Small target detection method and system based on improved YOLOv8 model

By improving the backbone network, neck network and detection head of the YOLOv8 model, combined with WIOU loss function and data enhancement, the problem of low detection accuracy of small objects is solved, and higher detection accuracy and robustness are achieved.

CN120495839APending Publication Date: 2025-08-15ANHUI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510564669.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The detection accuracy of small and medium-sized objects in the prior art is low, especially in complex contexts, and the model does not fully learn small objects.

Method used

By improving the YOLOv8 model, the CBAM module is introduced to improve the backbone network to enhance feature extraction capabilities; the neck network is improved to better integrate multi-scale features; the detection head is added to detect large-scale features; the loss function is replaced as a WIOU loss function, and data augmentation training is performed.

Benefits of technology

Without adding model parameters, the accuracy and robustness of small object detection is improved, noise and occlusion can be better handled, and the detection ability of small objects is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495839A_ABST
    Figure CN120495839A_ABST
Patent Text Reader

Abstract

The invention discloses a small target detection method and system based on an improved YOLOv8 model, and belongs to the field of target detection. The method comprises the following steps: improving a YOLOv8 model: introducing a CBAM module to improve an original backbone network; the original neck network is improved; an original detection head is improved; the original loss function of the YOLOv8 model is replaced; a public data set is adopted to train the improved YOLOv8 model; obtaining a to-be-tested picture; and inputting a to-be-detected picture into the improved YOLOv8 model after training to perform small target detection, wherein a detection result comprises category information, confidence and a target position. According to the invention, the detection precision of small target detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of target detection, and specifically relates to a small target detection method and system based on an improved YOLOv8 model. Background Art

[0002] Small object detection is a key research topic in computer vision, involving the identification and localization of relatively small objects in videos. Small object detection is crucial in many practical applications, including but not limited to: identifying small roadside objects such as pedestrians, animals, and traffic signs in autonomous driving; detecting small suspicious objects in surveillance systems; identifying small objects such as ships and vehicles in remote sensing images for environmental monitoring; and detecting small lesions or tumors in medical imaging analysis.

[0003] However, small target detection also faces a series of technical challenges: small targets in the scale problem usually occupy fewer pixels in the image, and the amount of information is insufficient, making it difficult for the model to learn effective features. In the case of complex backgrounds, small targets are often obscured by the complex background and are easily misdetected or missed. In the training dataset, the number of small targets is usually far less than that of large targets, resulting in insufficient model learning of small targets. Small targets in the image may be difficult to detect due to blur or occlusion by other objects, and many other difficulties. Small target detection is an important and challenging research direction in the field of computer vision. With the continuous development of technology, research on small target detection will continue to advance in the future to promote its application in various fields.

[0004] To solve the problems of the prior art, the present invention proposes a small target detection method and system based on an improved YOLOv8 model. Summary of the Invention

[0005] The present invention aims to overcome the shortcomings of the existing technology and proposes a small target detection method and system based on an improved YOLOv8 model to achieve the following objectives: improve the detection accuracy of small target detection.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is: a small target detection method based on an improved YOLOv8 model, the method comprising the following steps:

[0007] Step S1: Improve the YOLOv8 model, including:

[0008] Introducing the CBAM module to improve the original backbone network;

[0009] Based on the improved backbone network, the original neck network is improved;

[0010] Based on the improved neck, the original detection head is improved;

[0011] Replace the original loss function of the YOLOv8 model;

[0012] Step S2: Use the public dataset to train the improved YOLOv8 model;

[0013] Step S3, obtaining the image to be tested;

[0014] Step S4: Input the image to be tested into the trained improved YOLOv8 model to perform small target detection. The detection results include category information, confidence level, and target location.

[0015] Preferably, the CBAM module is introduced to improve the original backbone network. The improved backbone network includes CBS1, CBS2, C2F1, CBAM1, CBS3, C2F2, CBAM2, CBS4, C2F3, CBAM3, CBS5, C2F4, and SPPF connected in series in sequence, wherein the outputs of CBAM1, CBAM2, CBAM3, and SPPF, as well as the outputs of C2F2 and C2F3, are respectively used as features extracted by the improved backbone network for subsequent feature fusion.

[0016] Preferably, the original neck network is improved based on the improved backbone network, and the improved neck network includes:

[0017] Upsampling structure: including Upsample1, Concat1, C2F5, Upsample2, Concat2, C2F6, Upsample3, Concat3, C2F7 connected in series;

[0018] Downsampling structure: including CBS6, Concat4, C2F8, CBS7, Concat5, C2F9, CBS8, Concat6, C2F10 connected in series; C2F7 is connected to CBS6;

[0019] Among them, Concat1 is connected to CBAM3; Concat2 is connected to CBAM2; Concat3 is connected to CBAM1; Concat4 is connected to C2F2 and C2F6 respectively; Concat5 is connected to C2F5 and C2F3 respectively; SPPF is connected to Upsample1 and Concat6 respectively.

[0020] Preferably, based on the improved neck, the original detection head is improved, that is, a detection head for detecting large-scale features is added. The detection head includes H1, H2, H3, and H4 in descending order according to the size of the processed feature scale, among which H1 is connected to C2F7, H2 is connected to C2F8, H3 is connected to C2F9, and H4 is connected to C2F10.

[0021] Preferably, the loss function of the original YOLOv8 model is replaced with the WIOU loss function.

[0022] Preferably, the step S2 includes performing data enhancement on the adopted public dataset to expand the dataset, and the data enhancement includes image translation, scaling, flipping, and mosaicking.

[0023] Preferably, the public dataset adopts the Visdrone2019 dataset.

[0024] The present invention also proposes a small target detection system based on an improved YOLOv8 model. The system includes a camera, a processor, and a display screen. The camera is connected to the processor; the processor is connected to the display screen. The camera is configured to capture images to be tested and send them to the processor. The processor is configured to deploy a trained improved YOLOv8 model and apply it to small target detection in the images to be tested, obtaining detection results and sending them to the display screen. The display screen is configured to display the detection results. The processor uses an Orange Pi 5 Plus development board.

[0025] The technical effect of the present invention is as follows: in response to the problem of low small target detection accuracy in the prior art, the present invention improves the existing YOLOv8 model to increase the accuracy of small target detection without introducing too many parameters, and deploys it on the Orange Pi 5 Plus to verify its effectiveness, showing strong application potential. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a schematic diagram of the traditional YOLOv8 model structure;

[0027] Figure 2 A schematic diagram of the improved YOLOv8 model structure provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The following is a further detailed description of the specific embodiments of the present invention through the description of the embodiments with reference to the accompanying drawings. The purpose is to help those skilled in the art have a more complete, accurate and in-depth understanding of the inventive concept and technical solution of the present invention and to facilitate its implementation. It should be noted that the terms "first" and "second" described in this application are only used to facilitate the description of the technical solution to distinguish different components and are not intended to limit this application. To make the technical solution of the present invention clearer, the present invention is explained through the following embodiments.

[0029] As mentioned above, the existing YOLOv8 model includes a backbone network (backbone), a neck network (neck), and a detection head (head). Its structure is as follows Figure 1 As shown in the figure, it consists of seven "CBS" modules, eight "C2f" modules, one "SPPF" module, two "Unsample" modules, four "Concat" modules, and three detection heads (H1, H2, and H3). Although the YOLOv8 model has excellent object detection capabilities, its detection accuracy for small objects can no longer meet the increasingly high demand.

[0030] To improve the detection accuracy of small targets, this embodiment provides a small target detection method based on an improved YOLOv8 model, which includes the following steps:

[0031] Step S1: Improve the YOLOv8 model, including:

[0032] Introducing the CBAM (hybrid attention mechanism) module to improve the original backbone network;

[0033] Based on the improved backbone network, the original neck network is improved;

[0034] Based on the improved neck, the original detection head is improved;

[0035] Replace the original loss function of the YOLOv8 model;

[0036] Step S2: Use the public dataset to train the improved YOLOv8 model;

[0037] Step S3, obtaining the image to be tested;

[0038] Step S4: Input the image to be tested into the trained improved YOLOv8 model to perform small target detection. The detection results include category information, confidence level, and target location.

[0039] Specifically, in step S1, this embodiment improves the existing YOLOv8 model structure so that it can better learn the characteristics of small targets and achieve more accurate small target detection. The improved YOLOv8 model structure is as follows: Figure 2 shown.

[0040] First, the CBAM module is introduced to improve the original backbone network. The improved backbone network consists of CBS1, CBS2, C2F1, CBAM1, CBS3, C2F2, CBAM2, CBS4, C2F3, CBAM3, CBS5, C2F4, and SPPF, which are connected in series. The outputs of CBAM1, CBAM2, CBAM3, and SPPF, as well as the outputs of C2F2 and C2F3, are also used as features extracted by the improved backbone network for subsequent feature fusion.

[0041] Compared to the existing YOLOv8 model, this embodiment adds a CBAM module after the first three C2F modules in the backbone network. This effectively focuses on the most important information in the input data, thereby improving the model's ability to identify key features. By combining different types of attention mechanisms (such as spatial attention and channel attention), hybrid attention can extract features at different scales, enhancing the model's ability to detect small objects or details. The hybrid attention mechanism makes the model more robust in the face of challenges such as noise, blur, or occlusion, and can better handle incomplete or interfering information.

[0042] CBAM (Convolutional Block Attention Module) is a lightweight attention module designed to enhance the representation capabilities of convolutional neural networks (CNNs). CBAM improves the quality of feature representation by applying channel attention and spatial attention in parallel. The channel attention module aims to adjust the importance of each channel based on global information, while the spatial attention module aims to adjust the importance of each position based on spatial information. Specifically:

[0043] Assume that the input feature vector F∈R C*H*W , C represents dimension, H represents height, and W represents width. The channel attention module (ChannelAttention) process is as follows:

[0044] Average pooled feature vector:

[0045] Max pooling feature vector: F max =max i,j F i,j ;

[0046] Shared MLP: M C =σ(W2δ(W1F avg +W1F max ));

[0047] Reweighting: F c=M c ⊙F;

[0048] Where W1 and W2 are the weights of the shared MLP, δ is the ReLU activation function, σ is the Sigmoid activation function, and ⊙ is the element-wise multiplication.

[0049] The SpatialAttentionModule process is as follows:

[0050] Average pooling feature map:

[0051] Max pooling feature map:

[0052] Fusion features:

[0053] Convolution operation: M S =σ(Conv((F spatial )));

[0054] Reweighting: F CS =M S ⊙F c ;

[0055] Among them, Conv is a 1×1 convolution layer, F c is the feature map output by the channel attention module, M S is the spatial attention map, and ⊙ is the element-wise multiplication.

[0056] Next, this embodiment improves the original neck network based on the improved backbone network. The improved neck network includes:

[0057] Upsampling structure: including Upsample1, Concat1, C2F5, Upsample2, Concat2, C2F6, Upsample3, Concat3, C2F7 connected in series;

[0058] Downsampling structure: including CBS6, Concat4, C2F8, CBS7, Concat5, C2F9, CBS8, Concat6, C2F10 connected in series; C2F7 is connected to CBS6;

[0059] Among them, Concat1 is connected to CBAM3; Concat2 is connected to CBAM2; Concat3 is connected to CBAM1; Concat4 is connected to C2F2 and C2F6 respectively; Concat5 is connected to C2F5 and C2F3 respectively; SPPF is connected to Upsample1 and Concat6 respectively.

[0060] The neck network of the existing YOLOv8 model is a multi-scale fusion pyramid PAFPN structure, which includes two layers of upsampling structure and two layers of downsampling structure, namely:

[0061] In PAFPN, the calculation formula of the PA module is as follows:

[0062]

[0063] Among them, F PA is the feature map output by the PA module, F i (x) is the input feature map of different resolutions, m is the number of feature maps, α i is the weight of the feature map. The weight calculation formula of the PA module is:

[0064]

[0065] Among them, w i It is the weight coefficient calculated according to the size and resolution of the feature map.

[0066] Finally, the formula of PAFPN structure is obtained:

[0067] F PAFPN (x) = F PA (F Pi (x))

[0068] Among them, F Pi (x) is the i-th layer feature map output by the FPN module, F PAFPN (x) is the feature map output by the PAFPN module.

[0069] In this embodiment, in order to strengthen the learning of relatively low-level features and perform fusion so as to be suitable for the detection of small targets, on the basis of the neck network of the existing YOLOv8 model, a corresponding upsampling structure is added, namely, Upsample3, Concat3, and C2F7 connected in series, and a downsampling structure, namely, CBS6, Concat4, and C2F8 connected in series. At the same time, the input features of Concat3 include the output of CBAM1, that is, in the multi-scale feature fusion process of this embodiment, the large-scale feature map input by CBAM1 is introduced to better learn the features of small targets.

[0070] At the same time, the existing YOLOv8 model feeds the features extracted by C2F and SPPF in the backbone network into the upsampling structure for feature fusion, and the upsampling output is then fed into the downsampling for fusion. The features extracted by the backbone network are not utilized in the downsampling stage. In contrast, this embodiment replaces the features extracted by C2F in the backbone network of the existing YOLOv8 model with the output features of the CBAM module. Specifically, the outputs of CBAM1, CBAM2, and CBAM3 are fed into the upsampling structure for feature fusion. This allows the CBAM module to better capture important feature information, thereby improving the model's feature representation capability. The upsampling outputs C2F5, C2F6, and C2F7 are then fed into the downsampling for fusion. At the same time, the features extracted by C2F2 and C2F3 are introduced into the downsampling structure. Among them, Concat4 and Concat5 of the downsampling structure respectively introduce the outputs of C2F2 and C2F3 in the process of feature fusion with the features output by the upsampling structure. That is, Concat4 fuses the output features of C2F2, C2F6, and CBS6, and Concat5 fuses the output features of C2F5, C2F3, and CBS7, thereby making up for some local features missed by the CBAM module to obtain higher detection accuracy and improve the robustness of the model.

[0071] Then, since the existing YOLOv8 model has only three detection heads and lacks a detection head for large-scale features, this embodiment improves the original detection head based on the improved neck, that is, adds a detection head for detecting large-scale features. The detection heads include H1, H2, H3, and H4 in descending order according to the size of the processed feature scale. In this embodiment, the corresponding scales are 160*160, 80*80, 40*40, and 20*20, respectively. Among them, H1 is connected to C2F7, H2 is connected to C2F8, H3 is connected to C2F9, and H4 is connected to C2F10.

[0072] Different from the existing YOLOv8 model, this embodiment adds a 160*160 detection head H1, which retains more image detail features and has higher spatial information integrity, and can better detect small targets.

[0073] Finally, the loss function of the original YOLOv8 model is replaced with the WIOU loss function to enhance the learning of small objects. WIOU is a weighted loss function based on IoU. It addresses the importance differences between different pixels by assigning different weights to each pixel. Specifically, WIOU is calculated as follows:

[0074] L WIOU =r×L w1 ;

[0075] Where r is the gradient gain factor, Lw1 is the WIoU v1 loss.

[0076]

[0077] Among them, β is the outlier degree, α and δ are hyperparameters; R w is distance attention; L IoU is the bounding box loss IoU.

[0078]

[0079] in, Represents the average sliding loss, x, y are the horizontal and vertical coordinates of the center point of the prediction box, respectively, x gt 、y gt are the horizontal and vertical coordinates of the center of the real frame, c w 、c h are the width and height of the minimum bounding rectangle of the predicted and true boxes, respectively. * is a separation operation in the computation graph, which turns it into a constant without gradient. IoU stands for intersection over union.

[0080] After improving the YOLOv8 model, it needs to be trained and validated. This example uses the Visdrone2019 drone vision open dataset to train the improved YOLOv8 model, dividing the dataset into training, validation, and test sets at an 8:1:1 ratio. The Visdrone2019 dataset was collected by the AISKYEYE team at Tianjin University's Machine Learning and Data Mining Laboratory to provide a large-scale benchmark for drone vision systems. The dataset includes pedestrians, people, bicycles, cars, vans, trucks, tricycles, awning tricycles, buses, and motorcycles. Before training, the data needs to undergo a series of enhancement operations. This is because training requires a large amount of data support, but the number of small objects in existing datasets is typically far less than that of large objects, resulting in insufficient model learning of small objects. Therefore, this example introduces data augmentation in step S2. This data augmentation includes image translation, scaling, flipping, and mosaicking. Training with this enhanced dataset provides sufficient information about small objects, thereby improving the robustness and generalization of the model's detection capabilities.

[0081] During the training process, the model's hyperparameters need to be set. In this example, the input image size is set to 640, the epoch is set to 300 rounds, the batch size is set to 16, and the learning rate is set to 0.01. The platform used is: Linux system, the processor is an Intel (R) Xeon (R) Platinum 8160 CPU 2.10G, and the graphics card is two GeForce RTX 4090ti 24GB. During training, the existing YOLOv8 pre-trained weight file is used for transfer training to speed up the training. At the end of the training, the weight file is generated and the weight file with the highest detection accuracy is saved.

[0082] After verification, the detection performance of different algorithms on the Visdrone2019 dataset is shown in Table 1.

[0083]

[0084] Table 1

[0085] As can be seen from Table 1, the mAP@0.5 and mAP@0.5:0.95 of the improved YOLOv8 model are 0.071 and 0.042 higher than the original model, respectively. After changing the loss function to WIOU, the mAP@0.5 and mAP@0.5:0.95 of the model are 0.082 and 0.049 higher than the original model, respectively. Therefore, the present invention improves the accuracy of small objects without significantly affecting the detection speed.

[0086] At the same time, according to the above-mentioned small target detection method based on the improved YOLOv8 model, the present invention also proposes a small target detection system based on the improved YOLOv8 model, the system including a camera, a processor, and a display screen, the camera being connected to the processor; the processor being connected to the display screen, wherein: the camera is used to acquire the image to be tested and send it to the processor; the processor is used to deploy the trained improved YOLOv8 model, and apply it to perform small target detection on the image to be tested, obtain the detection results and send them to the display screen; the display screen is used to display the detection results.

[0087] In this embodiment, the processor uses the Orange Pi 5 Plus development board. The Orange Pi 5 Plus captures the image to be tested using an external camera and performs small target detection using a modified YOLOv8 model deployed on the Orange Pi 5 Plus. The trained model's pt file is converted into an onnx file, which is then converted into an rknn file. A dedicated inference engine is then used to accelerate inference on the file using the NPU, ultimately deploying the model to the Orange Pi 5 Plus. The Orange Pi 5 Plus then sends the detection results to the display, which displays them in a graphical user interface (GUI), including candidate box coordinates, target category information, and confidence levels. The currently detected image can also be saved locally.

[0088] The present invention has been described above with reference to the accompanying drawings. Obviously, the specific implementation of the present invention is not limited to the above-described method. Any non-substantial improvements made using the method concepts and technical solutions of the present invention, or any direct application of the above-described concepts and technical solutions to other situations without modification, fall within the scope of protection of the present invention.

Claims

1. A small target detection method based on an improved YOLOv8 model, characterized by: The method comprises the following steps: Step S1: Improve the YOLOv8 model, including: Introducing the CBAM module to improve the original backbone network; Based on the improved backbone network, the original neck network is improved; Based on the improved neck, the original detection head is improved; Replace the original loss function of the YOLOv8 model; Step S2: Use the public dataset to train the improved YOLOv8 model; Step S3, obtaining the image to be tested; Step S4: Input the image to be tested into the trained improved YOLOv8 model to perform small target detection. The detection results include category information, confidence level, and target location.

2. A small target detection method based on an improved YOLOv8 model according to claim 1, characterized in that: The CBAM module is introduced to improve the original backbone network. The improved backbone network includes CBS1, CBS2, C2F1, CBAM1, CBS3, C2F2, CBAM2, CBS4, C2F3, CBAM3, CBS5, C2F4, and SPPF connected in series. The outputs of CBAM1, CBAM2, CBAM3, and SPPF, as well as the outputs of C2F2 and C2F3, are respectively used as the features extracted by the improved backbone network for subsequent feature fusion.

3. A small target detection method based on an improved YOLOv8 model according to claim 2, characterized in that: Based on the improved backbone network, the original neck network is improved. The improved neck network includes: Upsampling structure: including Upsample1, Concat1, C2F5, Upsample2, Concat2, C2F6, Upsample3, Concat3, C2F7 connected in series; Downsampling structure: including CBS6, Concat4, C2F8, CBS7, Concat5, C2F9, CBS8, Concat6, C2F10 connected in series; C2F7 is connected to CBS6; Among them, Concat1 is connected to CBAM3; Concat2 is connected to CBAM2; Concat3 is connected to CBAM1; Concat4 is connected to C2F2 and C2F6 respectively; Concat5 is connected to C2F5 and C2F3 respectively; SPPF is connected to Upsample1 and Concat6 respectively.

4. A small target detection method based on an improved YOLOv8 model according to claim 3, characterized in that: Based on the improved neck, the original detection head is improved, that is, a detection head for detecting large-scale features is added. The detection heads include H1, H2, H3, and H4 in descending order according to the size of the processed feature scale. Among them, H1 is connected to C2F7, H2 is connected to C2F8, H3 is connected to C2F9, and H4 is connected to C2F10.

5. A small target detection method based on an improved YOLOv8 model according to any one of claims 1 to 4, characterized in that: Replace the loss function of the original YOLOv8 model with the WIOU loss function.

6. The small target detection method based on the improved YOLOv8 model according to claim 1, characterized in that: The step S2 includes performing data enhancement on the adopted public dataset, and the data enhancement includes image translation, scaling, flipping, and mosaicking.

7. A small target detection method based on an improved YOLOv8 model according to claim 1 or 6, characterized in that: The public dataset uses the Visdrone2019 dataset.

8. A small target detection system based on an improved YOLOv8 model according to the method of any one of claims 1 to 7, characterized in that: The system includes a camera, a processor, and a display screen, wherein the camera is connected to the processor; The processor is connected to the display screen, wherein: the camera is used to obtain the image to be tested and send it to the processor; The processor is used to deploy the trained improved YOLOv8 model and apply it to perform small target detection on the image to be tested, obtain the detection results and send them to the display screen; the display screen is used to display the detection results.

9. The small target detection system based on the improved YOLOv8 model according to claim 8, characterized in that: The processor uses the Orange Pi 5Plus development board.