Target detection method under view angle of unmanned aerial vehicle
By improving the YOLOv8 network model, the number of backbone network layers is reduced, the deformable attention mechanism and RepNCSPELAN4 module are introduced, and the loss function is improved, which solves the difficulty of object detection from the perspective of the drone and achieves high-precision small object detection effect.
Patent Information
- Application Number
- CN202411991979.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-12-31
AI Technical Summary
When performing target detection from the perspective of a drone, the targets to be detected are densely distributed and the pixels are small, the number of samples in various categories in the data center is uneven, and the drone hardware scale limits the size of the model, making it difficult for the detection results to meet the actual application needs.
Improve the YOLOv8 object detection network model, including reducing the number of backbone network layers, increasing the size of feature maps to be detected, introducing a deformable attention mechanism to the spatial pyramid pooling layer, replacing the C2f module with the RepNCSPELAN4 module, and improving the boundary regression loss function to inner-ShapeIoU.
While reducing network depth, the accuracy and performance of the model for detecting small objects is improved, missed and missed detection, and a small amount of model parameters is maintained.
Smart Images

Figure CN119919639A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of target detection technology, and in particular, relates to a target detection method from the perspective of an unmanned aerial vehicle. Background Art
[0002] With the development of drone technology, target detection through deep learning algorithms has been widely used in related fields. However, target detection from the perspective of drones faces difficulties such as dense distribution of targets to be detected, small pixels, and uneven number of samples in each category in the corresponding data set. At the same time, the hardware scale of drones limits the size of the model, making it difficult for the detection results to meet the actual application requirements.
[0003] At present, deep learning-based target detection algorithms are mainly divided into two categories: one is the two-stage region proposal algorithm, such as R-CNN, Fast R-CNN and Faster R-CNN, etc.; the other is the one-stage regression algorithm, such as SSD, YOLO series and RetinaNet, etc. Although the two-stage algorithm has high detection accuracy, the detection speed is slow and is not suitable for detection tasks under the perspective of drones; while the one-stage algorithm has a fast detection speed but relatively low accuracy, especially in the detection of small targets, which still needs further optimization.
[0004] The YOLO series of models performs well in various detection tasks, especially in the field of target detection from the perspective of drones. The YOLOv8 target detection model has five versions from small to large: v8n, v8s, v8m, v8l, and v8x. It is mainly composed of three parts: backbone network (Backbone), neck (Neck) and detection head (Head). The Backbone part is mainly responsible for feature extraction, including Conv convolution module, C2f module and SPPF module evolved from SPP in YOLOv5 and other architectures; the Neck part is mainly responsible for feature fusion, using C2f module and feature pyramid structure combining PAN and FPN. The Head part is gradually developed and integrated from YOLO-Head of the YOLO series, including three detection heads with feature maps of different sizes, which are used to detect target objects of different sizes. However, since the basic YOLOv8 model is designed based on general natural scenes, directly applying the original model to targets photographed by drones will result in a large amount of feature information loss in the sampling process, and the detection effect for small targets is poor, making it difficult to achieve ideal results. Summary of the invention
[0005] The purpose of this application is to provide a target detection method from the perspective of a drone, so as to solve the difficulties of target detection from the perspective of a drone, such as dense distribution of targets to be detected, small pixels, and uneven number of samples in each category in the corresponding data set. At the same time, due to the limitations of the hardware scale of the drone itself, it is necessary to ensure that the network model can still maintain a high detection accuracy with a small number of parameters.
[0006] In order to achieve the above purpose, the technical solution of this application is as follows: A method for detecting a target from a drone's perspective, comprising: Build and train an improved YOLOv8 target detection network model, including an improved backbone network, neck network, and detection head; The image to be recognized is input into the improved backbone network to extract multi-scale features. The improved backbone network removes the 7th and 8th layers of the original backbone network and introduces a deformable attention mechanism in the spatial pyramid pooling layer; The extracted multi-scale features pass through the neck network and the detection head to obtain the recognition results.
[0007] Furthermore, the multi-scale features extracted by the improved backbone network are the output features of the 2nd layer, the 4th layer and the spatial pyramid pooling layer after the deformable attention mechanism is introduced.
[0008] Furthermore, the multi-scale features extracted by the improved backbone network have sizes of 160×160×128 pixels, 80×80×256 pixels, and 40×40×512 pixels, respectively.
[0009] Furthermore, the C2f modules in the backbone network and the neck network are replaced by RepNCSPELAN4 modules.
[0010] Furthermore, when training the improved YOLOv8 target detection network model, the boundary regression loss function adopts ShapeIoU, and the Inner-IoU loss replaces the intersection-over-union (IoU) in the ShapeIoU loss function.
[0011] This application proposes a method for target detection from the perspective of a drone. Taking into account the hardware limitations of the drone itself, the YOLOv8s network structure with a smaller scale and higher detection accuracy is used as a benchmark. By reducing the number of layers of the backbone network and increasing the size of the feature map to be detected, a feature layer containing rich semantic information of small targets is added, thereby reducing the network depth while enabling the model to focus more on the detection of small targets; the RepNCSPELAN4 structure based on the GELAN design concept is used to replace the C2f structure in the original model. This structure allows it to support multiple types of computing blocks, so that it can better adapt to different computing and hardware requirements, and effectively improves computing efficiency and performance; a deformable attention mechanism DAT is introduced in the spatial pyramid pooling layer SPPF, and SPPF_DAT is proposed to enable the network to dynamically select sampling points, so that the model can focus more on the most important area in the current task; the ShapeIoU function is improved using the idea of the inner-IoU loss function, and the original loss function is replaced by inner-ShapeIoU. The loss is calculated by paying attention to the shape and scale of the bounding box itself, thereby improving the algorithm's positioning and classification effect on small targets in complex environments. This application improves the small target detection performance without increasing the overhead of complex algorithms, thereby achieving a better balance between model complexity and accuracy; the detection results are more accurate, the model parameters are reduced, and there are fewer missed detections and false detections. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is a flow chart of the target detection method from the perspective of a drone for this application.
[0013] Figure 2 This is the YOLOv8 network structure diagram.
[0014] Figure 3 This is a diagram showing the structure of the target detection network model for improving YOLOv8 in the embodiments of the present application.
[0015] Figure 4 This is a comparison chart of the target detection network model proposed in this application and the performance indicators of YOLOv8. DETAILED DESCRIPTION
[0016] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0017] Embodiment 1, as Figure 1 As shown, a target detection method from the perspective of a drone is provided, including: Step S1: construct and train an improved YOLOv8 target detection network model, including an improved backbone network, a neck network, and a detection head.
[0018] The target detection network model constructed in this application is improved on the basis of YOLOv8. Figure 2 , including the backbone network (Backbone), the neck network (Neck) and the detection head (Head). The backbone network has 10 layers, from layer 0 to layer 9, layers 0, 1, 3, 5, and 7 are convolutional layers, layers 2, 4, 6, and 8 are C2f modules, and the 9th layer is the spatial pyramid pooling layer SPPF; the neck network is responsible for multi-scale feature fusion, and enhances the feature representation capability by fusing feature maps from different stages of the backbone network; the detection head is responsible for the final target detection.
[0019] It should be noted that YOLOv8 is a relatively mature technology in this field. Its backbone network and bottleneck network include multiple convolution and C2f modules respectively, which are not described here one by one.
[0020] This embodiment improves the original backbone network of YOLOv8, and the improved backbone network is used for multi-scale feature extraction.
[0021] Step S2: input the image to be recognized into the improved backbone network to extract multi-scale features. The improved backbone network removes the 7th and 8th layers of the original backbone network and introduces a deformable attention mechanism in the spatial pyramid pooling layer.
[0022] This embodiment improves the original backbone network of YOLOv8, with the aim of reducing the number of backbone network layers and increasing the size of the feature map to be detected on the basis of YOLOv8, and adding a feature layer containing rich semantic information of small targets, thereby reducing the network depth while enabling the model to focus more on the detection of small targets.
[0023] Specifically, the input image of the YOLOv8 model is 640×640×3 pixels, and the sizes of the three feature maps finally generated are 80×80×256 pixels, 40×40×512 pixels, and 20×20×1024 pixels. Its network structure is as follows: Figure 2 As shown in the figure, they correspond to the outputs of the 4th, 6th and 9th layers respectively. When the targets to be detected are mainly small targets, the minimum feature map size is compressed to 20×20×1024 pixels, which may lead to significant loss of small target feature information. In addition, after deep convolution, small target features are easily submerged, thus affecting the detection accuracy.
[0024] To this end, in the first step of this embodiment, the convolution (Conv) and C2f layers of stage 4 (Stage Layer 4) in the backbone network are removed to reduce the impact of deep convolution, that is, the 7th and 8th layers of the original backbone network are removed.
[0025] Then, the deformable attention mechanism DAT was introduced in the spatial pyramid pooling layer SPPF. The spatial pyramid pooling layer (SPPF_DAT) after the introduction of the deformable attention mechanism enables the network to dynamically select sampling points, allowing the model to focus more on the most important areas in the current task.
[0026] When the image reaches the SPPF layer, the result obtained by processing the original SPPF module through multiple pooling layers is first input into the 4×4 non-overlapping convolution with a stride of 4 in the DAT module, and then the size of the input image is compressed to 1 / 4 of the original through the normalization layer, and in the last two stages of the hierarchical feature pyramid, the information is aggregated locally through the window-based local attention through continuous local attention and deformable attention modules, and then the global relationship between local enhanced tags is modeled using deformable attention blocks. This alternative design of attention blocks with local and global receptive fields can collect more effective features without incurring excessive memory and computational costs.
[0027] The feature map size output by the SPPF_DAT module in this embodiment is 40×40×512 pixels.
[0028] In a specific embodiment, the multi-scale features extracted by the backbone network of this embodiment are the output features of the 2nd layer, the 4th layer and the spatial pyramid pooling layer after the deformable attention mechanism is introduced, and the output feature map sizes are 160×160×128 pixels, 80×80×256 pixels, and 40×40×512 pixels, respectively.
[0029] It is worth mentioning that due to the reduction in the number of layers of the backbone network, in order to match the output length and width of the feature map corresponding to the Concat connection layer in the neck network, the number of layers of the output feature map is changed from the original 4, 6, and 9 layers to 2, 4, and 7 layers.
[0030] This application enlarges the multi-scale feature size to 160×160×128 pixels, 80×80×256 pixels, and 40×40×512 pixels. Through this improvement, the model can better retain the feature information of small targets.
[0031] Step S3: The extracted multi-scale features are passed through the neck network and the detection head to obtain the recognition result.
[0032] After the backbone network extracts multi-scale features, they are input into the bottleneck network and then pass through the detection head to output the final detection results.
[0033] Example 2 provides a method for target detection from the perspective of a drone. The difference from Example 1 is that the C2f module in the backbone network and the neck network of this embodiment is replaced by the RepNCSPELAN4 module.
[0034] RepNCSPELAN4 is the feature extraction-fusion module in YOLOv9. It uses the GELAN design concept and supports multiple types of computing blocks, so that it can better adapt to different computing and hardware requirements and effectively improve computing efficiency and performance.
[0035] GELAN combines the design concepts of CSPNet and ELAN, draws on the segmentation and reorganization concepts of CSPNet, and introduces ELAN's hierarchical convolution processing method in each part, thereby reducing the amount of computation and memory usage, and improving the inference speed and accuracy.
[0036] Example 3 provides a target detection method from the perspective of a drone. Different from Example 1, in the detection head of the network model, inner-ShapeIoU is used instead of the original loss function to improve the algorithm's positioning and classification effects on small targets in complex environments.
[0037] The purpose of bounding box regression is to adjust the detection window output by the detector so that it is closer to the actual detection window. Since its proposal, the intersection over union (IoU) has become the main criterion for evaluating the prediction box loss in the field of object detection. Its calculation formula is shown in formula (1):
[0038] ⑴ in Represents the calculation result of the prediction box, Represents the calculation result of the true annotation box. IoU measures the degree of matching by calculating the ratio of the intersection and union between the predicted box and the true annotation box. The bounding box regression loss function based on IoU has undergone continuous iteration and development, and has derived multiple versions, such as DIoU, GIoU, CIoU, EIoU, and SIoU. These improved loss functions accelerate convergence by introducing new loss terms, but do not take into account the limitations of IoU itself.
[0039] Inner-IoU based on IoU loss of auxiliary borders proposes to use auxiliary borders to calculate IoU to improve the generalization ability of the model. The specific calculation process is shown in equations (2) to (6), where It represents the center point of the anchor point and the internal anchor point, while the width and height of the anchor point are represented by w and h, and the scale factor ratio is used to adjust the size of the auxiliary border.
[0040] (2) (3) By using equations (2) and (3), the center point of the detection box can be transformed to obtain the corner points of the auxiliary detection box. At the same time, the predicted box and the real box output by the model are transformed accordingly, using and Represents the calculation results of the real box and the predicted box, and uses min and max to obtain the minimum and maximum values in brackets.
[0041] (4) (5) (6) Therefore, as shown in equations (4) to (6), inner-IoU actually calculates the IoU between auxiliary bounding boxes. ratio∈[0.5,1.5] and When , the size of the auxiliary border is smaller than the actual border, which causes the effective range of regression to be smaller than the IoU loss, but the absolute value of the obtained gradient is larger than the gradient of the IoU loss, thereby accelerating the convergence of high IoU samples. When , the size of the auxiliary bounding box is larger than the actual box, which expands the effective range of regression and helps the regression of low IoU samples.
[0042] ShapeIoU takes into account the influence of the shape and scale of the bounding box itself, and samples itself in the bounding box regression, making the bounding box regression more accurate. The specific calculation process is shown in equations (7) to (10).
[0043] (7) (8) (9) (10) in , Represents the weight coefficients in the vertical and horizontal directions respectively, and its value is related to the value of the GT border. Scale is the proportional factor. Represents the shape cost.
[0044] This embodiment adopts the idea of inner-IoU to transform ShapeIoU, and effectively improves the detection effect by replacing the IoU calculation part. The inner-ShapeIoU calculation formula is shown in formula (11): (11) in, ∈ R ∈ ∈ ℓ ℓ , ∈ ∈ ℓ ℓ , represents the inner-ShapeIoU loss function of this embodiment, which is used in the detection head for bounding box regression. Represents the Inner-IoU loss function.
[0045] like Figure 3 As shown, the target detection network model constructed by this application is referred to as DAT-YOLO, that is, a deformable attention mechanism DAT is introduced on the basis of YOLOv8. Specifically, on the basis of YOLOv8, the number of layers of the backbone network is reduced and the size of the feature map to be detected is increased, and a feature layer containing rich semantic information of small targets is added, so that the model can be more focused on the detection of small targets while reducing the network depth. The RepNCSPELAN4 structure using the GELAN design concept replaces the C2f structure in the original model. This structure allows it to support multiple types of computing blocks, so that it can better adapt to different computing and hardware requirements, and effectively improves computing efficiency and performance. The deformable attention mechanism DAT is introduced in the spatial pyramid pooling layer SPPF, and SPPF_DAT is proposed to enable the network to dynamically select sampling points, so that the model can focus more on the most important areas in the current task. The idea of the inner-IoU loss function is used to improve the ShapeIoU function. The inner-ShapeIoU loss function is used to replace the original loss function. The loss is calculated by focusing on the shape and scale of the bounding box itself, which improves the algorithm's positioning and classification of small targets in complex environments.
[0046] When training the constructed target detection network model, collect the image dataset taken by the drone and preprocess it, including converting the annotation file format and applying data enhancement methods. The image size in the training set is adjusted to 640×640 as the input of the network model, and the Mosaic data enhancement method is used to expand the dataset. Set the training epoch to 600 and the patience to 30 to ensure that the model training stops automatically within 30 rounds without better results.
[0047] After the training is completed, the corresponding weight model and the corresponding training results are saved. To evaluate the performance of the improved model, P (precision), R (recall), mAP@50 (multi-category average precision at 50% IoU threshold), and mAP@50:95 (multi-category average precision at 50%-95% IoU threshold) are used as evaluation indicators, and the results of YOLOv8 and DAT-YOLO network training are compared. The results are shown in the figure. Figure 4As shown in the figure, it can be seen that DAT-YOLO has better overall performance and the mAP curve is smoother.
[0048] Finally, the model obtained after training the YOLOv8, PVswin-YOLOv8, GAM-YOLO and DAT-YOLO networks proposed in this application is used to detect objects in the same images. Under different lighting conditions, object numbers and sizes, the models proposed in this application can obtain good detection results. Several typical scenarios are listed below.
[0049] The first scene: a sparse small target scene with sufficient lighting.
[0050] Experimental results show that GAM-YOLO and DAT-YOLO with a small target detection layer can recognize the truck in the distance, while the original model and PVswin-YOLOV8s cannot.
[0051] The second scene: a multi-category target scene with sufficient lighting.
[0052] The original model misclassified the roadblock as a pedestrian and only identified the pedestrian nearby; however, DAT-YOLO did not make any misclassifications and not only identified the pedestrian but also detected the bicycle under the pedestrian.
[0053] The third scene: a scene with dense small targets at night.
[0054] Compared with the recognition results of the original model, PVswin-YOLOv8s and GAM-YOLO, DAT-YOLO can also more accurately detect more small targets, such as smaller vehicles and bicycles in the distance, effectively reducing missed detections and false detections.
[0055] From the results, it can be seen that compared with other mainstream models, this application is more suitable for UAV target detection tasks.
[0056] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.
Claims
1. A method for detecting a target from the perspective of a drone, characterized in that: The target detection method from the perspective of the drone includes: Build and train an improved YOLOv8 target detection network model, including an improved backbone network, neck network, and detection head; The image to be recognized is input into the improved backbone network to extract multi-scale features. The improved backbone network removes the 7th and 8th layers of the original backbone network and introduces a deformable attention mechanism in the spatial pyramid pooling layer; The extracted multi-scale features pass through the neck network and the detection head to obtain the recognition results.
2. The target detection method from the perspective of a drone according to claim 1, characterized in that: The multi-scale features extracted by the improved backbone network are the output features of the 2nd layer, the 4th layer and the spatial pyramid pooling layer after the deformable attention mechanism is introduced.
3. The target detection method from the perspective of a drone according to claim 2, characterized in that: The multi-scale features extracted by the improved backbone network have sizes of 160×160×128 pixels, 80×80×256 pixels, and 40×40×512 pixels, respectively.
4. The target detection method from the perspective of a drone according to claim 1, characterized in that: The C2f modules in the backbone network and the neck network are replaced by the RepNCSPELAN4 modules.
5. The target detection method from the perspective of a drone according to claim 1, characterized in that: When training the improved YOLOv8 target detection network model, the boundary regression loss function adopts ShapeIoU, and the Inner-IoU loss replaces the intersection-over-union (IoU) in the ShapeIoU loss function.
Citation Information
Patent Citations
Lane line detection method based on deformable attention mechanism
CN116524449A
Small target detection method for images acquired by unmanned aerial vehicle based on improved YOLOv8 algorithm
CN118628939A
Improved YOLOv8 tower foundation target detection method and device
CN119048736A