Two-uav two-stage target detection method based on improved yolov8
By improving the dual-UAV two-stage target detection method of YOLOv8, and utilizing dual-UAV perspective image fusion and feature pyramid network, the problems of missed detection and false detection of UAV target detection in complex scenarios are solved, and efficient and accurate target detection is achieved.
Patent Information
- Application Number
- CN202511328150.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-17
AI Technical Summary
Existing UAV target detection methods struggle to accurately capture target features in complex scenarios, resulting in missed detections and false detections. Furthermore, they suffer from high computational complexity, low detection efficiency, and an inability to adaptively adjust the processing flow.
A two-stage target detection method based on improved YOLOv8 is adopted, which uses images from different perspectives captured by two drones for fusion, combines feature matching and multi-scale feature pyramid network, and adjusts the processing flow according to the image complexity through a two-stage detection strategy. Lightweight network and feature attention mechanism are used to improve detection accuracy and efficiency.
It improves the accuracy and efficiency of target detection, can adapt to complex conditions in different scenarios, reduces missed detections and false detections, and enhances the detection effect in complex environments such as occlusion and changes in lighting.
Smart Images

Figure CN120833566B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of unmanned aerial vehicle vision detection, and particularly relates to a two-stage target detection method based on improved YOLOv8 for double unmanned aerial vehicles. BACKGROUND
[0002] With the rapid development of unmanned aerial vehicle technology, unmanned aerial vehicles have been widely used in geographic mapping, environmental monitoring, security inspection, disaster rescue and other fields. Using unmanned aerial vehicles to carry cameras for target detection is one of the important application directions. Through analyzing the images taken by the unmanned aerial vehicle, the position, quantity, state and other information of the target can be quickly obtained. However, in actual application, the images taken by the unmanned aerial vehicle often face complex scene changes, such as unstable lighting conditions, large target size differences, and many background interferences, which bring great challenges to target detection.
[0003] Traditional monocular unmanned aerial vehicle target detection methods only rely on the analysis of images taken by a single camera. Due to the lack of multi-view information, it is difficult to accurately capture the characteristics of the target when dealing with complex scenes, and it is easy to miss detection and false detection. Although the target detection method based on multi-camera can obtain more image information, there are problems of high computational complexity and low detection efficiency in the image fusion and feature extraction process. In addition, most of the existing target detection algorithms adopt a fixed detection process, which cannot adaptively adjust according to the complexity of the image, resulting in waste of computing resources when dealing with simple images, and insufficient detection accuracy when dealing with complex images. Therefore, there is an urgent need for a target detection method that can combine the advantages of multi-camera, intelligently process according to the complexity of the image, and improve the detection efficiency and accuracy. SUMMARY
[0004] The purpose of the present application is to overcome the defects of the prior art and provide a two-stage target detection method based on improved YOLOv8 for double unmanned aerial vehicles, which uses visible light photos taken by two unmanned aerial vehicle cameras and combines a two-stage detection strategy to achieve efficient processing of large and small images of different targets and improve the accuracy and efficiency of target detection.
[0005] The technical scheme provided by the present application is as follows:
[0006] A two-stage target detection method based on improved YOLOv8 for double unmanned aerial vehicles, comprising the following steps:
[0007] Step 1: Simultaneously taking visible light images of different angles of the target area through the cameras on the two unmanned aerial vehicles to obtain a first angle image and a second angle image of the target area;
[0008] Among them, the first angle image and the second angle image have an overlapping area;
[0009] Step two, fusing the first view image and the second view image to obtain a fused image;
[0010] Step three, inputting the fused image into a first target detection model for target detection; determining the area S of each detection target frame in the fused image, and comparing the area S of the detection target frame with a target frame area threshold T_simple;
[0011] If the area of each detection target frame in the fused image satisfies S>T_simple, the detection result output by the first target detection model is used as the target detection result;
[0012] If there is a detection target frame area S≤T_simple in the fused image, the fused image is input into a second target detection model for target detection, and the target detection result is output by the second target detection model;
[0013] The first target detection adopts a YOLOv8 model, and the second target detection model adopts an improved YOLOv8 model.
[0014] Preferably, the method for obtaining the second target detection model is:
[0015] The PANet module in the YOLOv8 model is replaced by a BiFPN module, and a CBAM attention mechanism is embedded at the nodes of the BiFPN module.
[0016] Preferably, in the step one, the two unmanned aerial vehicles always maintain a distance of 5-20 meters for flight, and a time synchronization module of the unmanned aerial vehicle is used to ensure that the difference between the shooting moments of the two cameras is less than 10 milliseconds.
[0017] Preferably, in the step two, fusing the first view image and the second view image includes the following steps:
[0018] Step 1, matching the feature points in the first view image and the second view image by using a SIFT feature matching algorithm, and removing the mis-matched feature points by using a RANSAC algorithm;
[0019] Step 2, establishing a spatial transformation mathematical transformation model by using the matched feature points, and converting the first view image and the second view image into the same spatial coordinate system according to the mathematical transformation model to obtain the registered first view image and second view image;
[0020] Step 3, fusing the pixels in the overlapping area of the registered first view image and second view image by using a weighted average method, and retaining the corresponding pixels in the non-overlapping area;
[0021] Step 4, the Gaussian blur algorithm is used to smooth the fusion region boundary to obtain a fused image.
[0022] Preferably, in the step 1, the image to be matched is divided into 4*4 different regions, and feature point matching is performed on each region separately.
[0023] Preferably, in the step 2, the least square method is used to calculate the perspective transformation matrix, and the perspective transformation matrix is used to perform perspective transformation on the first view image, so that the transformed first view image and the second view image are aligned in the same coordinate system to obtain the registered first view image and the second view image.
[0024] Preferably, in the step 3, the formula for weighted fusion of the pixels in the overlapping region is:
[0025] ;
[0026] wherein, 、 are the pixel values of the pixel points to be fused in the first view image and the second view image respectively; is the distance from the pixel point to be fused to the edge of the first view image, is the distance from the pixel point to be fused to the edge of the second view image.
[0027] The beneficial effects of the present application are:
[0028] Improve detection accuracy: by using the double-view images taken by two unmanned aerial vehicle cameras, more target information is obtained; the correlation and fusion of multiple targets can accurately correlate the target information under multiple views and integrate the detection results of multiple views to further improve the accuracy of target detection; combined with feature enhancement, matching and multi-scale image fusion, the target features are more accurately extracted, and the missed detection and false detection situations are reduced.
[0029] Improve detection efficiency: the two-stage detection strategy can automatically adjust the processing flow according to the target size of the image, avoid using complex detection algorithms when processing simple images, waste computing resources, and effectively improve the overall detection efficiency; at the same time, the application of lightweight target detection network further speeds up the detection speed of simple images.
[0030] Strong adaptability: the present application can adapt to target detection tasks in different scenes, whether the light condition is unstable, the target scale difference is large or the background interference is complex, the feature pyramid can be used for fusion to highlight the target features and suppress the background interference, accurate target detection can be realized, and the present application has strong adaptability and universality. BRIEF DESCRIPTION OF DRAWINGS
[0031] Figure 1This is a flowchart of the dual-UAV two-stage target detection method based on improved YOLOv8 described in this invention.
[0032] Figure 2 This is a flowchart of the dual UAV image fusion method described in this invention.
[0033] Figure 3 This is a flowchart of the method for determining large and small targets as described in this invention.
[0034] Figure 4 This is a schematic diagram of the two-stage small target detection process described in this invention. Detailed Implementation
[0035] The present invention will now be described in further detail with reference to the accompanying drawings, so that those skilled in the art can implement it based on the description.
[0036] like Figure 1 As shown, this invention provides a two-stage target detection method for dual UAVs based on an improved YOLOv8, and the specific implementation process is as follows.
[0037] S1. Use two drones simultaneously to capture visible light images of the target area from different perspectives, obtaining first-view and second-view images of the target area.
[0038] Using a combination of the UAV's inertial navigation system (INS) and GPS for positioning, the UAVs undergo high-precision positioning calibration using an RTK base station before flight, ensuring that the images captured by the two UAVs have an accurate correspondence in the geographic coordinate system. During flight, the two UAVs maintain a distance of 5 to 20 meters, and the time synchronization module of the UAVs ensures that the difference in the shooting time of the two cameras is less than 10 milliseconds, avoiding positional deviations of moving targets due to time differences. The acquired image data is stored locally on an SD card, and low-resolution preview images are transmitted back to the ground station in real time for operators to monitor the shooting quality.
[0039] Dual-drone imaging effectively overcomes the limitations of a single drone's field of view. In complex scenarios such as occlusion, insufficient lighting, and low resolution, it provides a crucial prerequisite for solving image parallax and occlusion problems through complementary information from multiple perspectives. This project utilizes dual-drone image acquisition to construct a precise multi-target feature matching and fusion mechanism, thereby achieving efficient integration of dual-view detection results.
[0040] Step S2: Use the SIFT algorithm to extract feature points from the first-view image and the second-view image.
[0041] like Figure 2 As shown, in this embodiment, the first-view image and the second-view image are respectively divided into 4... 4 regions, and the SIFT algorithm is used for feature point extraction in each region to make the number of extracted feature points moderate and uniformly distributed. Then, the RANSAC algorithm is used to match the extracted feature points and eliminate the incorrectly matched feature points.
[0042] Calculate the transformation matrix: a mathematical transformation model of the unified space is established based on the matched feature points to realize the unified coordinate transformation of the images. In this embodiment, the least squares method is used to calculate the perspective transformation matrix, which is a 3x3 matrix. According to the calculated perspective transformation matrix, the first view image (image L) is subjected to perspective transformation, so that the first view image (image L) and the second view image (image R) are aligned in the same coordinate system, and the registered image is obtained.
[0043] Region division: the registered first view image (image L) and the second view image (image R) are divided into overlapping regions and non-overlapping regions, and the overlapping regions are the intersection part of the two images.
[0044] Overlapping region fusion: the weighted average method is used for fusion in the overlapping region, and the weight is determined according to the distance of the pixel point to the edge of the overlapping region. The closer the distance, the greater the weight.
[0045] For a certain pixel point in the overlapping region, if the distance of the pixel point to the edge of the first image is , and the distance of the pixel point to the edge of the second image is , then the pixel value of the pixel point after fusion is:
[0046] ;
[0047] wherein, , are the pixel values of the pixel point in the first image and the second image, respectively.
[0048] For the non-overlapping region, the pixel information of the corresponding image is directly retained.
[0049] The Gaussian blur algorithm is used to smooth the boundary of the fused region, and the standard deviation of the Gaussian filter is 0.5, so that the boundary transition is more natural, and the final fused image is obtained.
[0050] Step S3, input the fused image into the YOLOv8 model for target detection
[0051] For example, Figure 3As shown, the YOLOv8 model (first target detection model) outputs the bounding box information of each detection target, that is, the coordinates of the detection target in the image, which is usually represented by the left upper corner and the right lower corner coordinates, the left upper corner coordinates being (x1, y1) and the right upper corner coordinates being (x2, y2). Based on these bounding box data, the area of the target can be further calculated, and the specific method is as follows:
[0052] At the image pixel level, if the bounding box is a rectangle, its pixel (target frame) area can be calculated by the formula:
[0053] ;
[0054] This area reflects the pixel proportion of the detection target in the image, which can be used to judge the size of the target relative to the image.
[0055] If it is necessary to convert the pixel area into the actual physical area, the parameters of the image acquisition device (such as focal length, shooting distance, sensor size, etc.) or the size of the known reference object in the scene need to be combined to derive the result through geometric conversion formula.
[0056] After calculating the area of each detection target frame in the fusion image, a judgment module is introduced. The calculated area S of each target frame is compared with the preset target frame area threshold T_simple, and large and small targets are judged. If the area S of all target frames in the fusion image satisfies S>T_simple, it is judged that all detection targets in the fusion image are large targets, and the result after detection by the YOLOv8 model is directly used. If there is at least one target frame area S in the fusion image that satisfies S≤T_simple, it is judged that there is a small target in the detection target of the fusion image, and the fusion image enters the second stage complex processing.
[0057] Step S4, second stage complex processing (for complex small target image)
[0058] The second stage adopts a second target detection model (improved YOLOv8 model) for target detection. As shown, Figure 4 The original PANet module of the YOLOv8 model is removed, and a BiFPN module is used instead of the PANet module as the core fusion structure, and the feature extraction capability of the backbone network is retained. Based on the {P3, P4, P5} feature layers output by the backbone network, a P2 layer (with a resolution of 2 times that of P3) is added through shallow feature mapping, forming a four-level feature system {P2, P3, P4, P5}, wherein P2 is responsible for capturing the fine spatial information of small targets.
[0059] The fusion image in which small targets are detected in the first stage is input into the second target detection model (improved YOLOv8 model) as the initial feature for the second stage processing.
[0060] Top-down path: starting from the highest level P5, the feature map size is enlarged to match P4 by upsampling, and the new feature P5→P4 is generated after the weighted fusion of P4 original features; similarly, P4→P3 and P3→P2 are generated in turn, and the high-level semantic information is passed to the low layer.
[0061] Bottom-up path: starting from the P2 fusion feature, the size is reduced to match P3 by downsampling (such as max pooling), and the P2→P3 is generated after the weighted fusion of P3 original features and P4→P3 fusion features; similarly, P3→P4 and P4→P5 are generated in turn, and the low-level spatial details are supplemented to the high layer.
[0062] Repeated stacking: according to the accuracy requirement, 2-3 layers of complete bidirectional path are stacked, and each layer is based on the output features of the previous layer for iterative fusion, gradually strengthening the cross-scale feature correlation.
[0063] The BiFPN bidirectional feature pyramid network is adopted, and through the bidirectional feature interaction from top to bottom and from bottom to top, the multi-scale feature fusion can be more effective, the detection ability of small targets can be enhanced, and the network structure is simplified.
[0064] CBAM feature attention mechanism is added after the node of BiFPN module for feature fusion of different scales, and CBAM attention mechanism is through channel attention module and spatial attention module respectively.
[0065] The channel attention mechanism is to compress the feature map in the spatial dimension to obtain a one-dimensional vector and then operate. When compressing in the spatial dimension, not only the average value pooling (Average Pooling) is considered, but also the maximum value pooling (Max Pooling). Average pooling and maximum pooling can be used to aggregate spatial information of feature mapping, and sent to a shared network to compress the spatial dimension of input feature map, and element-wise sum is merged to produce a channel attention map. Just for a picture, the channel attention pays attention to which content on the picture is important.
[0066] The channel attention mechanism can be expressed as:
[0067]
[0068] where F is the input feature map, respectively represent the global average pooling and maximum pooling operation, represents a multi-layer perceptron, represents a Sigmoid activation function.
[0069] The spatial attention mechanism of CBAM (Convolutional Block Attention Module) is an important component, which is mainly used to weight the feature map in the spatial dimension to highlight key spatial positions. The main goal of the spatial attention module is to explicitly model the dependency between spatial positions, generate a spatial attention map, and thus focus on important spatial position features.
[0070] The spatial attention mechanism can be expressed as:
[0071] ;
[0072] Wherein, F is the input feature map, is a sigmoid function, is a 7 size of the convolution kernel. Indicates that the average pooling and maximum pooling results are spliced along the channel axis.
[0073] The overall attention process can be summarized as:
[0074] ;
[0075] ;
[0076] Wherein, indicates the input feature map; is a channel attention generation function for generating channel attention weights; indicates element-wise multiplication, is the feature map adjusted by the channel attention; is a spatial attention generation function for generating spatial attention weights; is the final feature map adjusted by the channel and spatial attention.
[0077] The feature map enhanced by BiFPN and CBAM is transmitted to the detection head, and through boundary box regression and class prediction, the accurate coordinates of small targets, class labels and confidence are output, and the two-stage detection process is completed.
[0078] After obtaining the detection frame of the small target, each detection target appears as a single detection frame, which can be visualized in the image after spatial unification.
[0079] The above embodiments show in detail the complete implementation process of the application in complex scenes, from device parameter selection to algorithm optimization details, which have been tested and verified, proving the effectiveness and practicality of the method in target detection in complex environments. In practical applications, parameters and algorithm configurations can be further adjusted according to specific scenes to achieve optimal detection results.
[0080] The improved YOLOv8-based two-unmanned aerial vehicle two-stage target detection method provided by the application has significant advantages in many aspects through innovative technical architecture and algorithm design.
[0081] In terms of detection accuracy, different view images are obtained by cameras on two unmanned aerial vehicles, and the feature pyramid network is used to extract different scale features, so that target information can be fully captured. The two-stage detection strategy can more accurately locate and identify small targets for complex images, and can effectively solve the problems of occlusion and similar background interference by realizing accurate multi-target association and fusing multi-view detection results.
[0082] In terms of efficiency, the two-stage detection mechanism intelligently allocates processing flow according to the complexity of the image. For simple images, a lightweight network and a fast fusion strategy are used to greatly improve the speed while ensuring detection effect. For small target images, the improved YOLOv8 model is used, which uses a bidirectional feature pyramid and a CBAM attention mechanism to control the computational overhead while ensuring accuracy, and the overall detection process achieves a good balance between real-time performance and accuracy.
[0083] In terms of scene adaptability, the application fully considers the characteristics of different application scenarios and effectively deals with challenges such as occlusion and light changes. This flexibility enables the application to be widely used in geographic mapping, security inspection, disaster rescue and other fields, and exhibits strong versatility.
[0084] Although the embodiments of the application have been disclosed as above, they are not limited to the application listed in the specification and embodiments, and can be fully applied to various fields suitable for the application, and additional modifications can be easily realized by those skilled in the art, therefore the application is not limited to specific details and the figures shown and described herein, without departing from the general concept defined by the claims and the equivalent scope.
Claims
1. A two-stage target detection method for dual UAVs based on improved YOLOv8, characterized in that, Includes the following steps: Step 1: Simultaneously capture visible light images of the target area from different perspectives using cameras on two drones to obtain first-view and second-view images of the target area. The first viewpoint image and the second viewpoint image have overlapping areas; Step 2: Fuse the first viewpoint image and the second viewpoint image to obtain a fused image; Step 3: Input the fused image into the first target detection model for target detection; determine the area S of each detected target box in the fused image, and compare the area S of the detected target box with the target box area threshold T_simple; If the area of each detected target box in the fused image satisfies S>T_simple, then the detection result output by the first target detection model is used as the target detection result. If there is a detection target box area S≤T_simple in the fused image, then the fused image is input into the second target detection model for target detection, and the target detection result is output through the second target detection model; Wherein, the first object detection uses the YOLOv8 model, and the second object detection model uses an improved YOLOv8 model; The method for obtaining the second object detection model is as follows: The PANet module in the YOLOv8 model is replaced with a BiFPN module, and the CBAM attention mechanism is embedded at the nodes of the BiFPN module.
2. The dual-UAV two-stage target detection method based on improved YOLOv8 according to claim 1, characterized in that, In step one, the two drones maintain a distance of 5 to 20 meters while flying, and the time synchronization module of the drones is used to ensure that the difference in the shooting time of the two cameras is less than 10 milliseconds.
3. The dual-UAV two-stage target detection method based on improved YOLOv8 according to claim 1 or 2, characterized in that, In step two, the first-view image and the second-view image are fused, including the following steps: Step 1: Use the SIFT feature matching algorithm to match feature points in the first-view image and the second-view image, and use the RANSAC algorithm to remove mismatched feature points. Step 2: Establish a spatial mathematical transformation model using matching feature points, and transform the first-view image and the second-view image into the same spatial coordinate system according to the mathematical transformation model to obtain the registered first-view image and second-view image. Step 3: The pixels in the overlapping areas of the registered first-view and second-view images are fused using a weighted average method, while the corresponding pixels in the non-overlapping areas are retained. Step 4: Use the Gaussian blur algorithm to smooth the boundaries of the fused regions to obtain the fused image.
4. The dual-UAV two-stage target detection method based on improved YOLOv8 according to claim 3, characterized in that, In step 1, the image to be matched is divided into 4*4 different regions, and feature point matching is performed separately for each region.
5. The dual-UAV two-stage target detection method based on improved YOLOv8 according to claim 4, characterized in that, In step 2, the least squares method is used to calculate the perspective transformation matrix, and the perspective transformation matrix is used to perform perspective transformation on the first-view image so that the transformed first-view image and the second-view image are aligned in the same coordinate system, thus obtaining the registered first-view image and the second-view image.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle cooperative target detection method
CN115294484A
Target detection method, device and system for hundred million-level pixel fusion video
CN119091109A