Unmanned aerial vehicle small target data set construction and evaluation method

By constructing the GA-Fly dataset, we address the lack of standard datasets in the field of drone detection, provide multi-factor analysis, and improve the performance and adaptability of drone detection algorithms, especially for small targets and complex backgrounds.

CN120673190APending Publication Date: 2025-09-19ZHEJIANG UNIV CITY COLLEGE +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510618160.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The existing technology lacks widely recognized open source large-scale drone visual detection datasets and multi-factor analysis, resulting in the insufficient performance of deep learning methods in the field of drone detection.

Method used

A GA-Fly dataset was constructed, containing 10,800 4K resolution images covering a variety of shooting angles, backgrounds, lighting conditions, and flight altitudes. The performance of drone detection algorithms, shooting angles, grayscale sensitivity, target scale, and image resolution were evaluated through annotation and evaluation methods.

Benefits of technology

A diverse dataset is provided for evaluating drone detection algorithms, revealing key challenges in ground-to-air detection and providing effective suggestions for improvement, thereby improving the accuracy and adaptability of drone detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673190A_ABST
    Figure CN120673190A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle small target data set construction and evaluation method, the data set construction comprises at least 10, 800 images and corresponding. Txt annotation files, each image comprises a target unmanned aerial vehicle, the image resolution is 3840 * 2160 pixels, the images are sampled from a section of video shot at 25 frames per second according to the frequency of 10 frames per second, and the sampled images are stored in a database; frames which are not provided with unmanned aerial vehicles or are poor in quality due to shooting problems are eliminated, and all the images are labeled by using Labelimg. According to the method, a ground-to-air unmanned aerial vehicle detection data set focusing on small target detection is constructed, and compared with an existing data set, the method has higher pertinence. Based on the data set, the performance of seven representative deep learning algorithms is evaluated, and the influence of the shooting angle, the gray scale, the image resolution and the target scale on the performance of the deep learning network is analyzed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image dataset construction, and relates to a method for constructing and evaluating a small target dataset of an unmanned aerial vehicle. Background Art

[0002] Existing machine learning methods typically involve a two-step process, while deep learning methods directly output detection results through an end-to-end neural network. Leveraging the powerful feature extraction capabilities of CNNs, they avoid the complexity of manual feature design, resulting in greater versatility. However, CNNs are data-intensive and require large datasets for effective training. This study employs deep learning methods, focusing on drone visual detection in ground-to-air scenarios.

[0003] While deep learning methods have demonstrated impressive performance in various object detection tasks, their potential in drone detection has yet to be fully explored and validated. In recent years, most research has focused on ground-to-air drone detection. However, this research area still lacks a mature and comprehensive deep learning-based drone visual detection framework. A review reveals two major issues with previous research:

[0004] (1) There is a lack of a widely recognized open source large-scale dataset (such as COCO, VOC, etc.) as a standard.

[0005] (2) There is a lack of multi-factor analysis based on standard datasets, such as key factors such as lighting conditions, background complexity, target scale, and detection algorithm type. Summary of the Invention

[0006] To solve the above problems, the present invention provides a method for constructing and evaluating a small drone target dataset. The dataset construction includes at least 10,800 images and corresponding .txt annotation files. Each image contains a target drone, and the image resolution is 3840×2160 pixels. These images are sampled at a frequency of 10 frames per second from a video shot at 25 frames per second, and frames without drones or with poor quality due to shooting problems are eliminated. All images are annotated using Labelimg.

[0007] Preferably, the viewing angles of the images in the dataset are divided into three groups: 0°-30°, 30°-60°, and 60°-90°, with each group accounting for one-third of the entire dataset; wherein, the scenes in the 0°-30° range include houses, trees, and grass, the scenes in the 30°-60° range include tall buildings and large trees, and the scenes in the 60°-90° range include clear skies or layers of clouds.

[0008] Preferably, the area occupied by the targets in one-third of the images in the dataset is less than 1 / 160 of the entire image; the scale of the targets in one-half of the images is between 1 / 160 and 1 / 80 of the entire image; the area of ​​the targets in one-sixth of the images exceeds 1 / 80 of the entire image; when the height and width of the target drone are both less than 10% of the image, it is considered a small target.

[0009] Preferably, the dataset evaluates the illumination condition of the image by analyzing the average grayscale value of the image, and the grayscale value is calculated as follows:

[0010]

[0011] Among them, the grayscale value Gray of the image is calculated based on the intensity values ​​of the red R, green G, and blue B layers. The value range of each layer is 0 to 255. At least 85% of the images in the dataset have grayscale values ​​between 100 and 142, which represents the grayscale range of daytime shots. At least 50% of the images in the dataset have grayscale values ​​between 105 and 125.

[0012] Preferably, the evaluation method includes algorithm performance evaluation, shooting angle evaluation, grayscale sensitivity evaluation, target scale evaluation and image resolution evaluation.

[0013] Preferably, the algorithm performance evaluation includes evaluating RTMDet, YOLOv8, SSD, YOLOv3, Cascade R-CNN, Faster R-CNN and FPN using the constructed data set, using non-maximum suppression to eliminate overlapping bounding boxes to ensure that each detected object is surrounded by only one bounding box, and using the intersection-over-union ratio between the predicted boxes to evaluate the degree of overlap; and using precision, recall and average precision mean to evaluate model performance.

[0014] Preferably, the shooting angle evaluation includes calculating the average precision mean of images at three shooting angles: 0°-30°, 30°-60°, and 60°-90°.

[0015] Preferably, the grayscale sensitivity assessment comprises calculating the mean average precision for images with grayscale values ​​less than 102.3, 102.3-110.9, 110.9-118, 118-127.3 and greater than 127.3.

[0016] Preferably, the target scale evaluation includes calculating an average precision mean of three target scale images, wherein the three target scale images include: the area occupied by the target in the image is less than 1 / 160 of the entire image; the target scale in the image is between 1 / 160 and 1 / 80 of the entire image; and the target area in the image exceeds 1 / 80 of the entire image.

[0017] Preferably, the image resolution evaluation comprises calculating the mean average precision for images with 4K resolution and 640×640 resolution.

[0018] The beneficial effects of the present invention include at least:

[0019] (1) We constructed a dataset of drones captured by high-definition ground cameras (GA-Fly). The dataset contains 10,800 4K resolution images, covering a variety of shooting angles, backgrounds, lighting conditions, and flight altitudes. The drones were located 5 to 200 meters away from the camera and flew at altitudes ranging from 0 to 100 meters, ensuring the diversity and comprehensiveness of the data.

[0020] (2) Based on seven deep learning algorithms (such as RTMDet, YOLOv8, and SSD), image grayscale tests, shooting angle tests, target scale tests, and resolution tests were conducted to reveal the impact of various factors on detection performance. The experiment also compared the differences between ground-to-air and air-to-air datasets, highlighting the key challenges in different detection scenarios.

[0021] This dataset can be used to evaluate various drone detection algorithms. Experimental results reveal key challenges in ground-to-air detection and provide effective suggestions for future improvements. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 The mAP value graph of the seven algorithms in different grayscale value intervals for the construction and evaluation method of the UAV small target dataset of the present invention;

[0023] Figure 2 This is the mAP value diagram of the seven algorithms for the construction and evaluation method of the UAV small target dataset of the present invention at different target sizes;

[0024] Figure 3 This is a graph of mAP values ​​for seven algorithms in the method for constructing and evaluating a small target drone dataset according to a specific embodiment of the present invention under the condition of mismatched training and test image resolutions;

[0025] Figure 4 This is a graph of the mAP values ​​of seven algorithms for the method for constructing and evaluating a small target drone dataset according to a specific embodiment of the present invention on the GA-Fly and Det-Fly datasets. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0027] On the contrary, the present invention covers any alternatives, modifications, equivalents, and solutions that fall within the spirit and scope of the present invention as defined by the claims. Furthermore, to facilitate a better understanding of the present invention, certain specific details are described in detail below in the detailed description of the present invention. Those skilled in the art will be able to fully understand the present invention without these details.

[0028] In this embodiment of the present invention, a dataset constructed from at least 10,800 images and corresponding .txt annotation files is named GA-Fly, a ground-to-air drone dataset. Each image contains a target drone, and the image resolution is 3840 × 2160 pixels. These images are sampled at 10 frames per second from a video shot at 25 frames per second, with frames without drones or with poor quality due to filming issues removed. The video was shot with a Sony Alpha 6300 camera, using a Sony SEL18135 mid-to-telephoto zoom lens with an 18mm focal length (18mm-135mm). All images are annotated using Labelimg.

[0029] The GA-Fly dataset covers a variety of drone flight scenarios, simulating diverse conditions found in real-world environments, including varying viewpoints, backgrounds, relative distances, lighting conditions, and flight altitudes. Specifically, the viewpoints are divided into three groups: 0°-30°, 30°-60°, and 60°-90°, each accounting for one-third of the dataset. The backgrounds corresponding to different viewpoints also vary: scenes from 0°-30° primarily feature houses, trees, and grass, scenes from 30°-60° include tall buildings and large trees, and scenes from 60°-90° are mostly clear skies or layers of clouds. Regarding the size distribution of drone targets: in 3598 images, the target occupies less than 1 / 160 of the total image area; in 5310 images, the target size ranges from 1 / 160 to 1 / 80 of the total image size; and in 1892 images, the target size exceeds 1 / 80 of the total image size. Across all images, only 5 images have targets larger than 10% of the image area. A target drone is considered a small object when both its height and width are less than 10% of the image. Small object detection has always been a major challenge in visual inspection. Since the objects in most images in the dataset are relatively small, this study focused on small object detection, which is a common characteristic of drone detection scenarios. Notably, during data collection, the closest distance between the target drone and the camera was approximately 5 meters, indicating that most drone detection scenarios are actually small object detection scenarios. Therefore, the GA-Fly dataset provides an important resource and testing benchmark for studying ground-to-air small-target drone detection.

[0030] In the dataset, the lighting conditions of the image are evaluated by analyzing the average grayscale value of the image. The grayscale value is calculated as follows:

[0031]

[0032] The grayscale value of the image is calculated based on the intensity values ​​of the red (R), green (G), and blue (B) layers. The value of each layer ranges from 0 to 255. At least 85% of the images in the dataset have grayscale values ​​between 100 and 142, which represents the grayscale range of daytime images. At least 50% of the images in the dataset have grayscale values ​​between 105 and 125. Overall, the dataset exhibits a rich grayscale distribution, reflecting the diversity of drone detection scenarios under different lighting conditions in the real world.

[0033] There are several points to note about this dataset: First, each image in the dataset contains only one target drone (DJI Mini 4Pro), but the algorithm trained on this dataset has the ability to detect multiple targets and can generalize to a certain extent to drones of different models and similar appearances. Second, although this dataset includes a variety of shooting environments, it still cannot cover all practical application scenarios. Therefore, it is recommended to supplement it with a small amount of real-life data through transfer learning. Finally, to further improve the detection capabilities of drones across models and appearances, future research should consider building a more diverse drone dataset, which will require close cooperation between researchers. In addition, combining target motion perception technology can effectively improve drone detection results. We look forward to more innovative solutions to address these problems in the future.

[0034] Evaluation methods include algorithm performance evaluation, shooting angle evaluation, grayscale sensitivity evaluation, target scale evaluation and image resolution evaluation.

[0035] The algorithm performance evaluation included RTMDet, YOLOv8, SSD, YOLOv3, Cascade R-CNN, Faster R-CNN, and FPN using a constructed dataset. Non-maximum suppression was used to eliminate overlapping bounding boxes, ensuring that each detected object was surrounded by only one bounding box. The degree of overlap was assessed using the intersection-over-union ratio between predicted boxes. Model performance was evaluated using precision, recall, and mean average precision. Based on their detection strategies, these algorithms can be categorized as one-stage or two-stage networks. One-stage networks (RTMDet, YOLOv8, SSD, and YOLOv3) directly predict the location and category of objects in the input image end-to-end, avoiding the additional region proposal step. They typically use a convolutional neural network (CNN) as the backbone network, combining convolutional layers with a classification-regression head to simultaneously predict object categories and bounding boxes. These methods are fast and suitable for real-time detection scenarios. Two-stage networks (Cascade R-CNN, Faster R-CNN, and FPN) employ a two-stage detection approach. First, a region proposal network (RPN) is used to generate candidate regions. Then, these candidate regions are classified and fine-tuned for localization. Although such methods generally provide higher detection accuracy, their inference speed is slow and it is difficult to meet real-time requirements.

[0036] Table 1 shows the mAP values ​​of the seven algorithms. FPN achieved the highest score with an mAP of 0.894, demonstrating superior performance to other models. It is worth noting that the two-stage Faster R-CNN ranked third with 0.771, lower than the one-stage YOLOv3 (0.823). Cascade R-CNN performed poorly, ranking second from the bottom, while RTMDet (0.370), which performed well on the COCO dataset, performed the worst in this study. Although YOLOv8 is a newer model, its mAP value is 0.458, even lower than SSD (0.586). This unexpected result requires further analysis of the reasons for YOLOv8's poor performance.

[0037] The one-stage network's inference speed is significantly faster than the two-stage network. Cascade R-CNN had the longest inference time, at 44.1 milliseconds. Not only did it have the longest inference time, it also had the lowest mAP. While FPN achieved the highest accuracy with an mAP of 0.894, its inference time was also longer, at 41.1 milliseconds. Faster R-CNN also required a long inference time of 37.2 milliseconds.

[0038] Table 1 mAP of different algorithms on GA-Fly

[0039]

[0040] Among one-stage networks, RTMDet achieved an inference time of 16.6 milliseconds. While slower for a one-stage network, it was still more than twice as fast as Faster R-CNN. SSD achieved an inference time of 5.2 milliseconds, the second slowest among one-stage networks. Notably, YOLOv8 achieved the fastest inference time of just 1.4 milliseconds, demonstrating its dominance as a state-of-the-art network. Meanwhile, YOLOv3 followed closely behind in inference speed, taking 1.6 milliseconds and achieving a higher mAP, achieving a balance between speed and accuracy.

[0041] Shooting angle evaluation involves calculating the mean average precision for images shot at three angles: 0°-30°, 30°-60°, and 60°-90°. Shooting angle inherently reflects background complexity, which significantly impacts drone detection. Complex backgrounds often interfere with the network's detection capabilities, mistaking objects with similar appearances to drones, such as windows, surveillance cameras, and lighting, for the target drone. As the shooting angle increases, the background gradually simplifies from a complex ground scene to a skyscape.

[0042] The mAP value increases with increasing shooting angle, especially in the 0°-30° range, where the mAP value is relatively low. However, from 0°-30° to 30°-60°, the mAP value increases significantly. However, when the angle is further increased to 60°-90°, the increase in mAP is relatively small. This is because the 0-30° shooting angle includes a large amount of complex background; in the 30°-60° range, the background complexity is reduced, with only a few tall buildings and trees; and in the 60°-90° range, the background is mainly simple sky and cloud layers.

[0043] Different networks exhibit significant differences in their sensitivity to complex backgrounds, resulting in varying detection performance across different angle ranges. Although complex backgrounds inhibit the performance of all models within the 0°-30° angle range, FPN consistently achieves the highest mAP across all angle ranges, demonstrating its robustness against complex background interference. YOLOv3 and SSD, on the other hand, are more sensitive to background variations, experiencing significant fluctuations in their mAP values ​​as background complexity increases, indicating limitations in their ability to handle complex backgrounds.

[0044] Among the evaluated models, FPN consistently achieved the highest mAP, while Cascade R-CNN, RTMDet, and YOLOv8 performed poorly, lagging far behind FPN and YOLOv3 even on the simplest backgrounds. This consistently poor performance suggests that background complexity may not be the primary reason for the poor performance of R-CNN, RTMDet, and YOLOv8.

[0045] Grayscale sensitivity assessment involves calculating the mean average precision for images with grayscale values ​​less than 102.3, 102.3-110.9, 110.9-118, 118-127.3, and greater than 127.3. The grayscale value of an image also affects algorithm performance. Grayscale values ​​can be obtained through image processing techniques and indirectly reflect the exposure conditions at the time the image was captured. Higher grayscale values ​​generally indicate brighter images, while lower grayscale values ​​indicate underexposure or low ambient light. In drone detection tasks, both overexposure and underexposure can interfere with detection performance, potentially leading to unclear target boundaries and loss of texture detail. Moderate grayscale intensity often provides clearer and more balanced image details, helping to improve target recognition accuracy. Therefore, in practical applications, properly controlling image exposure conditions can effectively improve drone detection results.

[0046] See also Figure 1 , mAP values ​​vary significantly across grayscale values. Optimal detection performance occurs between grayscale values ​​of 102.3 and 110.9. When grayscale values ​​exceed this range, mAP decreases. This demonstrates that proper exposure conditions can improve drone detection accuracy, while environments with excessively bright or dark lighting can make target identification difficult, thus affecting detection results.

[0047] Different networks have different sensitivities to grayscale changes, resulting in differences in their performance under different lighting conditions. FPN shows high robustness in all grayscale ranges, with stable detection performance and insensitivity to grayscale changes, demonstrating its adaptability in the face of lighting interference. Cascade R-CNN, RTMDet, YOLOv8, and SSD are more sensitive to grayscale changes, and their mAP values ​​fluctuate greatly in different grayscale ranges, indicating that the performance of these models will be significantly limited under unfavorable lighting conditions. However, even in the most suitable grayscale range, the performance of Cascade R-CNN, RTMDet, and YOLOv8 is still far behind that of FPN and YOLOv3, which shows that grayscale value is not the main reason for their poor performance.

[0048] The object scale evaluation involves calculating the mean average precision for three object scale images: the object in the image occupies less than 1 / 160 of the entire image area; the object scale in the image is between 1 / 160 and 1 / 80 of the entire image area; and the object area in the image exceeds 1 / 80 of the entire image area. The size of the target drone in the image has a significant impact on the detection performance of the model. This experiment explores this by measuring the ratio of the width and height of the target drone to the entire image. Let w and h represent the width and height of the target box, respectively. Figure 2The mAP values ​​of all algorithms at different object scales are shown in the figure. As the figure clearly shows, the mAP of all algorithms significantly improves with increasing object scale, indicating that smaller objects are more challenging to detect. This result emphasizes the importance of algorithmic improvements and optimizations for small object detection to improve detection performance in real-world applications. Cascade R-CNN, RTMDet, and YOLOv8 are particularly affected by changes in object scale. The performance of RTMDet and YOLOv8 fluctuates significantly across different scenarios, which can be attributed to their detection heads and anchor-free mechanisms. Specifically, YOLOv8's detection head resolution is insufficient to accurately identify small objects, and its anchor-free design complicates the object bounding box regression process, resulting in poor performance. RTMDet, utilizing a similar detection head and anchor-free mechanism, faces the same challenges. In contrast, YOLOv3, which uses a detection head with the same resolution but relies on an anchor mechanism, performs relatively well. Cascade R-CNN also performs poorly for small object detection, primarily due to its architecture cascading multiple detectors with different IoU thresholds. When detecting small objects, even small pixel errors can have a significant impact on the IoU, resulting in a significant performance degradation.

[0049] Furthermore, Faster R-CNN and FPN also show decreased performance when detecting larger objects, likely due to the RPN mechanism in the two-stage network. Since GA-Fly primarily contains small objects, the system tends to generate candidate boxes optimized for these small objects, which is detrimental when generating candidate boxes for larger objects, thus affecting the final detection accuracy. This phenomenon emphasizes the importance of considering multiple object scales during training to improve the model's adaptability and accuracy in practical applications.

[0050] Image resolution evaluation involves calculating the mean average precision for images with 4K resolution and 640×640 resolution. As camera resolution increases, the size of the captured image also increases. However, not all inspection cameras have 4K resolution. Therefore, this study conducted cropping experiments to simulate the performance of inspection cameras at different resolutions. The experiments were conducted at 4K resolution and 640×640 resolution. Figure 3 , the model performs better when training and testing are performed at the same image size.,However, a mismatch between training and testing sizes can significantly,affect the performance of the model.

[0051] Specifically, while models trained on high-resolution images maintain a reasonable level of accuracy when tested on low-resolution images, the reverse is true: models trained on low-resolution images and tested on high-resolution images nearly lose detection capability. This phenomenon emphasizes the importance of image size consistency during training and testing. Notably, models trained and tested at 640×640 resolution achieve the highest mAP score, but their detection range is relatively limited. In contrast, models trained on 4K images achieve a wider detection range, albeit with slightly lower accuracy, which may result in increased equipment costs.

[0052] Experimental results show that to improve the model's practical application, the training process should prioritize using image sizes that match the actual camera resolution to ensure good detection performance at different resolutions. Given the limited ability of embedded platforms to process 4K image transmission in real time, subsequent YOLO-Drone algorithm training and embedded deployment will randomly crop 4K images to 1K size to meet real-time processing requirements without significantly affecting detection accuracy.

[0053] In the existing technology, Det-Fly is a large-scale and diverse drone dataset covering a variety of scenarios. Its authors point out that the shooting distance of this dataset exceeds 100 meters. However, since it mainly focuses on air-to-air recognition scenarios, there are significant differences compared to ground-to-air shooting conditions. The challenge of air-to-air detection lies mainly in the complex environmental background, which easily leads to false detections, while the difficulty of ground-to-air detection is more reflected in the fact that small drone targets are easily obscured by the vast sky background, resulting in missed detections. In addition, ground-to-air shooting is usually affected by light interference due to the shooting angle, such as overexposure and backlighting problems, while the shooting angle of air-to-air detection is usually smaller, the background is richer and the light is softer.

[0054] In order to highlight the difference between the two, the present invention conducted a comparative experiment on its own dataset, see Figure 4 (The left side of the bar graph of each algorithm is the GA-Fly dataset of the present invention, and the right side is the Det-Fly dataset.) It can be observed that on the dataset of the present invention, the indicators of most algorithms are more than 10% lower than those of the Det-Fly dataset. This result shows that although these classic deep learning algorithms perform reliably in the air-to-air drone detection environment, they are still insufficient in the ground-to-air detection environment. This further shows that the current deep learning network has a strong adaptability in dealing with complex background problems, but it is still insufficient in small target detection. This may require the introduction of additional auxiliary mechanisms or decision-making mechanisms to help deep learning networks solve small target problems more effectively.

[0055] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for constructing and evaluating a small drone target dataset, characterized in that: The dataset construction includes at least 10,800 images and corresponding .txt annotation files. Each image contains the target drone, and the image resolution is 3840×2160 pixels. These images are sampled at a frequency of 10 frames per second from a video shot at 25 frames per second, and frames without drones or with poor quality due to shooting problems are excluded. All images are annotated using Labelimg.

2. The method for constructing and evaluating a small drone target dataset according to claim 1, characterized in that: The viewing angles of the images in the dataset are divided into three groups: 0°-30°, 30°-60°, and 60°-90°, with each group accounting for one-third of the entire dataset; among them, the 0°-30° scenes include houses, trees, and grass, the 30°-60° scenes include tall buildings and large trees, and the 60°-90° scenes are clear skies or layers of clouds.

3. The method for constructing and evaluating a small target dataset of a UAV according to claim 1, characterized in that: In one-third of the images in the dataset, the target area is less than 1 / 160 of the entire image; in one-half of the images, the target size is between 1 / 160 and 1 / 80 of the entire image; in one-sixth of the images, the target area exceeds 1 / 80 of the entire image; when the height and width of the target drone are both less than 10% of the image, it is considered a small target.

4. The method for constructing and evaluating a small drone target dataset according to claim 1, characterized in that: In the dataset, the illumination condition of the image is evaluated by analyzing the average grayscale value of the image. The grayscale value is calculated as follows: Among them, the grayscale value Gray of the image is calculated based on the intensity values ​​of the red R, green G, and blue B layers. The value range of each layer is 0 to 255. At least 85% of the images in the dataset have grayscale values ​​between 100 and 142, which represents the grayscale range of daytime shots. At least 50% of the images in the dataset have grayscale values ​​between 105 and 125.

5. The method for constructing and evaluating a small target dataset of a UAV according to claim 1, characterized in that: Evaluation methods include algorithm performance evaluation, shooting angle evaluation, grayscale sensitivity evaluation, target scale evaluation and image resolution evaluation.

6. The method for constructing and evaluating a small target dataset of a UAV according to claim 5, characterized in that: The algorithm performance evaluation includes evaluating RTMDet, YOLOv8, SSD, YOLOv3, Cascade R-CNN, Faster R-CNN and FPN using the constructed dataset, using non-maximum suppression to eliminate overlapping bounding boxes to ensure that each detected object is surrounded by only one bounding box, and using the intersection-over-union ratio between predicted boxes to evaluate the degree of overlap; and using precision, recall and average precision to evaluate model performance.

7. The method for constructing and evaluating a small target dataset of a UAV according to claim 5, characterized in that: The shooting angle evaluation includes calculating the average precision mean of images at three shooting angles: 0°-30°, 30°-60°, and 60°-90°.

8. The method for constructing and evaluating a small target dataset of a UAV according to claim 5, characterized in that: The grayscale sensitivity evaluation includes calculating the average precision mean for images with grayscale values ​​less than 102.3, 102.3-110.9, 110.9-118, 118-127.3 and greater than 127.

3.

9. The method for constructing and evaluating a small drone target dataset according to claim 5, characterized in that: The target scale evaluation includes calculating the average precision mean of three target scale images, including: the area occupied by the target in the image is less than 1 / 160 of the entire image; the target scale in the image is between 1 / 160 and 1 / 80 of the entire image; and the target area in the image exceeds 1 / 80 of the entire image.

10. The method for constructing and evaluating a small target dataset of a UAV according to claim 5, characterized in that: The image resolution evaluation includes calculating the mean average precision for images with 4K resolution and 640×640 resolution.