Anti-unmanned aerial vehicle infrared small target detection system and method based on attention mechanism
Through the infrared small object detection system based on attention mechanism, the problems of low efficiency and feature distribution differences in UAV detection under infrared conditions are solved, and high-precision UAV target detection and safety warning are achieved, which improves the applicability and accuracy of the model.
Patent Information
- Application Number
- CN202510478982.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art drone detection under infrared conditions has problems such as low efficiency, different feature distribution and small target degradation, making it difficult to achieve efficient and accurate drone target detection.
The anti-UAV infrared small object detection system based on attention mechanism is adopted, including image acquisition, preprocessing, feature extraction, a progressive fusion feature pyramid, attention-guided object detection and shape loss function module. Through technical means such as efficient spatial coordinate attention unit and shape bounding box loss function, the loss weight is dynamically adjusted to improve detection accuracy.
It realizes high-precision drone target detection in low contrast and complex contexts, improves the generalization ability of the model, can promptly warn of possible safety problems caused by drones, and provides protection against drone abuse.
Smart Images

Figure CN120298935A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and in particular relates to an anti-drone infrared small target detection system and method based on an attention mechanism. Background Art
[0002] With the development of information technology, unmanned aerial vehicles (UAVs) have seen more and more development in the application field. More and more UAVs with smaller sizes and more stable performance have been invented and manufactured, giving rise to the emergence of civilian UAVs. The applications of civilian UAVs include aerial photography and videography, geographical mapping and surveying, disaster monitoring and rescue, agriculture and plant protection, construction and engineering, entertainment and competition, and other fields. UAVs have the advantages of being efficient and flexible, having a unique shooting perspective, low cost, low safety risk, rapid deployment, simple data collection and processing, and environmental friendliness, which has led to an increase in the popularity of UAVs. With the rapid development of sensor technology and intelligent control technology, UAVs are more user-friendly, easier to operate, and tend to be low-cost and low-energy-consuming. In addition, with the support of artificial intelligence and 5G communication technology, UAVs have the capabilities of autonomous decision-making and intelligence. Through deep learning and machine learning algorithms, UAVs can analyze environmental data, perceive obstacles, perform path planning and obstacle avoidance, and thus achieve autonomous flight and task execution. This enables UAVs to complete various tasks more intelligently and efficiently; in addition, 5G technology provides high-speed and low-latency data transmission capabilities, enabling UAVs to collect, transmit, and process a large amount of data in real time. This is very important for task execution, monitoring, and decision-making. Through rapid data transmission and processing, UAVs can timely obtain and utilize environmental data, and make accurate decisions and responses.
[0003] However, while drones are widely used in fields such as logistics distribution, agricultural plant protection, and film shooting, the potential safety hazards caused by their disorderly flight are becoming increasingly prominent. With the widespread application of drones in civilian and military fields, potential issues such as privacy infringement and airspace conflicts have drawn social attention. The extensive use of drones has brought about severe security challenges while promoting social development, and there is an urgent need for efficient and reliable detection technologies to regulate the use order of drones. Radar detection has an extremely low radar cross-section (RCS) for drones made of plastic / carbon fiber materials, and drones have characteristics such as slow speed and low altitude flight, resulting in a low detection probability in complex urban environments; the effective operating range of acoustic detection technology is theoretically up to 300m at most, which cannot meet the national security needs and is easily interfered by environmental noise. The infrared band has advantages such as penetrating smoke, perceiving weak light environments, and being sensitive to thermal radiation, making it irreplaceable in fields such as military reconnaissance, industrial non-destructive testing, medical diagnosis, and intelligent driving. For example, in complex battlefield environments, infrared imaging can avoid relying on visible light and achieve all-weather target tracking; in the monitoring of power equipment, thermal anomaly features can provide early warnings of mechanical failures. Target detection for infrared images captured by infrared sensors includes general target detection. And for small target detection in infrared scenes, the picture area is usually less than 30×30 pixels. This is very common in anti-drone detection tasks.
[0004] The traditional detection methods mainly include the following: First is the method based on two-dimensional least mean square. This method can optimize the adaptive filter based on the stochastic gradient descent method. By continuously iterating to optimize the filter weights, the optimization goal is to minimize the mean square error between the output and the expectation. Target detection is achieved through three steps: background modeling, background prediction, and background enhancement. Then is the method based on maximum mean and maximum median. This method uses the maximum mean and maximum median of pixels within a local window to effectively retain the edge structure, enhancing the robustness to impulse noise while extracting the target. Finally is the method based on morphology. This method lies in the design of dynamic structure elements and multi-scale erosion-dilation chains. By adaptively selecting the size of the structure element, while maintaining the background suppression ability, the computational complexity is greatly reduced. Although these traditional methods have made remarkable progress in the field of infrared target detection, their core is limited by the experience dependence of artificial feature design and the generalization bottleneck in complex scenarios. For example, changes in parameters have a greater impact on the results. The model overly relies on linear assumptions, etc.
[0005] Deep learning-based methods provide new ways to solve the infrared target detection problem through data-driven feature learning and end-to-end optimization paradigms. Deep learning often uses convolutional neural networks for adaptive feature extraction, and convolutional neural networks can automatically obtain target features across scales. The R-CNN model generates candidate regions through the selective search algorithm, then scales each candidate region to a fixed size for feature extraction through the CNN network, and finally completes the classification task through a support vector machine and uses a linear regressor for bounding box correction. However, this method is less efficient and not suitable for real-time scenarios. RetinaNet effectively alleviates the class imbalance problem by dynamically adjusting the loss weights of easy and hard samples. This model uses a feature pyramid network to strengthen the multi-scale feature fusion ability, but its performance is sensitive to hyperparameters and the inference speed is slower than subsequent lightweight models. The YOLO series of models, through end-to-end design, directly output the target location and category through a single forward pass, significantly reducing memory occupancy and being more suitable for real-time scenarios. However, when the above methods are directly transferred to the infrared scenario, they face problems such as differences in feature distribution and degradation of small targets. In addition, there is an inherent contradiction between the architecture design preferences of existing detectors and the characteristics of infrared data. Therefore, how to efficiently and accurately detect drones under infrared conditions has become a crucial challenge. Summary of the Invention
[0006] To solve the above technical problems, the present invention proposes an anti-drone infrared small target detection system and method based on an attention mechanism, which realizes high-precision detection of drone targets and can timely warn of potential safety problems caused by drones, providing guarantee for preventing the abuse of drones in military and civilian fields.
[0007] On the one hand, to achieve the above object, the present invention provides an anti-drone infrared small target detection system based on an attention mechanism, including:
[0008] An image acquisition module, an image preprocessing module, a feature extraction module, a progressive fusion feature pyramid module, an attention-guided target detection module, a shape loss function module, and a target detection module;
[0009] The image acquisition module is used to acquire the source image or video;
[0010] The image preprocessing module is used to preprocess the source image or video to obtain the preprocessed image;
[0011] The feature extraction module is used to extract features from the preprocessed image to obtain feature maps of different scales;
[0012] The progressive fusion feature pyramid module is used to fuse the feature maps of different scales to obtain the fused feature map;
[0013] The attention-guided object detection module is used to focus on a specific area of the fused feature map and perform classification and regression tasks on the focused feature map;
[0014] The shape loss function module is used to dynamically adjust the loss weight through the shape loss function according to the shape features of the target, guide the training of the anti-UAV infrared detection model, and obtain a trained anti-UAV infrared detection model;
[0015] The object detection module is used to obtain the source image or video of the target area, input the trained anti-UAV infrared detection model, and obtain the object detection result.
[0016] Optionally, preprocessing the source image or video to obtain the preprocessed image includes:
[0017] Performing video frame extraction, bounding box information extraction, pixel filling, and data augmentation on the source image or video to obtain the preprocessed image.
[0018] Optionally, performing feature extraction on the preprocessed image to obtain feature maps of different scales includes:
[0019] Performing feature extraction on the preprocessed image through CSPDarkNet53 to obtain features of three or more different scales.
[0020] Optionally, fusing the feature maps of different scales to obtain the fused feature map includes:
[0021] Performing upsampling and downsampling on the feature maps of different scales respectively, balancing the sizes of the feature maps at different levels, and fusing the balanced feature maps to obtain the fused feature map.
[0022] Optionally, the attention-guided object detection module includes: an efficient spatial coordinate attention unit;
[0023] The efficient spatial coordinate attention unit is expressed as:
[0024]
[0025]
[0026] y c (i,j) = x c (i,j) × σ(Conv(x c ) × [ρ(AP(x f )), ρ(MP(x f ))]);
[0027] Among them, xC Denote the original input features, where H and W represent the width and height of the original input features respectively, σ represents the Sigmond function, Conv represents the convolution operation, ρ represents the Softmax function, and AP and MP represent average pooling and max pooling respectively.
[0028] Optionally, the shape loss function includes a shape bounding box loss function, a classification loss function, and a distribution focal loss function.
[0029] Optionally, the shape bounding box loss function is expressed as:
[0030] Ω shape = ∑ t=w,h (1 - e -ωt ) θ ;
[0031] L shape = 1 - IoU + distance shape + 0.5 × Ω shape ;
[0032] where IoU represents the intersection over union of the predicted box and the actual box, and distance shape represents the arithmetic mean of the distances of the length and width in the Euclidean space.
[0033] On the other hand, to achieve the above object, the present invention also provides an anti - UAV infrared small target detection method based on an attention mechanism, including:
[0034] Obtain a source image or video;
[0035] Pre - process the source image or video to obtain a pre - processed image;
[0036] Extract features from the pre - processed image to obtain feature maps of different scales;
[0037] Fuse the feature maps of different scales to obtain a fused feature map;
[0038] Focus the fused feature map on a specific area, and perform classification and regression tasks on the focused feature map;
[0039] Dynamically adjust the loss weight according to the shape characteristics of the target, guide the training of the anti - UAV infrared detection model, and obtain a trained anti - UAV infrared detection model;
[0040] Obtain a source image or video of the target area, input it into the trained anti - UAV infrared detection model, and obtain the target detection result.
[0041] Technical effects of the present invention: The present invention discloses an anti-UAV infrared small target detection system and method based on an attention mechanism. Infrared images or videos are captured by a professional infrared sensor, and an anti-UAV infrared small target detection model based on the attention mechanism is used for target detection to automatically identify UAV targets. For the anti-UAV infrared small target detection model in the present invention, features are fused through progressive fusion of a feature pyramid, thereby reducing the semantic gap between adjacent features. Through an efficient spatial coordinate attention module, the interaction of feature information is ensured, thus solving the pain points and difficulties of small target detection. Through a shape bounding box loss function, the influence of UAVs with different shapes on the detection results is fully considered. The present invention preferably realizes the detection of UAV targets in infrared images with low contrast and complex backgrounds, improves the generalization ability of the model, and ensures the accurate identification of UAV targets under various conditions. Furthermore, high-precision detection of UAV targets is achieved, and early warnings can be issued in a timely manner for potential safety problems caused by UAVs, providing a guarantee for preventing the abuse of UAVs in military and civilian fields. Description of the Drawings
[0042] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0043] Figure 1 It is a schematic structural diagram of the anti-UAV infrared detection model in the anti-UAV infrared small target detection system based on the attention mechanism according to an embodiment of the present invention;
[0044] Figure 2 It is a schematic flowchart of the anti-UAV infrared small target detection method based on the attention mechanism according to an embodiment of the present invention;
[0045] Figure 3 It is a schematic diagram of the progressive fusion feature pyramid module according to an embodiment of the present invention;
[0046] Figure 4 It is a schematic diagram of the efficient spatial coordinate attention unit according to an embodiment of the present invention;
[0047] Figure 5 It is a comparison diagram of the effects of each method according to an embodiment of the present invention;
[0048] Figure 6 It is a visual comparison diagram of the objective evaluation indexes of each method according to an embodiment of the present invention;
[0049] Figure 7 It is a comparison diagram of the objective evaluation indexes of each method according to an embodiment of the present invention. Detailed Embodiments
[0050] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The following will describe the present application in detail with reference to the drawings and in combination with the embodiments.
[0051] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0052] As Figure 1 shown, in this embodiment, an anti-UAV infrared small target detection system based on an attention mechanism is provided, including:
[0053] an image acquisition module, an image preprocessing module, a feature extraction module, a progressive fusion feature pyramid module, an attention-guided target detection module, a shape loss function module, and a target detection module;
[0054] The image acquisition module is used to acquire a source image or video;
[0055] The image preprocessing module is used to preprocess the source image or video to obtain a preprocessed image;
[0056] The feature extraction module is used to extract features from the preprocessed image to obtain feature maps of different scales;
[0057] The progressive fusion feature pyramid module is used to fuse the feature maps of different scales to obtain a fused feature map;
[0058] The attention-guided target detection module is used to focus on a specific area of the fused feature map and perform classification and regression tasks on the focused feature map;
[0059] The shape loss function module is used to dynamically adjust the loss weight through a shape loss function according to the shape features of the target, guide the training of the anti-UAV infrared detection model, and obtain a trained anti-UAV infrared detection model;
[0060] The target detection module is used to obtain a source image or video of the target area, input the trained anti-UAV infrared detection model, and obtain a target detection result.
[0061] Furthermore, preprocessing the source image or video to obtain a preprocessed image includes:
[0062] Performing video frame extraction and bounding box information extraction, pixel filling, and data augmentation on the source image or video to obtain a preprocessed image.
[0063] Further, feature extraction is performed on the preprocessed image to obtain feature maps of different scales, including:
[0064] Feature extraction is performed on the preprocessed image through CSPDarkNet53 to obtain features of different scales with three or more layers.
[0065] Further, the feature maps of different scales are fused to obtain the fused feature map, including:
[0066] The feature maps of different scales are respectively subjected to upsampling and downsampling processes to balance the sizes of the feature maps at different levels, and the balanced feature maps are fused to obtain the fused feature map.
[0067] Further, the attention-guided object detection module includes: an efficient spatial coordinate attention unit;
[0068] The efficient spatial coordinate attention unit is expressed as:
[0069]
[0070] y c (i,j) = x c (i,j) × σ(Conv(x c ) × [ρ(AP(x f ))], ρ(MP(x f ))]);
[0071] Among them, x C represents the input original feature, H and W respectively represent the width and height of the input original feature, σ represents the Sigmond function, Conv represents the convolution operation, ρ represents the Softmax function, and AP and MP respectively represent average pooling and max pooling.
[0072] Further, the shape loss function includes a shape bounding box loss function, a classification loss function, and a distribution focal loss function.
[0073] Further, the shape bounding box loss function is expressed as:
[0074] Ω shape = ∑ t=w,h (1 - e -ωt ) θ ;
[0075] L shape = 1 - IoU + distance shape + 0.5 × Ω shape ;
[0076] Among them, IoU represents the intersection over union of the predicted box and the actual box, distanceshape It represents the arithmetic mean of the distances in the Euclidean space with respect to the length and width.
[0077] Furthermore, the classification loss function is expressed as:
[0078]
[0079] where p i represents the probability that sample i is predicted as the positive class; y i is the sign function, taking 1 when sample i is the positive class and 0 otherwise.
[0080] Furthermore, the distribution focal loss function is expressed as:
[0081]
[0082] where N represents the number of images; C represents the number of classifications; y ic is the true label of the i-th sample; p ic is the predicted probability that the i-th sample belongs to class c; α is the balance factor to balance the weights of positive and negative samples; γ is the focusing parameter to control the degree of attention to difficult samples.
[0083] In this embodiment, an anti-drone infrared small target detection method based on the attention mechanism is also provided, including:
[0084] Obtain the source image or video;
[0085] Preprocess the source image or video to obtain the preprocessed image;
[0086] Extract features from the preprocessed image to obtain feature maps of different scales;
[0087] Fuse the feature maps of different scales to obtain the fused feature map;
[0088] Focus the fused feature map on a specific area, and perform classification and regression tasks on the focused feature map;
[0089] Dynamically adjust the loss weights according to the shape features of the target, guide the training of the anti-drone infrared detection model, and obtain the trained anti-drone infrared detection model;
[0090] Obtain the source image or video of the target area, input it into the trained anti-drone infrared detection model, and obtain the target detection result.
[0091] Specifically, as Figure 2 shown, the implementation process of the anti-drone infrared small target detection method in this embodiment includes:
[0092] S1. The image or video is converted into single images through frame extraction. It is necessary to fill the pictures to the specified size and then reconstruct the images by flipping, translation, etc.
[0093] S2. Feature extraction is performed through CSPDarkNet53.
[0094] S3. Progressive fusion feature pyramid: Fuse feature maps of different scales to reduce the semantic gap existing between feature layers at a relatively large distance and effectively reduce the distortion of feature maps.
[0095] S3.1. First, perform adaptive feature fusion on the P5 and P4 features after feature extraction to reduce the semantic gap between P4 and P5.
[0096] S3.2. Fuse the two feature maps output by S3.1 with P3 again. The specific process of feature fusion is as Figure 3 shown.
[0097] Since the shallow features of the network have a relatively small receptive field, the semantic information of its low-order feature maps is usually used to detect small-sized targets, while high-order features are generally used for large-sized target recognition. In the traditional FPN architecture, the high-level features at the top of the pyramid usually do not directly perform cross-layer fusion with non-adjacent low-level features at the bottom layer, but adopt a top-down progressive layer-by-layer fusion strategy. This process may cause the gradual degradation of low-level features during propagation, and there are still significant semantic differences between the fused feature maps of non-adjacent levels. The progressive fusion feature pyramid of the present invention adaptively allows the network to fuse from high-order features to low-order features. This fusion method can retain the small target information of the network to the greatest extent and avoid the gradual disappearance of information during propagation.
[0098] S4. Attention-guided target detection module: Used to focus the fused feature maps on specific regions and perform classification and regression tasks on the focused feature maps to obtain the final bounding boxes and classification information.
[0099] To effectively solve the problems of small target size and vulnerability to background interference in infrared UAV images, the present invention proposes an efficient spatial coordinate attention unit based on the coordinate attention mechanism. See Figure 4 , which adopts a two-stage structure design: In the first stage, attention feature maps are extracted through 1×1 convolution in the coordinate dimension, and in the second stage, efficient spatial attention calculation is realized through the splitting of the spatial weights of the feature maps. The efficient spatial coordinate attention unit guides the model to quickly focus on the region of interest by enhancing the feature response of the target region.
[0100] S5, loss function: Calculate the loss of the coordinates output by the detection network and the label image and back propagate to optimize the network parameters. In the training phase, the number of cycles is set to 50, the batch is set to 64, the optimizer is SGD, and the learning rate is set to cosine annealing. The present invention uses the shape bounding box loss function as the loss function, and its equation can be expressed as follows, where scale represents the scaling factor and is one of the hyperparameters, and w gt ,h gt are the width and height of the real bounding box respectively. c ,h c are the width and height of the network prediction box respectively.
[0101]
[0102] Ω shape =∑ t=w,h (1-e -ω,t ) θ ;
[0103] L shape =1-IoU+distance shape +0.5×Ω shape ;
[0104] S6. If there are still image targets that need to be tracked, the images are input into the model for detection in fixed batches, and S0-S5 are cycled in sequence until there are no images that need to be tracked.
[0105] In this embodiment, the Anti-UAV dataset is used and infrared images are extracted from it. In this embodiment, the images are uniformly preprocessed to a size of 640×640 as input images. The training set contains a total of 149,528 images, and the validation set contains a total of 61,999 images. The images in the dataset are all randomly selected.
[0106] In order to verify the advancement and effectiveness of the present invention in the field of anti-UAV infrared small target detection. This embodiment uses 10 methods, RetinaNet, YOLOv3-SPP, YOLOv3-tiny, YOLOv5, YOLOv5-BiFPN, YOLO-FIRI, YOLOX, YOLOv6, YOLOv7-tiny, and YOLOv8 for comparative analysis. The codes of the above methods are all public and the parameters are not changed. This embodiment also makes quantitative and qualitative evaluations of the experiments. The structure of the quantitative evaluation can be found in Figure 6 and Figure 7 , where AP represents the average prediction accuracy, AP 75 It represents the average prediction accuracy above the intersection-over-union threshold of 75%, AP 50:95Indicates the average of all average prediction rates with the intersection over union threshold ranging from 50% to 95% and separated by a step of 5%. The experimental results show that the method proposed in the present invention surpasses all 10 methods in detection accuracy. Compared with the second-ranked YOLOv6, our model increases the AP 75 by 2.30%, increases the AP 50:95 by 0.99%, and simultaneously reduces the number of model parameters by 4.6M. Compared with the latest method YOLOv8, our model increases the AP by 2.85% 75 and increases the AP by 1.12% 50:95 . In summary, the method proposed in the present invention achieves a significant improvement in accuracy without a significant increase in scale, while taking into account the balance between model size and accuracy.
[0107] As Figure 5 shown, the results of the method proposed in the present invention are visualized with the results of other currently popular object detection methods. In this embodiment, the method successfully detects all five test images, and the results show that the method proposed in the present invention has the highest accuracy compared with other methods.
[0108] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An anti-drone infrared small target detection system based on an attention mechanism, characterized in that, Including: An image acquisition module, an image preprocessing module, a feature extraction module, a progressive fusion feature pyramid module, an attention-guided object detection module, a shape loss function module, and an object detection module; The image acquisition module is used to acquire a source image or video; The image preprocessing module is used to preprocess the source image or video to obtain a preprocessed image; The feature extraction module is used to extract features from the preprocessed image to obtain feature maps of different scales; The progressive fusion feature pyramid module is used to fuse the feature maps of different scales to obtain a fused feature map; The attention-guided object detection module is used to focus on a specific region of the fused feature map and perform classification and regression tasks on the focused feature map; The shape loss function module is used to dynamically adjust the loss weight through a shape loss function according to the shape features of the target, guide the training of the anti-drone infrared detection model, and obtain a trained anti-drone infrared detection model; The object detection module is used to obtain a source image or video of the target region, input the trained anti-drone infrared detection model, and obtain an object detection result.
2. The anti-drone infrared small target detection system based on the attention mechanism according to claim 1, wherein Preprocessing the source image or video to obtain a preprocessed image includes: Performing video frame extraction and bounding box information extraction, pixel filling, and data augmentation on the source image or video to obtain a preprocessed image.
3. The anti-UAV infrared small target detection system based on the attention mechanism according to claim 1, characterized in that, Extracting features from the preprocessed image to obtain feature maps of different scales includes: Performing feature extraction on the preprocessed image through CSPDarkNet53 to obtain features of three or more different scales.
4. The anti-UAV infrared small target detection system based on the attention mechanism according to claim 1, characterized in that, Fusing the feature maps of different scales to obtain a fused feature map includes: Performing upsampling and downsampling processing on the feature maps of different scales respectively, balancing the sizes of the feature maps at different levels, and fusing the balanced feature maps to obtain a fused feature map.
5. The anti-drone infrared small target detection system based on the attention mechanism according to claim 1, characterized in that, The attention-guided object detection module includes: an efficient spatial coordinate attention unit; The efficient spatial coordinate attention unit is expressed as: y c (i, j) = x c (i, j) × σ(Conv(x c ) × [ρ(AP(x f )), ρ(MP(x f ))]); Among them, x C represents the original input feature, H and W respectively represent the width and height of the original input feature, σ represents the Sigmond function, Conv represents the convolution operation, ρ represents the Softmax function, and AP and MP respectively represent average pooling and max pooling.
6. The anti-drone infrared small target detection system based on the attention mechanism according to claim 1, wherein The shape loss function includes a shape bounding box loss function, a classification loss function, and a distribution focal loss function.
7. The anti-UAV infrared small target detection system based on the attention mechanism according to claim 1, wherein The shape bounding box loss function is expressed as: Ω shape = ∑ t=w,h (1 - e -ωt ) θ ; L shape = 1 - IoU + distance shape + 0.5 × Ω shape ; Among them, IoU represents the intersection over union of the predicted bounding box and the actual bounding box, and distance shape represents the arithmetic mean of the distances of the length and width in the Euclidean space respectively.
8. A method for an anti-drone infrared small target detection system based on an attention mechanism according to claims 1-7, characterized in that, Including: Obtaining a source image or video; Preprocessing the source image or video to obtain a preprocessed image; Extracting features from the preprocessed image to obtain feature maps of different scales; Fusing the feature maps of different scales to obtain a fused feature map; Focusing on a specific region of the fused feature map and performing classification and regression tasks on the focused feature map; Dynamically adjusting the loss weight according to the shape features of the target, guiding the training of the anti-drone infrared detection model, and obtaining a trained anti-drone infrared detection model; Obtaining a source image or video of the target region, inputting the trained anti-drone infrared detection model, and obtaining an object detection result.