Unmanned aerial vehicle real-time detection algorithm
By introducing efficient feature pyramid networks and multi-attention mechanism networks into the UAV vehicle detection algorithm, the problems of weak discriminant features and information imbalance in weak vehicle detection are solved, and higher detection accuracy and real-time stability are achieved.
Patent Information
- Application Number
- CN202510249382.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-20
AI Technical Summary
When dealing with weak vehicles, existing drone vehicle detection algorithms have problems such as weak discriminant characteristics, unbalanced information, high computational complexity and inability to meet real-time detection requirements.
A real-time detection algorithm for drone vehicles is proposed, combining basic networks, efficient feature pyramid networks, multi-attention mechanism networks and shared detection networks. By introducing local information transmission, multi-attention mechanisms and shared detection networks, the feature representation ability of weak vehicles is enhanced, and information imbalance and computational complexity are reduced.
It effectively improves the discriminant characteristics of weak vehicles, reduces information imbalance and information loss problems, improves the accuracy and real-time stability of drone vehicle detection, and meets the needs of real-time detection.
Smart Images

Figure CN120182575A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle detection, and specifically to a real-time detection algorithm for drones and vehicles. Background Art
[0002] Vehicle detection technology is the core of drone intelligent scene monitoring. With the increasing number of vehicles, traffic congestion and traffic accidents are common, resulting in an increasingly serious traffic management task. By realizing the real-time detection, recognition and tracking of vehicles in the drone intelligent scene, a large amount of human resources can be saved to easily manage traffic tasks. At the same time, it plays an important role in aspects such as urban management and planning, real-time road condition acquisition, and highway patrol for the intelligentization of traffic management. The drone intelligent scene monitoring has a wide field of view and flexible perspectives. However, this makes the environment and lighting situation more complex, and the vehicles under the drone's perspective will have defects such as small size, low resolution, and large scale changes due to factors such as imaging conditions and imaging distances, resulting in problems of weak discriminative features and information imbalance for the vehicles.
[0003] Currently, the research methods mainly enhance the discriminative features of vehicles and reduce the information imbalance problem by introducing image or feature super-resolution, feature fusion models, and feature pyramid networks. The image or feature super-resolution method starts from the image or feature perspective and processes vehicle targets of different scales by using input images or feature images of different sizes, reflecting the multi-scale feature representation of vehicle targets, increasing the adaptability of the network to targets of different scales, and improving the discriminative features of small and weak vehicles. However, these methods will greatly increase the computational cost and inference time and cannot meet the real-time requirements of actual applications. The method based on feature fusion uses the feature combinations of different levels of the network and different connection methods to enhance the feature representation of small and weak vehicles. However, the relevant effective information extracted by these methods is still limited, and a large amount of redundant information will be introduced during the fusion and connection process, which is not conducive to the detection of small and weak vehicles. The method based on the feature pyramid network can alleviate the problem that the shallow features of the network have spatial information but no semantic information by transmitting global context semantic information. At the same time, it can also transmit global spatial information to reduce the problem that the deep features of the network have semantic information but no spatial information. However, whether it is global context semantic information or global spatial information, it will bring more interference information for small and weak vehicles under the drone's perspective, thus affecting the accurate detection of vehicles. Currently, similar drone vehicle detection algorithms lack local information beneficial to the detection and recognition of small target vehicles and a multi-attention mechanism for enhancing the features of small and weak vehicles.
[0004] In summary, the problems and disadvantages of the current drone vehicle detection algorithms are mainly as follows:
[0005] (1) Image or feature super-resolution methods improve the vehicle feature extraction ability by supplementing vehicle information, but they are prone to introducing incorrect information during the training process, which is disadvantageous for the detection of small and weak vehicles composed of a small number of pixels. At the same time, separate training for different-scale targets has a high computational complexity and cannot meet the application requirements of real-time detection.
[0006] (2) The method based on feature fusion still extracts limited relevant effective information and introduces a large amount of redundant information during the fusion and connection processes, bringing unnecessary interference to the detection of small and weak vehicles from the UAV perspective.
[0007] (3) The method based on the feature pyramid network only considers the transmission of global information during the information transmission process. However, most of the vehicles from the UAV perspective are small-sized targets, and the introduction of global information will be accompanied by the introduction of a large amount of background noise, which will damage the performance of small target detection and thus limit the improvement of vehicle detection accuracy. Summary of the Invention
[0008] The purpose of the present invention is to provide a real-time detection algorithm for UAV vehicles to solve the problems mentioned in the above background technology. That is, image or feature super-resolution methods improve the vehicle feature extraction ability by supplementing vehicle information, but they are prone to introducing incorrect information during the training process, which is disadvantageous for the detection of small and weak vehicles composed of a small number of pixels. At the same time, separate training for different-scale targets has a high computational complexity and cannot meet the application requirements of real-time detection. The method based on feature fusion still extracts limited relevant effective information and introduces a large amount of redundant information during the fusion and connection processes, bringing unnecessary interference to the detection of small and weak vehicles from the UAV perspective. The method based on the feature pyramid network only considers the transmission of global information during the information transmission process. However, most of the vehicles from the UAV perspective are small-sized targets, and the introduction of global information will be accompanied by the introduction of a large amount of background noise, which will damage the performance of small target detection and thus limit the improvement of vehicle detection accuracy.
[0009] To achieve the above purpose, the present invention provides the following technical solution: A real-time detection algorithm for UAV vehicles, including a real-time detection algorithm. The real-time detection algorithm includes a basic network, an efficient feature pyramid network, a multi-attention mechanism network, and a shared detection network. The basic network is connected to the efficient feature pyramid network, the efficient feature pyramid network is connected to the multi-attention mechanism network, and the multi-attention mechanism network is connected to the shared detection network.
[0010] The basic network has a powerful representation ability for the features of each convolutional layer, and its output is used as the input of the efficient feature pyramid network to further enhance the vehicle feature representation; the efficient feature pyramid network provides global and local context semantic information for the shallow features of the network, provides a rich receptive field for the vehicle, thereby enhancing the feature representation of small and weak vehicles, and alleviating to a certain extent the imbalance problem that the shallow features of the network have spatial information but no semantic information. The multi-attention mechanism network suppresses these redundant information, and the shared detection network provides global and local spatial information for the deep features of the network, minimizing the loss of local spatial information during the feature transmission process, and reducing the imbalance problem that the deep features of the network have semantic information but no spatial information.
[0011] As a preferred technical solution of the present invention, the basic network is ResNet-50.
[0012] As a preferred technical solution of the present invention, the efficient feature pyramid network includes an adaptive top-down propagation path and a bottom-up propagation path.
[0013] As a preferred technical solution of the present invention, the multi-attention mechanism network includes a multi-pooling channel attention mechanism and a multi-pooling spatial attention mechanism. Both the inside of the multi-pooling channel attention mechanism and the inside of the multi-pooling spatial attention mechanism are provided with average pooling and max pooling. The average pooling is used to calculate the average feature value of each channel and space in the feature map, effectively reducing background interference and redundant information, and the max pooling is used to calculate the maximum feature value of each channel in the feature map, so as to extract key foreground features.
[0014] As a preferred technical solution of the present invention, the operation steps of the multi-attention mechanism network are specifically as follows:
[0015] S1. The first attention operation: The multi-pooling channel attention mechanism performs an attention operation;
[0016] S2. The second attention operation: The multi-pooling spatial attention mechanism performs an attention operation;
[0017] S3. Discriminate features: Combine the two attention operations to obtain the discriminative features of the vehicle target.
[0018] As a preferred technical solution of the present invention, the specific operation of the multi-pooling channel attention mechanism in S1 is: the features respectively pass through global average pooling and global max pooling operations to obtain channel dimension features and Both of these features use 1x1 convolution operations to reduce the dimension to 1 / 16 of the original, and then pass through 1x1 convolution operations to restore to the original size to obtain features and Will and Add them together and pass the result through the Sigmoid activation function to obtain the channel weight F cam , and finally multiply the original feature F by the channel weight F cam to obtain the channel attention feature F c .
[0019] As a preferred technical solution of the present invention, the specific operation of the multi-pooling spatial attention mechanism in S2 is as follows: The features are respectively subjected to global average pooling and global maximum pooling operations to obtain spatial dimension features and Then these two features are concatenated, and then the concatenated features are subjected to 3x3 convolution and ReLU activation function operations to obtain features Next, and are added together and passed through 3x3 convolution and Sigmoid activation function to obtain the spatial weight F sam , and finally the original feature F is multiplied by the spatial weight F sam to obtain the spatial attention feature F s .
[0020] As a preferred technical solution of the present invention, the specific operation of the discriminative feature in S3 is as follows: Add the channel attention feature F c and the spatial attention feature F s to obtain the discriminative feature F of the vehicle target h .
[0021] As a preferred technical solution of the present invention, the specific operation steps of the shared detection network are as follows:
[0022] Step 1. Vehicle information classification: The classification branch continuously uses four 3×3 convolutions on the predicted feature map to obtain vehicle classification information;
[0023] Step 2. Vehicle information localization: The regression branch continuously uses four 3×3 convolutions on the predicted feature map to obtain the localization information of the vehicle bounding box, and the operation of the anchor-free mechanism is adopted to regress the bounding box of the UAV vehicle;
[0024] Step 3. Output the correct result: Use the classification loss function and the regression loss function to train the samples respectively, where the classification function uses the focal loss and the binary cross-entropy loss, and the regression function uses the GIoU loss. Predict the C-category number vector of the classification label and the 4-dimensional vector of the bounding box in the last layer, and finally obtain the correct output result.
[0025] As a preferred technical solution of the present invention, the expressions of the focal loss and the binary cross-entropy loss in Step 3 are:
[0026] CE(p, y) = -y i log(p) - (1 - y i )log(1 - p),
[0027] Where yi ∈ {0, 1} represents the label, p represents the predicted probability. To determine whether the anchor sample is easy to predict, the most direct way is to analyze p. The smaller the gap between the predicted value p and the true value, the easier it is to predict the sample. Therefore, for convenience, assume pt is as follows:
[0028]
[0029] Then the cross-entropy loss is expressed as:
[0030] CE(p, y) = -log(p t ),
[0031] Where the larger pt is, and when it is less than 1, the smaller the result after taking the log operation, the smaller the impact of this loss on the network, indicating that the gap between the predicted result and the true result is smaller, and the sample is easier to predict. Therefore, the size of pt is used to represent the difficulty level on one side of the sample;
[0032] The GIoU loss in step three is expressed as:
[0033]
[0034] Where IoU represents the intersection over union of the predicted box and the ground truth box, C represents the area of the smallest closed rectangle enclosing the predicted box and the ground truth box, and U represents the union area of the predicted box and the ground truth box.
[0035] Compared with the prior art, the beneficial effects of the present invention are:
[0036] 1. While transmitting global information, local information transmission is introduced, and an efficient feature pyramid is proposed based on the feature pyramid to provide appropriate global and local semantic and spatial information for small and weak vehicle targets to solve the information imbalance problem; at the same time, aiming at the interference of a large amount of redundant information and background noise;
[0037] 2. The present invention aggregates multiple attention mechanisms to highlight the key features of foreground vehicles while suppressing these interferences;
[0038] 3. The algorithm of the efficient feature pyramid aggregating multiple attention mechanisms designed by the present invention is also a lightweight network, facilitating the real-time detection of unmanned aerial vehicle vehicles;
[0039] 4. To enhance the discriminative features of small and weak vehicles, reduce the information imbalance and information loss problems, and improve the detection accuracy and real-time stability of unmanned aerial vehicle vehicles;
[0040] 5. By integrating the adaptive top-down propagation path and the effective bottom-up propagation path, an efficient bidirectional feature pyramid network is designed. This network provides global and local shallow spatial information for the deep features of the network, and at the same time provides global and local context semantic information for the shallow features of the network, thereby effectively reducing the imbalance problem between spatial information and semantic information at different levels;
[0041] 6. By integrating the multi-pooling channel attention mechanism and the multi-pooling spatial attention mechanism, a multi-attention mechanism network is designed to focus on the target attention in the spatial and channel dimensions, effectively extract finer features, and further enhance the discriminative features of small and weak vehicles;
[0042] 7. The proposed algorithm achieves state-of-the-art detection performance and real-time detection speed on three public datasets. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic diagram of the architecture of the real-time detection algorithm of the present invention;
[0044] Figure 2 It is a framework diagram of the UAV vehicle detection algorithm of the present invention;
[0045] Figure 3 It is an operation flow chart of the multi-attention mechanism network of the present invention;
[0046] Figure 4 It is a framework diagram of the multi-attention mechanism network of the present invention;
[0047] Figure 5 It is an operation flow chart of the shared detection network of the present invention;
[0048] Figure 6 It is a framework diagram of the efficient feature pyramid network of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work shall fall within the protection scope of the present invention.
[0050] Please refer to Figure 1-6 , the present invention provides a UAV vehicle real-time detection algorithm, including a real-time detection algorithm. The real-time detection algorithm includes a basic network, an efficient feature pyramid network, a multi-attention mechanism network, and a shared detection network. The basic network is connected to the efficient feature pyramid network, the efficient feature pyramid network is connected to the multi-attention mechanism network, and the multi-attention mechanism network is connected to the shared detection network;
[0051] The basic network has a powerful representation ability for the features of each convolutional layer, and its output is used as the input of the efficient feature pyramid network to further enhance the vehicle feature representation; the efficient feature pyramid network provides global and local context semantic information for the shallow features of the network, provides a rich receptive field for the vehicle, thereby enhancing the feature representation of small and weak vehicles, and alleviating to a certain extent the imbalance problem that the shallow features of the network have spatial information but no semantic information. The multi-attention mechanism network suppresses these redundant information, and the shared detection network provides global and local spatial information for the deep features of the network, minimizing the loss of local spatial information during the feature transmission process, and reducing the imbalance problem that the deep features of the network have semantic information but no spatial information.
[0052] The basic network is ResNet-50.
[0053] The efficient feature pyramid network includes an adaptive top-down propagation path and a bottom-up propagation path.
[0054] The multi-attention mechanism network includes a multi-pooling channel attention mechanism and a multi-pooling spatial attention mechanism. Both the inside of the multi-pooling channel attention mechanism and the inside of the multi-pooling spatial attention mechanism are equipped with average pooling and max pooling. The average pooling is used to calculate the average feature value of each channel and space in the feature map, effectively reducing background interference and redundant information, and the max pooling is used to calculate the maximum feature value of each channel and space in the feature map, so as to extract key foreground features.
[0055] The operation steps of the multi-attention mechanism network are specifically as follows:
[0056] S1. The first attention operation: The multi-pooling channel attention mechanism performs an attention operation;
[0057] S2. The second attention operation: The multi-pooling spatial attention mechanism performs an attention operation;
[0058] S3. Discriminant features: Combine the two attention operations to obtain the discriminant features of the vehicle target.
[0059] The specific operation of the multi-pooling channel attention mechanism in S1 is as follows: The features respectively pass through the global average pooling and global max pooling operations to obtain the channel dimension features and Both of these features use 1x1 convolution operations to reduce the dimension to 1 / 16 of the original, and then pass through 1x1 convolution operations to restore to the original size to obtain the features and Add and and pass the result through the Sigmoid activation function to obtain the channel weight F cam, finally, multiply the original feature F by the channel weight F cam to obtain the channel attention feature F c .
[0060] The specific operation of the multi-pooling spatial attention mechanism in S2 is as follows: the feature is respectively subjected to global average pooling and global maximum pooling operations to obtain the spatial dimension features and Then these two features are concatenated, and then the concatenated feature is subjected to 3x3 convolution and activation function ReLU operations to obtain the feature Next, and are added and subjected to 3x3 convolution and Sigmoid activation function to obtain the spatial weight F sam , finally, multiply the original feature F by the spatial weight F sam to obtain the spatial attention feature F s .
[0061] The specific operation of the discriminative feature in S3 is as follows: add the channel attention feature F c and the spatial attention feature F s to obtain the discriminative feature F h of the vehicle target.
[0062] The specific operation steps of the shared detection network are as follows:
[0063] Step 1. Vehicle information classification: The classification branch continuously uses four 3×3 convolutions on the predicted feature map to obtain vehicle classification information;
[0064] Step 2. Vehicle information localization: The regression branch continuously uses four 3×3 convolutions on the predicted feature map to obtain the localization information of the vehicle bounding box, and the operation of the anchor-free mechanism is adopted to regress the bounding box of the UAV vehicle;
[0065] Step 3. Output the correct result: Use the classification loss function and the regression loss function to train the samples respectively. The classification function uses the focal loss and the binary cross-entropy loss, and the regression function uses the GIoU loss. Predict the C-category number vector of the classification label and the 4-dimensional vector of the bounding box in the last layer, and finally obtain the correct output result.
[0066] The expressions of the focal loss and the binary cross-entropy loss in Step 3 are as follows:
[0067] CE(p,y)=-y i log(p)-(1-y i )log(1-p),
[0068] Among them, yi ∈ {0, 1} represents the label, p represents the predicted probability. To determine whether the anchor sample is easy to predict, the most direct way is to analyze p. The smaller the gap between the predicted value p and the true value, the easier the sample is to predict. Therefore, for convenience, assume pt is:
[0069]
[0070] Then the cross-entropy loss is expressed as:
[0071] CE(p, y) = -log(p t )
[0072] Among them, the larger pt is, when it is less than 1, the result of the corresponding log operation is smaller, the impact of this loss on the network is smaller, indicating that the gap between the predicted result and the true result is smaller, and the sample is easier to predict. Therefore, the size of pt is used to represent the difficulty level on one side of the sample;
[0073] In step three, the GIoU loss is expressed as:
[0074]
[0075] Among them, IoU represents the intersection over union of the predicted box and the true box, C represents the area of the smallest closed rectangle of the predicted box and the true box, and U represents the union area of the predicted box and the true box.
[0076] In the present invention, ResNet-50 is used as the basic network, which has a powerful representation ability for the features of each convolutional layer. Its output is used as the input of the efficient feature pyramid network to further enhance the vehicle feature representation. The efficient feature pyramid network is composed of an adaptive top-down propagation path, a multi-attention mechanism network, and an effective bottom-up propagation path. Among them, ATDP provides global and local context semantic information for the shallow features of the network, provides a rich receptive field for the vehicle, thereby enhancing the feature representation of small and weak vehicles, and alleviating to a certain extent the imbalance problem that the shallow features of the network have spatial information but no semantic information. Since ATDP will bring some redundant information to the network while introducing context information, the present invention connects the output of ATDP to the multi-attention mechanism network to suppress this redundant information. Subsequently, the output features of HAA-Net are input into the effective bottom-up propagation path. EBUP provides global and local spatial information for the deep features of the network, minimizes the loss of local spatial information during the feature transmission process, and reduces the imbalance problem that the deep features of the network have semantic information but no spatial information. Among them, the introduction of HAA-Net in the efficient feature pyramid network also reduces the training burden of EBUP. The present invention integrates the adaptive top-down propagation path and the effective bottom-up propagation path into the efficient feature pyramid network, effectively reducing the information imbalance between different levels of spatial and semantic information. In the efficient feature pyramid network, the introduction of global and local information of context semantic information and spatial information will inevitably bring interference from redundant and background information. Therefore, a multi-attention mechanism network is connected to the output of each layer of the efficient feature pyramid network to suppress this useless information, highlight the key information of the foreground vehicle, and improve the discriminative features of the vehicle target. The efficient feature pyramid network is integrated by an adaptive top-down propagation path and an effective bottom-up propagation path. The adaptive top-down propagation path provides global and local context semantic information for the shallow features, while the effective bottom-up propagation path provides global and local spatial information for the deep features. Therefore, the efficient feature pyramid network effectively reduces the imbalance between spatial information and semantic information at different levels. Secondly, the multi-attention mechanism network is designed by parallel integrating the multi-pooling channel attention mechanism and the multi-pooling spatial attention mechanism. It focuses on the target attention in the spatial and channel dimensions to reduce background interference and highlight foreground information. Specifically, the multi-pooling channel attention mechanism extracts important channels and their correlations, while the multi-pooling spatial attention mechanism captures important spatial relationships in the feature map. Therefore, the multi-attention mechanism network effectively extracts finer features and further enhances the discriminative features of small and weak vehicles. The latest deep learning technology is adopted, with reliable performance, meeting real-time detection, and the performance indicators reaching the international advanced level.
[0077] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A real-time detection algorithm for unmanned aerial vehicle, including a real-time detection algorithm, characterized in that: The real-time detection algorithm includes a basic network, an efficient feature pyramid network, a multi-attention mechanism network and a shared detection network, wherein the basic network is connected to the efficient feature pyramid network, the efficient feature pyramid network is connected to the multi-attention mechanism network, and the multi-attention mechanism network is connected to the shared detection network; The basic network has a strong representation capability for the features of each convolutional layer, and its output is used as the input of the efficient feature pyramid network to further enhance the vehicle feature representation; The efficient feature pyramid network provides global and local contextual semantic information for the shallow features of the network and provides a rich receptive field for the vehicle, thereby enhancing the feature representation of weak vehicles and, to a certain extent, alleviating the imbalance problem that the shallow features of the network have spatial information but no semantic information. The multi-attention mechanism network suppresses these redundant information and the shared detection network provides global and local spatial information for the deep features of the network, minimizing the loss of local spatial information during feature transmission and reducing the imbalance problem that the deep features of the network have semantic information but no spatial information.
2. The real-time detection algorithm for unmanned aerial vehicle according to claim 1 is characterized by: The basic network is ResNet-50.
3. The real-time detection algorithm for unmanned aerial vehicle according to claim 1 is characterized by: The efficient feature pyramid network includes an adaptive top-down propagation path and a bottom-up propagation path.
4. The real-time detection algorithm for unmanned aerial vehicle according to claim 1 is characterized in that: The multi-attention mechanism network includes a multi-pooling channel attention mechanism and a multi-pooling space attention mechanism. The multi-pooling channel attention mechanism and the multi-pooling space attention mechanism are both equipped with average pooling and maximum pooling. The average pooling is used to calculate the average eigenvalue of each channel and space in the feature map, effectively reducing background interference and redundant information. The maximum pooling is used to calculate the maximum eigenvalue of each channel and space in the feature map, thereby extracting key foreground features.
5. The real-time detection algorithm for unmanned aerial vehicle according to claim 1 is characterized by: The operation steps of the multi-attention mechanism network are specifically as follows: S1. First attention operation: multi-pooled channel attention mechanism performs attention operation; S2, the second attention operation: multi-pooling spatial attention mechanism performs attention operation; S3, discriminative features: Combining the two attention operations to obtain the discriminative features of the vehicle target.
6. The real-time detection algorithm for unmanned aerial vehicle according to claim 5 is characterized by: The specific operation of the multi-pooling channel attention mechanism in S1 is: the features are respectively subjected to global average pooling and global maximum pooling operations to obtain channel dimension features and Both features use 1x1 convolution operation to reduce the dimension to 1 / 16 of the original size, and then restore it to the original size through 1x1 convolution operation to obtain the features and Will and Add them and pass the result through the Sigmoid activation function to get the channel weight F cam Finally, the original feature F and the channel weight F cam Multiply to get the channel attention feature F c .
7. The real-time detection algorithm for unmanned aerial vehicle according to claim 5 is characterized by: The specific operation of the multi-pooling spatial attention mechanism in S2 is: the features are respectively subjected to global average pooling and global maximum pooling operations to obtain spatial dimension features and Then the two features are concatenated, and then the concatenated features are subjected to 3x3 convolution and activation function ReLU operation to obtain the feature Next will and Add and pass 3x3 convolution and Sigmoid activation function to get the spatial weight F sam Finally, the original feature F and the spatial weight F sam Multiply to get the spatial attention feature F s .
8. The real-time detection algorithm for unmanned aerial vehicle according to claim 5 is characterized by: The specific operation of the discriminative feature in S3 is: c and spatial attention feature F s Add the discriminative features F of the vehicle target h .
9. The real-time detection algorithm for unmanned aerial vehicle according to claim 1 is characterized by: The shared detection network operation steps are specifically as follows: Step 1: Vehicle information classification: The classification branch uses four 3×3 convolutions on the predicted feature map to obtain vehicle classification information; Step 2: Vehicle information positioning: The regression branch uses four 3×3 convolutions on the predicted feature map to obtain the positioning information of the vehicle bounding box, in which the anchor-free mechanism is used to regress the bounding box of the drone vehicle; Step 3. Output the correct result: Use the classification loss function and regression loss function to train the samples separately. The classification function uses focal loss and binary cross entropy loss, and the regression function uses GIoU loss. In the last layer, the C-dimensional vector of the classification label and the 4-dimensional vector of the bounding box are predicted to obtain the correct output result.
10. The real-time detection algorithm for unmanned aerial vehicle according to claim 9 is characterized in that: The expressions of focal loss and binary cross entropy loss in step 3 are: CE(p,y)=-y i log(p)-(1-y i )log(1-p), Among them, yi∈{0,1} represents the label, p represents the predicted probability. The most direct way to determine whether the anchor point sample is easy to predict is to analyze p. The smaller the difference between the predicted value p and the true value, the easier it is to predict the sample. Therefore, for convenience, we assume that pt is: The cross entropy loss is then expressed as: CE(p,y)=-log(p t ), The larger the pt is, when it is less than 1, the smaller the result after the corresponding log operation is, the smaller the impact of the loss on the network is, the smaller the gap between the surface prediction result and the actual result is, and the easier the sample is to predict. Therefore, the size of pt is used to indicate the difficulty of the sample side. The GIoU loss in step 3 is expressed as: Where loU represents the intersection-over-union ratio of the predicted box and the true box, C represents the minimum enclosed rectangular area of the predicted box and the true box, and U represents the union area of the predicted box and the true box.