Unmanned aerial vehicle detection and tracking method based on convolutional neural network

By combining convolutional neural networks and extended Kalman filters in the drone detection system, the real-time and cross-camera tracking issues on embedded devices are solved, high-precision drone detection and tracking is achieved, the deployment difficulty of edge devices is reduced, and the accuracy of trajectory tracking is improved.

CN120708005APending Publication Date: 2025-09-26HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510786373.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing drone detection and tracking systems are difficult to achieve real-time processing on embedded devices, have low cross-camera tracking accuracy, poor adaptability to edge devices, and traditional Kalman filtering algorithms are difficult to accurately predict the nonlinear motion trajectory of drones, which can easily cause tracking drift.

Method used

A drone detection method based on convolutional neural networks is adopted, combined with extended Kalman filtering and a lightweight target detection network, and progressive partial convolution and attention mechanisms are used to achieve cross-camera target matching and tracking, reduce computational complexity, and improve detection accuracy and real-time performance.

Benefits of technology

Real-time drone detection and tracking is achieved on embedded devices, which improves the accuracy of cross-camera tracking, reduces the difficulty of deploying edge devices, and improves trajectory tracking accuracy through nonlinear motion modeling, overcoming the tracking drift defects of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708005A_ABST
    Figure CN120708005A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle detection and tracking method based on a convolutional neural network, and belongs to the technical field of computer vision. According to the invention, the problems of contradiction between detection tracking precision and real-time performance, difficulty in realizing cross-camera tracking, easiness in generating tracking drift and difficulty in deploying edge equipment in the existing method are solved. According to the invention, based on the lightweight target detection network, the redundant feature generation energy consumption can be significantly reduced, the reasoning speed can be significantly improved, and the real-time bottleneck problem of an existing algorithm on embedded equipment is effectively solved while the detection precision is ensured; the difficulty of deploying the edge equipment can be reduced. And the tracking precision of the dynamic trajectory of the unmanned aerial vehicle is remarkably improved through nonlinear motion modeling based on extended Kalman filtering, and the tracking drift defect of a traditional method is overcome. And cross-domain tracking can be carried out, so that the accuracy of target re-identification is improved. The method can be applied to the field of unmanned aerial vehicle detection and tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a drone detection and tracking method based on convolutional neural networks. Background Art

[0002] Drone detection and tracking systems integrate target detection, multi-target tracking, and communication networking technologies. Current mainstream technologies rely on deep learning models (such as YOLO and Faster R-CNN) for target detection and combine them with data association algorithms (such as DeepSORT) for multi-target tracking. However, existing technologies still have the following shortcomings: 1. Conflict between accuracy and real-time performance: Since the high-precision detection and tracking of existing algorithms rely on complex network structures, it is difficult to achieve real-time processing on embedded devices (such as Jetson Orin).

[0003] 2. Limitations of cross-camera tracking: Traditional methods rely on a single camera perspective and lack the ability to align cross-domain features, resulting in low target re-identification accuracy.

[0004] 3. Poor adaptability of edge devices: The existing model has a large number of parameters and cannot adapt to IoT nodes with limited computing resources, making it difficult to meet the needs of multi-platform collaborative supervision.

[0005] 4. Insufficient nonlinear dynamic tracking: UAV motion has nonlinear characteristics. Traditional Kalman filter algorithms are difficult to accurately predict trajectories and are prone to tracking drift.

[0006] Therefore, it is necessary to propose a new drone detection and tracking method to solve the above problems. Summary of the Invention

[0007] The purpose of this invention is to solve the problems of existing methods such as the contradiction between detection and tracking accuracy and real-time performance, difficulty in achieving cross-camera tracking, easy tracking drift, and difficulty in edge device deployment, and propose a drone detection and tracking method based on convolutional neural networks.

[0008] The technical solution adopted by the present invention to solve the above technical problems is: a drone detection and tracking method based on convolutional neural network, the method specifically comprising the following steps: Step 1: According to the first The tracking results obtained from the frame images are subjected to extended Kalman filtering to predict the position of each UAV target; Step 2: Use the built target detection model to detect the first Frame images are used for drone target detection; Step 3: Match the predicted result of the drone target position with the detection result of step 2 to obtain the drone target tracking result; Step 4: Perform extended Kalman filtering based on the UAV target tracking results in step 3 to continue predicting the position of each UAV target; Step 5: , return to step 2 and continue until the entire tracking process is completed.

[0009] Furthermore, the target detection model includes a first convolutional layer, a first SFP2 module, a second SFP2 module, a first SFP1 module, a third SFP2 module, a second SFP1 module, a fourth SFP2 module, a third SFP1 module, a first attention module, a second convolutional layer, a first upsampling module, a first CSF module, a third convolutional layer, a second upsampling module, a second attention module, a second CSF module, a fourth convolutional layer, a third attention module, a third CSF module, a fifth convolutional layer, a fourth attention module, a fourth CSF module, a sixth convolutional layer, a seventh convolutional layer and an eighth convolutional layer, wherein: The input of the target detection model is used as the input of the first convolutional layer, and the output of the first convolutional layer is used as the input of the first SFP2 module; Use the output of the first SFP2 module as the input of the second SFP2 module, and then use the output of the second SFP2 module as the input of the first SFP1 module; Use the output of the first SFP1 module as the input of the third SFP2 module, and then use the output of the third SFP2 module as the input of the second SFP1 module; Use the output of the second SFP1 module as the input of the fourth SFP2 module, and then use the output of the fourth SFP2 module as the input of the third SFP1 module; The output of the third SFP1 module is used as the input of the first attention module, and the output of the first attention module is used as the input of the second convolutional layer; The output of the second convolutional layer is used as the input of the first upsampling module, and the output of the first upsampling module is concatenated with the output of the second SFP1 module, and the concatenated result is used as the input of the first CSF module; The output of the first CSF module is used as the input of the third convolutional layer, and the output of the third convolutional layer is used as the input of the second upsampling module. The output of the second upsampling module is then concatenated with the output of the first SFP1 module, and the concatenated result is used as the input of the second attention module. The output of the second attention module is used as the input of the second CSF module, and then the output of the second CSF module is used as the input of the fourth convolutional layer. The output of the third convolutional layer is spliced ​​with the output of the fourth convolutional layer, and the spliced ​​result is used as the input of the third attention module; The output of the third attention module is used as the input of the third CSF module, and then the output of the third CSF module is used as the input of the fifth convolutional layer. The output of the fifth convolutional layer is spliced ​​with the output of the second convolutional layer, and the spliced ​​result is used as the input of the fourth attention module; The output of the fourth attention module is used as the input of the fourth CSF module; Then use the output of the second CSF module as the input of the sixth convolutional layer, the output of the third CSF module as the input of the seventh convolutional layer, and the output of the fourth CSF module as the input of the eighth convolutional layer. The target detection results are output through the sixth, seventh, and eighth convolutional layers.

[0010] Furthermore, the working process of the first SFP1 module is as follows: The input of the first SFP1 module passes through the ninth convolutional layer and the first progressive partial convolutional layer in sequence, the output of the first progressive partial convolutional layer is spliced ​​with the input of the first SFP1 module, and then the splicing result is shuffled, and the shuffled result is used as the output of the first SFP1 module.

[0011] Furthermore, the working process of the first SFP2 module is as follows: The first SFP2 module includes two parallel branches. In one branch, the input of the first SFP2 module passes through the first depth-separable convolution layer and the tenth convolution layer in sequence. In the other branch, the input of the first SFP2 module passes through the eleventh convolution layer, the second depth-separable convolution layer and the twelfth convolution layer in sequence. After the outputs of the two branches are spliced, the splicing results are shuffled and the shuffled results are used as the output of the first SFP2 module.

[0012] Furthermore, the working process of the first CSF module is: The first CSF module includes two parallel branches; In one branch, the input of the first CSF module passes through the thirteenth convolutional layer; In another branch, the input of the first CSF module is first passed through the fourteenth convolutional layer, and then the output of the fourteenth convolutional layer is passed through the second progressive partial convolutional layer and the fifteenth convolutional layer in sequence. Finally, the output of the fourteenth convolutional layer and the output of the fifteenth convolutional layer are concatenated and added to obtain the concatenation result a; The splicing result a is spliced ​​with the output of the thirteenth convolutional layer to obtain the splicing result b, and then the splicing result b passes through the sixteenth convolutional layer, and the output of the sixteenth convolutional layer is used as the output of the first CSF module.

[0013] Furthermore, the first attention module, the second attention module and the third attention module are all NAM attention.

[0014] Furthermore, before the image captured by the camera is input into the target detection model, it needs to first be resized and pixel value normalized, and then the processed image is used as the input of the target detection model; The target detection model outputs the detection frame position of the drone target and the confidence of the detection frame. The detection frame position is then mapped to the original image captured by the camera to obtain the position of the detection frame in the original image. The detection frame in the original image is then subjected to NMS to obtain the final detection frame position of each drone target.

[0015] Furthermore, in step 3, the area covered by each camera is matched separately. The specific matching process of the area covered by any camera is as follows: Step 3.1. For prediction box, calculate the first detected box in the image The target detection box and Cosine distance of predicted boxes and Mahalanobis distance , and traverse all the prediction frames and target detection frames corresponding to the current camera; According to the cosine distance and Mahalanobis distance Calculate the cost matrix: in, Indicates the cost matrix Rank Elements of the column, is a hyperparameter; Step 32: Use the Hungarian algorithm and cost matrix to perform cascade matching, determine the matching solution that minimizes the total matching cost, and obtain the initial matching result between the prediction box and the detection box; Execute step 33 for the prediction frame and detection frame that are initially matched successfully, and execute step 34 for the prediction frame and detection frame that are not initially matched successfully; Step 3. Perform IOU matching on a pair of prediction boxes and detection boxes on the initial matching; If the IOU match between the prediction box and the detection box is successful, then determine whether the tracking results of the previous two frames of the current prediction box are all IOU matched successfully. If so, the matching result of the current prediction box and the detection box is confirmed. Otherwise, the matching result of the current prediction box and the detection box is tentative. If the IOU matching between the prediction box and the detection box is unsuccessful, then perform steps three and four on the prediction box and the detection box whose IOU matching is unsuccessful; Step 34: For the prediction boxes and detection boxes that are not initially matched, as well as the prediction boxes and detection boxes that fail to IOU matching in step 33, perform IOU matching on the detection boxes and prediction boxes in descending order of the confidence of the detection boxes, and then perform step 35 on the matching results; Step 35: For a pair of prediction frames and detection frames with successful IOU matching, if the tracking results of the previous two frames of the current prediction frame are both IOU-matched, the matching result of the current prediction frame and the detection frame is confirmed; otherwise, the matching result of the current prediction frame and the detection frame is tentative. For unmatched detection frames, if the detection frame is not located in the common coverage area of ​​adjacent cameras, the unmatched detection frame is considered a new target and initialized as a new track, and a unique ID is assigned to the new target. If the detection frame is located in the common coverage area of ​​adjacent cameras, a cross-camera matching and tracking process is required. For unsuccessful prediction frames, when the prediction frame is not located in the common coverage area of ​​adjacent cameras, if the tracking results of the first two frames of the current prediction frame do not have successful IOU matching, the current prediction frame will be deleted; otherwise, the current prediction frame will be retained; when the prediction frame is located in the common coverage area of ​​adjacent cameras, a cross-camera matching and tracking process is required.

[0016] Furthermore, the specific process of cross-camera matching and tracking is as follows: For the detection box: (1) If there is an unmatched prediction frame in the coverage area of ​​the adjacent camera, and the detection frame and the unmatched prediction frame are both in the common coverage area of ​​the two cameras, then the detection frame is matched with the unmatched prediction frame in the coverage area of ​​the adjacent camera at the same time. If the IOU match is successful, a pair of matched detection frame and prediction frame is obtained. If the IOU match is unsuccessful, the unmatched detection frame is regarded as a new target and initialized as a new track, and a unique ID is assigned to the new target. (2) If there is no unmatched prediction frame in the adjacent camera coverage area, the unmatched detection frame is regarded as a new target and initialized as a new trajectory, and a unique ID is assigned to the new target; For the prediction box: (1) If there is an unmatched detection frame in the coverage area of ​​the adjacent camera, and the predicted frame and the unmatched detection frame are both in the common coverage area of ​​the two cameras, the predicted frame is matched with the detection frame in the coverage area of ​​the adjacent camera at the same time. If the IOU match is successful, a pair of successfully matched detection frame and predicted frame is obtained. If the IOU match is unsuccessful and the tracking results of the first two frames of the current predicted frame are not successfully matched by IOU, the current predicted frame is deleted. Otherwise, the current predicted frame is retained. (2) If there is no unmatched detection frame in the adjacent camera coverage area, and the tracking results of the first two frames of the current prediction frame do not have an IOU match, the current prediction frame is deleted; otherwise, the current prediction frame is retained.

[0017] Furthermore, the working process of the first progressive partial convolutional layer is: The number of channels of the first progressive partial convolution layer’s input passing through the convolution kernel is c p The output of the seventeenth convolutional layer is then concatenated with the input of the first progressive partial convolutional layer, and the concatenated result is used as the output of the first progressive partial convolutional layer.

[0018] The beneficial effects of the present invention are: The method proposed in this paper proposes a lightweight target detection network based on progressive partial convolution and attention mechanisms. This lightweight target detection network significantly reduces the energy consumption of redundant feature generation and significantly improves inference speed. While ensuring detection accuracy, it effectively solves the real-time bottleneck problem of existing algorithms on embedded devices. Furthermore, the lightweight processing of the target detection network can reduce the difficulty of edge device deployment. Furthermore, through nonlinear motion modeling based on the extended Kalman filter, the tracking accuracy of the dynamic trajectory of the drone is significantly improved, overcoming the tracking drift defects of traditional methods. Furthermore, cross-domain tracking is possible, improving the accuracy of target re-identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a schematic structural diagram of the target detection model of the present invention; Figure 2 is a schematic structural diagram of the first SFP1 module; Figure 3 is a schematic structural diagram of the first SFP2 module; Figure 4 It is a structural diagram of the CSF module; Figure 5 This is a comparison chart of three types of convolution; (a) represents a depth-wise separable convolutional layer; (b) represents a normal convolutional layer; (c) represents a progressive partial convolutional layer; Figure 6 is a flow chart of the matching pursuit algorithm of the present invention; Figure 7 This is a flowchart of the Hungarian matching algorithm; Figure 8 is the average precision curve of the lightweight object detection model; Figure 9 is the test result of target detection; Figure 10 It is the image one of the tracking process; Figure 11 This is the second image of the tracking process. DETAILED DESCRIPTION

[0020] Specific embodiment 1: This embodiment describes a method for detecting and tracking drones based on a convolutional neural network, and the method specifically includes the following steps: Step 1: According to the first The tracking results obtained from the frame image are extended Kalman filtering to predict the position of each drone target (the position that needs to be predicted here is the position of the drone target after the first frame image is collected). The moment of the frame image, the position of each UAV target); Step 2: Use the built target detection model to detect the first Frame images are used for drone target detection; Step 3: Match the predicted result of the drone target position with the detection result of step 2 to obtain the drone target tracking result; Step 4: Perform extended Kalman filtering based on the UAV target tracking results in step 3 to continue predicting the position of each UAV target; Step 5: , return to step 2 and continue until the entire tracking process is completed.

[0021] Specific implementation method 2: Combination Figure 1 This embodiment is different from the first embodiment in that the target detection model includes a first convolutional layer, a first SFP2 module, a second SFP2 module, a first SFP1 module, a third SFP2 module, a second SFP1 module, a fourth SFP2 module, a third SFP1 module, a first attention module, a second convolutional layer, a first upsampling module, a first CSF module, a third convolutional layer, a second upsampling module, a second attention module, a second CSF module, a fourth convolutional layer, a third attention module, a third CSF module, a fifth convolutional layer, a fourth attention module, a fourth CSF module, a sixth convolutional layer, a seventh convolutional layer, and an eighth convolutional layer, wherein: The input of the target detection model is used as the input of the first convolutional layer, and the output of the first convolutional layer is used as the input of the first SFP2 module; Use the output of the first SFP2 module as the input of the second SFP2 module, and then use the output of the second SFP2 module as the input of the first SFP1 module; Use the output of the first SFP1 module as the input of the third SFP2 module, and then use the output of the third SFP2 module as the input of the second SFP1 module; Use the output of the second SFP1 module as the input of the fourth SFP2 module, and then use the output of the fourth SFP2 module as the input of the third SFP1 module; The output of the third SFP1 module is used as the input of the first attention module, and the output of the first attention module is used as the input of the second convolutional layer; The output of the second convolutional layer is used as the input of the first upsampling module, and the output of the first upsampling module is concatenated with the output of the second SFP1 module, and the concatenated result is used as the input of the first CSF module; The output of the first CSF module is used as the input of the third convolutional layer, and the output of the third convolutional layer is used as the input of the second upsampling module. The output of the second upsampling module is then concatenated with the output of the first SFP1 module, and the concatenated result is used as the input of the second attention module. The output of the second attention module is used as the input of the second CSF module, and then the output of the second CSF module is used as the input of the fourth convolutional layer. The output of the third convolutional layer is spliced ​​with the output of the fourth convolutional layer, and the spliced ​​result is used as the input of the third attention module; The output of the third attention module is used as the input of the third CSF module, and then the output of the third CSF module is used as the input of the fifth convolutional layer. The output of the fifth convolutional layer is spliced ​​with the output of the second convolutional layer, and the spliced ​​result is used as the input of the fourth attention module; The output of the fourth attention module is used as the input of the fourth CSF module; Then use the output of the second CSF module as the input of the sixth convolutional layer, the output of the third CSF module as the input of the seventh convolutional layer, and the output of the fourth CSF module as the input of the eighth convolutional layer. The target detection results are output through the sixth, seventh, and eighth convolutional layers.

[0022] Other steps and parameters are the same as those in the first embodiment.

[0023] The present invention proposes SFP1 and SFP2 modules, and stacks these two modules to form a new lightweight feature extraction network to solve the problems of small size, high maneuverability and complex background interference of UAV targets, and achieve a balance between accuracy and efficiency.

[0024] It should be noted that, in the present invention, except for the specifically specified depthwise separable convolutional layer and progressive partial convolutional layer, the other convolutional layers are conventional ordinary convolutional layers.

[0025] Specific implementation method three: Combination Figure 2 The difference between this embodiment and the first or second embodiment is that the working process of the first SFP1 module is as follows: The input of the first SFP1 module passes through the ninth convolutional layer and the first progressive partial convolutional layer (PConv) in sequence, the output of the first progressive partial convolutional layer is spliced ​​with the input of the first SFP1 module, and then the spliced ​​result is shuffled, and the shuffled result is used as the output of the first SFP1 module.

[0026] Other steps and parameters are the same as those in the first or second embodiment.

[0027] In the present invention, the working process of each SFP1 module is the same. Figure 5 As shown in the figure, due to the problem that small target features are easily lost in the drone detection process and the edge end has high requirements for real-time detection, the present invention proposes a progressive partial convolution for drone characteristics. Compared with ordinary convolution, progressive partial convolution can greatly reduce the amount of calculation and is more conducive to extracting the detection features of small target drones.

[0028] Where, h , w , c are the height, width and number of channels of the input feature map, respectively. k is the size of the convolution kernel, c p is the ratio of the number of channels of the convolution kernel to the number of channels of the feature map. From the calculation formula, it can be seen that the computational cost of ordinary convolution is twice that of progressive partial convolution (PConv). 1 / c p 2 The use of progressive partial convolution can greatly reduce the amount of calculation and achieve the overall lightweight detection model.

[0029] Specific implementation method four: Combination Figure 3 The difference between this embodiment and any one of the first to third embodiments is that the working process of the first SFP2 module is as follows: The first SFP2 module includes two parallel branches. In one branch, the input of the first SFP2 module passes through the first depth-separable convolution layer and the tenth convolution layer in sequence. In the other branch, the input of the first SFP2 module passes through the eleventh convolution layer, the second depth-separable convolution layer and the twelfth convolution layer in sequence. After the outputs of the two branches are spliced, the splicing results are shuffled and the shuffled results are used as the output of the first SFP2 module.

[0030] The other steps and parameters are the same as those in the first to third embodiments.

[0031] In the present invention, the working process of each SFP2 module is the same.

[0032] Specific implementation method five: Combination Figure 4 The difference between this embodiment and the first embodiment 1 to 4 is that the working process of the first CSF module is as follows: The first CSF module includes two parallel branches; In one branch, the input of the first CSF module passes through the thirteenth convolutional layer; In another branch, the input of the first CSF module is first passed through the fourteenth convolutional layer, and then the output of the fourteenth convolutional layer is passed through the second progressive partial convolutional layer and the fifteenth convolutional layer in sequence. Finally, the output of the fourteenth convolutional layer and the output of the fifteenth convolutional layer are concatenated and added to obtain the concatenation result a; The splicing result a is spliced ​​with the output of the thirteenth convolutional layer to obtain the splicing result b, and then the splicing result b passes through the sixteenth convolutional layer, and the output of the sixteenth convolutional layer is used as the output of the first CSF module.

[0033] The other steps and parameters are the same as those in the first to fourth embodiments.

[0034] The working processes of other CSF modules in the present invention are the same as those of the first CSF module. The present invention replaces the 3×3 ordinary convolution in BottleNeck with progressive partial convolution, which can dynamically enhance the feature responses of key components such as rotors and fuselage according to the variable characteristics of the UAV morphology. The excessive number of convolution layers in traditional networks causes the obtained feature maps to greatly lose spatial information. Although upsampling can increase channel information, it will also cause spatial information blurring, which has a significant impact on the feature extraction of small-target UAVs. The present invention significantly improves the detection accuracy and robustness of the network through the collaborative optimization of high-fidelity feature extraction-adaptive fusion-attention enhancement.

[0035] Specific embodiment six: This embodiment differs from any one of specific embodiments one to five in that the first attention module, the second attention module and the third attention module are all NAM attention.

[0036] The other steps and parameters are the same as those in the first to fifth embodiments.

[0037] Specific embodiment seven: This embodiment differs from any one of specific embodiments one to six in that the image captured by the camera needs to be resized and pixel-normalized before being input into the target detection model, and the processed image is then used as the input of the target detection model; The target detection model outputs the detection frame position of the drone target and the confidence of the detection frame. The detection frame position is then mapped to the original image captured by the camera to obtain the position of the detection frame in the original image. The detection frame in the original image is subjected to NMS (non-maximum suppression, retaining only the result with the highest probability) to obtain the final detection frame position of each drone target.

[0038] The other steps and parameters are the same as those in the first to sixth embodiments.

[0039] The following is an expanded description of this implementation method: Input image processing: Use the affine transformation matrix to scale the input image to a fixed size (640×640). Fill the remaining space with gray bars. Then center the image and normalize the pixel values ​​to 0–255.

[0040] Forward propagation: The input image passes through the target detection model and outputs detection frames at three scales. Each detection frame contains 8 pieces of information (x, y, w, h, conf, c). (x, y), w, h are the center coordinates, width, and height of the detection frame respectively. conf is the confidence level of the detection frame, and c is the probability that a drone target exists within the detection frame. Decode the predicted box: Convert the relative coordinates of the detection box to absolute coordinates (based on the input image size) and output the confidence of the detection box.

[0041] Post-processing: Apply NMS (non-maximum suppression, retaining only the results with the highest probability) to obtain the final detection results.

[0042] Specific implementation method eight: combination Figure 6 and Figure 7 This embodiment is different from the first to seventh embodiments in that the specific process of step 3 is as follows: The tracking process of the detection frame in the image captured by any camera is used as an example to illustrate. After obtaining the tracking frame based on the previous frame image captured by this camera, the tracking frame is subjected to extended Kalman filtering to obtain a prediction frame, and then the prediction frame is matched and tracked with the detection frame in the current frame image captured by this camera.

[0043] Step 3.1. For prediction box, calculate the first detected box in the image The target detection box and Cosine distance of predicted boxes and Mahalanobis distance , and traverse all the prediction boxes and target detection boxes corresponding to this camera; It should be noted that both cosine distance and Mahalanobis distance are calculated based on features. The features of the detection frame are composed of the center point coordinates of the detection frame, the aspect ratio of the detection frame, and the height of the detection frame; the features of the prediction frame are composed of the center point coordinates of the prediction frame, the aspect ratio of the prediction frame, and the height of the prediction frame. According to the cosine distance and Mahalanobis distance Calculating the cost matrix can reduce the number of ID jumps: in, Indicates the cost matrix Rank Elements of the column, is a hyperparameter; Step 32: Use the Hungarian algorithm and cost matrix to perform cascade matching, determine the matching solution that minimizes the total matching cost, and obtain the initial matching result between the prediction box and the detection box; Execute step 33 for the prediction frame and detection frame that are initially matched successfully, and execute step 34 for the prediction frame and detection frame that are not initially matched successfully; Step 3. Perform IOU matching on the pair of prediction boxes and detection boxes on the initial match (the intersection-over-union ratio threshold is set to 0.5. If the IOU is higher than the threshold, the two are considered to be matched successfully); If the IOU match between the prediction box and the detection box is successful, then determine whether the tracking results of the previous two frames of the current prediction box are all IOU matched successfully. If so, the matching result of the current prediction box and the detection box is confirmed. Otherwise, the matching result of the current prediction box and the detection box is tentative. If the IOU matching between the prediction box and the detection box is unsuccessful, then perform steps three and four on the prediction box and the detection box whose IOU matching is unsuccessful; Step 34: For the prediction boxes and detection boxes that are not initially matched, as well as the prediction boxes and detection boxes that fail to IOU matching in step 33, perform IOU matching on the detection boxes and prediction boxes in descending order of the confidence of the detection boxes, and then perform step 35 on the matching results; that is, first match the detection box with the highest confidence with the prediction box that is not initially matched and the prediction box with unsuccessful IOU matching in step 33, and then match the detection box with the second highest confidence with the prediction box, and so on; Step 35: For a pair of prediction frames and detection frames with successful IOU matching, if the tracking results of the previous two frames of the current prediction frame are both IOU-matched, the matching result of the current prediction frame and the detection frame is confirmed; otherwise, the matching result of the current prediction frame and the detection frame is tentative. For unmatched detection frames, when the detection frame is not located in the common coverage area of ​​adjacent cameras, the unmatched detection frame is regarded as a new target and initialized as a new track. The new target is assigned a unique ID (initialize the EKF state vector and covariance matrix). When the detection frame is located in the common coverage area of ​​adjacent cameras, a cross-camera matching and tracking process is required. For unsuccessful prediction frames, when the prediction frame is not located in the common coverage area of ​​adjacent cameras, if the tracking results of the first two frames of the current prediction frame do not have successful IOU matching, the current prediction frame will be deleted; otherwise, the current prediction frame will be retained; when the prediction frame is located in the common coverage area of ​​adjacent cameras, a cross-camera matching and tracking process is required.

[0044] The other steps and parameters are the same as those in the first to seventh embodiments.

[0045] The trajectory tracking result of the UAV target is obtained based on all the confirmed states obtained in the matching tracking process.

[0046] Specific embodiment 9: This embodiment differs from any one of specific embodiments 1 to 8 in that the specific process of cross-camera matching and tracking is as follows: For the detection box: (1) If there is an unmatched prediction frame in the coverage area of ​​the adjacent camera, and the detection frame and the unmatched prediction frame are both in the common coverage area of ​​the two cameras, the detection frame is IOU matched with the unmatched prediction frame in the coverage area of ​​the adjacent camera at the same time (in the present invention, each camera first performs the tracking process within its own coverage area, and then performs cross-camera tracking after the end). If the IOU match is successful, a pair of successfully matched detection frame and prediction frame is obtained (if the tracking results of the first two frames of the current prediction frame are both IOU matched successfully, the matching result of the current prediction frame and the detection frame is confirmed, otherwise, the matching result of the current prediction frame and the detection frame is tentative). If the IOU match is unsuccessful, the unmatched detection frame is regarded as a newly appeared target and initialized as a new track, and a unique ID is assigned to the newly appeared target; (2) If there is no unmatched prediction frame in the adjacent camera coverage area, the unmatched detection frame is regarded as a new target and initialized as a new trajectory, and a unique ID is assigned to the new target; For the prediction box: (1) If there is an unmatched detection frame in the coverage area of ​​the adjacent camera, and the predicted frame and the unmatched detection frame are both in the common coverage area of ​​the two cameras, the predicted frame is IOU matched with the detection frame in the coverage area of ​​the adjacent camera at the same time. If the IOU match is successful, a pair of successfully matched detection frames and predicted frames is obtained (if the tracking results of the first two frames of the current predicted frame are both IOU matched successfully, the matching result of the current predicted frame and the detection frame is confirmed, otherwise, the matching result of the current predicted frame and the detection frame is tentative). If the IOU match is unsuccessful and the tracking results of the first two frames of the current predicted frame are both IOU matched successfully, the current predicted frame is deleted, otherwise, the current predicted frame is retained. (2) If there is no unmatched detection frame in the adjacent camera coverage area, and the tracking results of the first two frames of the current prediction frame do not have an IOU match, the current prediction frame is deleted; otherwise, the current prediction frame is retained.

[0047] The other steps and parameters are the same as those in Specific Embodiments 1 to 8.

[0048] The specific process of step four is: For the prediction frame and detection frame pair in the tentative or confirmed state, the prediction frame and detection frame are fused according to the Kalman filter gain to obtain the tracking frame. The extended Kalman filter is continued based on the obtained tracking frame to obtain the next prediction frame; The detection frame of the newly appeared target is used as the tracking frame, and the tracking frame is subjected to extended Kalman filtering to obtain the next prediction frame; The prediction frame that was not successfully matched and retained is used as the tracking frame, and the tracking frame is subjected to extended Kalman filtering to obtain the next prediction frame; The obtained prediction box is the prediction result of the target position in the next frame.

[0049] The Extended Kalman Filter (EKF) expands the nonlinear function through Taylor series and discards its high-order terms, approximating the nonlinearity to a linear system. The specific iterative formula is as follows: Where: Indicates the k -1 frame's true state; Indicates the k The actual state of the frame; is a nonlinear state transfer function based on the estimated state at the previous moment and the current control input Predict the current state; F For function The Jacobian matrix of the function exist The linear approximation at is the process noise.

[0050] For the moment k The observation state, For the moment k The prior state prediction of (i.e., the predicted state without observation correction), h is a nonlinear observation function, and the predicted state Mapped to the observation space, H is the observation matrix, representing h In the forecast state The linear approximation at F similar, is the observation noise, which represents the uncertainty of sensor measurement.

[0051] By using Taylor expansion to approximate the nonlinear motion model, the problem of trajectory prediction deviation caused by high-speed maneuvers of UAVs can be effectively solved. Compared with traditional Kalman filtering, it can significantly reduce the position prediction error.

[0052] Specific embodiment ten: This embodiment differs from any one of specific embodiments one to nine in that the working process of the first progressive partial convolutional layer is as follows: The number of channels of the first progressive partial convolution layer’s input passing through the convolution kernel is c p The output of the seventeenth convolutional layer is then concatenated with the input of the first progressive partial convolutional layer, and the concatenated result is used as the output of the first progressive partial convolutional layer.

[0053] The other steps and parameters are the same as those in Specific Embodiments 1 to 9.

[0054] Application scenario: Airport drone real-time monitoring and countermeasure system 1. Hardware configuration and deployment plan Node hardware: Core equipment: 6 NVIDIA Jetson Orin NANO modules, equipped with 1024 CUDA cores, 32 Tensor cores and 6-core ARM CPU, with a computing power of 40TOPS and support for INT8 quantized inference.

[0055] Camera: Each node is equipped with a Sony IMX477 sensor camera with a resolution of 1920 × 1080 @ 60 FPS, a horizontal field of view of 120°, support for HDR mode, and a sensitivity of 0.1 lux in low-light environments.

[0056] Communication module: Built-in QuectelRM500Q 5G module, supports the Sub-6GHz frequency band, has an uplink bandwidth of 500Mbps, end-to-end latency of less than 10ms, and builds a wireless mesh network through the IEEE802.11s protocol.

[0057] Physical layout: The nodes are deployed around the airport in a regular hexagonal topology, with a spacing of 200 meters, an installation height of 8 meters, a coverage radius of 250 meters, and an overlapping area greater than or equal to 30% to ensure seamless target handover.

[0058] The central control server is deployed in the tower and is equipped with dual-core Xeon Silver 4310 processors and NVIDIA A100 GPU for global trajectory fusion and threat analysis.

[0059] 2. System operation process Target detection stage: Image preprocessing: The input video stream undergoes image preprocessing (size, pixel normalization) and is then passed to the object detection model.

[0060] Lightweight model inference: The input resolution is adjusted to 640×640, the model inference time is less than or equal to 15ms / frame, and the confidence threshold of the output detection box is set to 0.7.

[0061] Using Det-Fly data to pre-train a YOLOv5s variant with the C3F module, the AP@0.5 for drone detection reached 94.5% after transfer learning. The average precision curve of the lightweight object detection model is shown in the figure below. Figure 8 As shown in Figure 2, the target detection test results are as follows: Figure 9 shown.

[0062] Cross-camera tracking and matching: When the target leaves the coverage area of ​​CameraNode 1, the Mesh network broadcasts the target features to adjacent nodes, and CameraNode 2 completes re-identification within 0.2 seconds.

[0063] Early warning and countermeasure stage: If a drone enters a no-fly zone (such as 500 meters around a runway), the system automatically marks it as a "high-risk target."

[0064] 3. Effect verification and performance testing Test environment: Scene: On a clear day, a drone is flying at low speed with a mountain forest in the background.

[0065] Test results: The detection accuracy of the video test results reached 94.7%, and the tracking results are as follows: Figure 10 and Figure 11As shown in the figure, the end-to-end delay of the entire process of detection → tracking → countermeasure is less than 220ms.

[0066] The above examples are merely illustrative of the calculation model and process of the present invention and are not intended to limit the embodiments of the present invention. Persons skilled in the art will readily appreciate that other variations or modifications based on the above description are possible. This list of embodiments is not exhaustive; however, any obvious variations or modifications derived from the technical solution of the present invention remain within the scope of protection of the present invention.

Claims

1. A drone detection and tracking method based on convolutional neural network, characterized in that: The method specifically comprises the following steps: Step 1: According to the first The tracking results obtained from the frame images are subjected to extended Kalman filtering to predict the position of each UAV target; Step 2: Use the built target detection model to detect the first Frame images are used for drone target detection; Step 3: Match the predicted result of the drone target position with the detection result of step 2 to obtain the drone target tracking result; Step 4: Perform extended Kalman filtering based on the UAV target tracking results in step 3 to continue predicting the position of each UAV target; Step 5: , return to step 2 and continue until the entire tracking process is completed.

2. The method for detecting and tracking drones based on convolutional neural networks according to claim 1, wherein: The target detection model includes a first convolutional layer, a first SFP2 module, a second SFP2 module, a first SFP1 module, a third SFP2 module, a second SFP1 module, a fourth SFP2 module, a third SFP1 module, a first attention module, a second convolutional layer, a first upsampling module, a first CSF module, a third convolutional layer, a second upsampling module, a second attention module, a second CSF module, a fourth convolutional layer, a third attention module, a third CSF module, a fifth convolutional layer, a fourth attention module, a fourth CSF module, a sixth convolutional layer, a seventh convolutional layer and an eighth convolutional layer, wherein: The input of the target detection model is used as the input of the first convolutional layer, and the output of the first convolutional layer is used as the input of the first SFP2 module; Use the output of the first SFP2 module as the input of the second SFP2 module, and then use the output of the second SFP2 module as the input of the first SFP1 module; Use the output of the first SFP1 module as the input of the third SFP2 module, and then use the output of the third SFP2 module as the input of the second SFP1 module; Use the output of the second SFP1 module as the input of the fourth SFP2 module, and then use the output of the fourth SFP2 module as the input of the third SFP1 module; The output of the third SFP1 module is used as the input of the first attention module, and the output of the first attention module is used as the input of the second convolutional layer; The output of the second convolutional layer is used as the input of the first upsampling module, and the output of the first upsampling module is concatenated with the output of the second SFP1 module, and the concatenated result is used as the input of the first CSF module; The output of the first CSF module is used as the input of the third convolutional layer, and the output of the third convolutional layer is used as the input of the second upsampling module. The output of the second upsampling module is then concatenated with the output of the first SFP1 module, and the concatenated result is used as the input of the second attention module. The output of the second attention module is used as the input of the second CSF module, and then the output of the second CSF module is used as the input of the fourth convolutional layer. The output of the third convolutional layer is spliced ​​with the output of the fourth convolutional layer, and the spliced ​​result is used as the input of the third attention module; The output of the third attention module is used as the input of the third CSF module, and then the output of the third CSF module is used as the input of the fifth convolutional layer. The output of the fifth convolutional layer is spliced ​​with the output of the second convolutional layer, and the spliced ​​result is used as the input of the fourth attention module; The output of the fourth attention module is used as the input of the fourth CSF module; Then use the output of the second CSF module as the input of the sixth convolutional layer, the output of the third CSF module as the input of the seventh convolutional layer, and the output of the fourth CSF module as the input of the eighth convolutional layer. The target detection results are output through the sixth, seventh, and eighth convolutional layers.

3. The method for detecting and tracking drones based on convolutional neural networks according to claim 2, wherein: The working process of the first SFP1 module is as follows: The input of the first SFP1 module passes through the ninth convolutional layer and the first progressive partial convolutional layer in sequence, the output of the first progressive partial convolutional layer is spliced ​​with the input of the first SFP1 module, and then the splicing result is shuffled, and the shuffled result is used as the output of the first SFP1 module.

4. The method for detecting and tracking drones based on convolutional neural networks according to claim 2, wherein: The working process of the first SFP2 module is as follows: The first SFP2 module includes two parallel branches. In one branch, the input of the first SFP2 module passes through the first depth-separable convolution layer and the tenth convolution layer in sequence. In the other branch, the input of the first SFP2 module passes through the eleventh convolution layer, the second depth-separable convolution layer and the twelfth convolution layer in sequence. After the outputs of the two branches are spliced, the splicing results are shuffled and the shuffled results are used as the output of the first SFP2 module.

5. The method for detecting and tracking drones based on convolutional neural networks according to claim 2, wherein: The working process of the first CSF module is: The first CSF module includes two parallel branches; In one branch, the input of the first CSF module passes through the thirteenth convolutional layer; In another branch, the input of the first CSF module is first passed through the fourteenth convolutional layer, and then the output of the fourteenth convolutional layer is passed through the second progressive partial convolutional layer and the fifteenth convolutional layer in sequence. Finally, the output of the fourteenth convolutional layer and the output of the fifteenth convolutional layer are concatenated and added to obtain the concatenation result a; The splicing result a is spliced ​​with the output of the thirteenth convolutional layer to obtain the splicing result b, and then the splicing result b passes through the sixteenth convolutional layer, and the output of the sixteenth convolutional layer is used as the output of the first CSF module.

6. The method for detecting and tracking drones based on convolutional neural networks according to claim 2, wherein: The first attention module, the second attention module and the third attention module are all NAM attention.

7. The method for detecting and tracking drones based on convolutional neural networks according to claim 2, wherein: Before the image captured by the camera is input into the target detection model, it needs to be resized and pixel normalized in sequence, and then the processed image is used as the input of the target detection model; The target detection model outputs the detection frame position of the drone target and the confidence of the detection frame. The detection frame position is then mapped to the original image captured by the camera to obtain the position of the detection frame in the original image. The detection frame in the original image is then subjected to NMS to obtain the final detection frame position of each drone target.

8. The method for detecting and tracking drones based on convolutional neural networks according to claim 1, wherein: In step 3, the area covered by each camera is matched separately. The specific matching process of the area covered by any camera is as follows: Step 3.

1. For prediction box, calculate the first detected box in the image The target detection box and Cosine distance of predicted boxes and Mahalanobis distance , and traverse all the prediction frames and target detection frames corresponding to the current camera; According to the cosine distance and Mahalanobis distance Calculate the cost matrix: in, Indicates the cost matrix Rank Elements of the column, is a hyperparameter; Step 32: Use the Hungarian algorithm and cost matrix to perform cascade matching, determine the matching solution that minimizes the total matching cost, and obtain the initial matching result between the prediction box and the detection box; Execute step 33 for the prediction frame and detection frame that are initially matched successfully, and execute step 34 for the prediction frame and detection frame that are not initially matched successfully; Step 3. Perform IOU matching on a pair of prediction boxes and detection boxes on the initial matching; If the IOU match between the prediction box and the detection box is successful, then determine whether the tracking results of the previous two frames of the current prediction box are all IOU matched successfully. If so, the matching result of the current prediction box and the detection box is confirmed. Otherwise, the matching result of the current prediction box and the detection box is tentative. If the IOU matching between the prediction box and the detection box is unsuccessful, then perform steps three and four on the prediction box and the detection box whose IOU matching is unsuccessful; Step 34: For the prediction boxes and detection boxes that are not initially matched, as well as the prediction boxes and detection boxes that fail to IOU matching in step 33, perform IOU matching on the detection boxes and prediction boxes in descending order of the confidence of the detection boxes, and then perform step 35 on the matching results; Step 35: For a pair of prediction frames and detection frames with successful IOU matching, if the tracking results of the previous two frames of the current prediction frame are both IOU-matched, the matching result of the current prediction frame and the detection frame is confirmed; otherwise, the matching result of the current prediction frame and the detection frame is tentative. For unmatched detection frames, if the detection frame is not located in the common coverage area of ​​adjacent cameras, the unmatched detection frame is considered a new target and initialized as a new track, and a unique ID is assigned to the new target. If the detection frame is located in the common coverage area of ​​adjacent cameras, a cross-camera matching and tracking process is required. For unsuccessful prediction frames, when the prediction frame is not located in the common coverage area of ​​adjacent cameras, if the tracking results of the first two frames of the current prediction frame do not have successful IOU matching, the current prediction frame will be deleted; otherwise, the current prediction frame will be retained; when the prediction frame is located in the common coverage area of ​​adjacent cameras, a cross-camera matching and tracking process is required.

9. The method for detecting and tracking drones based on convolutional neural networks according to claim 8, wherein: The specific process of cross-camera matching and tracking is as follows: For the detection box: (1) If there is an unmatched prediction frame in the coverage area of ​​the adjacent camera, and the detection frame and the unmatched prediction frame are both in the common coverage area of ​​the two cameras, then the detection frame is matched with the unmatched prediction frame in the coverage area of ​​the adjacent camera at the same time. If the IOU match is successful, a pair of matched detection frame and prediction frame is obtained. If the IOU match is unsuccessful, the unmatched detection frame is regarded as a new target and initialized as a new track, and a unique ID is assigned to the new target. (2) If there is no unmatched prediction frame in the adjacent camera coverage area, the unmatched detection frame is regarded as a new target and initialized as a new trajectory, and a unique ID is assigned to the new target; For the prediction box: (1) If there is an unmatched detection frame in the coverage area of ​​the adjacent camera, and the predicted frame and the unmatched detection frame are both in the common coverage area of ​​the two cameras, the predicted frame is matched with the detection frame in the coverage area of ​​the adjacent camera at the same time. If the IOU match is successful, a pair of successfully matched detection frame and predicted frame is obtained. If the IOU match is unsuccessful and the tracking results of the first two frames of the current predicted frame are not successfully matched by IOU, the current predicted frame is deleted. Otherwise, the current predicted frame is retained. (2) If there is no unmatched detection frame in the adjacent camera coverage area, and the tracking results of the first two frames of the current prediction frame do not have an IOU match, the current prediction frame is deleted; otherwise, the current prediction frame is retained.

10. The method for detecting and tracking drones based on convolutional neural networks according to claim 3, wherein: The working process of the first progressive partial convolutional layer is: The number of channels of the first progressive partial convolution layer’s input passing through the convolution kernel is c p The output of the seventeenth convolutional layer is then concatenated with the input of the first progressive partial convolutional layer, and the concatenated result is used as the output of the first progressive partial convolutional layer.