Multi-target tracking method under monitoring of unmanned aerial vehicle

By adopting EfficientNet network and dynamic matching strategy in the multi-objective tracking algorithm, combined with extended Kalman filtering, the difficulty of identifying and tracking of traditional multi-objective tracking algorithms in complex environments is solved, and higher tracking accuracy and robustness are achieved.

CN120182322APending Publication Date: 2025-06-20CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510325371.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

Traditional multi-objective tracking algorithms are difficult to accurately identify and track targets in complex environments, especially in the case of light changes, occlusion, rapid movement and target pose changes, and there are problems of mismatch and tracking failure.

Method used

The EfficientNet network is used for target feature extraction, and the target tracking is carried out in combination with dynamic matching strategy and extended Kalman filtering (EKF). The dynamic matching strategy uses the EIoU loss function for interchange and comparison calculation, and integrates the target's position information, speed information, appearance characteristics and historical trajectory information. EKF is used to process the nonlinear motion state of the target and outputs a continuous trajectory.

Benefits of technology

The tracking effect of the model is improved, the mismatch rate and tracking failure rate are reduced, the robustness of complex lighting and multi-scale targets is enhanced, and the low latency requirements of drones are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182322A_ABST
    Figure CN120182322A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-target tracking method under monitoring of an unmanned aerial vehicle, and belongs to the technical field of intelligent monitoring. According to the method, a traditional DeepSORT algorithm is optimized from three aspects of feature extraction, a matching strategy and motion modeling aiming at the problems of small target, frequent shielding, nonlinear motion and the like in an unmanned aerial vehicle scene. The method comprises the following steps: firstly, replacing an original feature extraction module with a lightweight OfficientNet network, and improving the small target feature discrimination through a composite scaling strategy and an MBConv module; secondly, providing a dynamic matching strategy based on EIoU, fusing position, shape and motion information, and reducing the mismatching rate of dense scenes; and finally, introducing extended Kalman filtering to process nonlinear motion trail prediction. According to the method, the MOTA index on a VisDrone data set reaches 85.6% and is improved by 12.3% compared with an original algorithm, and the tracking robustness in a complex scene is remarkably improved while the real-time performance is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent monitoring and relates to a multi-target tracking method under UAV monitoring. Background Art

[0002] With the vigorous development of intelligent monitoring technology, multi-target tracking technology plays a crucial role in many fields such as security, traffic management, and public safety. Multi-target tracking technology can monitor the trajectories of multiple targets in a scene in real time and provide key data support for various decisions. However, in a complex environment, multi-target tracking faces many challenges:

[0003] Lighting changes at different times and under different weather conditions can affect the recognition and tracking of targets.

[0004] Occlusions between people or objects can cause targets to disappear temporarily, increasing the difficulty of tracking.

[0005] Fast-moving targets may be difficult to accurately detect and track.

[0006] Changes in the target pose can affect the accuracy of feature extraction, resulting in false matching.

[0007] Traditional multi-target tracking algorithms have limitations in dealing with the above problems and are difficult to meet the requirements of practical applications.

[0008] In recent years, multi-target tracking algorithms based on deep learning have made significant progress, and the DeepSORT algorithm performs particularly well. The DeepSORT algorithm combines object detection and tracking algorithms, can effectively handle problems such as small targets, complex target motions, and occlusion situations, and achieves a good balance between accuracy and real-time performance. However, the DeepSORT algorithm still has some deficiencies:

[0009] Traditional feature extraction networks may be difficult to extract sufficiently discriminative features, resulting in false matching.

[0010] Traditional IoU matching algorithms overly rely on the geometric overlap degree between the detection box and the tracked target. For targets with complex or irregular shapes, as well as situations where the target is partially occluded or its shape has undergone significant deformation, it may not accurately reflect the matching degree, resulting in incorrect associations or matching failures.

[0011] The traditional DeepSORT algorithm defaults to using linear Kalman filtering for target motion prediction, which is difficult to accurately describe the non-linear motion of targets from the UAV perspective, resulting in tracking failures.

[0012] To solve the above problems, the present invention proposes a multi-target tracking method under UAV monitoring, aiming to improve the tracking effect of the model and enable it to better meet various requirements in practical applications. Summary of the Invention

[0013] In view of this, the purpose of the present invention is to provide a multi-target tracking method under UAV monitoring.

[0014] To achieve the above purpose, the present invention provides the following technical solutions:

[0015] A multi-target tracking method under UAV monitoring includes the following steps:

[0016] S1: Extract features of the targets in the video frames collected by the UAV through the EfficientNet network to generate target feature vectors;

[0017] S2: Based on the dynamic matching strategy, associate and match the detection boxes of the current frame with the historical tracking targets, where the dynamic matching strategy uses the EIoU loss function to calculate the intersection over union ratio and fuses the position information, speed information, appearance features, and historical trajectory information of the targets;

[0018] S3: Predict and update the non-linear motion state of the targets through the Extended Kalman Filter EKF and output continuous trajectories.

[0019] Further, the EfficientNet network adjusts the depth, width, and resolution of the network through the compound scaling method, specifically including:

[0020] Cooperatively scale the depth, width, and resolution according to the following formulas:

[0021] depth: d = α φ (1)

[0022] width: w = β φ (2)

[0023] resolution: r = γ φ (3)

[0024] where α, β, and γ are constants determined by grid search on a small dataset, respectively controlling the scaling ratios of the depth, width, and resolution; φ represents the scaling factor.

[0025] Further, the EIoU loss function includes the following components:

[0026] (a) Intersection over union loss, calculating the area ratio of the overlapping region and the union region between the predicted box and the ground truth box;

[0027] (b) Distance loss, which calculates the square of the Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box;

[0028] (c) Aspect ratio loss, which is achieved by minimizing the difference in width and height between the predicted bounding box and the ground truth bounding box. The specific calculation formula is:

[0029]

[0030] where, L IOU represents the intersection over union (IoU) loss, L dis represents the distance loss, L asp represents the aspect ratio loss, IoU represents the intersection over union; w and h are the width and height of the predicted bounding box, w gt and h gt are the width and height of the target bounding box, w c and h c are the width and height of the smallest closed region containing the predicted bounding box and the target bounding box, ρ 2 (b, b gt ) is the square of the Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box; b represents the position of the center point of the predicted bounding box, bg t represents the position of the center point of the ground truth bounding box.

[0031] Furthermore, the steps of the extended Kalman filter (EKF) include:

[0032] (a) Perform first-order Taylor expansion linearization on the non-linear state transition function and the observation function, and calculate the state transition Jacobian matrix F k and the observation Jacobian matrix H k respectively;

[0033] (b) Perform prior state estimation and error covariance prediction through the formulas and ; where, represents the prior state estimation value at time step k based on the state prediction at time step k - 1; represents the optimal estimation at time step k - 1; represents the prior state estimation at time step k; P k|k-1 represents the uncertainty of the predicted state;

[0034] (c) Update the posterior state and error covariance based on the formulas and P k|k =(I - K k H k )P k|k-1 ; where, is the difference between the observation and the prediction; represents the posterior state estimation value at time step k after fusing the observation data, which is the optimal state estimation corrected by combining the prediction value and the observation value; K kdenotes the Kalman gain at time step k, which is used to balance the weights of the prior prediction and the observed data and determine the influence degree of the observed data on the state update; z k denotes the actual observed value at time step k; denotes the predicted observed value based on the prior state estimate at time step k, P k|k denotes the posterior error covariance matrix after fusing the observed data at time step k, which reflects the uncertainty of the posterior state estimate; I denotes the identity matrix to ensure the matching of matrix operation dimensions; H k denotes the observation matrix at time step k, which describes the mapping relationship from the state variable to the observation space; P kk denotes the prior error covariance matrix at time step k, which represents the uncertainty of the prior state estimate.

[0035] Furthermore, in the dynamic matching strategy, the extraction of appearance features is based on the MBConv module, and the MBConv module includes:

[0036] a dilated convolutional layer, a depthwise separable convolutional layer, an SE channel attention module, and a projection convolutional layer, where the SE module generates channel attention weights through the formulas s = GAP(F dw ); and e = σ(W2·δ(W1…));

[0037] where, GAP represents global average pooling; s represents the output result of global average pooling, and e represents the channel attention weights generated after processing s through two fully connected layers; r is the dimensionality reduction factor, δ is the ReLU activation function, σ is the sigmoid activation function, and F dw denotes the feature map,

[0038] Furthermore, in the compound scaling coefficients of the EfficientNet, the depth scaling coefficient is 1.2, the width scaling coefficient is 1.1, and the resolution scaling coefficient is 1.15.

[0039] Furthermore, in the aspect ratio loss, the width w c and height h c of the closed region and ρ 2 (b, b gt ) are calculated through the minimum bounding rectangle enclosing the predicted box and the ground truth box.

[0040] Furthermore, the state transition Jacobian matrix F k and the observation Jacobian matrix H k are dynamically updated in each frame.

[0041] Furthermore, the dataset of the method uses VisDrone - MOT, and the resolution of the video frames from the drone's perspective is 1920×1080.

[0042] Furthermore, in the dynamic matching strategy, when the EIoU value exceeds the dynamically adjusted threshold, it is determined that the detection box matches the tracking target successfully, and the threshold is adaptively adjusted according to the scene complexity.

[0043] The beneficial effects of the present invention are as follows:

[0044] (1) The EfficientNet-B0 lightweight network is adopted, with the number of parameters reduced by 41% compared to the original ResNet-18, and the accuracy of feature similarity comparison increased by 18.5%, effectively distinguishing targets with similar appearances.

[0045] (2) The EIoU matching strategy comprehensively considers shape, displacement, and motion information, reducing the number of ID Switch times to 32 times on the VisDrone test set (the original algorithm was 89 times), and the false matching rate decreased by 64%.

[0046] (3) The Extended Kalman Filter (EKF) reduces the trajectory prediction error by 42% in the vehicle turning scenario, and the MOTA index is increased to 85.6%.

[0047] (4) Through the MBConv depthwise separable convolution and model pruning, real-time processing at 35 FPS is achieved on the NVIDIA Jetson TX2 edge device, meeting the low-latency requirements of drones.

[0048] (5) Verified on datasets such as UAVDT and DukeMTMC, the MOTA exceeds 80% in all cases, proving the robustness of the algorithm to complex lighting and multi-scale targets.

[0049] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0051] Figure 1 is the structural diagram of the EEE-DeepSORT model;

[0052] Figure 2 is the structural diagram of ResNet-18;

[0053] Figure 3 is the structural diagram of EfficientNet;

[0054] Figure 4It is a schematic diagram of EIoU. Specific implementation manners

[0055] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only schematically illustrate the basic concept of the present invention. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0056] Among them, the attached drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the attached drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted.

[0057] In the attached drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, it is based on the orientation or positional relationship shown in the attached drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the attached drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0058] 1. EEE-DeepSORT multi-object tracking algorithm structure

[0059] The specific network structure of EEE-DeepSORT proposed in this chapter is as Figure 1 shown, where the red module is the improved module or structure in this chapter. Based on DeepSORT, this algorithm uses the deeper and lightweight convolutional neural network EfficientNet in the feature extraction network module to extract high-level features of the target, improving the discrimination of the features. In the Hungarian matching module, by using E-IoU for dynamic threshold adjustment and multi-scale matching, the accuracy of matching is significantly improved. In the Kalman filter module, by introducing the Extended Kalman Filter (EKF), complex non-linear target tracking problems can be handled, better predicting and updating the target state, and improving the accuracy and robustness of tracking.

[0060] 2. Feature extraction network

[0061] In the DeepSORT algorithm, the feature extraction network usually adopts a convolutional neural network. Commonly, it is the ResNet-18 network pre-trained on a large-scale dataset. Its structural diagram is as shown in Figure 2 the figure. Its main function is to extract features from the detected target image regions, convert them into fixed-length feature vectors, and these feature vectors can be used for subsequent target matching and tracking. Especially when the target is occluded or its appearance changes, the similarity of features is used to assist in determining the consistency of the target in different frames. However, for targets with similar appearances, traditional feature extraction networks may be difficult to extract sufficiently discriminative features, resulting in incorrect matches during the tracking process, which will reduce the accuracy of the tracking results. Especially in complex scenarios, it may lead to frequent switching of target identities, increasing the probability of incorrect tracking. When the target reappears after occlusion or deformation, it is difficult to be correctly recognized and matched, resulting in tracking loss or incorrect tracking. And some high-performance feature extraction network structures are complex and have numerous parameters, requiring a large amount of computing resources and time for feature extraction. Especially in real-time tracking systems, this may lead to processing delays.

[0062] To solve the above problems, this section improves the feature extraction network of DeepSORT and uses the pre-trained network EfficientNet with better performance. This network can extract more discriminative features at different scales and resolutions. EfficientNet consists of multiple MBConv modules and some conventional convolutional layers and pooling layers. As shown in Figure 3 the figure, the overall architecture of the network presents a hierarchical structure. Starting from the input layer, through a series of feature extraction layers, the features of the image are gradually abstracted into high-level semantic features, and finally the prediction results are output through the fully connected layer or the global average pooling layer. Different versions of EfficientNet vary in the number of layers, the number of modules, and parameter settings, but all follow the basic design principles and achieve different performance and computing resource requirements through reasonable combination and scaling.

[0063] EfficientNet adopts a compound scaling method to uniformly adjust the depth, width, and resolution of the network. Traditional network scaling often only focuses on a single dimension, such as increasing the number of network layers (depth) or increasing the number of channels per layer (width). However, EfficientNet can reasonably scale the depth, width, and resolution simultaneously through a compound coefficient, significantly improving performance while maintaining model efficiency. The scaling of depth is shown in formula (1), the scaling of width is shown in formula (2), and the scaling of resolution is shown in formula (3).

[0064] depth: d = α φ (1)

[0065] width: w = β φ (2)

[0066] resolution: r = γ φ (3)

[0067] Among them, α, β, and γ are constants determined by grid search on a small dataset, controlling the scaling ratios of depth, width, and resolution respectively. In the design of EfficientNet, α ≈ 1.2, β ≈ 1.1, γ ≈ 1.15, and the overall scaling factor of the network is determined according to the required computational resources and performance balance, obtaining different variants through different values of φ.

[0068] The calculation of the computational complexity (FLOPS, Floating Point Operations Per Second) of the network is shown in formula (4).

[0069] FLOPS ≈ depth × width 2 × resolution 2 (4)

[0070] This co-scaling method enables the model to adaptively adjust its own structure under different computational resource constraints, thus achieving the effect of the algorithm.

[0071] MBConv (Mobile Inverted Bottleneck Convolution) is the core module of EfficientNet. It is a structure based on depthwise separable convolution and inverted residual structure, combining the advantages of depthwise separable convolution and SE (Squeeze-and-Excitation) module, achieving efficient feature extraction and expression.

[0072] In the MBConv structure, first, the feature expression ability is enhanced by increasing the number of channels. The input feature map is subjected to 1×1 convolution to expand the number of channels from C in to C exp = α × C in , where α is the expansion factor, usually taking a value of 6. The output feature map The specific operation is shown in formula (5).

[0073] F exp = W exp × F in (5)

[0074] Among them, Wexp is the dilated convolution kernel. Then, spatial feature extraction is performed independently for each channel to reduce the computational amount. For F exp , perform a k×k depthwise separable convolution to output the feature map as shown in Equation (6).

[0075] F dw = W dw × F exp (6)

[0076] where W dw is the depthwise separable convolution kernel, and k is usually taken as 3. Then, channel reweighting is performed to enhance important features. As shown in Equation (7), perform global average pooling on F dw to obtain the channel descriptor Then, as shown in Equation (8), generate the channel attention weight through two fully connected layers

[0077] s = GAP(F dw ) (7)

[0078] e = σ(W2·δ(W1…)) (8)

[0079] where r is the dimensionality reduction factor, usually taken as 4, δ is the ReLU activation function, and σ is the sigmoid activation function. Then, apply e to F dw to obtain the feature map Specifically, as shown in Equation (9).

[0080]

[0081] where denotes element-wise multiplication. Finally, reduce the number of channels to lower the computational complexity. As shown in Equation (10), perform a 1×1 convolution on F se to project the number of channels from C exp to C out , and output the feature map

[0082] F out = W proj × F se (10)

[0083] where W proj is the projection convolution kernel.

[0084] Through the combination of dilated convolutional layers and depthwise separable convolutional layers, MBConv can reduce the computational cost while maintaining the feature representation ability. The SE module enables MBConv to dynamically adjust the channel weights according to the importance of features, enhancing key features and suppressing unimportant features. This structure enhances the feature extraction ability while reducing the computational cost. The MBConv module can adapt to different model scales and computational resource limitations by adjusting the dilation factor α and the dimensionality reduction factor r. When processing large-scale image data, the MBConv module can quickly extract representative features, improving the running efficiency of the model.

[0085] 3. IoU Matching Improvement

[0086] In the DeepSORT algorithm, IoU Match (Intersection over Union Matching) is a crucial step in the object tracking process for associating detection boxes with existing tracked objects. In each frame of the image, the object detection module outputs a series of detection boxes, and the tracking module needs to determine the correspondence between these detection boxes and the previously tracked objects. IoU Match measures the overlap degree between the detection box and the tracked object based on the Intersection over Union metric. For each detection box, its IoU values with all existing tracked objects are calculated. If the IoU value of a detection box with a certain tracked object exceeds a pre-set threshold (usually around 0.5, but can be adjusted according to the actual situation), it is considered that they may be the same object, and thus the detection box is matched with the corresponding tracked object. However, the traditional Intersection over Union only considers the area ratio of the overlapping region and the union region between the detection box and the tracked object, which may have limitations for objects with complex or irregular shapes. At the same time, when the object is partially occluded or the object shape undergoes a large deformation, the traditional IoU calculation may not accurately reflect the matching degree between the two objects, ultimately resulting in incorrect associations or matching failures.

[0087] To address the above problems, this section improves the IoU in the DeepSORT Hungarian matching by using the loss function EIoU, which is more sensitive to overlap and shape differences. It splits the loss term of the aspect ratio, enabling the model to converge faster during training and improving the detection accuracy for small objects, especially suitable for small object tracking from the perspective of drones. By dividing the loss function into three parts: IOU loss, distance loss, and aspect ratio loss, its loss function is calculated as shown in Equation (11). In this way, the beneficial characteristics of the CIOU loss can be retained. At the same time, the EIoU loss directly minimizes the width and height differences between the target box and the anchor box, thus achieving a faster convergence speed and better localization results.

[0088]

[0089] where w and h are the width and height of the predicted box, wgt and h gt are the width and height of the target bounding box, w c and h c are the width and height of the smallest closed region containing the predicted bounding box and the target bounding box, ρ 2 (b, b gt ) is the square of the Euclidean distance between the centers of the predicted bounding box and the ground truth bounding box, as Figure 4 shown.

[0090] 4 Extended Kalman Filter

[0091] In the update and prediction phase of the traditional DeepSORT algorithm, linear Kalman Filter (KF) is defaultly used for target motion prediction. However, in the multi-target tracking task from the UAV perspective, targets (such as vehicles and pedestrians) often show curvilinear motion (turning, detouring) or acceleration changes (sudden stop, acceleration) from the UAV perspective. The motion model of the target is non-linear, and the traditional linear motion model is difficult to accurately describe.

[0092] To solve the above problems, in this section, the filtering algorithm of DeepSORT is replaced with a non-linear extended Kalman filter algorithm, which can more accurately predict the motion state of the target. Traditional Kalman filter is an optimal recursive least squares estimation algorithm based on linear systems, used to estimate the state of a dynamic system. This algorithm assumes that both the state transition and observation models of the system are linear, and then recursively estimates the system state through two steps: prediction and update. Extended Kalman filter is an extended version of the traditional Kalman filter, used to handle non-linear systems. It approximates the non-linear system as a linear system by linearizing the non-linear function at each time step (i.e., calculating the Jacobian matrix), and then applies it to the Kalman filter algorithm.

[0093] The core of the extended Kalman filter is to minimize the mean square error (MSE) of the state estimate in a recursive manner under the condition that both the system state equation and the observation equation are non-linear. Assume that the state space model of the non-linear dynamic system is as shown in Equation (12).

[0094]

[0095] where, x k ∈R n is the state vector at time k, u k ∈R m is the control input, f is the non-linear state transition function, h is the non-linear observation function, w k is the process noise, v k is the observation noise, and the noises both follow a zero-mean Gaussian distribution.

[0096] The EKF performs local linearization of the nonlinear function through first-order Taylor expansion. Equation (13) is the state transition function linearized at the state estimation point after expansion.

[0097]

[0098] where F k is the state transition Jacobian matrix.

[0099] Similarly, Equation (14) is the observation function expanded at the predicted state point after expansion.

[0100]

[0101] where H k is the observation Jacobian matrix.

[0102] First, state prediction is performed. Equation (15) is the prior state estimate, and Equation (16) is the prior error covariance prediction.

[0103]

[0104]

[0105] where P kk-1 represents the uncertainty of the predicted state. Then, as shown in Equation (17), the predicted observation value is calculated based on the linearized observation model.

[0106]

[0107] Then, Equation (18) is used to calculate the Kalman gain, and the calculated result is used to balance the confidence levels of prediction and observation.

[0108]

[0109] Subsequently, the state can be updated. Equation (19) is the update calculation formula for the posterior state estimate, and Equation (20) is the posterior error covariance update formula.

[0110]

[0111]

[0112] where is the difference between the observation and the prediction.

[0113] By calculating the Jacobian matrix through linearizing the nonlinear function at each time step, the nonlinear system is approximated as a linear system, and thus the Kalman filter algorithm can be applied to alleviate the problem of tracking failure caused by the nonlinear motion of the UAV during tracking.

[0114] 5. Experimental Environment and Datasets

[0115] (1) Experimental Environment

[0116] Hardware: NVIDIA RTX 3090 GPU (24GB video memory) is used at the training end, and Jetson TX2 embedded platform is used at the inference end;

[0117] Software: PyTorch 1.9.0, CUDA 11.1, and TensorRT acceleration is used during deployment;

[0118] Parameter settings: The dimension of the EKF state vector is 7 (position, velocity, aspect ratio), the input resolution of EfficientNet is 640×640, the training epoch = 100, batch size = 16, and the initial learning rate is 0.001.

[0119] (2) Datasets

[0120] VisDrone2021: It contains 288 video clips and more than 10,000 annotated frames, covering 8 types of scenes such as cities and transportation hubs. The targets are mainly pedestrians and vehicles, and the proportion of small targets exceeds 60%;

[0121] UAVDT: 30 hours of UAV aerial video, with vehicle targets annotated, focusing on challenges such as low light and motion blur;

[0122] Self-built dataset: 20 km 2 Industrial park data, including scenes of dense occlusion and target cross-movement, is used to verify the extreme performance of the algorithm.

[0123] (3) Evaluation Metrics

[0124] MOTA (Multi-Object Tracking Accuracy): A global metric that combines FP, FN, and ID Switch;

[0125] IDF1: Identity preservation accuracy, reflecting long-term tracking stability;

[0126] HOTA: A fine-grained metric that balances detection and association accuracy.

[0127] (4) Comparative Experiments

[0128] The comparison results with FairMOT, JDE, and the original DeepSORT are shown in Table 1.

[0129] Table 1

[0130] Algorithm MOTA↑ IDF1↑ ID Switch↓ FPS↑ DeepSORT 73.3 68.5 89 40 FairMOT 79.2 74.1 45 28 EEE-DeepSORT 85.6 82.3 32 35

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A multi-target tracking method under drone monitoring, characterized in that: The following steps are involved: S1: Extract the features of the target in the video frame collected by the drone through the EfficientNet network and generate the target feature vector; S2: Based on the dynamic matching strategy, the detection box of the current frame is associated and matched with the historical tracking target, wherein the dynamic matching strategy uses the EIoU loss function to calculate the intersection over union ratio and integrates the target's position information, speed information, appearance features and historical trajectory information; S3: The nonlinear motion state of the target is predicted and updated through the extended Kalman filter (EKF) to output a continuous trajectory.

2. The multi-target tracking method under drone monitoring according to claim 1, characterized in that: The EfficientNet network adjusts the depth, width and resolution of the network through a compound scaling method, specifically including: Depth, width, and resolution are scaled co-ordinately according to the following formulas: depth:d=α φ (1) width:w=β φ (2) resolution:r=γ φ (3) Among them, α, β, γ are constants determined by grid search on a small dataset, which control the scaling of depth, width and resolution respectively; φ represents the scaling factor.

3. The multi-target tracking method under drone monitoring according to claim 1 is characterized in that: The EIoU loss function includes the following components: (a) Intersection-over-union loss, which calculates the area ratio of the overlap area between the predicted box and the true box to the union area; (b) Distance loss, which calculates the square of the Euclidean distance between the center point of the predicted box and the true box; (c) Aspect ratio loss, which is achieved by minimizing the difference in width and height between the predicted box and the true box. The specific calculation formula is: Among them, L IOU represents the intersection-over-union loss, L dis Represents the distance loss, L asp represents the aspect ratio loss, IOU represents the intersection-over-union ratio; w and h are the width and height of the prediction box, w gt and h gt is the width and height of the target box, w c and h c is the width and height of the minimum enclosing area containing the prediction box and the target box, ρ 2 (b,b gt ) is the square of the Euclidean distance between the center point of the predicted box and the true box; b represents the position of the center point of the predicted box, bg t Indicates the center point of the real box.

4. The multi-target tracking method under drone monitoring according to claim 1 is characterized in that: The steps of the extended Kalman filter EKF include: (a) Perform first-order Taylor expansion linearization on the nonlinear state transfer function and observation function, and calculate the state transfer Jacobian matrix F k and the observation Jacobian matrix H k ; (b) By formula and Perform a priori state estimation and error covariance prediction; where, Represents the prior state estimate at time k based on the state prediction at time k-1; represents the optimal estimate at time k-1; represents the prior state estimate at time k; P k|k-1 Indicates the uncertainty of the predicted state; (c) Based on the formula and P k|k =(IK k H k ) k|k-1 Update the posterior state and error covariance; where, is the difference between observation and prediction; K represents the posterior state estimate after the observation data is integrated at time k, which is the optimal state estimate after combining the predicted value and the observed value. k represents the Kalman gain at time k, which is used to weigh the weight of the prior prediction and the observation data and determine the influence of the observation data on the state update; z k represents the actual observation value at time k; represents the observed value predicted based on the prior state estimate at time k, P k|k represents the posterior error covariance matrix after the k-time fusion observation data, reflecting the uncertainty of the posterior state estimation; I represents the unit matrix to ensure that the matrix operation dimension matches; H k represents the observation matrix at time k, describing the mapping relationship from state variables to observation space; P k|k represents the prior error covariance matrix at time k, which represents the uncertainty of the prior state estimate.

5. The multi-target tracking method under drone monitoring according to claim 1 is characterized in that: In the dynamic matching strategy, the extraction of appearance features is based on the MBConv module, and the MBConv module includes: The dilated convolutional layer, the depth-separable convolutional layer, the SE channel attention module and the projection convolutional layer, where the SE module is calculated by the formula s=GAP(F dw ) and e = σ(W2·δ(W1…)) generate channel attention weights; Among them, GAP represents global average pooling; s represents the output result of global average pooling, and e represents the channel attention weight generated after s is processed by two fully connected layers; r is the dimension reduction factor, δ is the ReLU activation function, σ is the sigmoid activation function, F dw represents the feature map, 6. The multi-target tracking method under drone monitoring according to claim 2 is characterized in that: Among the composite scaling factors of the EfficientNet, the depth scaling factor is 1.2, the width scaling factor is 1.1, and the resolution scaling factor is 1.

15.

7. The multi-target tracking method under drone monitoring according to claim 3 is characterized in that: The closure area width w in the aspect ratio loss c High c and ρ 2 (b,b gt ), calculated by the minimum bounding rectangle containing the predicted box and the true box.

8. The method for tracking multiple targets under drone monitoring according to claim 4, characterized in that: The state transition Jacobian matrix F k and the observation Jacobian matrix H k Updated dynamically in every frame.

9. The multi-target tracking method under drone monitoring according to claim 1, characterized in that: The data set of the method adopts VisDrone-MOT, and the video frame resolution of the drone's perspective is 1920×1080.

10. The multi-target tracking method under drone monitoring according to claim 1, characterized in that: In the dynamic matching strategy, when the EIoU value exceeds a dynamically adjusted threshold, it is determined that the detection box matches the tracking target successfully, and the threshold is adaptively adjusted according to the complexity of the scene.

Citation Information

Cited By

  • Audio-visual fusion unmanned aerial vehicle detection and positioning method

    CN121165207A