Multi-target tracking method for low-altitude airspace vehicles
By improving the FairMOT architecture and BYTE data association module, and combining the coordinate attention mechanism and ArcFace Loss, the problem of low detection accuracy caused by target occlusion and small size in low-altitude airspace aircraft clusters is solved, and higher accuracy multi-target tracking is achieved.
Patent Information
- Application Number
- CN202311104604.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-29
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2043-08-29
AI Technical Summary
In complex scenarios, due to factors such as mutual occlusion between low-altitude airspace vehicle clusters and small target size, existing multi-target tracking methods have low detection and tracking accuracy, making it difficult to effectively track low-altitude airspace vehicle clusters.
An improved FairMOT architecture is adopted, introducing a coordinate attention mechanism and a BYTE data association module. The Encoder-Decoder network extracts the target features of the aircraft, and ArcFace Loss is used to train the target detection and re-identification network. Combined with Kalman filter, multi-target tracking is performed.
It improves the accuracy of multi-target identification and tracking of low-altitude airspace aircraft clusters, reduces the probability of target loss and detection drift, and enhances the consistency of visual detection and tracking.
Smart Images

Figure CN117079165B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multi-target tracking, and in particular to a multi-target tracking method for low-altitude airspace vehicles. BACKGROUND
[0002] Currently, low-altitude airspace vehicle technologies such as unmanned aerial vehicles are developing rapidly and are widely used in various industries of national production. In the military field, low-altitude airspace vehicle cluster combat represented by unmanned aerial vehicles is developing rapidly towards intelligentization and actual combat, which will pose a great threat to future battlefields. Therefore, it is imperative to conduct research on visible light detection and tracking of countermeasures for low-altitude airspace vehicle clusters.
[0003] However, due to the small geometric size of low-altitude vehicles, small radar scattering cross section, and weak infrared features, it is difficult for ordinary air defense systems to timely detect and intercept targets, and detection and tracking mainly rely on low-altitude search radars, infrared and visible light detection systems.
[0004] In complex scenes, the low-altitude airspace vehicle cluster has problems such as mutual occlusion between targets and small target size, which leads to tracking drift, target loss, and other situations, thereby reducing the accuracy of visual detection and tracking. Therefore, multi-target tracking technology based on vision has become one of the key technologies for countering low-altitude airspace vehicle clusters.
[0005] The existing multi-target tracking method FairMOT is a typical method based on the Joint Detection and Embedding (JDE) framework in the field of multi-target tracking, which can be used for real-time multi-target tracking. While locating multiple targets of interest, it maintains the identity number of each target and records the continuous motion trajectory. FairMOT includes an Encoder-Decoder network, a target detection network, a re-identification network, and a data association module. However, this method is proposed for pedestrian datasets, and it is not suitable for the characteristics of low-altitude airspace vehicle clusters. Therefore, the present application improves FairMOT to make it more suitable for tracking low-altitude airspace vehicle clusters. SUMMARY
[0006] In view of the above problems of the prior art, the present application provides a multi-target tracking method for low-altitude airspace vehicles, which aims to solve the problem of low detection and tracking accuracy caused by mutual occlusion between targets and small target size in complex scenes, and improve the accuracy of multi-target recognition and tracking.
[0007] To solve the above technical problems, the technical solution adopted by the present application includes the following steps:
[0008] Step 1: Obtain a video sequence containing position information and identity number of an aircraft target, and construct a training data set;
[0009] Step 2: Input the training data set into an Encoder-Decoder network based on the FairMOT architecture for feature extraction, and input the extracted aircraft target features into a coordinate attention mechanism to obtain enhanced aircraft target features;
[0010] Further, the method of inputting the extracted aircraft target features into the coordinate attention mechanism to obtain enhanced aircraft target features comprises the following steps:
[0011] Step 2.1: Perform pooling operations on the extracted aircraft target features along the horizontal direction and the vertical direction respectively to obtain feature maps in the width and height directions;
[0012] Step 2.2: Concatenate the obtained feature maps in the two directions, input the concatenated feature map into a convolution layer with a convolution kernel of 1x1 for convolution transformation, and then perform batch normalization processing and Sigmoid nonlinear activation function to obtain a feature map F;
[0013] Step 2.3: Perform convolution operations on the feature map F along the horizontal direction and the vertical direction respectively with a convolution kernel of 1x1, and then obtain attention weights g h and g w in the height and width directions respectively after Sigmoid nonlinear activation function;
[0014] Step 2.4: Perform multiplication weighting calculation on the extracted aircraft target features and the attention weights g h and g w of the feature map F in the height and width directions to obtain enhanced aircraft target features;
[0015] Step 3: Input the enhanced aircraft target features into a target detection network based on the FairMOT architecture to obtain target detection results, input the enhanced aircraft target features into a re-identification network based on the FairMOT architecture to obtain appearance features of the target, and train the target detection network and the re-identification network to obtain a trained target detector;
[0016] Further, the target detector comprises an Encoder-Decoder network, a coordinate attention mechanism, a target detection network, and a re-identification network;
[0017] Further, the method of training the target detection network and the re-identification network is to use a loss function L detection for training the target detection network and a loss function L ArcFaceThe network total loss function trains the target detection network and the re-identification network, and the network total loss function is:
[0018]
[0019] Ltotal=w1Ltarget+w2Lreid detection Ltarget represents a loss function for training the target detection network, L ArcFace Lreid represents a loss function for training the re-identification network; w1 and w2 are weight coefficients of the target detection network and the re-identification network, respectively.
[0020] Step 4: Obtain a video sequence containing an aircraft target to be detected, and input the video sequence containing the aircraft target to be detected into a target detector to obtain a target detection result and an appearance feature of the target;
[0021] The target detection result includes a detection frame and a detection score.
[0022] Step 5: Input the target detection result and the appearance feature of the target into a BYTE data association module for tracking, and output a multi-target tracking result.
[0023] The multi-target tracking result includes a boundary frame of the target and an identity serial number.
[0024] Step 5.1: Set a detection score threshold τ, and divide a detection frame with a detection score higher than the detection score threshold into a high-score detection frame D high , and otherwise into a low-score detection frame D low ; regard the boundary frame and the identity serial number of the target in the first frame as a target track, establish a track set T according to the target tracks in all frames in the video sequence, and initialize the track set T, a first-time unsuccessfully matched detection frame set D remain , a first-time unsuccessfully matched track set T remain , a second-time unsuccessfully matched track set T re-remain , and a track set T lost that is unsuccessfully matched both times to be empty; for each frame in the video sequence, repeat steps 5.2-5.5;
[0025] Step 5.2: predict a new target track of the aircraft target to be detected in the current frame using a Kalman filter, and obtain a prediction frame of the target track.
[0026] Step 5.3: perform a first-time matching between D high and T, reserve a high-score detection frame that is unsuccessfully matched to D remain , and reserve a target track that is unsuccessfully matched to T remain .
[0027] The unsuccessfully matched detection frame is reserved to Dhigh The first matching between D and T is performed by using D high D is calculated based on the intersection over union between the prediction boxes of the new target trajectories of D and T in the current frame and the appearance features of the targets extracted by the re-identification network high The similarity of the prediction boxes of the new target trajectories of D and the to-be-detected aircraft targets in the current frame is calculated, and D is completed based on the similarity using the Hungarian algorithm high The matching between D and T is completed.
[0028] Step 5.4: The second matching between D low and T remain is performed, the unsuccessfully matched target trajectories are retained to T re-remain , and all the unsuccessfully matched low-score detection boxes in the current frame are deleted.
[0029] The second matching between D low and T remain is performed by using D low and T remain The intersection over union between the prediction boxes of the new target trajectories of D and T in the current frame is calculated to obtain D low The similarity of the prediction boxes of the new target trajectories of D and the to-be-detected aircraft targets in the current frame is calculated, and D is completed based on the similarity using the Hungarian algorithm low The matching between D and T remain is completed.
[0030] Step 5.5: The trajectories in T re-remain are retained to T lost , and if the trajectories still cannot be matched to the detection boxes for 30 consecutive frames, the trajectories are deleted from T and T lost , and the detection boxes in D remain are used to initialize new target trajectories and put the new target trajectories into T.
[0031] Step 5.6: Until the matching of all the frames of the video sequence is completed, the trajectory set T is output as the multi-target tracking result, including the positions and identity serial numbers of the targets in each frame.
[0032] The technical scheme has the beneficial effects that:
[0033] The method of the application embeds the position information of the aircraft target into the channel attention by introducing the coordinate attention mechanism into the Encoder-Decoder network, considers the relationship between the channels of the image and the position information, and helps the network to better locate and identify the target; in the training process, the ArcFace Loss is used as the loss function in the re-identification network, so that more accurate appearance features can be extracted, the discrimination ability is enhanced, and the accuracy of visual detection and tracking is improved; the BYTE data association module is selected for target tracking in the method of the application, the similarity between the detection frame and the tracking track is used, the background is removed from the low-score detection result while the high-score detection result is reserved, and the real object is mined, so as to reduce the missed detection and improve the continuity of the track. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 A flow chart of the multi-target tracking method for low-altitude airspace aircraft in the embodiment;
[0035] Figure 2 A principle diagram of the multi-target tracking method for low-altitude airspace aircraft in the embodiment;
[0036] Figure 3 A principle diagram of the coordinate attention mechanism in the embodiment. DETAILED DESCRIPTION
[0037] In order to facilitate the understanding of the present application, the specific embodiments of the application are further described in detail below in combination with the drawings and embodiments. The following embodiments are used to illustrate the application, but are not used to limit the scope of the application. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.
[0038] The core idea of the method of the application is: training a target detector using a training data set and completing aircraft multi-target tracking using the target detector and a BYTE data association module. The process of training the target detector includes: inputting the training data set into an Encoder-Decoder network based on the FairMOT architecture for feature extraction, introducing the extracted aircraft target features into the coordinate attention mechanism to obtain enhanced aircraft target features; inputting the enhanced aircraft target features into a target detection network and a re-identification network based on the FairMOT architecture to obtain target detection results; in the process of training the target detector, the loss function L detection and L ArcFaceThe target detection network and the re-identification network are trained together to obtain the target detector. The process of using the target detector and the BYTE data association module to complete the multi-target tracking of the aircraft includes: sending a video sequence containing the target aircraft to be detected into the target detector to obtain the target detection result; inputting the target detection result into the BYTE data association module to output the multi-target tracking result, including the target's bounding box and identification number.
[0039] This embodiment provides a multi-target tracking method for low-altitude airspace vehicles, such as... Figure 1 As shown, it includes the following steps:
[0040] Step 1: Obtain video sequences containing the location information and identification number of the aircraft target, and construct a training dataset;
[0041] In this embodiment, there is no limitation on the method of obtaining the training dataset. The training dataset can be obtained from an existing database, or a pre-collected and manually labeled video sequence can be used as the training dataset. The method is as follows: collect a video sequence containing multiple drone targets, convert it into images frame by frame, and label the location information of the targets, namely the bounding box coordinates and identification number.
[0042] Step 2: Input the training dataset into the Encoder-Decoder network based on the FairMOT architecture for feature extraction, and input the extracted aircraft target features into the coordinate attention mechanism to obtain enhanced aircraft target features;
[0043] In this embodiment, such as Figure 2 As shown, the training dataset is input into an Encoder-Decoder network based on the FairMOT architecture for feature extraction. An existing Encoder-Decoder network based on the FairMOT architecture is selected, and DLA-34 is used as the backbone network for feature extraction. This Encoder-Decoder network includes six levels from level 0 to level 5. Level 0 and level 1 consist of 3×3 convolutions with strides of 1 and 2, respectively, followed by processing through four tree-like levels from level 2 to level 5. Within each level, downsampling and concatenation operations are used to achieve residual structures, and there are larger residual connections between different levels than within the internal residual structures. These residual connections between different levels enable the detection of targets at different scales, allowing the network to dynamically adjust its receptive field based on the target's size and pose.
[0044] Further, the coordinate attention mechanism comprises: coordinate information embedding and coordinate attention generation, wherein the coordinate information embedding comprises two parallel average pooling layers, and the coordinate attention generation comprises a splicing and convolution layer, a batch normalization and Sigmoid nonlinear layer, two parallel convolution layers, and a Sigmoid nonlinear layer.
[0045] The method for inputting the extracted aircraft target feature into the coordinate attention mechanism to obtain an enhanced aircraft target feature comprises the following steps:
[0046] Step 2.1: performing a pooling operation on the extracted aircraft target feature along a horizontal direction and a vertical direction respectively to obtain feature maps in the width and height directions;
[0047] Step 2.2: splicing the obtained feature maps in the two directions, inputting the spliced feature map into a convolution layer with a convolution kernel of 1×1 for convolution transformation, and then performing batch normalization processing and Sigmoid nonlinear activation function to obtain a feature map F;
[0048] Step 2.3: performing a convolution operation on the feature map F along the horizontal direction and the vertical direction respectively with a convolution kernel of 1×1, and then obtaining attention weights g h and g w of the feature map F in the height and width directions respectively after Sigmoid nonlinear activation function;
[0049] Step 2.4: performing multiplication and weighting calculation on the extracted aircraft target feature and the attention weights g h and g w of the feature map F in the height and width directions to obtain an enhanced aircraft target feature;
[0050] Step 3: inputting the enhanced aircraft target feature into a target detection network based on the FairMOT architecture to obtain a target detection result, inputting the enhanced aircraft target feature into a re-identification network based on the FairMOT architecture to obtain an appearance feature of the target, and training the target detection network and the re-identification network to obtain a trained target detector;
[0051] In this embodiment, as shown in Figure 3 , the aircraft target feature extracted through the Encoder-Decoder network is input into the coordinate attention mechanism, is subjected to an X-direction average pooling layer and a Y-direction average pooling layer respectively, is subjected to splicing and convolution transformation through a splicing and convolution layer, is input into a batch normalization and Sigmoid nonlinear layer to obtain a feature map F, and is subjected to convolution transformation along the X direction and the Y direction respectively and is subjected to a Sigmoid nonlinear layer to obtain attention weights g h and gw The extracted aircraft target feature is multiplied by the attention weight g h and g w to obtain an enhanced aircraft target feature.
[0052] The target detector comprises an Encoder-Decoder network, a coordinate attention mechanism, a target detection network and a re-identification network.
[0053] In the embodiment, the target detection network comprises three parallel branches of a heatmap head, a center offset head and a box size head, each branch being composed of a 3x3 convolution layer with 256 channels and a 1x1 convolution layer.
[0054] The heatmap head is responsible for estimating the position of the target center, and outputs a feature layer with a size of 1xHxW; the center offset head is responsible for more accurate positioning of the object, and outputs a feature layer with a size of 2xHxW; and the box size head is responsible for estimating the height and width of the target bounding box at each anchor point, and outputs a feature layer with a size of 2xHxW; wherein H and W are the height and width of the enhanced aircraft target feature, respectively.
[0055] The re-identification network comprises a Re-ID embedding head branch composed of a 3x3 convolution layer with 256 channels and a 1x1 convolution layer, which is used to generate appearance features capable of distinguishing different aircraft targets for identity recognition of different targets. The re-identification network outputs a feature layer with a size of 128xHxW, wherein H and W are the height and width of the enhanced aircraft target feature, respectively.
[0056] The target detection result comprises a detection box and a detection score.
[0057] Further, the method for training the target detection network and the re-identification network comprises: training the target detection network and the re-identification network by using a total loss function composed of a loss function L detection for training the target detection network and a loss function L ArcFace for training the re-identification network, wherein the total loss function for training the network is:
[0058]
[0059] wherein L detection represents the loss function for training the target detection network, and L ArcFacerepresents a loss function for training the re-identification network; w1 and w2 are weight coefficients of the target detection network and the re-identification network respectively, used to balance the two branch tasks of target detection and re-identification;
[0060] In the embodiment, the training effect of the target detector is judged by using the multi-target tracking accuracy MOTA in the multi-target tracking evaluation index. Since the value range of MOTA is (-∞, 1), in the value range, the greater the value of MOTA, the better the training effect of the target detector. In addition, there are many evaluation indexes for judging the training effect of the target detector, for example, the number of training rounds is specified according to the number of training data and the structure of the network during training. The evaluation index for judging the training effect of the target detector is not limited in the method.
[0061] Step 4: Obtain a video sequence containing an aircraft target to be detected, and input the video sequence containing the aircraft target to be detected into the target detector to obtain a target detection result and an appearance feature of the target;
[0062] Step 5: Input the target detection result and the appearance feature of the target into the BYTE data association module for tracking, and output a multi-target tracking result, including a bounding box and an identity number of the target;
[0063] Step 5.1: Set a detection score threshold τ, and divide the detection boxes with a detection score higher than the detection score threshold into high-score detection boxes D high , and otherwise into low-score detection boxes D low ; the bounding box and the identity number of the target in the first frame are regarded as a target track, a track set T is established according to the target tracks in all frames in the video sequence, the track set T is initialized as empty, and a first-time unsuccessfully matched detection box set D remain , a first-time unsuccessfully matched track set T remain , a second-time unsuccessfully matched track set T re-remain , and a track set T lost that is unsuccessfully matched twice are initialized as empty; for each frame in the video sequence, steps 5.2-5.5 are repeated;
[0064] In the embodiment, according to the experimental results, the detection score threshold is set as τ = 0.6, at which time the tracking effect on the target is better;
[0065] Step 5.2: Use a Kalman filter to predict a new target track of the aircraft target to be detected in the current frame, and obtain a prediction box of the target track;
[0066] Step 5.3: Perform first-time matching between D high and T, and reserve the high-score detection boxes that are unsuccessfully matched to D remainThe unmatched target trajectory is saved to T. remain ;
[0067] Using D high D is calculated using the intersection-over-union (IoU) ratio between the predicted bounding boxes of the new target trajectory and T in the current frame, and the appearance features of the target extracted by the re-identification network. high The similarity between the predicted bounding box of the target aircraft and the new target trajectory in the current frame is used; based on this similarity, the Hungarian algorithm is used to complete the D... high Matching between T;
[0068] Step 5.4: In D low and T remain A second matching is performed between them, and the target trajectories that did not match are saved to T. re-remain And delete all low-scoring detection boxes that did not match successfully in the current frame;
[0069] Using D low and T remain The cross-union ratio D is calculated among the predicted bounding boxes of the new target trajectory in the current frame. low The similarity between the predicted bounding box and the new target trajectory in the current frame is used to perform D-scanning using the Hungarian algorithm. low and T remain Matching between;
[0070] In this embodiment, step 5.4 uses only the cross-union ratio to calculate similarity without adding appearance features, because low-scoring detection boxes usually contain severe occlusion or motion blur, in which case appearance features are unreliable;
[0071] Step 5.5: Place T re-remain The trajectory in T is preserved. lost During this process, if a detection box is still not matched after 30 consecutive observations, it is removed from T and T's. lost Delete it, and use D at the same time. remain The detection box in the middle initializes the new target trajectory and puts it into T;
[0072] In this embodiment, D remain The first detection box in step 5.3 that failed to match the aircraft trajectory, and whose detection score is high enough, indicates that D... remain The target aircraft corresponding to the detection box in the image should be a newly appearing target in this frame, and it needs to be detected using D. remain In other words, the detection box establishes a new trajectory for newly appearing targets.
[0073] Step 5.6: Until all frames of the video sequence match end, the output track set T is the multi-target tracking result, including the position and identity sequence number of the target in each frame.
[0074] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, but not limited to them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope defined by the claims of the present application.
Claims
1. A multi-target tracking method for low-altitude airspace vehicles, characterized by, The method comprises the following steps: Step 1: Obtain a video sequence containing position information and identity number of an aircraft target, and construct a training data set; Step 2: input the training data set into an Encoder-Decoder network based on the FairMOT architecture for feature extraction, and input the extracted aircraft target features into a coordinate attention mechanism to obtain enhanced aircraft target features; Step 3: input the enhanced aircraft target features into a target detection network based on the FairMOT architecture to obtain a target detection result, input the enhanced aircraft target features into a re-identification network based on the FairMOT architecture to obtain an appearance feature of the target, and train the target detection network and the re-identification network to obtain a trained target detector; The target detection result comprises a detection box and a detection score. Step 4: obtain a video sequence containing a to-be-detected aircraft target, and input the video sequence containing the to-be-detected aircraft target into the target detector to obtain a target detection result and an appearance feature of the target; Step 5: input the target detection result and the appearance feature of the target into a BYTE data association module for tracking, and output a multi-target tracking result; Step 5.1: Set the detection score threshold detection boxes with detection scores higher than the detection score threshold are classified into high-score detection boxes , and vice versa, into low-score detection boxes ; the bounding boxes and identity numbers of the target in the first frame are regarded as a target track, and a track set is established according to the target tracks in all frames in the video sequence ; the track set is initialized as empty, and a set of detection boxes that are unsuccessfully matched for the first time is initialized ; a set of tracks that are unsuccessfully matched for the first time is initialized ; a set of tracks that are unsuccessfully matched for the second time is initialized ; and a set of tracks that are unsuccessfully matched for both times is initialized ; and the sets are empty; for each frame in the video sequence, steps 5.2-5.5 are repeated; Step 5.2: use a Kalman filter to predict a new target trajectory of the to-be-detected aircraft target in the current frame, and obtain a prediction box of the target trajectory; Step 5.3: Perform the first match between and , keep the high-score bounding boxes that do not successfully match to , and keep the target trajectories that do not successfully match to ; The method for performing the first matching between the and is to use and to calculate the intersection over union between the prediction boxes of the new target track in the current frame and the appearance features of the target extracted by the re-identification network and the similarity of the prediction boxes of the new target track in the current frame of the to-be-detected aircraft target; based on the similarity, the Hungarian algorithm is used to complete the matching between the and Step 5.4: Perform the second matching between and , keep the target trajectories that are not successfully matched to , and delete all low-score bounding boxes that are not successfully matched in the current frame; The method for the second time matching between the and is to use and The intersection over union calculation between the prediction boxes of the new target track in the current frame and the prediction box similarity of the new target track in the current frame, based on the similarity, the Hungarian algorithm is used to complete the matching between and Step 5.5: If the trajectory in T is not matched to a detection box in D, then remove it from T and initialize a new trajectory with the detection box in D and put it into T. Step 5.6: until the matching of all frames of the video sequence is completed, output a track set T as the multi-target tracking result, comprising the position and identity number of the target in each frame. 2.The multi-target tracking method for low-altitude airspace vehicles according to claim 1, wherein, The method for obtaining enhanced aircraft target features by inputting the extracted aircraft target features into a coordinate attention mechanism in step 2 comprises the following steps: Step 2.1: perform a pooling operation on the extracted aircraft target features along the horizontal direction and the vertical direction respectively to obtain feature maps in the width and height directions; Step 2.2: splice the obtained feature maps in the two directions, input the spliced feature map into a convolution layer with a convolution kernel of 1×1 for convolution transformation, and then perform batch normalization processing and Sigmoid nonlinear activation function to obtain a feature map F; Step 2.3: The feature map F is respectively subjected to a convolution operation with a 1x1 convolution kernel along the horizontal direction and the vertical direction, and then respectively subjected to a Sigmoid nonlinear activation function to obtain the attention weight g of the feature map F in the height and width directions h and g w ; Step 2.4: Multiply the extracted aircraft target features with the attention weights g on the height and width of the feature map F h and g w to get the enhanced aircraft target features. 3.The multi-target tracking method for low-altitude airspace vehicles according to claim 1, wherein, The target detector comprises an Encoder-Decoder network, a coordinate attention mechanism, a target detection network and a re-identification network. 4.The multi-target tracking method for low-altitude airspace vehicles according to claim 1, wherein, The method for training the target detection network and the re-identification network is: training the target detection network and the re-identification network by using a network total loss function composed of a loss function for training the target detection network and a loss function for training the re-identification network The network total loss function is: ; wherein, represents a loss function for training the target detection network, represents a loss function for training the re-identification network; and are a weight coefficient of the target detection network and a weight coefficient of the re-identification network, respectively.
5. The multi-target tracking method for low-altitude airspace vehicles according to claim 1, wherein, The multi-target tracking result comprises a bounding box and an identity number of the target.
Citation Information
Patent Citations
Unmanned aerial vehicle video multi-target tracking method based on attention feature fusion
CN113807187A
Chain type multi-target tracking method of secondary correlation low-resolution detection frame
CN114724059A