A Target Tracking Method from the Perspective of a Low-Altitude Aircraft

By designing a specific network framework and feature fusion strategy, the problem of missing target tracking methods under short-term occlusion in low-altitude aircraft perspective is solved, achieving higher tracking accuracy and accuracy, ensuring accurate positioning of target positions.

CN115880332BActive Publication Date: 2025-07-22NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211453407.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-21
Publication Date
2025-07-22
Estimated Expiration
2042-11-21

AI Technical Summary

Technical Problem

The existing target tracking method at the perspective of low-altitude aircraft is prone to losing targets under short-term occlusion, and the target feature extraction and matching are not accurate enough, resulting in uncertain location.

Method used

Design a specific network framework for feature extraction and matching, use twin networks and Transformer structures to fusion, combine attention mechanisms to strengthen the similarity between the target and the template, and estimate the position after the occlusion through the target motion information before occlusion to avoid template updates.

Benefits of technology

It improves the accuracy and accuracy of target tracking, reduces the probability of target loss after short-term occlusion, reduces background interference and error accumulation, and achieves accurate re-tracking in the occlusion situation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115880332B_ABST
    Figure CN115880332B_ABST
Patent Text Reader

Abstract

The present invention relates to a target tracking method from the perspective of a low-altitude aircraft. In view of the analysis of the characteristics of aviation target tracking data, a network framework for feature extraction and feature matching under specific conditions is designed. At the same time, during the target tracking process, when a short-term foreground occlusion occurs to a moving target, the movement trajectory of the target under the window before occlusion is calculated, and the position estimation of the target object under occlusion is maintained to achieve accurate re-tracking of the target after occlusion. Through the analysis of the spatio-temporal sequence data from the perspective of a low-altitude aircraft, a backbone network adapted to the characteristics of its data is designed to obtain low coupling degree and rich target spatial and semantic information. During the matching process of target features and template features, an attention mechanism is adopted to strengthen the similarity between the target and the template, reduce background interference, and further ensure the accuracy of tracking through two-layer feature fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of artificial intelligence, optoelectronic engineering, computer vision, and target tracking, and relates to a target tracking method from the perspective of a low-altitude aircraft. Background Art

[0002] The artificial intelligence industry has been fully penetrated into all walks of life in human society and has profoundly changed people's working methods and living habits. Among them, computer vision is a simulation technology of biological vision using computers and related devices. The modeling of vision systems requires the guidance of technologies such as optoelectronic engineering. Target tracking in computer vision technology has a wide range of applications. Target tracking is the process of continuously positioning a selected object in an image sequence and is an important research topic in the field of computer vision. As an important research direction of target tracking, the single-object tracking task completes the subsequent target positioning process on the premise of a given target in the initial few frames. Tracking models are divided into generative models and discriminant models. The template information is updated in real time according to the tracking results, and the accuracy and robustness of the tracker can be improved to a great extent through appropriate template update strategies.

[0003] Target tracking from the perspective of low-altitude aircraft (such as drones, low-altitude manned aircraft) has a wide range of applications in military strategic strikes and defenses, drone positioning, border patrols, traffic control, etc. Since target tracking is the process of continuously positioning a selected object in an image sequence, it can be understood as an object recognition behavior in the time series. A series of difficult problems will arise in target tracking tasks due to completely different targets and constraint conditions. In addition to some common challenges in other computer vision tasks, such as objects under different viewpoints, lighting, intra-class variations, etc., the challenges of target tracking from an aerial perspective also include, but are not limited to, the following aspects: rotation and scaling changes of objects (such as small objects), precise positioning of objects, dense and occluded detection of objects, detection speed, etc. At the same time, aerial single-object tracking is a continuous positioning process of a specific single object in the time series. Therefore, the difficult problems faced also include error accumulation in long-term tracking, tracking drift caused by interference from intra-class similar targets, tracking failure problems caused by occlusion during the tracking process, and precise tracking of objects under scale changes. These difficult problems have also promoted the research of aerial single-object tracking algorithms. The present invention focuses on the difficult problems in target tracking from the perspective of low-altitude aircraft, and proposes a set of accurate and precise target tracking solutions for low-altitude aircraft perspectives by designing a network framework, improving the target search strategy, and formulating a template update method.

[0004] The feature of the single-object tracking algorithm based on the Siamese network is that it uses a network with shared weights to extract features of the target and the search domain, simplifies the feature processing of object tracking, and improves the tracking speed. Its general process is to extract features of the target template and the search domain, and complete the estimation task of the target position and size through the interaction of deep features. Bertinetto et al. proposed a simple and efficient single-object tracking algorithm in the literature "L. Bertinetto, J. Valmadre, J. Henriques, A. Vedaldi, and P. Torr, “Fully-Convolutional Siamese Networks for Object Tracking,” in Proc. European Conference on Computer Vision, 2016, pp. 850-865." On the one hand, the fully convolutional network is used in the backbone network part of feature extraction, and on the other hand, a simple cross-correlation operation is used in similarity measurement, which greatly improves the accuracy and speed of object tracking. However, the complex background changes interfere with the local target features, and relying on a single cross-correlation operation will affect the accuracy of target feature matching.

[0005] The search strategies for object tracking are divided into local search strategies and global search strategies. The local search strategy is to search and track the object within the local area of the current frame based on the position information of the object in the previous frame. Wang et al. proposed a simple and effective sampling strategy to break the limitation of spatial invariance and a cross-layer feature aggregation structure to aggregate multi-scale feature maps in the literature "B.Li, W.Wu, Q.Wang, F.Zhang, J.Xing, and J.Yan, “SiamRPN++: Evolution of Siamese Visual Tracking with Very Deep Networks,” in Proc. IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 4277-4286." However, since these methods follow the locality assumption and cannot re-track the object after the tracking object is lost, these tracking methods only detect locally, resulting in not meeting the requirements for long-term tracking. The global search strategy is to directly search and track the object globally in the current frame. Huang et al. eliminated the local position assumption in the literature "L.Huang, X.Zhao, and K.Huang, “GlobalTrack: A Simple and Strong Baseline for Long-Term Tracking,” in Proc. AAAI Conference on Artificial Intelligence, 2020, pp. 11037-11044." and proposed a two-stage-based global instance search tracker that can search for the object at any position and scale, thus avoiding the cumulative error in the long-term tracking process. However, the problem brought by the global search strategy is that the model processing time is slow and the real-time response is not strong. Summary of the Invention

[0006] Technical Problems to be Solved

[0007] To avoid the deficiencies of the prior art, the present invention proposes an object tracking method from the perspective of a low-altitude aircraft, mainly alleviating the problem of object tracking loss under short-term occlusion, and at the same time improving the accuracy and precision of object tracking from the perspective of a low-altitude aircraft.

[0008] The purpose of the present invention is to improve the following aspects.

[0009] 1. The loss of the object after short-term occlusion in the existing aviation object tracking methods.

[0010] 2. The effective extraction and mining of the features of the object in the object tracking method based on a low-altitude aircraft.

[0011] 3. Target tracking method from the perspective of low-altitude aircraft, the accuracy of matching between target features and template features.

[0012] 4. Target tracking method from the perspective of low-altitude aircraft, the movement at the target end and the tracking end respectively will cause the uncertainty of the object position.

[0013] Technical solution

[0014] A target tracking method from the perspective of low-altitude aircraft, characterized in that the steps are as follows:

[0015] Step 1: Use the drone to collect data in the real scene, simulate the target tracking data from the perspective of a low-altitude aircraft in the virtual scene, and finally annotate the data on darklabel to establish a real-virtual aerial remote sensing data set;

[0016] Step 2: Modify the total stride of the last two convolutional blocks of the Siamese network to be reduced to 8, and add an additional 1×1 convolutional layer to each convolutional block output to reduce the number of feature channels to 256;

[0017] Use the existing general target tracking data set as the training data of the network model. The input of the improved Siamese network is the RGB images of the target template and the current search frame, and the feature maps of the target template and the search frame are obtained;

[0018] Step 3: Compress and process the feature maps of the target template and the current frame extracted by the Siamese backbone network into one-dimensional representations, and then obtain q and k through position embedding encoding; the one-dimensional representation information without processing is v, and the input feature fusion module strengthens the features during the matching stage of the target and the template;

[0019] The feature fusion module is divided into temporal feature fusion and spatial feature fusion, specifically:

[0020] Input the feature map of the target template in the previous frame and the information of the target template in the current frame into the first Transformer structure. Among them, the information of the target template in the current frame is input into the decoder module of the Transformer, and then the k and v of the feature of the target template in the previous frame are input for feature matching;

[0021] Input the feature map of the current search frame and the information of the previous search frame feature map into the second Transformer structure. Among them, the feature map of the current search frame is input into the decoder module of the Transformer, and then the k and v of the information of the previous search frame feature map are input for feature matching;

[0022] Input the current search frame feature map and the current frame target template information into the third Transformer structure, where the current frame target template information is input into the encoder module of the Transformer, and the current search frame feature map is input into the decoder module of the Transformer for feature matching;

[0023] Input the information fused by the first Transformer structure and the information fused by the second Transformer structure into the first multi-head cross-attention module for fusion, then input the information fused by the second Transformer structure into the first sum regularization module for fusion, and input it into the first feed-forward neural network for fusion; The output of the first feed-forward neural network and the output of the first sum regularization module are simultaneously input into the second sum regularization module for fusion; The output of the second sum regularization module and the output of the third Transformer structure are simultaneously input into the second multi-head cross-attention module for fusion, and then the output of the second sum regularization module and the output of the third Transformer structure are simultaneously input into the third sum regularization module for fusion, and input into the second feed-forward neural network for fusion; The output of the second feed-forward neural network and the output of the third sum regularization module are simultaneously input into the fourth sum regularization module for fusion to obtain the output of the feature fusion information;

[0024] Step 4: Input the result obtained in Step 3 into the classification network and the regression network. Obtain the center position of the target through the classification network and the size of the target through the regression network;

[0025] Step 5: Use the publicly available dataset to perform end-to-end training on the network framework in Steps 2 to 4;

[0026] Step 6: In the online learning stage, estimate the target motion information through the target positions of several frames before the current frame to implement the target tracking method from the perspective of a low-altitude aircraft.

[0027] The classification network in Step 4 outputs a feature map of w*h*2, where w and h are the length and width of the feature map respectively, that is, each point contains a 2D vector as the foreground and background corresponding position scores for classification.

[0028] The regression network in Step 4 outputs a feature map of w*h*4, and each point contains a four-dimensional vector representing the distances of the four corners to determine the target.

[0029] After the end-to-end training in Step 5 is completed, fine-tune the obtained network framework on the data we collected.

[0030] In the target tracking in Step 6, in the case of no occlusion, use the target matching result to determine the target position, and at the same time correct the estimation of the target motion and update the target template.

[0031] In the target tracking in step 6, in the case of occlusion, the target position is estimated using the target motion information, the target search area is maintained, the template is not updated, the motion trajectory of the target under short-term occlusion is predicted, and when it is determined that the target is no longer occluded, the target is re-tracked relying on the template.

[0032] Beneficial effects

[0033] A target tracking method from the perspective of a low-altitude aircraft proposed by the present invention analyzes the characteristics of aviation target tracking data and designs a network framework for feature extraction and feature matching under specific conditions. At the same time, during the target tracking process, when a short-term foreground occlusion occurs to a moving target, by calculating the moving trajectory of the target under the window before occlusion and maintaining the position estimation of the target object under occlusion, accurate target re-tracking after occlusion is achieved.

[0034] The present invention has the following advantages compared with the existing aviation target tracking methods:

[0035] 1. The edge information of target feature extraction is richer and the feature expression is more accurate. The present invention is a target tracking method from the perspective of a low-altitude aircraft. By analyzing the spatio-temporal sequence data from the perspective of a low-altitude aircraft, a backbone network adapted to its data characteristics is designed to obtain low-coupling, rich target spatial and semantic information.

[0036] 2. The target feature matching is more accurate. The present invention uses an attention mechanism in the process of matching target features and template features to strengthen the similarity between the target and the template, reduce background interference, and further ensure the accuracy of tracking through two-layer feature fusion.

[0037] 3. The target re-tracking after short-term occlusion is more accurate. When a short-term foreground occlusion occurs to the target in the present invention, the target position after occlusion is estimated using the target motion information of the previous several frames, realizing the target re-tracking after short-term occlusion and reducing the probability of target tracking failure or drift.

[0038] 4. The information pollution of the target template is reduced. The present invention does not update the target template after the target is occluded, reducing the interference of foreground information on the template. At the same time, using the target motion information to assist in guiding the target tracking can obtain a more accurate estimation of the target position and size, making the target template information more accurate and reducing the error accumulation over time. Description of the drawings

[0039] Figure 1 : Flow chart of the overall framework

[0040] Figure 2 : Flow chart of target motion information estimation

[0041] Figure 3 : Network framework of the present invention Detailed implementation manners

[0042] The present invention will be further described below in conjunction with embodiments and the accompanying drawings:

[0043] The present invention provides a target tracking method from the perspective of a low-altitude aircraft. By analyzing the characteristics of aviation target tracking data, a network framework for feature extraction and feature matching under specific conditions is designed. At the same time, during the target tracking process, when a short-term foreground occlusion occurs to a moving target, the movement trajectory of the target under the window before occlusion is calculated, and the position estimation of the target object under occlusion is maintained to achieve accurate target re-tracking after occlusion. The technical solution is as follows:

[0044] 1. Target feature extraction with an adaptive receptive field: In terms of extracting target feature information, a convolutional model is used to build a Siamese-based backbone network. By adjusting the convolutional kernel parameters to be suitable for the receptive field of aviation remote sensing targets, the sensitivity of the model to target edge information is improved, and the ability of the model to capture feature changes of the target during movement and learn the overall features of the target is enhanced. Using the idea of grouped convolution, information mixing is performed in the spatial dimension, and by bringing the depth wise convolutional layer forward, the network parameters are reduced to a certain extent. Through the overall design of the backbone network, the target feature extraction work from the aviation perspective is completed.

[0045] 2. Multi-level target feature matching and fusion strategy: During the matching process of the target features and template features in the Siamese backbone network, a two-layer serial feature fusion network is used to strengthen the feature matching effect. In each layer, first, the self-attention mechanism is used to perform self-attention feature processing on the template features and the target features in the current frame respectively to strengthen the semantic information. Secondly, the processed template features and target features are used for cross-attention feature fusion to strengthen the template information on the target features in the current frame and improve the target matching effect. In the serial feature fusion network, the method of multi-layer fusion is used to further strengthen the target feature information in the current frame and minimize the interference of noise and background to the greatest extent.

[0046] 3. Target re-tracking and template selection and update method based on moving target information: It is determined whether the target is occluded according to the target feature confidence and target movement information, and the target movement information is estimated based on the target positions in several frames before the current frame. In the case of non-occlusion, the target position is determined using the target matching result, and at the same time, the estimation of the target movement is corrected and the target template is updated. In the case of occlusion, the target position is estimated using the target movement information, the target search area is maintained, the template is not updated, the movement trajectory of the target under short-term occlusion is predicted, and when it is determined that the target is no longer occluded, the target is re-tracked relying on the template.

[0047] Specific embodiment: Refer to Figure 1, the implementation steps of the present invention are as follows:

[0048] Step 1, in order to obtain a low-altitude aircraft perspective target tracking model with strong robustness, first establish a rich aerial remote sensing tracking data set. Use drones to collect data in real scenarios and simulate target tracking data from the perspective of low-altitude aircraft in virtual scenarios. Finally, perform data annotation on darklabel to complete the establishment of the real-virtual aerial remote sensing data set.

[0049] Step 2, use the existing general target tracking data set as the training data for the network model. The input of the network is the target template and the RGB image of the current search frame. For the two input RGB images, a Siamese network with weight sharing is adopted to obtain the feature maps of the target and the search frame. Among them, the Siamese network uses ConvNext, which performs well in classification tasks. The spatial information it obtains is crucial for the accurate positioning of the target, so it is used in our tracking task. We modified the stride of the last two convolutional blocks to reduce the total stride to 8 and increased the receptive field through dilated convolution. An additional 1×1 convolutional layer is added to the output of each convolutional block to reduce the feature channels to 256. Use appropriate convolutional kernels to form a highly adaptable receptive field, use grouped convolution to strengthen the understanding of spatial dimension information cross, and adopt an improved residual network to extract features with low coupling degree and strong representativeness, enhancing the network's ability to extract and distinguish target features.

[0050] Step 3, in the feature fusion stage, for the template and the feature maps of the current frame extracted by the Siamese backbone network, use the feature fusion module to strengthen the features during the matching stage of the target and the template, and at the same time to slow down the gradient dispersion problem caused by the deepening of the network. At the same time, use temporal background context information to strengthen the connection between spatio-temporal target feature information to complete the design of the feature fusion module. The feature fusion module is divided into temporal feature fusion and spatial feature fusion. In temporal feature fusion, the feature map output from the previous frame is used, and the feature map enhanced by self-attention is input into the cross-attention module respectively for attention matching between the template feature and the target feature. In spatial feature fusion, the template and the current frame features are directly input into the Transformer structure for feature matching. Finally, the outputs obtained from the two modules are input into spatio-temporal feature fusion. As Figure 3 , and finally the output is obtained.

[0051] Specifically: compress and process the target template and the feature maps of the current frame extracted by the Siamese backbone network into one-dimensional representations, and then obtain q and k through positional embedding encoding; the one-dimensional representation information without processing is v, and input it into the feature fusion module to strengthen the features during the matching stage of the target and the template;

[0052] The feature fusion module is divided into temporal feature fusion and spatial feature fusion, specifically as follows:

[0053] Input the previous frame target template feature map and the current frame target template information into the first Transformer structure. Among them, the current frame target template information is input into the decoder module of the Transformer, and then the k and v of the previous frame target template features are input for feature matching;

[0054] Input the current search frame feature map and the previous search frame feature map information into the second Transformer structure. Among them, the current search frame feature map is input into the decoder module of the Transformer, and then the k and v of the previous search frame feature map information are input for feature matching;

[0055] Input the current search frame feature map and the current frame target template information into the third Transformer structure. Among them, the current frame target template information is input into the encoder module of the Transformer, and the current search frame feature map is input into the decoder module for feature matching;

[0056] Input the information fused by the first Transformer structure and the information fused by the second Transformer structure into the first multi-head cross-attention module for fusion, then input it into the first sum regularization module for fusion with the information fused by the second Transformer structure, and input it into the first feed-forward neural network for fusion; The output of the first feed-forward neural network and the output of the first sum regularization module are simultaneously input into the second sum regularization module for fusion; The output of the second sum regularization module and the output of the third Transformer structure are simultaneously input into the second multi-head cross-attention module for fusion, and then input into the third sum regularization module for fusion with the output of the third Transformer structure, and input into the second feed-forward neural network for fusion; The output of the second feed-forward neural network and the output of the third sum regularization module are simultaneously input into the fourth sum regularization module for fusion to obtain the output of the feature fusion information;

[0057] Step 4: Classify and regress the result obtained in Step 3 to respectively determine the position and size of the target. At the same time, synchronously optimize the classification and regression to generate an output with consistent precision, thus completing the overall framework design of the network. (Firstly, in the target tracking process, classification is to determine the central position of the target, and regression is to determine the size of the target, which is also a part of the network. After obtaining the response map, it is divided into a classification branch and a regression branch. The main manifestation is that the classification branch outputs a feature map of w*h*2, where w and h are the length and width of the feature map respectively, that is, each point contains a 2D vector representing the foreground and background corresponding position scores for classification. The regression branch outputs a feature map of w*h*4, and each point contains a four-dimensional vector representing the distances of the four corners to achieve the determination of the target).

[0058] Step 5: Conduct end-to-end training on the network framework on the publicly available dataset, and at the same time complete the fine-tuning work on one's own dataset.

[0059] Step 6: In the online learning stage, estimate the target motion information based on the target positions of several frames before the current frame. In the case of no occlusion, use the target matching result to determine the target position, and at the same time correct the estimation of the target motion and update the target template. In the case of occlusion, use the target motion information to estimate the target position, maintain the target search area, do not update the template, predict the next motion trajectory of the target under short-term occlusion, and when it is judged that the target is no longer occluded, continue to re-track the target relying on the template. This ensures that the approximate position of the target can be determined without global search in the case of occlusion, guarantees the real-time performance of the model, and realizes the target tracking method from the perspective of a low-altitude aircraft.

[0060] Generally speaking, the present invention is a target tracking method from the perspective of a low-altitude aircraft, which guarantees the accuracy and precision of aerial target tracking. When the target is short-term occluded, estimate the target trajectory after occlusion based on the target position before occlusion, and at the same time, it can also be used as auxiliary information guidance in the target tracking process to prevent target drift and interference from similar targets. Obtain target information with low coupling degree through feature extraction, and obtain an accurate target tracking result from the perspective of a low-altitude aircraft through the matching of the target and the template.

Claims

1. A target tracking method from the perspective of a low-altitude aircraft, characterized in that The steps are as follows: Step 1: Use a drone to collect data in the real scene, simulate the target tracking data from the perspective of a low-altitude aircraft in the virtual scene, and finally annotate the data on darklabel to establish a real-virtual aerial remote sensing dataset; Step 2: Modify the total stride of the last two convolutional blocks of the Siamese network to be reduced to 8, and add an additional 1×1 convolutional layer to each convolutional block output to reduce the feature channels to 256; Use the existing general target tracking dataset as the training data for the network model. The input of the improved Siamese network is the RGB images of the target template and the current search frame, and the feature maps of the target template and the search frame are obtained; Step 3: Compress and process the feature maps of the target template and the current frame extracted by the Siamese backbone network into one-dimensional representations, and then obtain q and k through positional embedding encoding; the one-dimensional representation information without processing is v, and the input feature fusion module strengthens the features during the matching stage of the target and the template; The feature fusion module is divided into temporal feature fusion and spatial feature fusion, specifically: Input the feature map of the previous frame target template and the current frame target template information into the first Transformer structure. Among them, the current frame target template information is input into the decoder module of the Transformer, and then the k and v of the previous frame target template features are input for feature matching; Input the feature map of the current search frame and the information of the previous search frame feature map into the second Transformer structure. Among them, the feature map of the current search frame is input into the decoder module of the Transformer, and then the k and v of the previous search frame feature map information are input for feature matching; Input the feature map of the current search frame and the information of the current frame target template into the third Transformer structure. Among them, the current frame target template information is input into the encoder module of the Transformer, and the feature map of the current search frame is input into the decoder module for feature matching; Input the information fused by the first Transformer structure and the information fused by the second Transformer structure into the first multi-head cross-attention module for fusion, then input the information fused by the second Transformer structure into the first sum regularization module for fusion, and input it into the first feed-forward neural network for fusion; the output of the first feed-forward neural network and the output of the first sum regularization module are simultaneously input into the second sum regularization module for fusion; The output of the second sum regularization module and the output of the third Transformer structure are simultaneously input into the second multi-head cross-attention module for fusion, then input into the third sum regularization module for fusion with the output of the third Transformer structure, and input into the second feed-forward neural network for fusion; the output of the second feed-forward neural network and the output of the third sum regularization module are simultaneously input into the fourth sum regularization module for fusion to obtain the output of the feature fusion information; Step 4: Input the result obtained in Step 3 into the classification network and the regression network. Obtain the central position of the target through the classification network and the size of the target through the regression network; Step 5: Use the publicly available dataset to perform end-to-end training on the network framework of Steps 2 to 4; Step 6: In the online learning stage, estimate the target motion information through the target positions of several frames before the current frame to implement the target tracking method from the perspective of a low-altitude aircraft.

2. The target tracking method from the perspective of a low-altitude aircraft according to claim 1, characterized in that: The classification network in Step 4 outputs a feature map of w*h*2, where w and h are the length and width of the feature map respectively, that is, each point contains a 2D vector of foreground and background corresponding position scores for classification.

3. The target tracking method from the perspective of a low-altitude aircraft according to claim 1, characterized in that: The regression network in Step 4 outputs a feature map of w*h*4, and each point contains a four-dimensional vector which are the distances of the four corners respectively to determine the target.

4. The target tracking method from the perspective of a low-altitude aircraft according to claim 1, wherein: After the end-to-end training in Step 5 is completed, fine-tune the obtained network framework on the data we collected.

5. The target tracking method from the perspective of a low-altitude aircraft according to claim 1, characterized in that: The target tracking method from the perspective of a low-altitude aircraft according to Claim 1, wherein in the target tracking of Step 6, in the case of non-occlusion, use the target matching result to determine the target position, and at the same time correct the estimation of the target motion and update the target template.

6. The target tracking method from the perspective of a low-altitude aircraft according to claim 1, characterized in that: In the target tracking of Step 6, in the case of occlusion, use the target motion information to estimate the target position, maintain the target search area, do not update the template, predict the motion trajectory of the target under short-term occlusion next, and when it is determined that it is no longer occluded, continue to re-track the target relying on the template.

Citation Information

Patent Citations

  • Twin network structure target tracking method fusing target re-identification

    CN113963032A

  • Robust online learning ship tracking method based on twin network

    CN115272405A