An air target tracking method, device, medium and product

Through the self-attention mechanism neural network model combines visual features and historical motion information, it predicts the future motion trajectory of aerial targets, solving the problem of insufficient accuracy in aerial target tracking and achieving higher tracking accuracy and robustness.

CN119648750BActive Publication Date: 2025-07-04YUNNAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510179528.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-07-04
Estimated Expiration
2045-02-19

AI Technical Summary

Technical Problem

Existing aerial target tracking methods are difficult to maintain high accuracy when facing fast movement, occlusion and complex environments, especially when repositioning after targets are lost.

Method used

The self-attention mechanism neural network model is adopted, combining the visual feature information of the target and historical motion information, and adjust the search area by predicting the future motion trajectory of the target to improve tracking accuracy.

Benefits of technology

By predicting the future motion trajectory of the target, reducing the search range, improving the accuracy and robustness of aerial target tracking, overcoming the problem of severe changes in the movement of aerial targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119648750B_ABST
    Figure CN119648750B_ABST
Patent Text Reader

Abstract

The present invention discloses an air target tracking method, device, medium and product, relating to the technical field of target tracking. The method includes: obtaining a video sequence of a tracking target from a preset historical moment to the current moment; extracting initial feature information; obtaining the position information of the tracking target; centering on the position information of the tracking target, cropping the tracking frame image to obtain a search area; extracting the visual feature information of the search area; determining the historical motion information and the visual feature information as the first input information; determining the initial feature information and the visual feature information as the second input information; inputting the first input information and the second input information into a self-attention mechanism neural network model to obtain the tracking frame position information of the tracking target; traversing all the tracking frame images to obtain the position information of the tracking target at the next moment. The present invention can improve the accuracy of air target tracking and solve the tracking problem in the case where the air target is occluded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target tracking, and particularly to an air target tracking method, device, medium and product. Background Art

[0002] In this digital age, unmanned aerial vehicles (UAVs) have become powerful tools for urban management and services. UAV technology has provided new opportunities and possibilities for the construction and operation of smart cities. However, the large-scale commercial and civilian use of UAVs will inevitably bring problems such as illegal flights and private flights of UAVs, and it is necessary to efficiently supervise and effectively manage UAVs, and adopt innovative technical and regulatory means to prevent the abuse of UAVs.

[0003] Therefore, research on the precise tracking of air targets represented by UAVs is carried out to provide support for formulating comprehensive and effective UAV supervision policies, and at the same time promote the innovation of related technologies to balance the reasonable application of UAV technology and the needs of social security.

[0004] Currently, most mainstream trackers adopt a standard detection and tracking framework, and independently detect each captured frame to achieve target tracking. Among these trackers, the discriminative correlation filter (DCF)-based tracker is widely used in air platforms due to its high processing efficiency and low demand for computing resources. However, when the target has rapid movement, appearance changes, occlusion, etc., the performance of these DCF-based trackers will be greatly reduced.

[0005] In recent years, the Siamese neural network-based tracker has demonstrated excellent performance in multiple competitions, and its efficiency is also amazing, and it has been widely used. The Siamese neural network-based tracker believes that the target will not have a large displacement between adjacent frames, and the position of the target in the current frame is within the local neighborhood of the center point of the previous frame. Therefore, the object is detected and tracked based on the locality principle. Only when the target is lost due to occlusion, violent movement, camera shake, etc., will it switch to the global search method to detect and re-associate the target. This alternating use strategy of local and global search is widely adopted because it can effectively save computing resources and improve tracking efficiency, meeting the requirements of high-performance tracking in complex scenarios.

[0006] However, when the tracking background shifts to the air, due to the more complex movement trajectories and faster moving speeds of airborne targets, and the fact that they are not restricted by driving rules such as "lanes", the tracking difficulty increases significantly. In addition, possible occluders such as clouds, trees, and buildings, as well as adverse factors such as light changes and camera jitter, all make the tracking of airborne targets particularly difficult, especially in the problem of repositioning after the target is lost. The global search method that performs well in ground tracking tasks may fail in the air due to sudden changes in the movement direction and speed of the target, and the target is likely to escape from the global search range, resulting in tracking failure. Summary of the Invention

[0007] The object of the present invention is to provide an airborne target tracking method, device, medium and product, which can improve the accuracy of tracking airborne targets.

[0008] To achieve the above object, the present invention provides the following solutions:

[0009] An airborne target tracking method, the method includes:

[0010] Obtain a video sequence of a tracking target from a preset historical moment to the current moment; the video sequence includes multiple frames of images arranged in chronological order; the multiple frames of images include a first frame image and multiple tracking frame images starting from the second frame image;

[0011] Extract the target feature information of the first frame image as the initial feature information;

[0012] Obtain the position information of the tracking target in the first frame image as the position information of the tracking target at the 0th iteration;

[0013] Let the tracking frame serial number i = 1;

[0014] Centering on the position information of the tracking target at the (i - 1)th iteration, crop the ith tracking frame image to obtain the ith search area;

[0015] Extract the visual feature information of the ith search area;

[0016] Determine the historical motion information of the tracking target according to the historical frame images before the ith tracking frame image; the historical motion information includes historical running speed, historical motion acceleration, and historical motion direction;

[0017] Determine the historical motion information and the visual feature information as the first input information at the ith iteration;

[0018] Determine the initial feature information and the visual feature information as the second input information at the ith iteration;

[0019] Input the first input information at the i-th iteration and the second input information at the i-th iteration into the self-attention mechanism neural network model to obtain the (i + 1)-th tracking frame position information of the tracking target, and use the (i + 1)-th tracking frame position information of the tracking target as the center of the position information of the tracking target at the i-th iteration;

[0020] Increment the tracking frame serial number by 1 and return to the step "Crop the i-th tracking frame image centered on the position information of the tracking target at the (i - 1)-th iteration to obtain the i-th search area" until all tracking frame images are traversed to obtain the position information of the tracking target at the next moment.

[0021] Optionally, use the backbone network of Transformer to extract the target feature information of the first frame image and the visual feature information of the i-th search area.

[0022] Optionally, the self-attention mechanism neural network model includes a first interaction module, a second interaction module, and a prediction module;

[0023] The first input information at the i-th iteration is used as the input of the first interaction module; the output of the first interaction module is connected to the input of the prediction module; the output of the first interaction module is the interaction result of motion information and local visual features;

[0024] The second input information at the i-th iteration is used as the input of the second interaction module; the output of the second interaction module is connected to the input of the prediction module; the output of the second interaction module is the interaction result of local visual features and template visual features;

[0025] The output of the prediction module is the position information of the tracking target at the next moment.

[0026] Optionally, the first interaction module includes a splicing module, a first two-dimensional convolutional module, a multi-head self-attention module, and a multi-head cross self-attention module;

[0027] The motion information of the tracking target is used as the input of the splicing module; the output of the splicing module is connected to the input of the multi-head self-attention module; the output of the multi-head self-attention module is connected to the input of the multi-head cross self-attention module; the visual feature information of the i-th search area is used as the input of the first two-dimensional convolutional module; the output of the first two-dimensional convolutional module is connected to the input of the multi-head cross self-attention module; the output of the multi-head cross self-attention module is the interaction result of motion information and local visual features.

[0028] Optionally, the second interaction module includes a second two-dimensional convolutional module, a third two-dimensional convolutional module, a first masked multi-head attention module, a first residual and normalization module, a second masked multi-head attention module, a second residual and normalization module, a first feed-forward neural network module, and a third residual and normalization module;

[0029] The target feature information of the first frame is used as the input of the third two-dimensional convolutional module; the output of the third two-dimensional convolutional module is respectively connected to the input of the first masked multi-head attention module and the input of the first residual and normalization module; the output of the first masked multi-head attention module is connected to the input of the first residual and normalization module; the output of the first residual and normalization module is respectively connected to the input of the second masked multi-head attention module and the input of the second residual and normalization module; the output of the second masked multi-head attention module is connected to the input of the second residual and normalization module; the output of the second residual and normalization module is respectively connected to the input of the first feed-forward neural network module and the input of the third residual and normalization module; the output of the first feed-forward neural network module is connected to the input of the third residual and normalization module; the output of the third residual and normalization module is the interaction result of the local visual feature and the template visual feature; the visual feature information of the i-th search area is used as the input of the second two-dimensional convolutional module; the output of the second two-dimensional convolutional module is connected to the input of the second masked multi-head attention module.

[0030] Optionally, the prediction module includes a third masked multi-head attention module, a fourth residual and normalization module, a fourth masked multi-head attention module, a fifth residual and normalization module, a second feed-forward neural network module, a sixth residual and normalization module, a fully-connected layer, and a multi-layer perceptron;

[0031] The motion information and the interaction result of the local visual feature are used as the input of the fourth masked multi-head attention module;

[0032] The interaction results of the local visual features and the template visual features are respectively used as the inputs of the third masked multi-head attention module and the fourth residual and normalization module; the output of the third masked multi-head attention module is used as the input of the fourth residual and normalization module; the output of the fourth residual and normalization module is respectively used as the input of the fourth masked multi-head attention module and the input of the fifth residual and normalization module; the output of the fourth masked multi-head attention module is connected to the input of the fifth residual and normalization module; the output of the fifth residual and normalization module is connected to the input of the second feed-forward neural network module; the output of the second feed-forward neural network module is connected to the input of the sixth residual and normalization module; the output of the sixth residual and normalization module is connected to the input of the fully-connected layer; the output of the fully-connected layer is connected to the input of the multi-layer perceptron; the output of the multi-layer perceptron is the position information of the tracking target at the next moment.

[0033] Optionally, determining the historical motion information of the tracking target according to the historical frame images before the i-th tracking frame image specifically includes:

[0034] Obtain the historical frame images before the i-th tracking frame image; when i is greater than the preset number of frames of the historical frame images, the historical frame images are arranged in chronological order and the (i - 1)-th tracking frame image is the last frame image; the number of the historical frame images is the preset number of frames; when i is less than or equal to the preset number of frames, the (i - 1)-th tracking frame image is used as the historical frame image;

[0035] When i is greater than the preset number of frames of the historical frame images, determine the historical motion information of the tracking target according to the position information of the tracking target in each frame image of the historical frame images;

[0036] When i is less than or equal to the preset number of frames, determine the historical motion information of the tracking target according to the position information of the (i - 1)-th tracking frame image.

[0037] A computer device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the air target tracking method described in any one of the above.

[0038] A computer-readable storage medium, on which a computer program is stored, wherein the computer program, when executed by a processor, implements the steps of the air target tracking method described in any one of the above.

[0039] A computer program product includes a computer program, wherein the computer program, when executed by a processor, implements the steps of the air target tracking method described in any one of the above.

[0040] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0041] An air target tracking method, device, medium and product provided by the present invention designs a complete air target tracking framework. By fully utilizing the visual feature information of the target, the historical motion information of the target and the search background information, the future motion trajectory of the object is predicted. Since air targets represented by unmanned aerial vehicles usually have high mobility, the concept of a motion factor is introduced, so that the prediction result can better fit objects with different motion speeds and has better robustness. In addition, the prediction result is used as an auxiliary factor in the search process, and the search range is further narrowed by using the prediction result. Compared with a single local principle, the search area cropped according to the prediction result not only has a smaller range, but also can overcome the problem of drastic changes in the motion of air targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 It is a structure diagram of the interaction between motion information and visual information provided in Embodiment 1 of the present invention; Figure 1 Part (a) in it is a schematic structural diagram of the first interaction module; Figure 1 Part (b) in it is a schematic structural diagram of the second interaction module;

[0044] Figure 2 It is a schematic structural diagram of the prediction module;

[0045] Figure 3 It is a schematic diagram of the overall framework of the system of the present invention;

[0046] Figure 4 It is a flowchart of the air target tracking method of the present invention;

[0047] Figure 5 It is an internal structure diagram of a computer device. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0049] The purpose of the present invention is to provide an air target tracking method, device, medium and product, which can improve the accuracy of air target tracking.

[0050] The present invention improves the accuracy of air target tracking by adding a trajectory prediction method during the target tracking process. The idea of adding prediction in the target tracking task has existed for a long time, but most of the prediction methods add a predictor based on the Kalman filter after the tracker, and predict the possible future movement trajectory of the object by mathematically modeling the movement information of the object. This method is effective, but the rich and easily obtained visual feature information of the tracker is ignored, including the appearance information of the object and the environmental information that plays a key role in the movement trajectory of the object. The present invention designs a complete air target tracking framework, and predicts the future movement trajectory of the object by making full use of the visual feature information of the target, the historical movement information of the target and the search background information. At the same time, considering the problem of computational efficiency and the timeliness of movement information, the historical observations are reduced from a long time series to three frames, and the target trajectory is predicted only based on the effective information captured in the past three frames.

[0051] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0052] Embodiment 1

[0053] As Figures 1 to 4 shown, the air target tracking method in this embodiment includes:

[0054] Step S1: Obtain a video sequence of a tracking target from a preset historical moment to the current moment; the video sequence includes multiple frames of images arranged in chronological order; the multiple frames of images include a first frame of image and multiple tracking frame images starting from the second frame of image.

[0055] Step S2: Extract the target feature information of the first frame of image as the initial feature information.

[0056] Specifically, the backbone network of Transformer is used to extract the target feature information of the first frame of image.

[0057] Step S3: Obtain the position information of the tracking target in the first frame image as the position information of the tracking target at the 0th iteration.

[0058] Step S4: Let the tracking frame sequence number i = 1.

[0059] Step S5: Center on the position information of the tracking target at the (i - 1)th iteration, and crop the ith tracking frame image to obtain the ith search area.

[0060] Step S6: Extract the visual feature information of the ith search area.

[0061] Specifically, use the backbone network of Transformer to extract the visual feature information of the ith search area.

[0062] Step S7: Determine the historical motion information of the tracking target according to the historical frame images before the ith tracking frame image; the historical motion information includes historical running speed, historical motion acceleration, and historical motion direction.

[0063] S7 specifically includes:

[0064] Step S71: Obtain the historical frame images before the ith tracking frame image; when i is greater than the preset number of frames of the historical frame images, the historical frame images are arranged in chronological order and the (i - 1)th tracking frame image is the last frame image; the number of the historical frame images is the preset number of frames; when i is less than or equal to the preset number of frames, use the (i - 1)th tracking frame image as the historical frame image.

[0065] Step S72: When i is greater than the preset number of frames of the historical frame images, determine the historical motion information of the tracking target according to the position information of the tracking target in each frame image of the historical frame images.

[0066] Step S73: When i is less than or equal to the preset number of frames, determine the historical motion information of the tracking target according to the position information of the (i - 1)th tracking frame image.

[0067] Step S8: Determine the historical motion information and the visual feature information as the first input information at the ith iteration.

[0068] Step S9: Determine the initial feature information and the visual feature information as the second input information at the ith iteration.

[0069] Step S10: Input the first input information at the i-th iteration and the second input information at the i-th iteration into the self-attention mechanism neural network model to obtain the (i + 1)-th tracking frame position information of the tracking target, and use the (i + 1)-th tracking frame position information of the tracking target as the center of the position information of the tracking target at the i-th iteration.

[0070] Specifically, the self-attention mechanism neural network model includes a first interaction module, a second interaction module, and a prediction module.

[0071] The first input information at the i-th iteration is used as the input of the first interaction module; the output of the first interaction module is connected to the input of the prediction module; the output of the first interaction module is the interaction result of motion information and local visual features.

[0072] The second input information at the i-th iteration is used as the input of the second interaction module; the output of the second interaction module is connected to the input of the prediction module; the output of the second interaction module is the interaction result of local visual features and template visual features.

[0073] The output of the prediction module is the position information of the tracking target at the next moment.

[0074] Among them, the first interaction module includes a splicing module, a first two-dimensional convolutional module, a multi-head self-attention module, and a multi-head cross self-attention module.

[0075] The motion information of the tracking target is used as the input of the splicing module; the output of the splicing module is connected to the input of the multi-head self-attention module; the output of the multi-head self-attention module is connected to the input of the multi-head cross self-attention module; the visual feature information of the i-th search area is used as the input of the first two-dimensional convolutional module; the output of the first two-dimensional convolutional module is connected to the input of the multi-head cross self-attention module; the output of the multi-head cross self-attention module is the interaction result of motion information and local visual features.

[0076] Among them, the second interaction module includes a second two-dimensional convolutional module, a third two-dimensional convolutional module, a first masked multi-head attention module, a first residual and normalization module, a second masked multi-head attention module, a second residual and normalization module, a first feed-forward neural network module, and a third residual and normalization module.

[0077] The target feature information of the first frame is used as the input of the third two-dimensional convolutional module; the output of the third two-dimensional convolutional module is respectively connected to the input of the first masked multi-head attention module and the input of the first residual and normalization module; the output of the first masked multi-head attention module is connected to the input of the first residual and normalization module; the output of the first residual and normalization module is respectively connected to the input of the second masked multi-head attention module and the input of the second residual and normalization module; the output of the second masked multi-head attention module is connected to the input of the second residual and normalization module; the output of the second residual and normalization module is respectively connected to the input of the first feed-forward neural network module and the input of the third residual and normalization module; the output of the first feed-forward neural network module is connected to the input of the third residual and normalization module; the output of the third residual and normalization module is the interaction result of the local visual feature and the template visual feature; the visual feature information of the i-th search area is used as the input of the second two-dimensional convolutional module; the output of the second two-dimensional convolutional module is connected to the input of the second masked multi-head attention module.

[0078] Wherein, the prediction module includes a third masked multi-head attention module, a fourth residual and normalization module, a fourth masked multi-head attention module, a fifth residual and normalization module, a second feed-forward neural network module, a sixth residual and normalization module, a fully connected layer, and a multi-layer perceptron.

[0079] The interaction result of the motion information and the local visual feature is used as the input of the fourth masked multi-head attention module.

[0080] The interaction result of the local visual feature and the template visual feature is respectively used as the input of the third masked multi-head attention module and the fourth residual and normalization module; the output of the third masked multi-head attention module is used as the input of the fourth residual and normalization module; the output of the fourth residual and normalization module is respectively used as the input of the fourth masked multi-head attention module and the input of the fifth residual and normalization module; the output of the fourth masked multi-head attention module is connected to the input of the fifth residual and normalization module; the output of the fifth residual and normalization module is connected to the input of the second feed-forward neural network module; the output of the second feed-forward neural network module is connected to the input of the sixth residual and normalization module; the output of the sixth residual and normalization module is connected to the input of the fully connected layer; the output of the fully connected layer is connected to the input of the multi-layer perceptron; the output of the multi-layer perceptron is the position information of the tracking target at the next moment.

[0081] Step S11: Increment the tracking frame sequence number by 1 and return to step S5 until all tracking frame images are traversed to obtain the position information of the tracking target at the next moment.

[0082] As a specific implementation, N is 3. Centering on the position information of the previous moment, crop the first tracking frame to obtain the first search area. Extract the target feature information of the first frame and the visual feature information of the first search area. Input the first input information and the second input information into the self-attention mechanism neural network model to obtain the position information of the tracking target in the first tracking frame; the first input information is the motion information of the tracking target and the visual feature information of the first search area; the second input information is the target feature information of the first frame and the visual feature information of the first search area.

[0083] Centering on the position information of the first tracking frame, crop the second tracking frame to obtain the second search area. Extract the target feature information of the first frame and the visual feature information of the second search area. Input the third input information and the fourth input information into the self-attention mechanism neural network model to obtain the position information of the tracking target in the second tracking frame; the third input information is the motion information of the tracking target and the visual feature information of the second search area; the fourth input information is the target feature information of the first frame and the visual feature information of the second search area.

[0084] Centering on the position information of the second tracking frame, crop the third tracking frame to obtain the third search area; extract the target feature information of the first frame and the visual feature information of the third search area. Input the fifth input information and the sixth input information into the self-attention mechanism neural network model to obtain the position information of the tracking target in the third tracking frame; the fifth input information is the motion information of the tracking target and the visual feature information of the third search area; the sixth input information is the target feature information of the first frame and the visual feature information of the third search area.

[0085] Take the position information of the tracking target in the third tracking frame as the position information of the tracking target at the next moment.

[0086] In the present invention, the first two-dimensional convolution, the second two-dimensional convolution, and the third two-dimensional convolution are two-dimensional convolutions with the same structure. The first masked multi-head attention module, the second masked multi-head attention module, the third masked multi-head attention module, and the fourth masked multi-head attention module are all masked multi-head attention modules with the same structure. The first feed-forward neural network and the second feed-forward neural network are both feed-forward neural networks with consistent structures. The first residual and normalization module, the second residual and normalization module, the third residual and normalization module, the fourth residual and normalization module, the fifth residual and normalization module, and the sixth residual and normalization module are all residual and normalization modules with the same structure.

[0087] The visual feature information of the i-th search region is the predicted search region feature of the corresponding tracking frame. The target feature information of the first frame is the template feature.

[0088] In practical applications, based on the airborne target tracking method disclosed in the present invention, in order to achieve the same function, an airborne target tracking system is provided, including: the backbone of the Transformer, a tracking module, a motion module, and a prediction module.

[0089] In the prior art, predictors based on filters are not learnable, so neither the existing large-scale datasets nor the large amount of visual feature information obtained during the detection process is utilized. Different from the prior art, the tracking framework MPAT provided by the present invention integrates historical motion information and visual feature information, and realizes more powerful and accurate tracking and prediction in an end-to-end manner.

[0090] Specifically, the video sequence captured by the camera is sent into a Transformer-based backbone network. By extracting features frame by frame from the images in the video sequence, the visual feature information of the video sequence is obtained, including the target feature information Z (Template Feature) in the first frame and the visual feature information X (Search Feature) in the subsequent tracking frames. Each frame of the image contains the object to be tracked and other background information. In the first frame of the image, the object to be tracked is specified (by giving the coordinates of the object), so the visual feature information in the first frame is called the target feature information. For the other frames after the first frame, the coordinates of the object to be tracked are not given, and it is necessary to find the object specified in the first frame. Therefore, the frames after the first frame are collectively called tracking frames. The target feature information is the feature information of the object to be tracked extracted from the first frame, which is very specific and only contains the feature information of the object to be tracked. For the visual feature information of the tracking frames, since the specific position of the object to be tracked in the tracking frame is unknown, we cannot, like in the first frame, only extract the feature information of the target, but extract the visual feature information of a range.

[0091] The tracking module (Tracker Model) crops the tracking frame according to the predicted position information output by the prediction module , and tracks the target according to the target feature information Z (Template Feature) obtained from the first frame and the predicted search feature PSF (Prediction Search Feature) centered on the predicted coordinates, that is, classifies and regresses the target according to the predicted search feature. Classification means identifying the target, such as identifying a drone, and regression means positioning the target.

[0092] Specifically, assuming the size of a frame of the image is 512×512, the prediction module will output the position coordinates (x, y) of the target being tracked at the predicted time. The tracking module will crop the image (512×512) with (x, y) as the center point according to the predicted position coordinates (x, y) to obtain a slice (equivalent to cutting out a part of an image). The size of this cut-out part (may be 224×224, anyway smaller than the original 512×512). Feature extraction is performed on this slice, and the tracking target is searched in this slice. What is obtained after the feature extraction of this slice is the PSF. Since it is cropped according to the predicted position coordinates, it is called the Prediction Search Feature.

[0093] The Motion Model extracts the historical motion information of the target based on the past K frames (in the present invention, K is default set to 3), including the motion speed, motion acceleration, motion direction, etc. of the target. Considering the high mobility of aerial targets, the timeliness of their motion information within a continuous time period is very short. Therefore, the historical observation time is shortened from a long time series to 3 frames.

[0094] The Prediction Model interacts the motion information obtained from the Motion Model with the visual feature information of the current frame to learn the influence of the local environment on the target's motion trend; it interacts the template features obtained from the Backbone with the visual information of the current frame to cope with the possible pose changes of aerial targets. Both are used as inputs to predict the motion trajectory of the target at future moments. Among them, the first frame is special because the first frame defines the tracking target. All frames other than this are called tracking frames. The current frame is the frame at the current moment in the tracking frames. Since the first frame defines the tracking target, only the features of the tracking target are extracted in the first frame, and this feature is used as a template.

[0095] Currently, mainstream trackers generally believe that the target being tracked will not undergo drastic displacement in adjacent frames of the video. Therefore, during the tracking process, the tracking frame is usually cropped using the positioning principle, and only the local neighborhood of the target center point in the previous frame is searched and tracked. Therefore, during the tracking process, the tracking frame is usually cropped using the locality principle, and only the local neighborhood of the target center point in the previous frame is searched and tracked. However, the high mobility and flexibility of aerial targets, as well as the complexity of the aerial environment, make it easy for these targets to be occluded, lose sight, or directly escape from the local neighborhood, etc. In such cases, the local search strategy may cause the tracker to completely fail.

[0096] Based on this, in order to reduce the impact of target loss, especially when the target is severely occluded or out of the field of view. The present invention readjusts the cropping of the search area according to the output result of the Predictor (i.e., the Prediction Module). In the present invention, the observed visual feature information and the historical motion information of the target are comprehensively utilized to jointly predict the position coordinates of the target at time point t. Given the "observation deformation" that may be caused by the change in the flight angle of aerial targets, the prediction result in this invention is based on the center point of the target's location.

[0097] Formally, the output position center point coordinates are formulated as:

[0098] R = (X1, Y1), (X2, Y2), …, (X T , YT ).

[0099] Among them, T is the length of the prediction time range, and (X i , Y i ) respectively represent the coordinates of the target in the x-direction and y-direction at the i-th moment.

[0100] According to the coordinates of the target center point (X i , Y i ) at the predicted moment output by the Prediction Module, and the initialized template size (INSTANCE_SIZE), a search area is re-cropped for local detection and tracking. The initialized template size is a hyperparameter that can be specified according to the size of the target to be tracked.

[0101] (1)

[0102] The local detection centered on the prediction result successfully overcomes the challenges brought by the complex motion trajectory of aerial targets, thus significantly improving the target tracking accuracy in case of occlusion.

[0103] The historical motion state of the target is an important factor for predicting the future trajectory of the target. However, aerial targets represented by unmanned aerial vehicles mostly have high mobility and flexibility, and can quickly change their own motion states, such as flight speed, flight angle, flight height, etc. within a short period of time. Therefore, the present invention believes that the timeliness of the motion information of aerial targets is very short. Based on this, in the present invention, an attempt is made to reduce the observed historical video frames from a long time series to three frames. Observe the motion state of the target from these three frames, including speed , acceleration , motion direction , etc. As Figure 1 shown, these motion information can all be easily calculated through the following formulas.

[0104] The calculation formula for the running speed of the tracked target is:

[0105] (2)

[0106] Among them, is the running speed of the tracked target; is the displacement of the tracked target on the X-axis within the time interval of k frames; is the displacement of the tracked target on the Y-axis within the time interval of k frames; is the frame rate.

[0107] The calculation formula for the motion acceleration of the tracked target is:

[0108] (3)

[0109] Wherein, is the motion acceleration of the tracked target; is the frame rate; represents the average motion speed of the tracked target from i - k + 1 to the intermediate time j; represents

[0110] the motion direction The calculation formula is:

[0111]

[0112] Wherein, atan2() represents the function for calculating the azimuth angle.

[0113] If i < k, that is, there are not enough historical frames to obtain historical motion information, the position information of the target in the last frame can be used for filling.

[0114] Specifically, k represents the kth historical frame. i < k means that there are not enough k historical frames in the front. Then the ith frame is the last frame. For example, in the present invention, it is set to predict the position of the next frame based on the motion information and visual information of the previous three frames. Then when there are only two frames of historical information, not enough for three frames, the second historical frame is the last frame, and the second frame is used for filling to make up three frames.

[0115] Further, the last frame filled with the position information of the target in the last frame is the last frame among the obtained k historical frames. For example, k = 1, that is, there is only one frame of historical frame, but the present invention requires three historical frames, then the position information of this one frame of historical frame is repeated three times.

[0116] The present invention proposes a prediction module (PredictionModule) based on end - to - end design and integrating multi - modal data. In addition to the historical motion state of the target, just as the influence of the ground lanes, obstacles, etc. on the motion trajectory of the vehicle, the motion trajectory of the aerial target will also be affected by the aerial environmental factors. Especially the local environment around the target directly affects the action trajectory and motion state of the target. Therefore, in the present invention, after splicing and encoding and sorting out the motion information of the target, it interacts with the visual feature information (PSF) obtained after the Backbone feature extraction of the prediction search area to learn the influence of local environmental factors on the motion trajectory of the target. The interaction between the two can be expressed formulaically as:

[0117] (4)

[0118] (5)

[0119] Among them, concatenation is cat(m i , …, m i-k ), encoding is Encoder(cat(m i , …, m i-k )); MHCA and MHSA respectively represent the multi-head cross-attention module (Multi-head Cross-Attention) and the multi-head self-attention module (Multi-Head Self-Attention). In the present invention, first, the historical motion information extracted from the Motion Module is subjected to Cat and Encoder operations. After concatenation, the historical motion information is mapped into a high-dimensional space to obtain richer features of the data and reduce the sparsity of the data. The Multi-Head Self-Attention (MHSA) helps the prediction model capture long-range dependencies when processing sequential motion information, consider the correlation between the motion information of an object at different times, and allows the prediction model to simultaneously focus on the motion information of different parts and learn the weights between them. The Multi-Head Cross-Attention (MHCA) establishes a global connection between the motion information and the visual feature information of the search area, improving the fusion effect of the interaction and the generalization ability of the model. At the same time, considering the possible pose changes of the tracking target during the motion process, as well as the influence of factors such as illumination and camera jitter, in the research, the local search features of the current frame are interacted with the template feature Z given from the initial frame to assist in prediction. To improve efficiency and avoid network complexity, the visual feature information directly extracted by the backbone network is used: Template Feature (TF) and Prediction Search Feature (PSF). The implementation details are as shown in part (b) of Figure 1 , and the specific interaction formula can be expressed as:

[0120] (6)

[0121] Among them, TS is the result of visual feature interaction; TF is the template feature.

[0122] The interaction result of the historical motion information and the local search feature , together with the visual feature interaction result and the prediction step t, are jointly used as the input of the Prediction Module. The structure of the prediction module is as shown in Figure 2As shown in the figure, the Prediction Module consists of N layers, and each layer obtains two types of interactions: Motion-PSF and PSF-TF. The prediction module consists of N stacked Motion-PSF and TF-PSF interactions. The overall structure is constructed using a standard Transformer Decoder. The final output is processed by an MLP to obtain the prediction result.

[0123] It is difficult to fit the direct prediction of the trajectory coordinates of the target at future moments, and it is difficult to achieve the expected goal in different test sets and real environments. To solve this problem, the concept of a motion factor is introduced. Different from directly predicting the future trajectory coordinates of an aerial target, the output of the predictor P is defined as the relative offset value of the position coordinates based on the average moving speed of the target in the past K frames. The specific formula is as follows:

[0124] (7)

[0125] (8)

[0126] Among them, represents the average motion speed of the target obtained based on the historical motion information of K frames; represents the trajectory coordinates obtained when the target is moving at a constant speed at the prediction time (i + t); represents element-wise addition. The output of the prediction module (Prediction Module) is added element-wise to to obtain the final prediction result.

[0127] By introducing the concept of a motion factor, the present invention enables the predictor to have good generalization ability and performs well in tracking objects with different motion speeds. In addition, the prediction module (Prediction Module) predicts the multi-modal future motion of the tracking target, that is, the motion trajectory of the target within the next X frames. This prediction method generates a single forward-propagated target trajectory coordinate in the frame, greatly saving the computational cost of aligning the entire scene with the target coordinates.

[0128] The present invention has the following advantages:

[0129] (1)By deeply interacting and modeling visual features and target motion information, a simple, concise, and efficient aerial target tracker is proposed.

[0130] (2)By introducing the concept of a motion factor in the present invention, the tracker of the present invention can better fit objects with different motion speeds.

[0131] (3) According to the prediction results, the present invention further corrects the local search ability in the tracking stage, reduces the inference time, and improves the accuracy.

[0132] (4) The prediction module proposed by the present invention can be easily applied to different trackers, plug-and-play, and the effect is significantly improved.

[0133] Embodiment 2

[0134] A computer device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the steps of the air target tracking method in Embodiment 1.

[0135] Embodiment 3

[0136] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the air target tracking method in Embodiment 1 are implemented.

[0137] Embodiment 4

[0138] A computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the air target tracking method in Embodiment 1 are implemented.

[0139] Embodiment 5

[0140] A computer device, which can be a database, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store transactions to be processed. The input / output interface of the computer device is used for the processor to exchange information with external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the air target tracking method in Embodiment 1 is implemented.

[0141] It should be noted that the object information (including but not limited to object device information, object personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present invention are all information and data authorized by the object or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions.

[0142] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided by the present invention can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided by the present invention can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0143] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0144] In this text, specific examples are used to illustrate the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. An air target tracking method, characterized in that, The method includes: Obtaining a video sequence of a tracking target from a preset historical moment to the current moment; the video sequence includes multiple frames of images arranged in chronological order; the multiple frames of images include a first frame image and multiple tracking frame images starting from the second frame image; Extracting the target feature information of the first frame image as the initial feature information; Obtaining the position information of the tracking target in the first frame image as the position information of the tracking target at the 0th iteration; Let the tracking frame serial number i = 1; Centering on the position information of the tracking target at the (i - 1)th iteration, cropping the i-th tracking frame image to obtain the i-th search region; Extracting the visual feature information of the i-th search region; Determining the historical motion information of the tracking target according to the historical frame images before the i-th tracking frame image; the historical motion information includes historical running speed, historical motion acceleration, and historical motion direction; Determining the historical motion information and the visual feature information as the first input information at the i-th iteration; Determining the initial feature information and the visual feature information as the second input information at the i-th iteration; Inputting the first input information at the i-th iteration and the second input information at the i-th iteration into the self-attention mechanism neural network model to obtain the (i + 1)-th tracking frame position information of the tracking target, and using the (i + 1)-th tracking frame position information of the tracking target as the center at the i-th iteration; Increasing the tracking frame serial number by 1 and returning to the step "Centering on the position information of the tracking target at the (i - 1)th iteration, cropping the i-th tracking frame image to obtain the i-th search region" until all tracking frame images are traversed to obtain the position information of the tracking target at the next moment; The self-attention mechanism neural network model includes a first interaction module, a second interaction module, and a prediction module; The first input information at the i-th iteration is used as the input of the first interaction module; the output of the first interaction module is connected to the input of the prediction module; the output of the first interaction module is the interaction result of motion information and local visual features; The second input information at the i-th iteration is used as the input of the second interaction module; the output of the second interaction module is connected to the input of the prediction module; the output of the second interaction module is the interaction result of local visual features and template visual features; The output of the prediction module is the position information of the tracking target at the next moment; The first interaction module includes a splicing module, a first two-dimensional convolutional module, a multi-head self-attention module, and a multi-head cross self-attention module; The motion information of the tracked target is used as the input of the stitching module; the output of the stitching module is connected to the input of the multi-head self-attention module; the output of the multi-head self-attention module is connected to the input of the multi-head cross-attention module; the visual feature information of the i-th search region is used as the input of the first two-dimensional convolutional module; the output of the first two-dimensional convolutional module is connected to the input of the multi-head cross-attention module; the output of the multi-head cross-attention module is the interaction result of the motion information and the local visual features.

2. The air target tracking method according to claim 1, wherein Apply the backbone network of Transformer to extract the target feature information of the first frame image and the visual feature information of the i-th search region.

3. The air target tracking method according to claim 1, characterized in that The second interaction module includes a second two-dimensional convolutional module, a third two-dimensional convolutional module, a first masked multi-head attention module, a first residual and normalization module, a second masked multi-head attention module, a second residual and normalization module, a first feed-forward neural network module, and a third residual and normalization module; The target feature information of the first frame is used as the input of the third two-dimensional convolutional module; the output of the third two-dimensional convolutional module is respectively connected to the input of the first masked multi-head attention module and the input of the first residual and normalization module; the output of the first masked multi-head attention module is connected to the input of the first residual and normalization module; the output of the first residual and normalization module is respectively connected to the input of the second masked multi-head attention module and the input of the second residual and normalization module; the output of the second masked multi-head attention module is connected to the input of the second residual and normalization module; the output of the second residual and normalization module is respectively connected to the input of the first feed-forward neural network module and the input of the third residual and normalization module; the output of the first feed-forward neural network module is connected to the input of the third residual and normalization module; the output of the third residual and normalization module is the interaction result of the local visual features and the template visual features; the visual feature information of the i-th search region is used as the input of the second two-dimensional convolutional module; the output of the second two-dimensional convolutional module is connected to the input of the second masked multi-head attention module.

4. The air target tracking method according to claim 1, wherein The prediction module includes a third masked multi-head attention module, a fourth residual and normalization module, a fourth masked multi-head attention module, a fifth residual and normalization module, a second feed-forward neural network module, a sixth residual and normalization module, a fully-connected layer, and a multi-layer perceptron; The interaction result of the motion information and the local visual features is used as the input of the fourth masked multi-head attention module; The interaction results of the local visual features and the template visual features are respectively used as the inputs of the third masked multi-head attention module and the fourth residual and normalization module; the output of the third masked multi-head attention module is used as the input of the fourth residual and normalization module; the output of the fourth residual and normalization module is respectively used as the input of the fourth masked multi-head attention module and the input of the fifth residual and normalization module; the output of the fourth masked multi-head attention module is connected to the input of the fifth residual and normalization module; the output of the fifth residual and normalization module is connected to the input of the second feed-forward neural network module; the output of the second feed-forward neural network module is connected to the input of the sixth residual and normalization module; the output of the sixth residual and normalization module is connected to the input of the fully-connected layer; the output of the fully-connected layer is connected to the input of the multi-layer perceptron; the output of the multi-layer perceptron is the position information of the tracking target at the next moment.

5. The air target tracking method according to claim 1, characterized in that Determining the historical motion information of the tracking target according to the historical frame images before the i-th tracking frame image specifically includes: Obtaining the historical frame images before the i-th tracking frame image; when i is greater than the preset number of frames of the historical frame images, the historical frame images are arranged in chronological order and the (i - 1)-th tracking frame image is the last frame image; the number of the historical frame images is the preset number of frames; when i is less than or equal to the preset number of frames, the (i - 1)-th tracking frame image is used as the historical frame image; When i is greater than the preset number of frames of the historical frame images, determining the historical motion information of the tracking target according to the position information of the tracking target in each frame image of the historical frame images; When i is less than or equal to the preset number of frames, determining the historical motion information of the tracking target according to the position information of the (i - 1)-th tracking frame image.

6. A computer device, comprising: A memory and a processor with a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the air target tracking method according to any one of claims 1-5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the air target tracking method according to any one of claims 1-5.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the air target tracking method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Autoregressive visual target tracking algorithm based on token fusion

    CN119477974A