Target tracking method and device based on satellite video, equipment and storage medium

By performing spatiotemporal feature fusion in satellite video and using calibration network to process spatiotemporal fusion features, the accuracy problem of the target tracking algorithm in satellite video under background motion interference is solved, and higher tracking accuracy is achieved.

CN120259374AActive Publication Date: 2025-07-04北京观微科技有限公司
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510743301.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

The target tracking algorithm in satellite videos is insufficient in accuracy under background motion interference, and the prior art is difficult to effectively filter noise responses, resulting in inaccurate tracking.

Method used

By performing space-time fusion processing on the spatial and temporal features in satellite videos, and calibrating the spatial and fusion features using calibration networks, filtering the noise response caused by background motion and improving tracking accuracy.

Benefits of technology

It improves the accuracy of single target tracking in satellite videos, can filter background motion interference more effectively, and obtain more accurate target tracking results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259374A_ABST
    Figure CN120259374A_ABST
Patent Text Reader

Abstract

The invention provides a target tracking method, device and equipment based on a satellite video and a storage medium, and is applied to the technical field of satellite target tracking, and the method comprises the steps: carrying out the spatial feature extraction processing of a template image and a search image corresponding to a target in the satellite video, and determining spatial features; obtaining a first autoregressive query information sequence corresponding to the target in the current frame image, and performing decoding processing in combination with the spatial features to determine time features corresponding to the target; performing spatio-temporal information fusion processing on the spatial feature and the time feature to determine a first spatio-temporal fusion feature; carrying out calibration processing on the first space-time fusion feature by adopting a calibration network, and determining a second space-time fusion feature; determining a corresponding tracking result of the target in the current frame image according to the second space-time fusion feature; the tracking result comprises corresponding position information and scale information of the target in the current frame image. By adopting the technical scheme of the invention, the accuracy of tracking the single target in the satellite can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of satellite target tracking, and in particular, to a target tracking method, device, equipment and storage medium based on satellite video. Background Art

[0002] Satellite videos often have large-scale background movements, such as the movements of clouds, ground traffic, etc. These background movements may cause misdetection when the target tracking algorithm processes the target, thus affecting the subsequent tracking effect. Therefore, it is necessary to accurately detect and track the target.

[0003] In related technologies, when performing target tracking, most often the current coordinates are predicted through the historical coordinates of the target to achieve the tracking of the target.

[0004] However, the above technologies still have certain limitations in the target tracking of satellite videos. Summary of the Invention

[0005] The present invention provides a target tracking method, device, equipment and storage medium based on satellite video, which is used to solve the defect that there are still certain limitations in the target tracking of satellite videos in the prior art, and realizes the calibration processing of the spatio-temporal fusion features after spatio-temporal fusion processing of the spatial features and temporal features in the satellite video, and determines the tracking result of the target in the current frame image through the calibrated spatio-temporal fusion features, so as to filter the noise response caused by the background disturbance of satellite movement and improve the accuracy of single-target tracking in satellite videos.

[0006] The present invention provides a target tracking method based on satellite video, including: Obtain a template image corresponding to a target in a satellite video and a search image corresponding to the target in the current frame image, and input the template image and the search image into a spatial encoder for spatial feature extraction processing to determine a joint spatial feature corresponding to the target and a spatial feature corresponding to the target in the search image; Obtain a first autoregressive query information sequence corresponding to the target in the current frame image, and input the first autoregressive query information sequence and the joint spatial feature into a temporal decoder for decoding processing to determine a target query information corresponding to the target in the current frame image and a temporal feature corresponding to the target; the above target query information is used to characterize the state of the target in the current frame image, and the first autoregressive query information sequence is determined according to a historical query information sequence corresponding to the target in the current frame image and a query information to be learned corresponding to the target in the current frame image; Perform spatio-temporal information fusion processing on the spatial feature and the temporal feature to determine a first spatio-temporal fusion feature; The first spatio-temporal fusion feature is calibrated by a calibration network to determine the second spatio-temporal fusion feature; the calibration network is pre-trained according to a plurality of sample first spatio-temporal fusion features and corresponding sample second spatio-temporal fusion features; The tracking result corresponding to the target in the current frame image is determined according to the second spatio-temporal fusion feature; the tracking result includes the position information and scale information corresponding to the target in the current frame image.

[0007] According to a target tracking method based on satellite video provided by the present invention, the calibration network includes: a first convolutional layer, a non-linear activation function layer, and a second convolutional layer. The process of calibrating the first spatio-temporal fusion feature by using the calibration network to determine the second spatio-temporal fusion feature includes: The first spatio-temporal fusion feature is input into the first convolutional layer for feature extraction processing to determine the first intermediate feature corresponding to the target; the first intermediate feature includes features of the target in multiple scales and multiple directions; The first intermediate feature is input into the non-linear activation function layer to filter the noise in the first intermediate feature to determine the second intermediate feature corresponding to the target; The second intermediate feature is input into the second convolutional layer to extract local calibration information to determine the second spatio-temporal fusion feature.

[0008] According to a target tracking method based on satellite video provided by the present invention, the first convolutional layer is a deformable convolutional layer.

[0009] According to a target tracking method based on satellite video provided by the present invention, the second convolutional layer is a depthwise separable convolutional layer. The depthwise separable convolutional layer includes a depth convolutional layer and a pointwise convolutional layer. The process of inputting the second intermediate feature into the second convolutional layer to extract local calibration information to determine the second spatio-temporal fusion feature includes: The second intermediate feature is input into the depth convolutional layer for local feature extraction processing to determine the local feature corresponding to the target; The local feature is input into the pointwise convolutional layer for feature fusion processing to determine the second spatio-temporal fusion feature.

[0010] According to a target tracking method based on satellite video provided by the present invention, the historical query information sequence includes at least one historical query information, and the method further includes: The target query information corresponding to the target in the current frame image is added to the end of the historical query information sequence to obtain a new historical query information sequence; The historical query information that is farthest from the current time in the new historical query information sequence is removed according to a preset sequence length to obtain the historical query information sequence corresponding to the target in the next frame image; Determine the second autoregressive query information sequence corresponding to the target in the next-frame image based on the historical query information sequence corresponding to the target in the next-frame image and the query information to be learned corresponding to the target in the next-frame image.

[0011] According to a target tracking method based on satellite video provided by the present invention, the above-mentioned determining the tracking result corresponding to the target in the current-frame image according to the second spatio-temporal fusion feature includes: Perform element-wise product fusion processing on the second spatio-temporal fusion feature and the spatial feature to determine the enhanced feature corresponding to the target in the current-frame image; Use the prediction head to perform target detection processing on the enhanced feature to determine the tracking result corresponding to the target in the current-frame image.

[0012] The present invention also provides a target tracking device based on satellite video, including the following modules: A spatial encoding module, configured to obtain a template image corresponding to a target in a satellite video and a search image corresponding to the target in the current-frame image, and input the template image and the search image into a spatial encoder for spatial feature extraction processing to determine the joint spatial feature corresponding to the target and the spatial feature corresponding to the target in the search image; A temporal decoding module, configured to obtain the first autoregressive query information sequence corresponding to the target in the current-frame image, and input the first autoregressive query information sequence and the joint spatial feature into a temporal decoder for decoding processing to determine the target query information corresponding to the target in the current-frame image and the temporal feature corresponding to the target; the above-mentioned target query information is used to characterize the state of the target in the current-frame image, and the above-mentioned first autoregressive query information sequence is determined according to the historical query information sequence corresponding to the target in the current-frame image and the query information to be learned in the current-frame image; A spatio-temporal fusion module, configured to perform spatio-temporal information fusion processing on the spatial feature and the temporal feature to determine the first spatio-temporal fusion feature; A calibration module, configured to perform calibration processing on the first spatio-temporal fusion feature by using a calibration network to determine the second spatio-temporal fusion feature; the above-mentioned calibration network is pre-trained according to a plurality of sample first spatio-temporal fusion features and corresponding sample second spatio-temporal fusion features; A tracking result determination module, configured to determine the tracking result corresponding to the target in the current-frame image according to the second spatio-temporal fusion feature; the above-mentioned tracking result includes the position information and scale information corresponding to the target in the current-frame image.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, it implements the target tracking method based on satellite video as described in any one of the above.

[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the target tracking method based on satellite video as described in any one of the above.

[0015] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the target tracking method based on satellite video as described in any one of the above.

[0016] The target tracking method, device, equipment and storage medium based on satellite video provided by the present invention obtain a template image corresponding to a target in a satellite video and a search image corresponding to the target in the current frame image, and input the template image and the search image into a spatial encoder for spatial feature extraction processing to determine the joint spatial feature corresponding to the target and the spatial feature corresponding to the target in the search image. Then, a first autoregressive query information sequence corresponding to the target in the current frame image is obtained, and the first autoregressive query information sequence and the joint spatial feature are input into a temporal decoder for decoding processing to determine the target query information corresponding to the target in the current frame image and the temporal feature corresponding to the target. Then, spatio-temporal information fusion processing is performed on the spatial feature and the temporal feature to determine a first spatio-temporal fusion feature, and a calibration network is used to perform calibration processing on the first spatio-temporal fusion feature to determine a second spatio-temporal fusion feature. Then, a tracking result including position information and scale information of the target in the current frame image is determined according to the second spatio-temporal fusion feature; wherein, the above target query information is used to characterize the state of the target in the current frame image, the first autoregressive query information sequence is determined according to the historical query information sequence corresponding to the target in the current frame image and the query information to be learned corresponding to the target in the current frame image, and the calibration network is pre-trained according to a plurality of sample first spatio-temporal fusion features and corresponding sample second spatio-temporal fusion features. In this method, since the spatio-temporal fusion feature of the target in the satellite video can be calibrated by the calibration network, noise responses caused by satellite background motion disturbances can be filtered, so that more accurate spatio-temporal fusion features can be obtained. Furthermore, the satellite target can be tracked by the more accurate spatio-temporal fusion features, and a more accurate satellite target tracking result can be obtained, that is, the accuracy of satellite target tracking can be improved. Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1It is one of the flow schematic diagrams of the object tracking method based on satellite video provided by the present invention.

[0019] Figure 2 It is the model architecture block diagram of the AQA-Track algorithm.

[0020] Figure 3 It is the architecture block diagram of the object tracking model provided by the present invention.

[0021] Figure 4 It is the second flow schematic diagram of the object tracking method based on satellite video provided by the present invention.

[0022] Figure 5 It is the third flow schematic diagram of the object tracking method based on satellite video provided by the present invention.

[0023] Figure 6 It is the structural schematic diagram of the object tracking device based on satellite video provided by the present invention.

[0024] Figure 7 It is the structural schematic diagram of the electronic device provided by the present invention. Detailed implementation manners

[0025] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.

[0026] In satellite video tracking, since satellite videos often have large-scale background motion (such as clouds and ground traffic), the algorithm may be misdetected due to background interference. Especially when the texture of the target is similar to that of the background, the target also faces irregular motion. This poses a huge challenge to single-object tracking in satellite video scenarios. Most trackers perform well by introducing backbone networks with strong spatial feature extraction capabilities in detection tasks or natural language processing (NLP) tasks. Luca Bertinetto et al.'s Siamese Fully Convolutional Network (SiamFC) designed a Siamese network framework, using AlexNet as the backbone network to extract features from the template image and the search image, achieving good performance in terms of speed and accuracy. In recent years, Transformer-based algorithms introduced into object recognition have shown amazing global modeling capabilities. Therefore, visual object tracking has also started to use Transformer. Initially, some trackers used the attention mechanism or Transformer for feature extraction or fusion during tracking. For example, Chen et al. proposed TransT, designing two modules for feature interaction based on the attention mechanism. Bin Yan et al. proposed STARK, using the Transformer structure as a fusion module. Later, due to the excellent performance of Transformer, Cui et al. used Transformer as the backbone network. In addition, some researchers proposed trackers based on the full Transformer, combining feature extraction and fusion, greatly improving the performance. Spatiotemporal information is crucial for the model to capture the target state changes and motion trends. Therefore, many mainstream studies have explored spatiotemporal information in visual object tracking. A commonly used method is to update the appearance representation. Song et al. adopted a dynamic template to update the target appearance to capture changes, making the matching between the template image and the search image more accurate. Lin et al. focused on learning a feature that describes the target's previous state or motion information. Another common method to explore spatiotemporal information is to integrate historical appearance information. Fu et al. proposed a Spatiotemporal Memory Network STMTrack to make full use of historical information. Zhang et al. proposed UpdateNet, using a set of historical appearance information to estimate the optimal template for the next frame. However, most of these methods require some artificial rule designs, and spatiotemporal information is far from being fully exploited. Recently, Wei et al. proposed an autoregressive tracker ARTrack, predicting the current coordinates autoregressively based on historical coordinates. The above methods have significantly improved the accuracy and robustness of general video object tracking, but there are still certain limitations for satellite video object tracking.

[0027] The multi-object tracking paradigm for satellite videos is basically the same as that for natural videos, mainly including the detection-based tracking paradigm, the joint detection and tracking paradigm, and the multi-object tracking paradigm based on the attention mechanism. The following is an explanation of each: Detection-based tracking paradigm: This is a traditional paradigm that decomposes the tracking task into two steps: detection and data association. First, a target detection algorithm is used to identify targets in video frames, and then the similarity between targets, such as the intersection over union (IoU) or appearance features, is calculated to associate targets in consecutive frames. Representative algorithms of this method include SORT and DeepSORT, which perform target matching and tracking through techniques such as Kalman filtering and the Hungarian algorithm.

[0028] Joint detection and tracking paradigm: This paradigm attempts to integrate target detection and tracking into a unified framework to improve the robustness and accuracy of the system. For example, the JDE (Joint Detection and Embedding) algorithm uses a feature pyramid network to detect targets at different scales and a shared network for embedding learning and matching.

[0029] Multi-object tracking paradigm based on the attention mechanism: With the rise of deep learning, especially Transformer models, tracking methods based on the attention mechanism have begun to receive attention. These methods use the self-attention mechanism of Transformer to capture long-range dependencies between frames, such as algorithms like TransTrack and TrackFormer.

[0030] From the above description, it can be seen that due to the relatively small size of satellite targets in satellite videos, as well as characteristics such as complex background interference and irregular movement of satellite targets, there are still certain limitations in using the above technologies for satellite target tracking, resulting in inaccurate tracking of satellite targets. Based on this, the embodiments of the present invention provide a target tracking method, device, equipment, and storage medium based on satellite videos, which can solve the above technical problems.

[0031] It should be noted that the execution subject of the embodiments of the present invention can be a target tracking device based on satellite videos, or an electronic device including a target tracking device based on satellite videos, or other devices, equipment, or systems, etc. There is no specific limitation here. The following embodiments will take an electronic device as the execution subject for illustration. The electronic device here can be a terminal or a server, and there is no specific limitation on the specific form of the terminal or the server.

[0032] Figure 1 is one of the flow diagrams of the target tracking method based on satellite videos provided by the present invention. As Figure 1 shown, the method includes the following steps: Step 102: Obtain the template image corresponding to the target in the satellite video and the search image corresponding to the target in the current frame image, and input the template image and the search image into the spatial encoder for spatial feature extraction processing to determine the joint spatial feature corresponding to the target and the spatial feature corresponding to the target in the search image.

[0033] In this step, the target can be imaged by an image acquisition device (such as a camera, a video camera, etc.) in the satellite to obtain a satellite video including the target. The satellite video may include one or more satellite images acquired in sequence. The target to be tracked in the satellite video can be a pre-set target, such as an airplane, a vehicle, etc. In this embodiment, mainly a single target in the satellite video is tracked, and the single target here can be, for example, an airplane.

[0034] During the process of tracking the target, the regional image corresponding to the target (i.e., the image of the area where the target is located) can be cropped from the first satellite image or a specific satellite image of the satellite video. For example, the area including the initial appearance of the target can be cropped from a certain frame, and the cropped regional image is used as the template image corresponding to the target. The template image includes the initial state of the target, and can be matched and compared with the search area / search image (i.e., the image corresponding to the search area) corresponding to the target in subsequent frames during the target tracking process to help the target tracking model / algorithm identify the position and appearance changes of the target. The template image can be saved after being obtained for subsequent use.

[0035] For the search image corresponding to the target, after obtaining the satellite video, the current frame image (i.e., the satellite image at the current moment) can be obtained from the satellite video, and then the current frame image is processed by a target tracking algorithm to determine the area where the target may appear in the current frame image, denoted as the search image. The search image can be an area formed by expanding a certain range centered on the position of the target in the previous frame image in the current frame image. The size of the search image can usually be dynamically adjusted according to the size and movement speed of the target to ensure that the target is within the search image.

[0036] Furthermore, when tracking satellite targets, a target tracking algorithm can be used to achieve target tracking. Optionally, the target tracking algorithm in this embodiment selects the AQA-Track algorithm (Autoregressive Queries for Adaptive Tracking with Spatio-Temporal Transformers), which uses simple autoregressive queries to effectively learn spatio-temporal information without many manually designed components. First, a set of learnable autoregressive queries is introduced into the target tracking model corresponding to this target tracking algorithm to capture the instantaneous changes in the target appearance in a sliding window manner; then, a novel attention mechanism is designed for the interaction between existing queries to generate a new query in the current frame; finally, based on the initial target template and the learned autoregressive queries, a spatio-temporal information fusion module (STM) is designed for spatio-temporal information aggregation to locate the target object. Refer to Figure 2 the model architecture block diagram of the AQA-Track algorithm shown. The template image and search image (Template&Search) and the learnable autoregressive queries (Queries) can be input into the target tracking model Model for processing to obtain the tracking result Result, and this operation is performed on the timeline Timeline, and the tracking of the target can be achieved. Thanks to STM, this target tracking algorithm can effectively combine the static appearance and instantaneous changes of the target, thus guiding robust target tracking.

[0037] Based on the target tracking model corresponding to the above AQA-Track algorithm, this embodiment improves the spatio-temporal information fusion module STM in it. Refer to Figure 3 the architecture block diagram of the target tracking model shown. This target tracking model can include a spatial encoder Spatial Encoder, a temporal decoder Temporal Decoder, a spatio-temporal information fusion module STM, and a detection head Head. Among them, the spatio-temporal information fusion module STM can include a spatio-temporal fusion module, a calibration module, and a feature enhancement and fusion module. The calibration module is composed of a calibration network.

[0038] Based on the above architecture of the target tracking model, the following describes the specific target tracking process.

[0039] After obtaining the template image corresponding to the satellite target in the satellite video and the search image corresponding to the satellite target in the current frame image, the template image and the search image can be input into the spatial encoder, where the spatial feature extraction process can be performed on both the template image and the search image to obtain the spatial features corresponding to the target. Optionally, when performing the spatial feature extraction process, step-wise path embedding can be used to extract image features from the template image and the search image. The step-wise path embedding process refers to using multiple downsampling processes to avoid the loss of pixel correlation caused by a single downsampling process, which may affect subsequent tracking performance. Additionally, the step-wise path embedding process can be performed separately on the template image and the search image to obtain the template features corresponding to the template image and the search features corresponding to the search image. Then, the template features and the search features can be concatenated and input into a multi-layer Transformer encoder for encoding to obtain the normalized spatial features. The joint spatial features can be denoted as f zx , for example, it can be expressed as: ; where R is the set of real numbers, and represent the dimensions of the template features and the search features respectively, and D = 512.

[0040] Additionally, after obtaining the search features corresponding to the search image, the search features can be input into a multi-layer Transformer encoder for encoding to obtain the spatial features corresponding to the target in the search image.

[0041] Step 104: Obtain the first autoregressive query information sequence corresponding to the target in the current frame image, and input the first autoregressive query information sequence and the joint spatial features into the temporal decoder for decoding to determine the target query information corresponding to the target in the current frame image and the temporal features corresponding to the target; the target query information is used to represent the state of the target in the current frame image, and the first autoregressive query information sequence is determined based on the historical query information sequence corresponding to the target in the current frame image and the query information to be learned in the current frame image.

[0042] In this step, a sliding window mechanism is adopted to obtain the current temporal features of the satellite target. In this sliding window mechanism, a query information sequence with a length of m is maintained. Each frame of the image has its corresponding query information sequence, which is determined by its corresponding historical query information sequence and the query information to be learned at the current time. This query information sequence is an autoregressive query information sequence (Autoregressive Queries), which is used to capture the instantaneous changes in the appearance of the satellite target during the target tracking process and improve the tracking accuracy.

[0043] For the query information sequence corresponding to the current frame image, it can be denoted as the first autoregressive query information sequence. This first autoregressive query information sequence is composed of the historical query information sequence corresponding to the target in the current frame image (denoted as Q pre ) and the query information to be learned corresponding to the current frame image (denoted as Q cur ). For example, the historical query information sequence and the query information to be learned can be concatenated to obtain the first autoregressive query information sequence (denoted as Q all ). This first autoregressive query information sequence can be expressed as Q all =concat(Q pre ,Q cur ). The sequence length of this first autoregressive query information sequence is m, and m is greater than or equal to 1. Through experiments, it is found that the tracking effect is the best when m = 4. Optionally, m can be selected as 4.

[0044] For the query information to be learned corresponding to the target in the current frame image, it can be determined in the following way: First, extract the features of the target in the current frame image, and then fuse these features through a specific module (such as an attention module, etc.) to determine the query information to be learned corresponding to the target in the current frame image. For the historical query information sequence corresponding to the target in the current frame image, it can be obtained by combining the query information of the historical time before the current time of the target. For example, if m is 4, the length of the historical query information sequence corresponding to the target in the current frame image can be 3, and it can be obtained by combining the query information in the 3 frames of images before the current frame image of the target.

[0045] After determining the first autoregressive query information sequence corresponding to the target in the current frame image as described above, the first autoregressive query information sequence and the joint spatial features corresponding to the target can be input into the temporal decoder for decoding processing to generate the query information corresponding to the current frame image containing temporal information, denoted as the target query information. At the same time, the temporal features after temporal feature extraction of the target can be obtained.

[0046] Optionally, the above time decoder may include multiple decoding layers, and each decoding layer includes a temporal attention layer (TA), a multi-head attention layer (MHA), and a feed-forward network (FFN). Among them, the temporal attention layer uses a customized attention mechanism to take the to-be-learned query information Q corresponding to the target in the current frame image cur as the query Query of the time decoder, and interacts with the key Key and value Value of the first autoregressive query information sequence Q all to generate the target query information corresponding to the current frame image containing temporal information.

[0047] Step 106: Perform spatio-temporal information fusion processing on the spatial feature and the temporal feature to determine the first spatio-temporal fusion feature.

[0048] In this step, after obtaining the spatial feature and the temporal feature corresponding to the target in the search image, the spatial feature is actually a spatial feature matrix, which can be represented by the following formula: ; where represents the spatial feature matrix corresponding to the target in the search image, N x represents the dimension of the spatial feature matrix.

[0049] The temporal feature is also a temporal feature matrix, which can be represented by the following formula: ; where represents the spatial feature matrix corresponding to the target in the search image, and can also be denoted as where T represents time, N temporal represents the dimension of the temporal feature matrix.

[0050] Specifically, after obtaining the spatial feature matrix and the temporal feature matrix corresponding to the target in the search image, the spatio-temporal fusion module in the spatio-temporal information fusion module STM can perform spatio-temporal information fusion processing on these two features to obtain the first spatio-temporal fusion feature corresponding to the target. Optionally, the similarity map corresponding to these two features can be calculated. Specifically, the spatial feature matrix and the temporal feature matrix of the target can be subjected to a dot product calculation to obtain a dot product similarity map, and the dot product similarity map can be used as the first spatio-temporal fusion feature. The dot product similarity map can be represented as S: ; Through the above method, the first spatio-temporal fusion feature after spatio-temporal information fusion of the target can be obtained.

[0051] Step 108: Use a calibration network to perform calibration processing on the first spatio-temporal fusion feature to determine the second spatio-temporal fusion feature. The calibration network is pre-trained based on multiple sample first spatio-temporal fusion features and their corresponding sample second spatio-temporal fusion features.

[0052] In this step, the calibration network can be, for example, a convolutional neural network. The specific network architecture is not specifically limited here. This calibration network is mainly used to calibrate the first spatio-temporal fusion information of the target. More specifically, it performs calibration processing on the dot product similarity map of the target to filter out the noise response caused by satellite background motion perturbation, improve the accuracy of the finally obtained spatio-temporal fusion information, and obtain the final spatio-temporal fusion feature of the target, denoted as the second spatio-temporal fusion feature.

[0053] For the above calibration network, it can be pre-trained. Specifically, it can be trained based on multiple sample first spatio-temporal fusion features and the corresponding sample second spatio-temporal fusion features of each sample, or it can also be trained based on multiple sample first spatio-temporal fusion features and the corresponding labeled positions of each sample. The samples here can be any sample satellite targets. In addition, the above calibration network can be trained independently or jointly with the entire target tracking model.

[0054] Specifically, after obtaining the first spatio-temporal fusion feature of the target and the trained calibration network, the dot product similarity map corresponding to the first spatio-temporal fusion feature can be input into the calibration network. In the calibration network, operations such as feature extraction, activation, and pooling can be performed on the dot product similarity map to obtain the calibrated similarity map, denoted as S. calibrated This calibrated similarity map can be used as the final second spatio-temporal fusion feature of the target.

[0055] Step 110: Determine the tracking result of the target corresponding to the current frame image based on the second spatio-temporal fusion feature. The tracking result includes the position information and scale information of the target corresponding to the current frame image.

[0056] In this step, after obtaining the second spatio-temporal fusion feature of the target, the second spatio-temporal fusion feature can be directly input into the detection head for object detection processing, or it can also be further processed by the feature enhancement and fusion module in the spatio-temporal information fusion module and then input into the detection head for object detection processing, or it can also be that the second spatio-temporal fusion feature and other features are processed together by the feature enhancement and fusion module in the spatio-temporal information fusion module and then input into the detection head for object detection processing. In the detection head, a center point-based prediction network can be used to generate a classification score map and a bounding box offset, and then the bounding box with the highest score in the classification score map can be used as the bounding box corresponding to the object. At the same time, the bounding box can be corrected by combining the bounding box offset to obtain the final bounding box corresponding to the object. The final bounding box corresponding to the object can include the scale information and center point information of the bounding box, where the center point information can be the position information of the object in the current frame image, denoted as (x, y), and the scale information can be the scale information corresponding to the object in the current frame image, which can include information such as the width w, height h, and aspect ratio of the bounding box. Finally, the position information and scale information of the object in the current frame image can be combined as the tracking result corresponding to the object in the current frame image, denoted as (x, y, w, h).

[0057] In this embodiment, by obtaining the template image corresponding to the target in the satellite video and the search image corresponding to the target in the current frame image, and inputting the template image and the search image into the spatial encoder for spatial feature extraction processing to determine the joint spatial feature corresponding to the target and the spatial feature corresponding to the target in the search image, then obtaining the first autoregressive query information sequence corresponding to the target in the current frame image, and inputting the first autoregressive query information sequence and the joint spatial feature into the temporal decoder for decoding processing to determine the target query information corresponding to the target in the current frame image and the temporal feature corresponding to the target, then performing spatio-temporal information fusion processing on the spatial feature and the temporal feature to determine the first spatio-temporal fusion feature, and using a calibration network to perform calibration processing on the first spatio-temporal fusion feature to determine the second spatio-temporal fusion feature, and then determining the tracking result including the position information and the scale information of the target in the current frame image according to the second spatio-temporal fusion feature; wherein, the above target query information is used to characterize the state of the target in the current frame image, the first autoregressive query information sequence is determined according to the historical query information sequence corresponding to the target in the current frame image and the query information to be learned corresponding to the target in the current frame image, and the calibration network is pre-trained according to multiple sample first spatio-temporal fusion features and corresponding sample second spatio-temporal fusion features. In this method, since the spatio-temporal fusion feature of the target in the satellite video can be calibrated by the calibration network, noise responses caused by satellite background motion disturbances can be filtered, so that more accurate spatio-temporal fusion features can be obtained. Furthermore, the satellite target can be tracked by the more accurate spatio-temporal fusion features to obtain a more accurate satellite target tracking result, that is, the accuracy of satellite target tracking can be improved.

[0058] The following embodiments will describe the specific composition and specific calibration process of the above calibration network.

[0059] Figure 4 is the second flowchart of the target tracking method based on satellite video provided by the present invention. As Figure 4 shown, the above step 108 uses a calibration network to perform calibration processing on the first spatio-temporal fusion feature to determine the second spatio-temporal fusion feature, which may include the following steps: Step 202, input the first spatio-temporal fusion feature into the first convolutional layer for feature extraction processing to determine the first intermediate feature corresponding to the target; the above first intermediate feature includes features of the target in multiple scales and multiple directions.

[0060] Among them, the above calibration network may include: a first convolutional layer, a non-linear activation function layer, and a second convolutional layer.

[0061] After obtaining the first spatio-temporal fusion feature, the first spatio-temporal fusion feature can be input into the first convolutional layer in the calibration network, where multi-scale feature extraction processing, multi-directional feature adjustment processing, multi-scale feature fusion, multi-directional feature fusion processing, etc. are performed on the first spatio-temporal fusion feature to obtain the first intermediate feature of the target in multiple scales and multiple directions.

[0062] Further, since satellite targets are relatively small and move irregularly, this will further increase the difficulty of target tracking. To solve this problem, the first convolutional layer is improved in this embodiment. Optionally, the first convolutional layer is a deformable convolutional layer (Deformable Convolution). After the first spatio-temporal fusion feature is input into the deformable convolutional layer, deformable convolution can be applied to the dot product similarity map corresponding to the first spatio-temporal fusion feature to dynamically adjust the sampling position to capture irregular motion patterns and enhance the robustness to deformation and fast motion. In addition, for the number of layers of the deformable convolutional layer, it can be one layer or multiple layers, which can be specifically determined according to the target tracking performance and is not specifically limited here.

[0063] Step 204: Input the first intermediate feature into the non-linear activation function layer to filter the noise in the first intermediate feature and determine the second intermediate feature corresponding to the target.

[0064] In this step, after obtaining the first intermediate feature of the target, the first intermediate feature can be input into the non-linear activation function layer of the calibration network for non-linear activation processing. This non-linear activation processing is mainly used to filter the noise in the dot product similarity map to filter the noise response and obtain the second intermediate feature. Optionally, the non-linear activation function layer can be a ReLU non-linear activation function layer.

[0065] Step 206: Input the second intermediate feature into the second convolutional layer to extract local calibration information and determine the second spatio-temporal fusion feature.

[0066] In this step, after obtaining the second intermediate feature of the target, the second intermediate feature can be input into the second convolutional layer of the calibration network. In the second convolutional layer, local information in the dot product similarity map is calibrated, such as feature extraction processing, feature fusion processing, etc., to obtain local calibration information, and the final similarity map of the target is generated through the local calibration information, that is, the second spatio-temporal fusion feature is obtained.

[0067] Further, in order to reduce the computational complexity of the calibration network to improve the efficiency of calibrating the spatio-temporal fusion features of the target, and thus improve the efficiency of target tracking, the calibration network in this step can be a lightweight calibration network, which is specifically implemented by lightweight processing of the second convolutional layer. For this second convolutional layer, its type or architecture can be different from that of the first convolutional layer. Optionally, the second convolutional layer is a depthwise separable convolutional layer, and the depthwise separable convolutional layer includes a depth convolutional layer and a pointwise convolutional layer. Then, the process of determining the second spatio-temporal fusion feature in this step can include the following steps: Input the second intermediate feature into the depth convolutional layer for local feature extraction processing to determine the local feature corresponding to the target; Input the local feature into the pointwise convolutional layer for feature fusion processing to determine the second spatio-temporal fusion feature.

[0068] Among them, the second intermediate feature can be first input into the depth convolutional layer of the depthwise separable convolutional layer to extract multiple local features of the target, and then each local feature is input into the pointwise convolutional layer of the depthwise separable convolutional layer for feature fusion processing to generate the second spatio-temporal fusion feature. For the number of layers of the depthwise separable convolutional layer, it can be one layer or multiple layers. For the number of layers of the depth convolutional layer and the pointwise convolutional layer in the depthwise separable convolutional layer, both can be one layer or multiple layers, and can be specifically set according to the actual situation.

[0069] In this embodiment, by sequentially inputting the first spatio-temporal fusion feature into the first convolutional layer, the non-linear activation function layer, and the second convolutional layer of the calibration network for calibration processing, a more accurate second spatio-temporal fusion feature is obtained. In this way, by refining the network structure of the calibration network, the accuracy of calibrating the first spatio-temporal fusion feature can be improved, and thus the accuracy of subsequent target tracking can be improved. In addition, the first convolutional layer in the calibration network can be a deformable convolutional layer, which can enable the calibration network and the target tracking model to better adapt to the complex motion patterns of the target, such as irregular motion, and improve the accuracy of satellite target tracking. Further, the second convolutional layer in the calibration network can be a depthwise separable convolutional layer, and the depthwise separable convolutional layer includes a separated depth convolutional layer and a pointwise convolutional layer. By calibrating the spatio-temporal fusion feature of the target through this depthwise separable convolutional layer, the computational complexity of the calibration network can be reduced, and the efficiency of target tracking can be improved.

[0070] In the above embodiment, when the time decoder decodes the spatial feature of the target and the autoregressive query information, the target query information corresponding to the target in the current frame image can be obtained. For this target query information, it can be transmitted to the next frame image and continue to be used for tracking, thereby improving the accuracy of target tracking. The following embodiments will illustrate this process.

[0071] Figure 5 This is the third schematic flowchart of the object tracking method based on satellite video provided by the present invention. As Figure 5 shown, the above method may further include the following steps: Step 302: Add the object query information corresponding to the object in the current frame image to the end of the historical query information sequence to obtain a new historical query information sequence.

[0072] Step 304: Eliminate the historical query information that is farthest from the current time in the new historical query information sequence according to the preset sequence length, to obtain the historical query information sequence corresponding to the object in the next frame image.

[0073] Step 306: Determine the second autoregressive query information sequence corresponding to the object in the next frame image according to the historical query information sequence corresponding to the object in the next frame image and the query information to be learned corresponding to the object in the next frame image.

[0074] Among them, the historical query information sequence corresponding to the object in the current frame image includes at least one historical query information, that is, the historical query information sequence may include one or more historical query information. For example, for the current frame image, the historical query information sequence may include the historical query information corresponding to each of the previous m - 1 frames of the current frame.

[0075] After obtaining the object query information corresponding to the object in the current frame image through the time decoder, the object query information may be added to the end of the historical query information sequence corresponding to the current frame image as a historical query information to obtain a new historical query information sequence, and the sequence length of the new historical query information sequence is m + 1. Then, according to the sequence length m set by the sliding window, the earliest / farthest historical query information in the new historical query information sequence may be eliminated or discarded, so as to maintain a fixed sliding window length. The obtained historical query information sequence after elimination may be used as the historical query information sequence corresponding to the object in the next frame image, that is, the update of the historical query information sequence is realized.

[0076] When performing object tracking processing on the next frame image, the query information to be learned corresponding to the next frame image may be obtained, and then the historical query information sequence corresponding to the object in the next frame image and the query information to be learned corresponding to the next frame image are combined to obtain the autoregressive query information sequence corresponding to the next frame image, denoted as the second autoregressive query information sequence. For subsequent frames of the next frame, the autoregressive query information sequence can be updated in this way, which will not be elaborated here.

[0077] In this embodiment, by adding the target query information corresponding to the target in the current frame image to the end of the historical query information sequence and removing the farthest historical query information to keep the length of the historical query sequence as a fixed length, it is convenient for time decoding processing. At the same time, the historical query information sequence after removing the farthest historical query information can be combined with the query information to be learned corresponding to the next frame image to obtain the autoregressive query information sequence corresponding to the next frame image. In this way, the transmission of query information can be realized, so that the appearance changes of the satellite target over time can be better captured through the historical query information, and the accuracy of satellite target tracking can be improved.

[0078] The following embodiments illustrate an implementation manner of determining the corresponding tracking result of the target in the current frame image through the second spatio-temporal fusion feature of the target.

[0079] In one embodiment, the step 110 of determining the corresponding tracking result of the target in the current frame image according to the second spatio-temporal fusion feature may include the following steps: Perform element-wise product fusion processing on the second spatio-temporal fusion feature and the spatial feature to determine the enhanced feature corresponding to the target in the current frame image; Use the prediction head to perform target detection processing on the enhanced feature to determine the corresponding tracking result of the target in the current frame image.

[0080] Among them, after obtaining the second spatio-temporal fusion feature, that is, the calibrated similarity map, through the calibration network, the calibrated similarity map S calibrated and the spatial feature of the target in the search image of the current frame image are subjected to element-wise product (Element-wise-Product) fusion processing to suppress satellite background interference and enhance the response of the target area, and finally the enhanced feature corresponding to the target in the current frame image is obtained. Then, the enhanced feature can be input into the detection head for target detection processing to obtain the corresponding tracking result of the target in the current frame image.

[0081] In this embodiment, the enhanced feature is obtained by performing element-wise product fusion processing on the second spatio-temporal fusion feature of the target and the spatial feature in the search image, and the target detection processing is performed by inputting the enhanced feature into the detection head to obtain the corresponding tracking result of the target in the current frame image. In this way, satellite background interference can be suppressed and the response of the target area can be enhanced, further improving the accuracy of satellite target tracking.

[0082] The target tracking device based on satellite video provided by the present invention will be described below. The target tracking device based on satellite video described below can be mutually referred to the target tracking method based on satellite video described above.

[0083] Figure 6 This is a schematic structural diagram of the target tracking device based on satellite video provided by the present invention. Refer to Figure 6 As shown, the device may include: A spatial encoding module 410, configured to obtain a template image corresponding to a target in a satellite video and a search image corresponding to the target in the current frame image, and input the template image and the search image into a spatial encoder for spatial feature extraction processing to determine a joint spatial feature corresponding to the target and a spatial feature corresponding to the target in the search image; A temporal decoding module 420, configured to obtain a first autoregressive query information sequence corresponding to the target in the current frame image, and input the first autoregressive query information sequence and the joint spatial feature into a temporal decoder for decoding processing to determine a target query information corresponding to the target in the current frame image and a temporal feature corresponding to the target; the above target query information is used to characterize the state of the target in the current frame image, and the above first autoregressive query information sequence is determined according to a historical query information sequence corresponding to the target in the current frame image and a query information to be learned corresponding to the target in the current frame image; A spatio-temporal fusion module 430, configured to perform spatio-temporal information fusion processing on the spatial feature and the temporal feature to determine a first spatio-temporal fusion feature; A calibration module 440, configured to perform calibration processing on the first spatio-temporal fusion feature by using a calibration network to determine a second spatio-temporal fusion feature; the above calibration network is pre-trained according to a plurality of sample first spatio-temporal fusion features and corresponding sample second spatio-temporal fusion features; A tracking result determination module 450, configured to determine a tracking result corresponding to the target in the current frame image according to the second spatio-temporal fusion feature; the above tracking result includes position information and scale information corresponding to the target in the current frame image.

[0084] In one embodiment, the above calibration network includes: a first convolutional layer, a non-linear activation function layer, and a second convolutional layer. The above calibration module 440 is specifically configured to input the first spatio-temporal fusion feature into the first convolutional layer for feature extraction processing to determine a first intermediate feature corresponding to the target; the above first intermediate feature includes features of the target in multiple scales and multiple directions; input the first intermediate feature into the non-linear activation function layer to filter noise in the first intermediate feature to determine a second intermediate feature corresponding to the target; input the second intermediate feature into the second convolutional layer to extract local calibration information to determine the second spatio-temporal fusion feature.

[0085] Optionally, the above first convolutional layer is a deformable convolutional layer.

[0086] Optionally, the second convolutional layer is a depthwise separable convolutional layer. The depthwise separable convolutional layer includes a depth convolutional layer and a pointwise convolutional layer. The calibration module 440 is specifically configured to input the second intermediate feature into the depth convolutional layer for local feature extraction processing to determine the local feature corresponding to the target; input the local feature into the pointwise convolutional layer for feature fusion processing to determine the second spatio-temporal fusion feature.

[0087] In one embodiment, at least one historical query information is included in the historical query information sequence, and the apparatus further includes: A query information update module, configured to add the target query information corresponding to the target in the current frame image to the end of the historical query information sequence to obtain a new historical query information sequence; remove the historical query information with the farthest distance from the current time in the new historical query information sequence according to a preset sequence length to obtain the historical query information sequence corresponding to the target in the next frame image; determine the second autoregressive query information sequence corresponding to the target in the next frame image according to the historical query information sequence corresponding to the target in the next frame image and the to-be-learned query information corresponding to the target in the next frame image.

[0088] In one embodiment, the tracking result determination module 450 is specifically configured to perform element-wise product fusion processing on the second spatio-temporal fusion feature and the spatial feature to determine the enhanced feature corresponding to the target in the current frame image; use the prediction head to perform target detection processing on the enhanced feature to determine the tracking result corresponding to the target in the current frame image.

[0089] It should be noted here that the apparatus provided in the embodiments of the present invention can implement all the method steps implemented in the above method embodiments and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiments will not be specifically described in this embodiment.

[0090] Figure 7 An example of a schematic physical structure diagram of an electronic device is shown in Figure 7As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 may call logic instructions in the memory 530 to execute a target tracking method based on satellite video. The method includes: obtaining a template image corresponding to a target in the satellite video and a search image corresponding to the target in the current frame image, and inputting the template image and the search image into a spatial encoder for spatial feature extraction processing to determine a joint spatial feature corresponding to the target and a spatial feature corresponding to the target in the search image; obtaining a first autoregressive query information sequence corresponding to the target in the current frame image, and inputting the first autoregressive query information sequence and the joint spatial feature into a temporal decoder for decoding processing to determine a target query information corresponding to the target in the current frame image and a temporal feature corresponding to the target; the above target query information is used to characterize the state of the target in the current frame image, and the above first autoregressive query information sequence is determined according to a historical query information sequence corresponding to the target in the current frame image and a to-be-learned query information corresponding to the target in the current frame image; performing spatio-temporal information fusion processing on the spatial feature and the temporal feature to determine a first spatio-temporal fusion feature; using a calibration network to perform calibration processing on the first spatio-temporal fusion feature to determine a second spatio-temporal fusion feature; the above calibration network is pre-trained according to a plurality of sample first spatio-temporal fusion features and corresponding sample second spatio-temporal fusion features; determining a tracking result corresponding to the target in the current frame image according to the second spatio-temporal fusion feature; the above tracking result includes position information and scale information corresponding to the target in the current frame image.

[0091] In addition, when the logic instructions in the above memory 530 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.

[0092] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the target tracking method based on satellite video provided by each of the above methods. The method includes: obtaining a template image corresponding to a target in a satellite video and a search image corresponding to the target in the current frame image, and inputting the template image and the search image into a spatial encoder for spatial feature extraction processing to determine a joint spatial feature corresponding to the target and a spatial feature corresponding to the target in the search image; obtaining a first autoregressive query information sequence corresponding to the target in the current frame image, and inputting the first autoregressive query information sequence and the joint spatial feature into a temporal decoder for decoding processing to determine a target query information corresponding to the target in the current frame image and a temporal feature corresponding to the target; the above target query information is used to represent the state of the target in the current frame image, and the above first autoregressive query information sequence is determined according to the historical query information sequence corresponding to the target in the current frame image and the to-be-learned query information corresponding to the target in the current frame image; performing spatio-temporal information fusion processing on the spatial feature and the temporal feature to determine a first spatio-temporal fusion feature; using a calibration network to perform calibration processing on the first spatio-temporal fusion feature to determine a second spatio-temporal fusion feature; the above calibration network is pre-trained according to a plurality of sample first spatio-temporal fusion features and corresponding sample second spatio-temporal fusion features; determining a tracking result corresponding to the target in the current frame image according to the second spatio-temporal fusion feature; the above tracking result includes position information and scale information corresponding to the target in the current frame image.

[0093] In another aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a target tracking method based on satellite video provided by the above-mentioned various methods. The method includes: obtaining a template image corresponding to a target in a satellite video and a search image corresponding to the target in the current frame image, and inputting the template image and the search image into a spatial encoder for spatial feature extraction processing to determine a joint spatial feature corresponding to the target and a spatial feature corresponding to the target in the search image; obtaining a first autoregressive query information sequence corresponding to the target in the current frame image, and inputting the first autoregressive query information sequence and the joint spatial feature into a temporal decoder for decoding processing to determine a target query information corresponding to the target in the current frame image and a temporal feature corresponding to the target; the above-mentioned target query information is used to characterize the state of the target in the current frame image, and the above-mentioned first autoregressive query information sequence is determined according to a historical query information sequence corresponding to the target in the current frame image and a query information to be learned corresponding to the target in the current frame image; performing spatio-temporal information fusion processing on the spatial feature and the temporal feature to determine a first spatio-temporal fusion feature; using a calibration network to perform calibration processing on the first spatio-temporal fusion feature to determine a second spatio-temporal fusion feature; the above-mentioned calibration network is pre-trained according to a plurality of sample first spatio-temporal fusion features and corresponding sample second spatio-temporal fusion features; determining a tracking result corresponding to the target in the current frame image according to the second spatio-temporal fusion feature; the above-mentioned tracking result includes position information and scale information corresponding to the target in the current frame image.

[0094] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0095] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A target tracking method based on satellite video, characterized in that, Including: Obtain a template image corresponding to a target in a satellite video and a search image corresponding to the target in the current frame image, and input the template image and the search image into a spatial encoder for spatial feature extraction processing to determine the joint spatial feature corresponding to the target and the spatial feature corresponding to the target in the search image; Obtain a first autoregressive query information sequence corresponding to the target in the current frame image, and input the first autoregressive query information sequence and the joint spatial feature into a temporal decoder for decoding processing to determine the target query information corresponding to the target in the current frame image and the temporal feature corresponding to the target; The target query information is used to represent the state of the target in the current frame image, and the first autoregressive query information sequence is determined according to the historical query information sequence corresponding to the target in the current frame image and the query information to be learned corresponding to the target in the current frame image; Perform spatio-temporal information fusion processing on the spatial feature and the temporal feature to determine a first spatio-temporal fusion feature; Use a calibration network to perform calibration processing on the first spatio-temporal fusion feature to determine a second spatio-temporal fusion feature; the calibration network is pre-trained according to a plurality of sample first spatio-temporal fusion features and corresponding sample second spatio-temporal fusion features; Determine the tracking result corresponding to the target in the current frame image according to the second spatio-temporal fusion feature; the tracking result includes the position information and scale information corresponding to the target in the current frame image.

2. The object tracking method based on satellite video according to claim 1, characterized in that, The calibration network includes: a first convolutional layer, a non-linear activation function layer, and a second convolutional layer. The step of using the calibration network to perform calibration processing on the first spatio-temporal fusion feature to determine the second spatio-temporal fusion feature includes: Input the first spatio-temporal fusion feature into the first convolutional layer for feature extraction processing to determine a first intermediate feature corresponding to the target; the first intermediate feature includes the features of the target in multiple scales and multiple directions; Input the first intermediate feature into the non-linear activation function layer to filter the noise in the first intermediate feature to determine a second intermediate feature corresponding to the target; Input the second intermediate feature into the second convolutional layer to extract local calibration information to determine the second spatio-temporal fusion feature.

3. The object tracking method based on satellite video according to claim 2, characterized in that, The first convolutional layer is a deformable convolutional layer.

4. The object tracking method based on satellite video according to claim 2, wherein, The second convolutional layer is a depthwise separable convolutional layer. The depthwise separable convolutional layer includes a depth convolutional layer and a pointwise convolutional layer. The step of inputting the second intermediate feature into the second convolutional layer to extract local calibration information to determine the second spatio-temporal fusion feature includes: Input the second intermediate feature into the depth convolutional layer for local feature extraction processing to determine a local feature corresponding to the target; Input the local feature into the pointwise convolutional layer for feature fusion processing to determine the second spatio-temporal fusion feature.

5. The object tracking method based on satellite video according to claim 1, wherein, The historical query information sequence includes at least one historical query information, and the method further includes: Add the target query information corresponding to the target in the current frame image to the end of the historical query information sequence to obtain a new historical query information sequence; Eliminate the historical query information with the farthest distance from the current time in the new historical query information sequence according to a preset sequence length to obtain the historical query information sequence corresponding to the target in the next frame image; Determine the second autoregressive query information sequence corresponding to the target in the next frame image according to the historical query information sequence corresponding to the target in the next frame image and the to-be-learned query information corresponding to the target in the next frame image.

6. The method for target tracking based on satellite video according to any one of claims 1 to 5, characterized in that, The determining the tracking result corresponding to the target in the current frame image according to the second spatio-temporal fusion feature includes: Perform an element-wise product fusion process on the second spatio-temporal fusion feature and the spatial feature to determine the enhanced feature corresponding to the target in the current frame image; Use a prediction head to perform target detection processing on the enhanced feature to determine the tracking result corresponding to the target in the current frame image.

7. A target tracking device based on satellite video, characterized in that, including: A spatial encoding module, configured to obtain a template image corresponding to a target in a satellite video and a search image corresponding to the target in the current frame image, and input the template image and the search image into a spatial encoder for spatial feature extraction processing to determine the joint spatial feature corresponding to the target and the spatial feature corresponding to the target in the search image; A temporal decoding module, configured to obtain the first autoregressive query information sequence corresponding to the target in the current frame image, and input the first autoregressive query information sequence and the joint spatial feature into a temporal decoder for decoding processing to determine the target query information corresponding to the target in the current frame image and the temporal feature corresponding to the target; The target query information is used to characterize the state of the target in the current frame image, and the first autoregressive query information sequence is determined according to the historical query information sequence corresponding to the target in the current frame image and the to-be-learned query information corresponding to the target in the current frame image; A spatio-temporal fusion module, configured to perform spatio-temporal information fusion processing on the spatial feature and the temporal feature to determine a first spatio-temporal fusion feature; A calibration module, configured to perform calibration processing on the first spatio-temporal fusion feature by using a calibration network to determine a second spatio-temporal fusion feature; the calibration network is pre-trained according to a plurality of sample first spatio-temporal fusion features and corresponding sample second spatio-temporal fusion features; A tracking result determination module, configured to determine the tracking result corresponding to the target in the current frame image according to the second spatio-temporal fusion feature; the tracking result includes the position information and scale information corresponding to the target in the current frame image.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the satellite video-based target tracking method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the satellite video-based target tracking method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the satellite video-based target tracking method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Target tracking method and system for deep space-time information fusion based on CUDA

    CN108664935A

  • Attention-enhanced space-time Transform visual single-target tracking method

    CN117011342A

  • Satellite video single target tracking method and device

    CN117197192A

  • Target tracking method based on spatial-temporal feature fusion

    CN117423035A

  • Single-target tracking method and system, electronic equipment and medium

    CN119741339A