Satellite video single target tracking method based on appearance and motion feature fusion

By integrating appearance and motion characteristics in satellite video single target tracking and using Transformer encoder to process data, missed or misdetection problems caused by mixing small target features and background features are solved, improving the accuracy and efficiency of target tracking.

CN119919455BActive Publication Date: 2025-06-06PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510406867.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-06
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

In single-target tracking of satellite video, Transformer has too large receptive field of the self-attention mechanism, resulting in the mixing of small target features and background features, resulting in missed or missed detection.

Method used

Using a method based on the fusion of appearance and motion characteristics, tracking sequence data is generated by obtaining template image data, current frame image data, appearance prompt data, confidence mark data, historical position data and trajectory change data, and processing is performed using a preset Transformer encoder to extract the position and motion information of small targets for tracking.

Benefits of technology

It effectively reduces the missed detection or misdetection problems caused by the lack of obvious characteristics of small targets, improves the accuracy and efficiency of target tracking, and alleviates the problem of missing or drifting of target tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919455B_ABST
    Figure CN119919455B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of image data processing and analysis technology, and relates to a satellite video single target tracking method based on the fusion of appearance and motion features, which solves the problem of missed detection or false detection in small target tracking. The present invention includes: generating tracking sequence data of the current frame based on template image data, current frame image data, appearance prompt data, confidence mark data, historical position data and trajectory change data; the encoder tracks and processes the tracking sequence data to obtain the current frame position data, current frame appearance data and current frame confidence data of the single target; and tracks and processes the next frame image data based on the current frame position data, current frame appearance data, current frame confidence data and the encoder. Based on the historical position data and trajectory change data, the present invention uses the encoder to accurately extract the position information and motion information of the small target and perform target tracking, thereby improving the accuracy of target tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the rapid development of satellite remote sensing technology, satellite video target tracking has become an important research direction in remote sensing applications. Satellite video target tracking is conducive to expanding remote sensing applications and actual needs. It can be applied to scenes that are inaccessible in traditional target tracking and make up for the defects of traditional target tracking. Compared with traditional natural video target tracking, satellite video target tracking faces unique challenges, such as small target size, weak features, and complex background. In order to meet these challenges, researchers have proposed a variety of improved methods, such as correlation filtering-based methods, deep learning-based methods, multimodal fusion-based methods, reinforcement learning-based methods, and Transformer-based methods.

[0003] The Transformer-based satellite video single target tracking method has been a research hotspot in the field of computer vision in recent years. Transformer was originally used for natural language processing (NLP). Due to its powerful global modeling capabilities, it has gradually been introduced into visual tasks, including target tracking. The core of Transformer is the self-attention mechanism, which can capture the global dependencies between elements in a sequence. In the target tracking task, Transformer models the relationship between the target and the background through the self-attention mechanism, thereby improving the robustness of target tracking and thus improving the accuracy of target tracking.

[0004] However, the Transformer's self-attention mechanism will mix the features of small targets with those of the surrounding background because of its relatively large receptive field, resulting in obvious technical problems of missed detection or false detection during small target tracking. Summary of the invention

[0005] In order to solve the above problems in the prior art, the present invention provides a satellite video single target tracking method based on appearance and motion feature fusion, comprising:

[0006] Obtain template image data, current frame image data, appearance prompt data, confidence mark data, historical position data and trajectory change data;

[0007] Wherein, the image corresponding to the current frame image data contains a single target to be tracked;

[0008] Generate tracking sequence data of the current frame based on the template image data, the current frame image data, the appearance prompt data, the confidence mark data, the historical position data and the trajectory change data;

[0009] The tracking sequence data of the current frame is tracked and processed by a preset Transformer encoder to obtain the current frame position data, current frame appearance data and current frame confidence data of the single target corresponding to the current frame image data;

[0010] The next frame of image data is tracked based on the current frame position data, the current frame appearance data, the current frame confidence data and the Transformer encoder.

[0011] Further, the generating of the tracking sequence data of the current frame based on the template image data, the current frame image data, the appearance prompt data, the confidence mark data, the historical position data and the trajectory change data includes:

[0012] Performing cropping and reshaping processing on the template image data and the current frame image data in sequence, respectively, to obtain corresponding template image sequence data and current frame image sequence data;

[0013] Discretizing the historical position data and the trajectory change data respectively to obtain discrete historical position data and discrete trajectory change data;

[0014] Performing linear mapping processing on the template image sequence data, the current frame image sequence data, the appearance prompt data, the confidence mark data, the discrete historical position data and the discrete trajectory change data respectively to obtain one-dimensional sequence data composed of the processing units of the Transformer encoder;

[0015] A learnable position code is embedded in the one-dimensional sequence data to obtain tracking sequence data of the current frame.

[0016] Furthermore, the step of performing cropping and reshaping on the template image data and the current frame image data in sequence comprises:

[0017] Based on the image size data of the single target, the template image data and the current frame image data are respectively cropped to obtain respective corresponding image size data;

[0018] The two image size data are respectively divided into two-dimensional image blocks and the two-dimensional image blocks are reshaped to obtain template image sequence data and current frame image sequence data corresponding to the template image data and the current frame image data respectively.

[0019] Furthermore, the pixel size of the current frame image data is greater than the pixel size of the template image data.

[0020] Furthermore, the tracking process of the tracking sequence data of the current frame by using a preset Transformer encoder includes:

[0021] The Transformer encoder reshapes the tracking sequence data of the current frame to obtain a feature map of the current frame and generates a reference grid according to the feature map;

[0022] The Transformer encoder generates an offset according to a query vector corresponding to the tracking sequence data of the current frame, generates a deformation point grid according to the reference grid and the offset, and performs feature sampling at the deformation point position of the deformation point grid to obtain a Transformer encoder key and a Transformer encoder value;

[0023] The attention mechanism of the Transformer encoder generates weighted features based on the query vector, the key of the Transformer encoder, and the value of the Transformer encoder;

[0024] The Transformer encoder performs feature prediction on the weighted features to obtain current frame position data, current frame appearance data and current frame confidence data of a single target corresponding to the current frame image data.

[0025] Furthermore, the Transformer encoder reshapes the tracking sequence data of the current frame, including:

[0026] Perform a reshape operation on the tracking sequence data of the current frame to convert the tracking sequence data into a three-dimensional feature map.

[0027] Further, generating a reference grid according to the feature map includes:

[0028] Downsampling the feature map to obtain a grid size of a reference grid;

[0029] A uniform reference grid is generated according to the grid size.

[0030] Furthermore, the Transformer encoder generates an offset according to a query vector corresponding to the tracking sequence data of the current frame, including:

[0031] The query vector is offset through a dynamic offset network to obtain a learnable offset.

[0032] Further, generating a deformation point grid according to the reference grid and the offset and performing feature sampling at deformation point positions of the deformation point grid includes:

[0033] Adding the reference grid and the offset to obtain a deformation point grid;

[0034] Feature sampling is performed on the deformation point positions on the deformation point grid, and features obtained by the feature sampling are used as keys of the Transformer encoder and values ​​of the Transformer encoder.

[0035] Furthermore, the tracking process of the next frame image data based on the current frame position data, the current frame appearance data, the current frame confidence data and the Transformer encoder includes:

[0036] Updating the trajectory change data based on the current frame position data;

[0037] Generate tracking sequence data for the next frame according to the updated trajectory change data, the current frame position data, the current frame appearance data and the current frame confidence data;

[0038] The tracking sequence data of the next frame is tracked and processed through the preset Transformer encoder to obtain the next frame position data, the next frame appearance data and the next frame confidence data of the single target corresponding to the next frame image data.

[0039] Beneficial effects of the present invention:

[0040] (1) Based on historical position data and trajectory change data, the present invention uses a Transformer encoder to accurately extract the position information and motion information of small targets and perform target tracking, thereby reducing the technical problems of missed detection or false detection caused by unclear features of small targets, thereby improving the accuracy of target tracking.

[0041] (2) The present invention uses trajectory change data and appearance prompt data to track the target. In the case where the tracking result is missing due to occlusion, the target position can still be accurately tracked based on the change trajectory and appearance of the tracked target, and an accurate target tracking result can be obtained, thereby improving the tracking efficiency. It can also alleviate the problem of target tracking loss or drift, thereby improving the accuracy of target tracking.

[0042] (3) The present invention utilizes the parallel processing input data capability of the Transformer encoder to reduce computational complexity and improve the operational efficiency of target tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:

[0044] Figure 1It is a flowchart of a satellite video single target tracking method based on appearance and motion feature fusion provided by an embodiment of the present invention.

[0045] Figure 2 It is a data tracking schematic diagram of the Transformer encoder provided by an embodiment of the present invention.

[0046] Figure 3 It is a schematic diagram of the structure of a multi-layer perceptron provided by an embodiment of the present invention.

[0047] Figure 4 The present invention provides a flowchart of satellite video single target tracking based on appearance and motion feature fusion.

[0048] Figure 5 It is a structural diagram of a computer system of a server for implementing the method, system, and device embodiments of the present application. DETAILED DESCRIPTION

[0049] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the relevant invention are shown in the accompanying drawings.

[0050] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0051] The present invention provides a satellite video single target tracking method based on the fusion of appearance and motion features. In order to more clearly illustrate the satellite video single target tracking method based on the fusion of appearance and motion features of the present invention, the following is combined with Figure 1 Each step in the embodiment of the present invention is described in detail.

[0052] The satellite video single target tracking method based on the fusion of appearance and motion features of the first embodiment of the present invention includes steps S101 to S104, and each step is described in detail as follows:

[0053] S101: Acquire template image data, current frame image data, appearance prompt data, confidence mark data, historical position data and trajectory change data.

[0054] The image corresponding to the current frame image data contains a single target to be tracked.

[0055] In this step, the template image data is represented by T, and the template image data is the standard image data in target tracking. The template image data is the first frame or the first few frames of the satellite video. The target area of ​​interest is selected manually or by algorithm as the template image. For example: when monitoring ships at sea, the ship to be tracked can be selected in the starting frame of the video as a template. The ship is a single target to be tracked in the template image. The current frame image data is represented by S, and the current frame image data is the image information contained in a frame of the image currently being processed in the video stream or image sequence. Appearance prompt data is data related to the appearance features of the tracked target, which is used to prompt the appearance attributes and characteristics of the target. It can include color feature data, texture feature data, shape feature data, etc. Confidence mark data is data used to indicate the reliability of the tracking result. In the target tracking task, a confidence value will be given to the tracked target. Historical position data is data that records the position information of the tracked target at different times in the past. It contains the coordinate values ​​of the tracked target at each time point. Track change data is data that describes the change of the motion trajectory of the tracked target over time. It can be understood that in this step, trajectory change data is added to fully utilize motion changes to assist target tracking and improve tracking accuracy.

[0056] S102: Generate tracking sequence data of the current frame based on the template image data, the current frame image data, the appearance prompt data, the confidence mark data, the historical position data and the trajectory change data.

[0057] In this step, the template image data, current frame image data, appearance prompt data, confidence mark data, historical position data and trajectory change data are used to generate the tracking sequence data of the current frame. The tracking sequence data is convenient for input into the Transformer encoder. The Transformer encoder can process input data in parallel to improve the efficiency of target tracking.

[0058] It should be noted that the input of the Transformer encoder is usually vector data in the form of a sequence. Template image data and current frame image data can be converted into vector sequences through operations such as image segmentation and linear embedding. Features in appearance prompt data can be extracted as numerical vectors. Confidence tag data itself is numerical and can be used as a dimension of a vector. Historical position data and trajectory change data can be represented as position coordinate sequences and change amount sequences, respectively, which can naturally form input sequences.

[0059] Step S102 includes: performing cropping and reshaping processing on the template image data and the current frame image data in sequence, respectively, to obtain corresponding template image sequence data and current frame image sequence data.

[0060] Specifically, during the cropping process, the template image data and the current frame image data are cropped based on the image size data of the single target (tracking target) to obtain the image size data corresponding to each. It should be noted that the pixel size of the current frame image data is larger than the pixel size of the template image data. In this embodiment, the pixel size of the template image data and the pixel size of the current frame image data are 2 times and 4 times the pixel size of the tracking target, respectively.

[0061] The two image size data are divided into two-dimensional image blocks respectively and the two-dimensional image blocks are reshaped to obtain template image sequence data and current frame image sequence data corresponding to the template image data and the current frame image data respectively.

[0062] Specifically, the image size data is divided to realize dividing an image into multiple two-dimensional image blocks, that is, each image size data corresponds to multiple two-dimensional image blocks. Each image size data corresponds to multiple two-dimensional image blocks and is flattened to realize reshaping the two-dimensional image blocks into a one-dimensional image sequence, that is, the image sequence corresponding to the template image data is the template image sequence data, and the image sequence corresponding to the current frame image sequence data is the current frame image sequence data.

[0063] Discretization processing is performed on the historical position data and the trajectory change data respectively to obtain discrete historical position data and discrete trajectory change data.

[0064] Specifically, divide the value range of historical location data and trajectory change data into several intervals of equal width, and each interval corresponds to a discrete value. Ensure that the historical location data and trajectory change data exist in a suitable array form. Select equal-width discretization, equal-frequency discretization, or cluster-based discretization according to data characteristics and requirements. Pass the data to be discretized and related parameters (such as the number of intervals and the number of clusters) into the corresponding discretization function for processing. The discretization function returns the discretized data (x min ,y min ,x max ,y max ). The historical position (i.e., the target bounding box, with coordinates x min ,y min ,x max ,y max ), trajectory changes (i.e., the change value of the coordinates and ) is converted into a series of discrete processing units (tokens, tokens refer to the basic units of text processing, such as words or subwords.), and each continuous coordinate and change value is discretized into an integer.

[0065] It should be noted that before discretization, it is necessary to check whether the data to be discretized has missing values, outliers, etc., and process them if necessary.

[0066] The template image sequence data, the current frame image sequence data, the appearance prompt data, the confidence mark data, the discrete historical position data and the discrete trajectory change data are linearly mapped to obtain one-dimensional sequence data composed of the processing units of the Transformer encoder.

[0067] Specifically, the two one-dimensional image sequences mentioned above (i.e., template image sequence data, current frame image sequence data), appearance prompt data, confidence mark data, discrete historical position data, and discrete trajectory change data are linearly mapped and reshaped into a series of one-dimensional sequence data with multiple dimensions, namely, one-dimensional tokens.

[0068] The learnable position encoding is embedded in the one-dimensional sequence data to obtain the tracking sequence data of the current frame.

[0069] Specifically, learnable positional encoding is added or embedded into these one-dimensional sequence data to obtain the input sequence of the backbone network of the Transformer encoder. The input sequence is concatenated along the first dimension of the multi-dimensionality, and the concatenation result is the tracking sequence data of the current frame.

[0070] S103: Tracking the tracking sequence data of the current frame through a preset Transformer encoder to obtain the current frame position data, current frame appearance data and current frame confidence data of the single target corresponding to the current frame image data;

[0071] In this step, the parallel processing input data capability of the Transformer encoder is utilized to improve the efficiency of target tracking.

[0072] It should be noted that the Transformer encoder is able to process all tokens in parallel, abandon intra-frame autoregression while retaining the temporal autoregression characteristics, improve tracking efficiency, and use masked deformable attention to extract small target features in satellite videos.

[0073] In step S103, tracking processing is performed on the tracking sequence data of the current frame by using a preset Transformer encoder, including:

[0074] The Transformer encoder reshapes the tracking sequence data of the current frame. Before inputting the encoder, the tracking sequence data of the current frame is position-encoded. The purpose is to enable the encoder to perceive the order of each sequence, so as to correctly process the sequence data, obtain the feature map of the current frame, and generate a reference grid based on the feature map. It should be pointed out that after encoding the tracking sequence data of the current frame, the feature map of the current frame is still sequence data in N*D format. It is necessary to use a reshaping operation to convert it into a three-dimensional feature map in H*W*C format, where C is the number of channels, H and W are the length and width of the feature map. After converting it into a three-dimensional feature map, a reference grid is generated based on this feature map.

[0075] The Transformer encoder generates an offset according to the query vector corresponding to the tracking sequence data of the current frame, generates a deformation point grid according to the reference grid and the offset, and performs feature sampling at the deformation point position of the deformation point grid to obtain the key of the Transformer encoder and the value of the Transformer encoder;

[0076] The attention mechanism of the Transformer encoder generates weighted features based on the query vector, the key of the Transformer encoder, and the value of the Transformer encoder;

[0077] The Transformer encoder performs feature prediction on the weighted features to obtain the current frame position data, current frame appearance data and current frame confidence data of the single target corresponding to the current frame image data.

[0078] In this step, the tracking sequence data of the current frame is tracked based on the Transformer encoder to predict the position of the tracked target in the current video frame. The Transformer encoder introduces a masked deformable attention mechanism that can focus on small target features and reduce redundant calculations, thereby improving the robustness of tracking.

[0079] Specifically, see Figure 2 The data processing diagram of the Transformer encoder shown in the figure converts the tracking sequence data of the current frame ( Figure 2 The input vector in ( ) is input into the Transformer encoder, and the tracking sequence data of the current frame is reshaped by reshaping, that is, the tracking sequence data of the current frame is reshaped to convert the tracking sequence data into a three-dimensional feature map. The feature map is downsampled to obtain the grid size of the reference grid, and a uniform reference grid is generated according to the grid size.

[0080] It should be noted that the feature map is obtained through the reshape (reshape is an operation that transforms the shape of data in order to make the data conform to the input requirements of various parts of the model for better processing and calculation) deformation operation.

[0081] The query vector (query, abbreviated as q) corresponding to the tracking sequence data of the current frame is input into the Transformer encoder, and the dynamic offset network Perform offset processing to generate learnable offsets , the reference grid and the offset , add them together to get the deformation point grid; perform feature sampling on the deformation point position on the deformation point grid, and use the features obtained by feature sampling as the key (key, abbreviated as k) and value (value, abbreviated as v) of the Transformer encoder. The query vector q, key k and value v are input together into the masked multi-head attention mechanism of the Transformer encoder. Masking operation is performed through masked multi-head attention.

[0082] During the mask operation, the attention mechanism of the Transformer encoder generates weighted features based on the query vector q, the key k of the Transformer encoder, and the value v of the Transformer encoder. The attention mechanism first calculates the similarity between the query vector q and the key k, and outputs the similarity matrix (that is, the attention weight) through the softmax function based on the similarity. After masking unnecessary marks in the matrix, the weighted features are obtained by matrix multiplication with the value v.

[0083] The Transformer encoder predicts the weighted features. Specifically, the weighted features pass through a normalization layer, a jump connection layer, a multi-layer perceptron layer, a normalization layer, and a jump connection layer. The output of the last jump connection layer passes through multiple of the above Transformer encoders, and the last jump connection layer of the last Transformer encoder outputs the tracking result. The tracking result includes: the current frame position data, the current frame appearance data, and the current frame confidence data of the single target corresponding to the current frame image data.

[0084] S104: Tracking the next frame of image data based on the current frame position data, the current frame appearance data, the current frame confidence data and the Transformer encoder.

[0085] In this step, the tracking results of the current frame (current frame position data, current frame appearance data, current frame confidence data) are used to track the next frame of image data, which can avoid full-image target tracking for each frame, thereby greatly reducing the amount of calculation. In addition, targeted processing based on the known tracking results of the current frame can significantly speed up the overall tracking speed of the next frame of image data, which is more suitable for scenes with high real-time requirements, and further helps to make decisions and respond faster.

[0086] Based on the temporal continuity of image frames in the video, there is a certain correlation and similarity between the current frame and the next frame. The tracking result of the current frame can provide contextual information and prior knowledge for the tracking of the next frame, thereby improving the accuracy of target tracking. Even if the target is blocked or blurred in some frames, it can be judged more accurately with the help of the correlation information of the previous and next frames. In addition, the tracking result of the current frame can be used as a reference to verify and correct the tracking result of the next frame. If the tracking result of the next frame is too different from the tracking result of the current frame and does not conform to the normal change law, the tracking result of the next frame can be further analyzed and adjusted, thereby reducing the probability of erroneous tracking and improving the stability and reliability of tracking.

[0087] Specifically, tracking processing is performed on the next frame image data based on the current frame position data, the current frame appearance data, the current frame confidence data and the Transformer encoder, including:

[0088] Update the trajectory change data based on the current frame position data. The center coordinates of the bounding box of the tracked target in the previous N frames of tracking results Collect and save, and calculate the trajectory changes of the tracked target on the horizontal coordinate x and the vertical coordinate y between each frame and , the trajectory changes of the tracking target are fitted into two polynomial functions ( and ):

[0089] ;

[0090] When tracking the target position of the tth frame, the trajectory change data ( , ) is used as the input variable of the encoder to track the target.

[0091] The tracking sequence data of the next frame is generated according to the updated trajectory change data, the current frame position data, the current frame appearance data and the current frame confidence data. The current frame is input into the reconstruction decoder, and the reconstruction decoder realizes autoregressive appearance reconstruction based on the input data. The decoder is used to reconstruct the appearance change of the target in the current video frame, and maintain the target appearance when it is blocked, which is conducive to the effective fusion of motion information and appearance information, and facilitates the subsequent propagation to subsequent video frames.

[0092] The input is fed into a multilayer perceptron (MLP) to obtain the IOU (Intersection over Union) of the predicted bounding box and the true bounding box, which is used to indicate whether the appearance is in an evolving state or a maintained state. The multilayer perceptron consists of a three-layer perceptron, such as Figure 3 As shown in the figure, the three-layer perceptron consists of a linear layer, an activation layer, a linear layer, an activation layer, and a linear layer. The three-layer perceptron outputs the IOU value between the real bounding box and the predicted bounding box. Among them, if the IOU value is greater than the preset value, the corresponding confidence is marked as 1, which belongs to the maintenance state. If the IOU value is less than or equal to the preset value, the corresponding confidence is marked as 0, which belongs to the evolution state.

[0093] It is understandable that the encoder outputs a learnable confidence tag that interacts with all input information, and the confidence tag can guide the appearance tag about whether it is an evolving state or a preserved state. Using a multi-layer perceptron as an appearance evolution indicator to determine the appearance state of the input (evolving state or preserved state) helps solve the common occlusion problem in videos.

[0094] It should be noted that the generation of the tracking sequence data of the next frame also requires the first input template image data and the next frame image data. The specific method of generating the tracking sequence data of the next frame can refer to the method of generating the tracking sequence data of the current frame, which will not be repeated here.

[0095] The tracking sequence data of the next frame is tracked and processed through the preset Transformer encoder to obtain the next frame position data, the next frame appearance data and the next frame confidence data of the single target corresponding to the next frame image data.

[0096] As can be seen from the above description, the present invention is based on historical position data and trajectory change data, and accurately extracts the position information and motion information of small targets through the Transformer encoder and performs target tracking, thereby reducing the technical problems of missed detection or false detection caused by unclear features of small targets, thereby improving the accuracy of target tracking. By using trajectory change data and appearance prompt data to track targets, in the case of missing tracking results due to occlusion, the target position can still be accurately tracked based on the changing trajectory and appearance of the tracked target, and accurate target tracking results can be obtained, thereby improving tracking efficiency. It can also alleviate the problem of target tracking loss or drift, thereby improving the accuracy of target tracking. By utilizing the parallel processing input data capability of the Transformer encoder, it is possible to reduce computational complexity and improve the operational efficiency of target tracking.

[0097] It is understandable that although the various steps in the above embodiment are described in the above-mentioned order, those skilled in the art can understand that in order to achieve the effect of this embodiment, different steps do not have to be executed in such an order, and they can be executed simultaneously (in parallel) or in a reverse order. These simple changes are within the scope of protection of the present invention.

[0098] The embodiment of the present invention also provides a flow chart of a satellite video single target tracking method based on the fusion of appearance and motion features, as shown in FIG. Figure 4 As shown,

[0099] Step 1: Input the current frame image, template image, appearance prompt, confidence mark, historical position prompt and fitted trajectory change of the satellite video to be tracked, and use the parallel processing input data capability of the Transformer encoder to improve the running speed.

[0100] Step 2: Extract features and interact with the input data based on the Transformer encoder, where the appearance hint only interacts with the confidence mark and the current frame image. After executing this step M times, the target position of the current video frame is output. , appearance hint features, and confidence hint features.

[0101] Step 3: Input the appearance hint features obtained in step 2 into the reconstruction decoder, which implements autoregressive appearance reconstruction and outputs the appearance of the target in the current frame. Input the confidence hint features obtained in step 2 into the multi-layer perceptron and output the IOU value between the real bounding box and the predicted bounding box.

[0102] Step 4: The IOU value obtained in step 3 is used as an appearance evolution indicator. When occlusion occurs, the appearance prompt is not updated and the current state is maintained. Otherwise, the input corresponding to the encoder is updated for the appearance prompt obtained in step 3, that is, the confidence mark is updated to the evolution state.

[0103] Step 5: The Transformer encoder fits the trajectory change of the target center point between each frame based on the target position predicted by the previous N frames to generate the trajectory change ( and ) is input to the encoder to assist the model in accurately predicting the target position.

[0104] The above steps are repeated alternately to obtain the target position of the current frame and update the input of the encoder when predicting the next video frame until the last video frame is predicted.

[0105] A second embodiment of the present invention provides a system for tracking a single target in satellite video based on the fusion of appearance and motion features, the system comprising:

[0106] An acquisition unit, used to acquire template image data, current frame image data, appearance prompt data, confidence mark data, historical position data and trajectory change data;

[0107] Wherein, the image corresponding to the current frame image data contains a single target to be tracked;

[0108] A preprocessing unit, configured to generate tracking sequence data of a current frame based on the template image data, the current frame image data, the appearance prompt data, the confidence mark data, the historical position data and the trajectory change data;

[0109] A tracking unit is used to perform tracking processing on the tracking sequence data of the current frame through a preset Transformer encoder to obtain the current frame position data, current frame appearance data and current frame confidence data of the single target corresponding to the current frame image data;

[0110] An iterative unit is used to track the next frame of image data based on the current frame position data, the current frame appearance data, the current frame confidence data and the Transformer encoder.

[0111] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process and related instructions of the system described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0112] It should be noted that the above embodiment provides a system for tracking a single target in satellite video based on the fusion of appearance and motion features, and only uses the division of the above functional modules as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be decomposed or combined. For example, the modules in the above embodiments can be combined into one module, or further divided into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing the modules or steps, and are not regarded as improper limitations on the present invention.

[0113] An electronic device according to a third embodiment of the present invention includes:

[0114] at least one processor;

[0115] and a memory communicatively coupled to at least one of the processors;

[0116] The memory stores instructions that can be executed by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned system method for satellite video single target tracking based on the fusion of appearance and motion features.

[0117] A computer-readable storage medium according to a fourth embodiment of the present invention stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above-mentioned system method for satellite video single target tracking based on the fusion of appearance and motion features.

[0118] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process and related instructions of the storage device and processing device described above can refer to the corresponding process in the aforementioned method embodiment and will not be repeated here.

[0119] Those skilled in the art should be able to appreciate that the modules and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented with electronic hardware, computer software, or a combination of the two, and the programs corresponding to the software modules and method steps can be placed in random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the technical field. In order to clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described in the above description according to the function. Whether these functions are performed in electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0120] Reference below Figure 5 , which shows a schematic diagram of the structure of a computer system of a server for implementing the method, system, and device embodiments of the present application. Figure 5 The server shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0121] like Figure 5 As shown, the computer system includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage part 508 to the random access memory (RAM) 503. Various programs and data required for system operation are also stored in the RAM 503. The CPU 501, ROM 502 and RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0122] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that a computer program read therefrom is installed into the storage section 508 as needed.

[0123] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above-mentioned functions defined in the method of the present application are executed. It should be noted that the above-mentioned computer-readable medium of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, - but not limited to - a system, device or device of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection with one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, an apparatus or a device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate, or transmit a program for use by or in combination with an instruction execution system, an apparatus or a device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0124] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0125] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0126] The term "comprise" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus / device that includes a list of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to such process, method, article, or apparatus / device.

[0127] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. A satellite video single target tracking method based on the fusion of appearance and motion features, characterized in that: include: Obtain template image data, current frame image data, appearance prompt data, confidence mark data, historical position data and trajectory change data; Wherein, the image corresponding to the current frame image data contains a single target to be tracked; Generate tracking sequence data of the current frame based on the template image data, the current frame image data, the appearance prompt data, the confidence mark data, the historical position data and the trajectory change data; The tracking sequence data of the current frame is tracked and processed by a preset Transformer encoder to obtain the current frame position data, current frame appearance data and current frame confidence data of the single target corresponding to the current frame image data; The next frame of image data is tracked based on the current frame position data, the current frame appearance data, the current frame confidence data and the Transformer encoder.

2. The satellite video single target tracking method based on appearance and motion feature fusion according to claim 1 is characterized in that: The generating of the tracking sequence data of the current frame based on the template image data, the current frame image data, the appearance prompt data, the confidence mark data, the historical position data and the trajectory change data comprises: Performing cropping and reshaping processing on the template image data and the current frame image data in sequence, respectively, to obtain corresponding template image sequence data and current frame image sequence data; Discretizing the historical position data and the trajectory change data respectively to obtain discrete historical position data and discrete trajectory change data; Performing linear mapping processing on the template image sequence data, the current frame image sequence data, the appearance prompt data, the confidence mark data, the discrete historical position data and the discrete trajectory change data respectively to obtain one-dimensional sequence data composed of the processing units of the Transformer encoder; A learnable position code is embedded in the one-dimensional sequence data to obtain tracking sequence data of the current frame.

3. The satellite video single target tracking method based on appearance and motion feature fusion according to claim 2 is characterized in that: The step of performing cropping and reshaping on the template image data and the current frame image data in sequence comprises: Based on the image size data of the single target, the template image data and the current frame image data are respectively cropped to obtain respective corresponding image size data; The two image size data are respectively divided into two-dimensional image blocks and the two-dimensional image blocks are reshaped to obtain template image sequence data and current frame image sequence data corresponding to the template image data and the current frame image data respectively.

4. The satellite video single target tracking method based on appearance and motion feature fusion according to claim 3 is characterized in that: The pixel size of the current frame image data is greater than the pixel size of the template image data.

5. The satellite video single target tracking method based on appearance and motion feature fusion according to claim 1 is characterized in that: The tracking process of the tracking sequence data of the current frame by using a preset Transformer encoder includes: The Transformer encoder reshapes the tracking sequence data of the current frame to obtain a feature map of the current frame and generates a reference grid according to the feature map; The Transformer encoder generates an offset according to a query vector corresponding to the tracking sequence data of the current frame, generates a deformation point grid according to the reference grid and the offset, and performs feature sampling at the deformation point position of the deformation point grid to obtain a Transformer encoder key and a Transformer encoder value; The attention mechanism of the Transformer encoder generates weighted features based on the query vector, the key of the Transformer encoder, and the value of the Transformer encoder; The Transformer encoder performs feature prediction on the weighted features to obtain current frame position data, current frame appearance data and current frame confidence data of a single target corresponding to the current frame image data.

6. The satellite video single target tracking method based on appearance and motion feature fusion according to claim 5 is characterized in that: The Transformer encoder reshapes the tracking sequence data of the current frame, including: Perform a reshape operation on the tracking sequence data of the current frame to convert the tracking sequence data into a three-dimensional feature map.

7. The satellite video single target tracking method based on appearance and motion feature fusion according to claim 5 is characterized in that: The step of generating a reference grid according to the feature map comprises: Downsampling the feature map to obtain a grid size of a reference grid; A uniform reference grid is generated according to the grid size.

8. The satellite video single target tracking method based on appearance and motion feature fusion according to claim 5 is characterized in that: The Transformer encoder generates an offset according to a query vector corresponding to the tracking sequence data of the current frame, including: The query vector is offset through a dynamic offset network to obtain a learnable offset.

9. The satellite video single target tracking method based on appearance and motion feature fusion according to claim 5, characterized in that: Generating a deformation point grid according to the reference grid and the offset and performing feature sampling at deformation point positions of the deformation point grid includes: Adding the reference grid and the offset to obtain a deformation point grid; Feature sampling is performed on the deformation point positions on the deformation point grid, and features obtained by the feature sampling are used as keys of the Transformer encoder and values ​​of the Transformer encoder.

10. The satellite video single target tracking method based on appearance and motion feature fusion according to claim 1, characterized in that: The tracking process of the next frame image data based on the current frame position data, the current frame appearance data, the current frame confidence data and the Transformer encoder includes: Updating the trajectory change data based on the current frame position data; Generate tracking sequence data for the next frame according to the updated trajectory change data, the current frame position data, the current frame appearance data and the current frame confidence data; The tracking sequence data of the next frame is tracked and processed through the preset Transformer encoder to obtain the next frame position data, the next frame appearance data and the next frame confidence data of the single target corresponding to the next frame image data.

Citation Information

Patent Citations

  • Satellite video single target tracking method, system and device and storage medium

    CN116128926A

  • Video dynamic target tracking method and device, equipment and storage medium

    CN116503441A