A Correlation Filter Satellite Video Object Tracking Method Based on Trajectory Prediction

By constructing trajectory prediction network (TPN) and adaptive occlusion perception technology, the tracking difficulty problem when the target is blocked in satellite video is solved, and the tracking accuracy and success rate in occlusion situations are improved.

CN116977869BActive Publication Date: 2025-07-11NORTHWESTERN POLYTECHNICAL UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310759710.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-26
Publication Date
2025-07-11
Estimated Expiration
2043-06-26

AI Technical Summary

Technical Problem

The existing satellite video target tracking methods are difficult to continuously and accurately track when the target is blocked, especially when the target image characteristics are insufficient in satellite video, the existing methods lack adaptability and lack prediction capabilities in multi-frame occlusion.

Method used

Using a correlation filtering method based on trajectory prediction, timing features are extracted by constructing a trajectory prediction network (TPN), combined with adaptive occlusion perception and weight fusion technology, the target position is predicted and adaptive adjustments are performed during occlusion.

Benefits of technology

The target tracking accuracy and success rate in occlusion situations are improved, and the continuous and accurate tracking of targets in satellite video is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116977869B_ABST
    Figure CN116977869B_ABST
Patent Text Reader

Abstract

The present invention relates to a correlation filtering satellite video object tracking method based on trajectory prediction. The position information of the object is introduced into the tracking method, and a trajectory prediction network with an attention and temporal feature splicing mechanism is designed to fully extract temporal features and predict future positions. An occlusion perception index and an adaptive threshold are proposed to perceive whether the object in the satellite video is occluded. When occlusion occurs, the predicted trajectory with an adaptive weight and the correlation filtering tracking position result are fused, so as to effectively improve the continuous tracking ability of the object in the case of occlusion and enhance the tracking accuracy and success rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of satellite video target tracking, and relates to a satellite video target tracking method based on trajectory prediction and correlation filtering. Background Art

[0002] Video satellites acquire videos by continuously imaging a locked observation area for a long time. Compared with static observation images, satellite video data contains more spatio-temporal information. As one of the important technical means for dynamic monitoring, video target tracking technology studies the continuous positioning of specific targets in video image sequences. Due to the bird's-eye view and the complexity of the photographed scene, the target in the video frame may be occluded. When the target is occluded, the occluder will introduce samples that do not belong to the target into the detection and update process of the correlation filtering model. In satellite videos, the target occupies fewer pixels, and its image features are not sufficient to enable existing tracking methods to accurately track the target in subsequent frames. To meet the needs of continuous satellite earth observation, satellite video tracking methods should have the ability to continuously and accurately track the target even after the target is occluded.

[0003] Liu Yaosheng (Research on Video Satellite Target Tracking Algorithm Based on Kernel Correlation Filtering, Fire Control & Command Control, 2022, 47-02, 49-55) addressed the problem of target occlusion in satellite videos. Based on the KCF algorithm, the peak sidelobe ratio was used as an index to judge target occlusion, and the Kalman filter algorithm was used to predict the next position of the target when target occlusion was determined. However, this method uses a fixed threshold to judge whether the target is occluded, lacking adaptability to different video sequences. Moreover, the Kalman filter model used to predict the target position is only applicable to linear situations, and is limited in prediction when the target is occluded for multiple consecutive frames, making it difficult to achieve continuous and accurate tracking of satellite video targets. Summary of the Invention

[0004] Technical Problems to be Solved

[0005] To avoid the deficiencies of the prior art, the present invention proposes a satellite video target tracking method based on trajectory prediction and correlation filtering, aiming at the problem that the image features extractable from the target in satellite videos are few, making it difficult to maintain tracking after the correlation filtering model is contaminated by wrong samples.

[0006] Technical Solution

[0007] A satellite video target tracking method based on trajectory prediction and correlation filtering, characterized by the following steps:

[0008] Step 1: Construct a satellite video trajectory prediction data set:

[0009] Extract the bounding box sequence from the data set to obtain the trajectory data for training the trajectory prediction network;

[0010] The trajectory data in the x direction is

[0011] The trajectory data in the y direction is

[0012] Step 2: Use the trajectory data obtained in Step 1 to train the TPN trajectory prediction network using the gradient descent and backpropagation algorithms;

[0013] Step 3: Use the correlation filtering algorithm based on the TPN trajectory prediction network to track the satellite video target:

[0014] Step 3-1: Initialize the correlation filtering model using the first-frame image I1 of the satellite video, and set the continuous occlusion count variable count occ = 0; Crop the target area from I1 and extract the image feature samples to initialize the correlation filter h1. The target position is denoted as pos1 = (x1, y1), and this position is added to the position sequence set {pos t};

[0015] Step 3-2: Read the next frame image of the satellite video. Using the previous-frame target position pos t-1 as the center, extract the image feature sample f t from the current frame I t and perform a correlation operation with the correlation filter h t-1 to obtain the correlation response map R t , and the position where the maximum value in R t is located is the target position determined by the correlation filtering

[0016] Step 3-3: Input the correlation response map R t in Step 3-2 into the occlusion perception module, using the response confidence conf t and the occlusion perception threshold to determine whether the target is in an occluded state;

[0017] The response confidence conf t is the proportion of the part in the response map R t where the response value is less than 0.1 times the response peak:

[0018]

[0019] where card{·} is a function for calculating the set elements, and R max represents taking the maximum value of R;

[0020] Take the response map of the second frame of the video as the base response map, i.e., R base = R2; Calculate R baseThe response confidence is used as the base confidence, i.e., conf base = conf2;

[0021] Calculate the ratio of the response confidence of the current frame to the base confidence:

[0022] occ t = conf t / conf base

[0023] The occlusion perception threshold is the occlusion perception threshold adapted to different video sequences According to the similarity between the response maps R t and R base calculate:

[0024]

[0025]

[0026] where k is a constant controlling the threshold size, ρ(·) is the Pearson correlation coefficient, Cov and σ represent covariance and standard deviation;

[0027] If it is considered that no occlusion has occurred, set the continuous occlusion count variable count occ = 0, and take the target position obtained by correlation filtering as the target tracking position result pos of the current frame t and jump to step 3-5. Otherwise, it is considered that the target in the current frame is occluded, set the continuous occlusion count variable count occ = count occ + 1 and enter step 3-4;

[0028] Step 3-4: Use the TPN trajectory prediction network to predict the trajectory points of the current frame in the occluded state and use adaptive weights to determine the position of the target:

[0029]

[0030] where: is the decay weight;

[0031] Step 3-5: Extract scale samples centered on pos t = (x t , y t ), use the CSRDCF scale filter to obtain the width w t and height h t of the target in the current frame, and take pos t = (x t , y t)Extract samples centered around to update the CSRDCF filter;

[0032] Step 3-6: Add the current frame tracking position result pos t to the position sequence set {pos t}, and output the bounding box tracking result of the current frame as (x t - w t / 2, y t - h t / 2, w t , h t );

[0033] After this step, the bounding box tracking result corresponding to the current frame is obtained. Then, repeat steps 3-2 to 3-6 for the next consecutive frame until the end of the video sequence.

[0034] The dataset used is the SV248 dataset.

[0035] The process of extracting the bounding box sequence to obtain the trajectory data for training the trajectory prediction network is as follows: For each set of bounding box sequences {bbox t}, convert the bounding box data {bbox t} to the position coordinate sequence data {pos t} according to the following formula:

[0036] pos t = (bbox t [0] + bbox t [2] / 2, bbox t [1] + bbox t [3] / 2) = (x t , y t )

[0037] Separate the data in the x and y directions and obtain the position coordinate sequence data {x t} and {y t} in the x and y directions from the position coordinate sequence data {pos t}; The trajectory data in the x direction and the trajectory data in the y direction are calculated based on {x t} and {y t} as follows:

[0038] where the calculation methods of v x , a x , dir x , dist are as follows:

[0039]

[0040] where v y , a y , dir y , the calculation of dist is as follows:

[0041]

[0042] The TPN trajectory prediction network includes a trajectory segmentation module, an encoder network, a temporal splicing layer, a self-attention layer, a decoder network, and a fully connected layer; among them, the trajectory segmentation module divides the input data into 3 subsequences according to time series, and the encoder network includes 3 1D convolutional layers with convolutional channel numbers of 16, 64, and 32. The temporal splicing layer splices the encoded network features of the subsequences in chronological order. The self-attention layer adds the scale dot product attention value to the network features. The decoder consists of a long short-term memory network (LSTM) with 128 hidden units, and the fully connected layer maps the decoded network features to a 1D linear network layer.

[0043] The training steps of step 2 are as follows:

[0044] Step 2-1: Obtain m historical trajectory points from the trajectory data in step 1 in a sliding window manner as the input historical trajectory of the network Obtain the (m + 1)-th historical trajectory point as the training label for training the network; the historical trajectory is divided into three parts of sequence data in chronological order, namely and and these three parts of sequence data are respectively sent into the encoder network to obtain three groups of features tcf 1 , tcf 2 and tcf 3 ;

[0045] Step 2-2: Keep the feature channel dimension unchanged, splice the three groups of features obtained in step 2-1 in the time series dimension. The spliced feature tcf passes through a self-attention layer to further enable the network to focus on important features. Use a linear layer to map the feature tcf to a query matrix Q, a key matrix K, and a value matrix V, and calculate the self-attention matrix using the scale dot product attention mechanism according to the following formula:

[0046]

[0047] where softmax is a normalization function, d k is the feature channel dimension of tcf; after the self-attention matrix attn is added to tcf, the feature atf is obtained through a network layer with a residual connection form, and after being input into an LSTM network with 128 hidden neurons, it passes through a fully connected layer to obtain the network prediction output;

[0048] Step 2-3: Calculate the loss between the network prediction output obtained in Step 2-2 and the training labels obtained in Step 2-1 according to the mean square error loss function, and update the network weights; after the network loss converges, keep the network structure and weights unchanged to complete the training of the TPN trajectory prediction network.

[0049] When updating the network weights in Step 2-3, use gradient descent backpropagation and update the network weights using the Adam network optimizer with a learning rate of 0.001.

[0050] The process of using the TPN trajectory prediction network in Step 3-4 to predict the trajectory points of the current frame in the occlusion state and using the adaptive weights to determine the position of the target is as follows:

[0051] Process the historical position sequence data {pos t-1} according to the trajectory sequence construction method in Step 1 to obtain the trajectory sequence data and Input it into the trained TPN network in Step 2 to obtain the predicted positions in the x and y directions and Perform weighted fusion on the predicted trajectory position and the correlation filtering position to determine the target position. The calculation formula for the weight of the proposed predicted trajectory position is:

[0052]

[0053] If the continuous occlusion count variable count occ exceeds 15, attenuate the weight and set the attenuated part not to exceed 0.3:

[0054]

[0055] The target position pos t =(x t , y t ) is calculated according to the following formula:

[0056]

[0057] The constant k for controlling the threshold size in Step 3-3 takes a value of 1.3.

[0058] Beneficial effects

[0059] A correlation filtering satellite video target tracking method based on trajectory prediction proposed by the present invention introduces the position information of the target into the tracking method, designs a trajectory prediction network with an attention and temporal feature splicing mechanism to fully extract temporal features and predict future positions, proposes an occlusion perception index and an adaptive threshold to perceive whether the target in the satellite video is occluded, and fuses the predicted trajectory with an adaptive weight and the correlation filtering tracking position result when occlusion occurs, so as to effectively improve the continuous tracking ability of the target in the case of occlusion and enhance the tracking accuracy and success rate.

[0060] The technical effect of the present invention is that, compared with the existing correlation filtering tracking method, the present invention designs a trajectory prediction network with an attention and temporal splicing mechanism through Invention Step 2 to fully extract and utilize the position information, introduces an occlusion perception index and a threshold adaptive to the video sequence through Invention Step 3-3 to perceive whether the target in the satellite video is occluded, and uses an adaptive weight to fuse the correlation filtering tracking position result and the predicted trajectory to obtain the target position through Invention Step 3-4 when occlusion occurs, so as to effectively improve the prediction accuracy of the target position and the continuous tracking ability in the case of occlusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 It is a flowchart of the method of the present invention

[0062] Figure 2 It is a network structure diagram of the trajectory prediction network TPN, where the kernel size and channels size respectively represent the convolution kernel size and convolution channel size of the convolution layer, and Conv and LSTM respectively represent the convolutional network and the long short-term memory network

[0063] Figure 3 It is a schematic diagram of the tracking result (blue box) of the method of the present invention Figure 3 (a) is a video frame in the original occlusion case Figure 3 (b) and Figure 3 (c) are the tracking results of the CSRDCF method and the method of the present invention.

[0064] Figure 4 It is a performance comparison result diagram of the tracking results of the method of the present invention and other algorithms in the satellite video Figure 4 (a) is the success rate diagram Figure 4 (b) is the accuracy diagram. DETAILED DESCRIPTION OF THE INVENTION

[0065] The present invention will be further described in combination with the embodiments and the drawings as follows:

[0066] The correlation filtering satellite video tracking method based on trajectory prediction of the present invention is realized through the following technical solutions, and the specific steps are as follows:

[0067] Step 1: Construct a satellite video trajectory prediction dataset. Extract the bounding box sequences from the SV248 dataset to obtain the trajectory data for training the trajectory prediction network. For each set of bounding box sequences {bbox t}, convert the bounding box data {bbox t} into a set of position sequences {pos t} according to the following calculation formula:

[0068] pos t =(bbox t [0]+bbox t [2] / 2,bbox t [1]+bbox t [3] / 2)=(x t ,y t ) (1)

[0069] Separate the data in the x and y directions for processing. From the set of position sequences {pos t}, the sets of position sequences {x t} and {y t} in the x and y directions can be obtained. The trajectory data in the x direction and the trajectory data in the y direction are calculated according to {x t} and {y t}. Taking the x direction as an example, the trajectory data is where the calculation methods of v x , a x , dir x , dist are as follows:

[0070]

[0071] Similarly for the y direction, after processing in the above manner, the trajectory data in the y direction is obtained

[0072]

[0073]

[0074] Step 2: Construct a TPN trajectory prediction network and train the network using the gradient descent and backpropagation algorithms with the trajectory data obtained in Step 1, as follows:

[0075] Step 2-1: Build the TPN network framework (as shown in the appendix Figure 2As shown in the figure. The TPN network framework consists of a trajectory segmentation module, an encoder network, a temporal concatenation layer, a self-attention layer, a decoder network, and a fully connected layer. Among them, the trajectory segmentation module divides the input data into 3 subsequences according to time series. The encoder network includes 3 convolutional (1DConv) layers with convolutional channel numbers of 16, 64, and 32. The temporal concatenation layer concatenates the encoded network features of the subsequences in chronological order. The self-attention layer adds the scale dot product attention values to the network features. The decoder consists of a long short-term memory network (LSTM) with 128 hidden units. The fully connected layer is a linear network layer that maps the decoded network features to 1 dimension.

[0076] Step 2-2: Obtain m historical trajectory points from the trajectory data in Step 1 in a sliding window manner as the input historical trajectory of the network. Obtain the (m + 1)-th historical trajectory point as the training label for training the network. The historical trajectory is divided into three parts of sequence data in chronological order, namely and and these three parts of sequence data are respectively sent into the encoder network to obtain three groups of features tcf 1 、tcf 2 and tcf 3 .

[0077] Step 2-3: Keep the feature channel dimension unchanged, concatenate the three groups of features obtained in Step 2-1 in the temporal dimension. The concatenated feature tcf passes through a self-attention layer to further enable the network to focus on important features. Use a linear layer to map the feature tcf to a query matrix Q, a key matrix K, and a value matrix V. Calculate the self-attention matrix using the scale dot product attention mechanism according to the following formula:

[0078]

[0079] where softmax is the normalization function, d k is the feature channel dimension of tcf. After the sum of the self-attention matrix attn and tcf passes through a network layer with a residual connection form to obtain the feature atf, it is input into an LSTM network with 128 hidden neurons and then passes through a fully connected layer to obtain the network prediction output.

[0080] Step 2-4: Calculate the loss between the network prediction output obtained in Step 2-3 and the training label obtained in Step 2-1 according to the mean square error loss function. Use gradient descent backpropagation and use the Adam network optimizer with a learning rate of 0.001 to update the network weights. After the network loss converges, keep the network structure and weights unchanged to complete the training of the TPN trajectory prediction network.

[0081] Step 3: Use the correlation filtering algorithm based on the TPN trajectory prediction network to track the satellite video target. The details are as follows:

[0082] Step 3-1: Initialize the correlation filter model using the first frame of the satellite video image I1 and set the continuous occlusion count variable count occ = 0. For the satellite video tracking task, the position and size of the target to be tracked in the first frame image are given. The target area is cropped from I1 and the image feature samples are extracted to initialize the correlation filter h1. The target position is recorded as pos1 = (x1, y1), and the position is added to the position sequence set {pos t};

[0083] Step 3-2: Read the next frame of satellite video, using the target position pos of the previous frame t-1 Centered on the current frame I t Extract image feature samples f t and the correlation filter h t-1 Perform correlation operations to obtain the response graph R t , R t The location of the maximum value is the target location determined by the correlation filter

[0084] Step 3-3: Convert the response graph R in step 3-2 to t Input occlusion perception module. Propose a response confidence and occlusion perception threshold to determine whether the target is in an occluded state. Define response confidence conf t For the response graph R t The proportion of the response value that is less than 0.1 times the peak value of the response, that is:

[0085]

[0086] where card{·} is a function for calculating set elements, R max Indicates taking the maximum value of R.

[0087] The response graph of the second frame of the video is used as the basic response graph, that is, R base = R2; R base The response confidence calculated according to formula (4) is used as the basic confidence, that is, conf base =conf2. Calculate the ratio of the response confidence of the current frame to the basic confidence:

[0088] occ t =conf t / conf base (5)

[0089] Propose an occlusion perception threshold that is adaptive to different video sequences This threshold is calculated based on the response map R t and R base The similarity between them is calculated as follows:

[0090]

[0091] where k is a constant that controls the threshold size, and its value is 1.3; ρ(·) is the Pearson correlation coefficient. The closer this value is to 0, the weaker the similarity, and the occlusion perception threshold is larger; Cov and σ represent covariance and standard deviation. If it is considered that no occlusion has occurred, the continuous occlusion count variable count occ is set to 0, and the target position obtained by correlation filtering is used as the target tracking position result pos of the current frame t and jump to step 3-5. Otherwise, it is considered that the target in the current frame is occluded, and the continuous occlusion count variable count occ is set to count occ +1 and enter step 3-4;

[0092] Step 3-4: Use the TPN trajectory prediction network to predict the trajectory points of the current frame in the occluded state and use adaptive weights to determine the position of the target. Process the position sequence set {pos t-1} according to the trajectory sequence construction method in step 1 to obtain trajectory data and Input it into the TPN network trained in step 2 to obtain the predicted positions in the x and y directions and Perform weighted fusion on the trajectory prediction position and the correlation filtering position obtained in step 3-2 to determine the target position. The calculation formula for the weight of the proposed trajectory prediction position is:

[0093]

[0094] The Pearson correlation coefficient value involved in this weight has been calculated in step 3-3. If the continuous occlusion count variable count occ exceeds 15, the weight is attenuated and it is set that the attenuated part does not exceed 0.3:

[0095]

[0096] The target position pos t =(x t ,y t ) is calculated according to the following formula:

[0097]

[0098] The width w of the current frame target t and height h t are the same as the width w of the target in the previous frame t-1 and height h t-1 Keep consistent, and output the bounding box tracking result of the current frame as (x t -w t / 2, y t -h t / 2, w t , h t );

[0099] Step 3-5: Extract scale samples centered on pos t =(x t , y t ), use the CSRDCF scale filter to obtain the width w t and height h t of the current frame target, and extract samples centered on pos t =(x t , y t ) to update the CSRDCF filter;

[0100] Step 3-6: Add the current frame tracking position result pos t to the position sequence set {pos t}, and output the bounding box tracking result of the current frame as (x t -w t / 2, y t -h t / 2, w t , h t ).

[0101] After this step, the bounding box tracking result corresponding to the current frame is obtained. Then, repeat steps 3-2 to 3-6 for the next consecutive frame until the end of the video sequence.

[0102] The basic process of the improved visual background extraction method with time-domain interval reference added in the present invention is as Figure 1 shown. The specific implementation manner of the present invention is illustrated by examples, but the technical content of the present invention is not limited to the described scope. The specific implementation manner includes the following steps:

[0103] Step 1: Construct a satellite video trajectory prediction dataset. Obtain the bounding box sequence from the SV248 dataset to construct the training data of the trajectory prediction network. For each set of bounding box sequences {bbox t}, the bounding box is composed of the upper left corner coordinates and width and height of the target Convert the bounding box data {bbox t} into a position sequence set {post}:

[0104]

[0105] Separate the data in the x - direction and y - direction for processing. From the set of position sequences {pos t}, the sets of position sequences in the x - direction and y - direction {x t} and {y t} can be obtained. The trajectory data in the x - direction and the trajectory data in the y - direction are calculated according to {x t} and {y t}. Taking the x - direction as an example, the trajectory data is where v x , a x , dir x , and dist are calculated as follows:

[0106]

[0107] Similarly for the y - direction, according to the following formula:

[0108]

[0109] The trajectory data in the y - direction is obtained after the above processing

[0110] Step 2: Construct a TPN trajectory prediction network and use the trajectory data obtained in Step 1 to train the network using the gradient descent and backpropagation algorithms, as follows:

[0111] Step 2 - 1: Build the TPN network framework (as shown in the appendix Figure 2 ). The TPN network framework consists of a trajectory segmentation module, an encoder network, a temporal concatenation layer, a self - attention layer, a decoder network, and a fully - connected layer. Among them, the trajectory segmentation module divides the input data into 3 subsequences according to time series. The encoder network includes 3 convolutional (1DConv) layers with 16, 64, and 32 convolutional channels. The temporal concatenation layer concatenates the encoded network features of the subsequences in chronological order. The self - attention layer adds the scaled dot - product attention values to the network features. The decoder consists of a long short - term memory (LSTM) network with 128 hidden units. The fully - connected layer is a linear network layer that maps the decoded network features to 1 - D.

[0112] Step 2 - 2: Obtain m historical trajectory points from the trajectory data in Step 1 in a sliding window manner as the input historical trajectory of the network and obtain the (m + 1) - th historical trajectory point as the training label for training the network. The historical trajectory is divided into three parts of sequence data in chronological order, that is and and separately send these three parts of sequence data into the encoder network to obtain three groups of features tcf 1 , tcf 2 and tcf 3 .

[0113] Step 2-3: Keeping the feature channel dimension unchanged, concatenate the three groups of features obtained in Step 2-2 in the temporal dimension. The concatenated feature tcf passes through a self-attention layer to further enable the network to focus on important features. Use a fully connected layer to map the feature tcf into a query matrix Q, a key matrix K, and a value matrix V. The calculation formula for the self-attention matrix of the scaled dot-product attention mechanism is:

[0114]

[0115] where softmax is the normalization function, and d k is the feature channel dimension of tcf. After summing the self-attention matrix attn and tcf, the feature atf is obtained through a network layer with a residual connection form. After inputting it into an LSTM network with 128 hidden neurons and then passing through a fully connected layer, the network prediction output is obtained.

[0116] Step 2-4: Calculate the loss between the network prediction output obtained in Step 2-3 and the training label obtained in Step 2-1 according to the mean square error loss function and use gradient descent for backpropagation. Use the Adam network optimizer with a learning rate of 0.001 to update the network weights. After the network loss converges, keep the network structure and weights unchanged to complete the training of the TPN trajectory prediction network.

[0117] Step 3: Use the trained TPN trajectory prediction network and the CSRDCF correlation filtering algorithm to track satellite video targets. Specifically as follows:

[0118] Step 3-1: Use the first-frame image I1 of the satellite video to initialize the CSRDCF correlation filtering model and set the continuous occlusion count variable count occ = 0. For the satellite video tracking task, the position and size of the target to be tracked in the first-frame image are given. Crop the target area from I1, and extract the histogram of oriented gradients features and color features. Concatenate the features in the channel dimension to obtain the initial feature sample f0. Use the feature sample f0 to train and obtain the initialized correlation filter h1. The target position is denoted as pos1 = (x1, y1), and add this position to the position sequence set {pos t};

[0119] Step 3-2: Read the next frame image of the satellite video. The t-th frame image is denoted as I t . Use the previous-frame target position pos t-1Centered on the current frame I t Crop the image region and extract features in it to obtain the feature sample f t Then, for the feature sample f t Perform a correlation operation with the correlation filter h t-1 to obtain the response map R t For R t the position of the maximum value in it is the target position determined by the correlation filtering

[0120] Step 3-3: Input the response map R in Step 3-2 t into the occlusion perception module. A response confidence and an occlusion perception threshold are proposed to determine whether the target is in an occluded state according to the response map. For the response map R t , define the response confidence conf t as the proportion of the part in the response map R t where the response value is less than 0.1 times the response peak:

[0121]

[0122] where card{·} is a function for calculating the elements of a set, and R max represents taking the maximum value of R. Take the response map of the second frame of the video as the base response map, that is, R base = R2, and take the response confidence calculated from R base according to Equation (13) as the base confidence, that is, conf base = conf2. Calculate the ratio of the current frame response confidence to the base confidence:

[0123] occ t = conf t / conf base (14)

[0124] Propose an occlusion perception threshold adapted to different video sequences This threshold is calculated according to the similarity between the response map R t and R base :

[0125]

[0126]

[0127] where k is a constant controlling the threshold size, and it is recommended to take the value of 1.3; ρ(·) is the Pearson correlation coefficient, and the closer this value is to 0, the weaker the similarity and the larger the corresponding occlusion perception threshold; Cov and σ represent the covariance and standard deviation respectively. Compare occ t with the occlusion perception threshold If It is considered that there is no occlusion, and the continuous occlusion count variable count is set occ = 0, and the target position obtained by correlation filtering is used as the target tracking position result pos of the current frame t and jump to step 3-5. Otherwise, it is considered that the target in the current frame is occluded, and the continuous occlusion count variable count is set occ = count occ + 1 and enter step 3-4;

[0128] Step 3-4: After it is sensed in step 3-3 that the target is occluded, use the TPN trajectory prediction network to predict the trajectory points of the current frame and use adaptive weights to determine the target position. Process the collected position sequence set {pos t-1} according to the trajectory sequence construction method in step 1 to obtain trajectory data and Input it into the trained TPN network in step 2 to obtain the predicted positions in the x and y directions and Next, perform weighted fusion on the trajectory prediction position and the correlation filtering position obtained in step 3-2 to determine the target position. The calculation formula for the weight occupied by the proposed trajectory prediction position is:

[0129]

[0130] The Pearson correlation coefficient value involved in this weight has been calculated in step 3-3. If the continuous occlusion count variable count occ exceeds 15, it is considered that the target is gradually leaving the occlusion area and the weight is attenuated and it is set that the attenuated part does not exceed 0.3:

[0131]

[0132] The target position pos t = (x t , y t ) is calculated according to the following formula:

[0133]

[0134] The width w of the target in the current frame t and the height h t are the same as the width w t-1 and the height h t-1 of the target in the previous frame;

[0135] Step 3-5: With pos t = (x t , yt )Extract scale samples centered around t and use the CSRDCF scale filter to obtain the width w of the target in the current frame t and height h t =(x t ,y t ), and extract samples centered around pos

[0136] Step 3-6: Output the bounding box tracking result of the current frame as (x t -w t / 2,y t -h t / 2,w t ,h t ), add the current frame tracking position result pos t to the position sequence set {pos t}. If the video sequence frames have not been fully traversed, continue processing from Step 3-2, otherwise end.

Claims

1. A correlation filtering satellite video target tracking method based on trajectory prediction, characterized in that The steps are as follows: Step 1: Construct a satellite video trajectory prediction dataset: Extract the bounding box sequence from the dataset to obtain the trajectory data for training the trajectory prediction network; The trajectory data in the x direction is The trajectory data in the y direction is Step 2: Use the trajectory data obtained in Step 1 to train the TPN trajectory prediction network using the gradient descent and backpropagation algorithms; Step 3: Use the correlation filtering algorithm based on the TPN trajectory prediction network to track the satellite video target: Step 3-1: Initialize the correlation filtering model using the first frame image I1 of the satellite video, and set the continuous occlusion count variable count occ = 0; Crop the target area from I1 and extract image feature samples to initialize the correlation filter h1. The target position is denoted as pos1 = (x1, y1), and add this position to the position sequence set {pos t}; Step 3-2: Read the next frame image of the satellite video, and extract the image feature sample f in the current frame I centered on the target position pos of the previous frame t-1 and perform a correlation operation with the correlation filter h t to obtain the correlation response map R t The position of the maximum value in R t-1 is the target position determined by the correlation filter t t ​​ Step 3-3: Input the relevant response graph R in Step 3-2 t into the occlusion perception module to obtain the response confidence conf t and the occlusion perception threshold for determining whether the target is in an occluded state; The response confidence conf t is the proportion of the part in the response map R t where the response value is less than 0.1 times the response peak value: where card{·} is a function for calculating the elements of a set, and R max represents taking the maximum value of R; Use the response map of the second frame of the video as the base response map, i.e., R base = R2; Calculate the response confidence of R base as the base confidence, i.e., conf base = conf2; Calculate the ratio of the response confidence of the current frame to the base confidence: occ t = conf t / conf base The occlusion perception threshold is an occlusion perception threshold adapted to different video sequences calculated according to the similarity between the response maps R t and R base : where k is a constant that controls the threshold size, ρ(·) is the Pearson correlation coefficient, and Cov and σ represent covariance and standard deviation; If it is considered that there is no occlusion, and the continuous occlusion count variable count occ is set to 0, and the target position obtained by correlation filtering is used as the target tracking position result pos of the current frame t and jump to step 3-5. Otherwise, it is considered that the target in the current frame is occluded, and the continuous occlusion count variable count occ is set to count occ +1 and enter step 3-4; Step 3-4: Use the TPN trajectory prediction network to predict the trajectory points of the current frame in the occluded state and use adaptive weights to determine the position of the target: Wherein: is the attenuation weight; Step 3-5: Extract scale samples centered at pos t =(x t , y t ), use the CSRDCF scale filter to obtain the width w t and height h t of the target in the current frame, and extract samples centered at pos t =(x t , y t ) to update the CSRDCF filter; Step 3-6: Add the current frame tracking position result pos t to the position sequence set {pos t}, and output the bounding box tracking result of the current frame as (x t - w t / 2, y t - h t / 2, w t , h t ); After this step, the bounding box tracking result corresponding to the current frame is obtained. Then, repeat Steps 3-2 to 3-6 for the next consecutive frame until the video sequence ends.

2. The method for tracking satellite video targets based on correlation filtering using trajectory prediction according to claim 1, wherein: The dataset uses the SV248 dataset.

3. The method for tracking satellite video targets based on correlation filtering with trajectory prediction according to claim 1, characterized in that: The process of extracting the bounding box sequence to obtain trajectory data for training the trajectory prediction network is as follows: For each set of bounding box sequences {bbox t}, the bounding box data {bbox t} is converted into position coordinate sequence data {pos t} according to the following formula: pos t = (bbox t [0] + bbox t [2] / 2, bbox t [1] + bbox t [3] / 2) = (x t , y t ) Separate the data in the x - direction and y - direction for processing, and obtain the position coordinate sequence data in the x - direction and y - direction {x t} and {y t} from the position coordinate sequence data {pos t}; Calculate the trajectory data in the x - direction {Tr t x |Tr x =(x, v x , a x , dir x , dist)} and the trajectory data in the y - direction {Tr t y |Tr y =(y, v y , a y , dir y , dist)} according to {x t} and {y t}: where v x , a x , dir x , the calculation method of dist is as follows: where v y , a y , dir y , the calculation method of dist is as follows:

4. The method for satellite video target tracking based on correlation filtering with trajectory prediction according to claim 1, wherein: The TPN trajectory prediction network includes a trajectory segmentation module, an encoder network, a temporal splicing layer, a self-attention layer, a decoder network, and a fully connected layer; among them, the trajectory segmentation module divides the input data into 3 subsequences according to time series. The encoder network includes 3 1D convolutional layers with convolutional channel numbers of 16, 64, and 32. The temporal splicing layer splices the encoded network features of the subsequences in chronological order. The self-attention layer adds the scale dot product attention value to the network features. The decoder consists of a long short-term memory network (LSTM) with 128 hidden units. The fully connected layer is a linear network layer that maps the decoded network features to 1 dimension.

5. The method for tracking satellite video targets based on correlation filtering with trajectory prediction according to claim 1, wherein: The training steps of Step 2 are as follows: Step 2-1: Obtain m historical trajectory points from the trajectory data in Step 1 in a sliding window manner as the input historical trajectory of the network Obtain the (m + 1)-th historical trajectory point as the training label for training the network; the historical trajectory is divided into three parts of sequence data in chronological order, namely and and these three parts of sequence data are respectively fed into the encoder network to obtain three sets of features tcf 1 、tcf 2 and tcf 3 ; Step 2-2: Keep the feature channel dimension unchanged, splice the three groups of features obtained in Step 2-1 in the temporal dimension. The spliced feature tcf passes through a self-attention layer to further enable the network to focus on important features. Use a linear layer to map the feature tcf to a query matrix Q, a key matrix K, and a value matrix V. Calculate the self-attention matrix using the scale dot product attention mechanism according to the following formula: Among them, softmax is the normalization function, and d k is the feature channel dimension of tcf; after the self-attention matrix attn and tcf are summed, the feature atf is obtained through a network layer with a residual connection form, and after being input into an LSTM network with 128 hidden neurons, the network prediction output is obtained through a fully connected layer; Step 2-3: Calculate the loss between the network prediction output obtained in Step 2-2 and the training label obtained in Step 2-1 according to the mean squared error loss function, and update the network weights; after the network loss converges, keep the network structure and weights unchanged to complete the training of the TPN trajectory prediction network.

6. The method for tracking satellite video targets based on correlation filtering with trajectory prediction according to claim 5, wherein: When updating the network weights in Step 2-3, use gradient descent backpropagation and use the Adam network optimizer with a learning rate of 0.001 to update the network weights.

7. The method for tracking satellite video targets by correlation filtering based on trajectory prediction according to claim 1, wherein: The process of using the TPN trajectory prediction network to predict the trajectory points of the current frame in the occluded state and using adaptive weights to determine the position of the target in Step 3-4 is as follows: Construct the trajectory sequence data according to the trajectory sequence construction method in step 1 for the historical position sequence data {pos t-1} to obtain the trajectory sequence data and Input it into the trained TPN network in step 2 to obtain the predicted positions in the x and y directions and Perform weighted fusion on the predicted trajectory position and the correlation filtering position to determine the target position. The calculation formula for the weight of the proposed predicted trajectory position is as follows: If the continuous occlusion count variable count occ exceeds 15, then the attenuation weight is set such that the attenuated part does not exceed 0.3: Target position pos t =(x t , y t ) is calculated according to the following formula:

8. The method for tracking satellite video targets based on correlation filtering with trajectory prediction according to claim 1, characterized in that: The value of the constant k that controls the threshold size in Step 3-3 is 1.3.

Citation Information

Patent Citations

  • Satellite video dynamic target tracking method fusing correlation filter and motion estimation

    CN110956653A

  • Satellite video moving vehicle tracking method based on feature enhancement and position prediction

    CN115908484A