A Satellite Video Single Object Tracking Method Combining CNN and Decoder

By combining CNN and decoder methods, feature extraction and spatiotemporal information fusion of satellite videos are solved, and the problem of low target tracking accuracy in satellite videos is achieved, achieving higher tracking accuracy and robustness.

CN119919454BActive Publication Date: 2025-07-11PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510406597.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-11
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

In satellite videos, tracking targets usually occupy only a small number of pixels, with weak features and low contrast with backgrounds, making it difficult for existing feature extraction methods to distinguish tracking targets from backgrounds, reducing the accuracy of visual tracking, especially in complex backgrounds and occlusions.

Method used

Using a method combining CNN and decoder, the current video frame and template images of satellite video are extracted through the CNN feature extraction network, and spatial feature maps are generated, and feature interaction enhancement and spatial information fusion are performed through the decoder. The deep learning model is used for prediction, and the tracking results of a single target are generated.

Benefits of technology

It improves the accuracy and accuracy of single target tracking in satellite videos, can effectively deal with the problems of changing appearance and occlusion of targets in complex scenarios, and achieves more comprehensive target recognition and tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919454B_ABST
    Figure CN119919454B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of remote sensing tracking technology, and relates to a satellite video single-object tracking method combining a CNN and a decoder. The present invention includes: extracting respective spatial feature maps of the current video frame of the satellite video and the template image of the single object through a CNN feature extraction network; generating spatial features of the current video frame according to the spatial feature maps; generating temporal features of the current video frame based on the spatial features of the current video frame and the temporal features of the previous video frame by means of a decoder; generating spatio-temporal fusion features based on the spatial and temporal features of the current video frame; and generating a tracking result of the single object by using a prediction head of a deep learning model. The present invention extracts features of the target through a CNN feature extraction network, interacts with the extracted features through a decoder to generate spatio-temporal fusion features, and uses spatio-temporal information for tracking, so as to accurately identify the target in the satellite video, thereby improving the accuracy and precision of target tracking.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Visual tracking technology is an important branch in the field of computer vision. Visual tracking refers to detecting, extracting, recognizing, and tracking moving targets in an image sequence to obtain the motion parameters of the moving targets, such as position, speed, acceleration, and motion trajectory, etc., so as to perform further processing and analysis, realize the understanding of the behavior of the moving targets, and complete a higher-level detection task.

[0003] In visual tracking based on satellite video, since the tracking targets in satellite video usually only occupy a small number of pixels, and the details of the tracking targets are few, the features are weak, and the contrast is low compared with the complex background, it is difficult for existing feature extraction methods to effectively distinguish the tracking targets from the background, reducing the accuracy of visual tracking. In particular, the satellite video has a wide coverage range, presenting complex structural features and spatial distribution patterns. At the same time, it may also be affected by factors such as atmospheric conditions and solar angles, and the situation where the tracking targets are occluded is likely to occur, further interfering with target tracking and leading to a decrease in the accuracy of visual tracking.

[0004] Therefore, there is an urgent need for a method that can improve the accuracy of satellite video tracking. Summary of the Invention

[0005] To solve the above problems in the prior art, the present invention provides a satellite video single-object tracking method combining CNN and a decoder, including:

[0006] Performing feature extraction processing on the current video frame of the satellite video and the template image corresponding to the single object respectively through a preset CNN feature extraction network to obtain the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image;

[0007] Generating the spatial feature corresponding to the current video frame according to the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image;

[0008] Based on a preset decoder, performing feature interaction enhancement processing on the spatial feature corresponding to the current video frame and the temporal feature of the previous video frame to obtain the temporal feature corresponding to the current video frame;

[0009] Performing spatio-temporal information fusion processing on the spatial feature and the temporal feature corresponding to the current video frame respectively to obtain the spatio-temporal fusion feature corresponding to the current video frame;

[0010] Using the prediction head of the deep learning model to perform prediction processing based on the spatio-temporal fusion feature, and generating the tracking result of the single object in the current video frame based on the prediction processing result.

[0011] Further, the feature extraction process for the current video frame of the satellite video and the template image corresponding to a single target by the preset CNN feature extraction network includes:

[0012] Performing local feature extraction on the current video frame and the template image respectively through the convolutional layers of the CNN feature extraction network to obtain respective corresponding local feature maps;

[0013] Performing dimensionality reduction processing on the two local feature maps respectively through the max pooling layer of the CNN feature extraction network;

[0014] Performing target feature extraction on the two local feature maps after dimensionality reduction processing respectively through the convolutional blocks of the CNN feature extraction network to obtain the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image.

[0015] Further, the performing target feature extraction on the two local feature maps after dimensionality reduction processing respectively through the convolutional blocks of the CNN feature extraction network includes:

[0016] Extracting the local feature map through the convolutional layer in the convolutional block to obtain the first feature map corresponding to the local feature map;

[0017] Performing the first addition fusion of the local feature map and the first feature map corresponding to the local feature map to obtain a fusion feature map;

[0018] Extracting the fusion feature map through the deformable convolutional layer in the convolutional block to obtain the second feature map corresponding to the local feature map;

[0019] Performing the second addition fusion of the fusion feature map and the second feature map to obtain the spatial feature map corresponding to the local feature map.

[0020] Further, the performing the first addition fusion of the local feature map and the first feature map corresponding to the local feature map to obtain a fusion feature map includes:

[0021] Using a 1×1 convolutional kernel to perform downsampling on the local feature map, and performing element-wise addition of the downsampled local feature map and the first feature map to obtain the fusion feature map.

[0022] Further, the performing the second addition fusion of the fusion feature map and the second feature map to obtain the spatial feature map corresponding to the local feature map includes:

[0023] Performing element-wise addition of the fusion feature map and the second feature map to obtain the spatial feature map.

[0024] Further, generating the spatial feature corresponding to the current video frame based on the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image includes:

[0025] Performing a concatenation process on the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image to obtain a one-dimensional feature sequence;

[0026] Performing information interaction processing on the one-dimensional feature sequence in a cross-attention manner to obtain the spatial feature corresponding to the current video frame.

[0027] Further, the performing a concatenation process on the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image includes:

[0028] Performing linear projection and reshaping on the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image in sequence;

[0029] Performing a splicing process on the two reshaped spatial feature maps, and embedding position encoding in the one-dimensional spatial feature sequence obtained by the splicing process to obtain a one-dimensional feature sequence.

[0030] Further, the performing feature interaction enhancement processing on the spatial feature corresponding to the current video frame and the temporal feature of the previous video frame based on a preset decoder includes:

[0031] Determining respective corresponding query vectors according to the spatial feature corresponding to the current video frame and the temporal feature of the previous video frame;

[0032] Inputting the query vector corresponding to the current video frame and the query vector corresponding to the previous video frame into the first multi-head attention of the decoder, and the output vector of the first multi-head attention sequentially passes through the first normalization layer and the first skip connection of the decoder to generate the input vector of the second multi-head attention of the decoder;

[0033] Inputting the input vector of the second multi-head attention and the spatial feature corresponding to the current video frame into the second multi-head attention, and the output vector of the second multi-head attention sequentially passes through the second normalization layer, the fully connected feed-forward network layer, the third normalization layer and the second skip connection to obtain the temporal feature corresponding to the current video frame.

[0034] Further, the performing spatio-temporal information fusion processing on the spatial feature and the temporal feature respectively corresponding to the current video frame includes:

[0035] Adopting a feature fusion method of dot product to fuse the spatial feature and the temporal feature respectively corresponding to the current video frame to obtain the spatio-temporal fusion feature corresponding to the current video frame.

[0036] Further, the prediction head using the deep learning model performs prediction processing based on the spatio-temporal fusion feature, and generating the tracking result of a single target in the current video frame based on the prediction processing result includes:

[0037] Input the spatio-temporal fusion feature into the prediction head of the deep learning model, and output the bounding box, offset, and multiple classification results of a single target in the current video frame through the prediction head of the deep learning model;

[0038] Select the position corresponding to the classification result with the highest classification score from the multiple classification results as the tracking position of the single target;

[0039] Perform correction processing on the tracking position through the offset and the bounding box to obtain the tracking result of a single target in the current video frame.

[0040] Advantages of the present invention:

[0041] The present invention extracts features of a single target through a CNN feature extraction network. The CNN feature extraction network can better focus on the boundaries and detailed features of the single target, and adapt to changes in the appearance of the single target, so as to effectively extract the significant features of the single target in the satellite video, thereby improving the accuracy of target tracking.

[0042] The present invention interacts with the features extracted by the CNN feature extraction network through a decoder to generate spatio-temporal fusion features. The spatio-temporal fusion features can characterize the change trajectory of a single target in all directions and from multiple angles. Based on the spatio-temporal fusion features, the single target in the satellite video can be accurately identified, thereby improving the accuracy of target tracking. In addition, using spatio-temporal information for tracking can effectively address the problems of variable appearance and occlusion of a single target in complex scenarios, thereby improving the tracking accuracy. Description of the Drawings

[0043] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more apparent:

[0044] Figure 1 is a flowchart of a method for tracking a single target in a satellite video by combining a CNN and a decoder according to an embodiment of the present invention.

[0045] Figure 2 is a flowchart of the process of extracting features by the CNN feature extraction network according to an embodiment of the present invention.

[0046] Figure 3 is a flowchart of the processing process of the convolutional block in the CNN feature extraction network according to an embodiment of the present invention.

[0047] Figure 4 It is a schematic diagram of the processing flow of the decoder in the embodiment of the present invention.

[0048] Figure 5 It is a schematic diagram of the flow of the satellite video single-object tracking method in the embodiment of the present invention.

[0049] Figure 6 It is a schematic diagram of the structure of the computer system of the server for implementing the method, system, and device embodiments of the present application. Detailed implementation manners

[0050] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention and are not intended to limit the invention. In addition, it should be noted that only the parts related to the relevant invention are shown in the drawings for the convenience of description.

[0051] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and embodiments.

[0052] In order to more clearly illustrate a satellite video single-object tracking method combining CNN and a decoder of the present invention, each step in the embodiment of the present invention will be described in detail below with reference to the accompanying drawings.

[0053] A satellite video single-object tracking method combining CNN and a decoder according to the first embodiment of the present invention, as shown in Figure 1 shown, includes steps S101 - S105, and each step is described in detail as follows:

[0054] S101: Respectively perform feature extraction processing on the current video frame of the satellite video and the template image corresponding to the single object through a preset CNN feature extraction network to obtain the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image.

[0055] In this step, it is necessary to first obtain the template image of the single object to be tracked and send it to a preset CNN (Convolutional Neural Network) feature extraction network. Input the satellite video data containing the single object to be tracked into the preset CNN feature extraction network, and perform feature extraction processing on the current video frame of the satellite video and the template image corresponding to the single object through the CNN feature extraction network to obtain the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image.

[0056] The convolutional kernels of the CNN feature extraction network slide on the image for convolutional operations, enabling the CNN feature extraction network to focus on local regions of the image, extract detailed features of small targets, and not be interfered by surrounding background information. Thus, it can more comprehensively describe the features of small targets and improve the recognition accuracy of small targets.

[0057] It should be noted that the CNN feature extraction network in this step refers to a CNN model with a specific structure that has been trained or selected before satellite video target tracking. The parameters of this model (such as the weights and biases of the convolutional kernels) may have been learned and trained on a large amount of image data.

[0058] Furthermore, as shown in Figure 2 The CNN feature extraction network includes a convolutional layer, a max pooling layer, and a convolutional block. Among them, the convolutional layer of the CNN feature extraction network extracts local features from the current video frame and the template image respectively, obtaining their corresponding local feature maps. The convolutional layer can initially extract features and extract some detailed features of small targets, such as edges and corners. The convolutional layer combines these detailed features, which helps to recognize the overall small target. Through the feature extraction method of the convolutional layer, it can more comprehensively describe the features of small targets and improve the recognition accuracy of small targets.

[0059] The max pooling layer of the CNN feature extraction network performs dimensionality reduction processing on the two local feature maps respectively. The max pooling layer can downsample the local feature maps, reduce the resolution of the local feature maps, and reduce the data volume without losing important information. For small targets, this can, to a certain extent, solve the problem that small targets occupy a small proportion in the image and are easily ignored, enabling the network to perceive the features of small targets from a more macroscopic perspective and enhancing the robustness to small targets.

[0060] The convolutional block of the CNN feature extraction network extracts target features from the two local feature maps after dimensionality reduction processing respectively, obtaining the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image. By using the convolutional block of the CNN feature extraction network to extract target features from the two local feature maps after dimensionality reduction processing, it can effectively extract target features while reducing the amount of calculation and the number of parameters, improve the quality and distinguishability of the spatial feature map, and provide strong support for subsequent tasks such as target detection, tracking, and recognition.

[0061] It should be noted that the local feature maps output by the max pooling layer sequentially pass through multiple convolutional blocks. Since the sizes of small targets in satellite videos are small, if there are too many convolutional blocks, it is easy to lose the local boundary information or semantic information of small targets, and increase the complexity and computational amount of the CNN feature extraction network. Therefore, a smaller number of convolutional blocks is used. For example, the number of convolutional blocks is 3, which helps the CNN feature extraction network to quickly process small targets in complex scenes of satellite videos and improve the tracking efficiency.

[0062] When the convolutional blocks of the CNN feature extraction network respectively perform target feature extraction on the two downsampled local feature maps, the convolutional layer in the convolutional block is used to extract the local feature maps to obtain the first feature maps corresponding to the local feature maps; the local feature maps are added and fused with the first feature maps corresponding to the local feature maps for the first time to obtain the fused feature maps; the deformable convolutional layer in the convolutional block is used to extract the fused feature maps to obtain the second feature maps corresponding to the local feature maps; the fused feature maps are added and fused with the second feature maps for the second time to obtain the spatial feature maps corresponding to the local feature maps.

[0063] See Figure 3 As shown, each convolutional block sequentially includes a convolutional layer, a convolutional layer, a skip connection, a deformable convolutional layer, a deformable convolutional layer, and a skip connection. This structural design combines the advantages of traditional convolutional layers and deformable convolutional layers. First, preliminary feature extraction is performed through traditional convolutional layers, and then more adaptable features are further extracted using deformable convolutional layers. At the same time, taking advantage of the fact that deformable convolutions are robust to target appearance changes, it is beneficial to extract the features of small targets in satellite videos. The use of skip connections helps to alleviate problems such as gradient disappearance in the training of deep networks and ensure the effective transmission of information.

[0064] It should be noted that in satellite videos, small targets may have various pose, shape, and scale changes. Deformable convolutions can adaptively adjust the sampling positions, better capture the features of small targets, have more advantages than traditional convolutions, and improve the detection and recognition capabilities of small targets.

[0065] In addition, during the skip connection process, a 1×1 convolutional kernel is used for downsampling, which can change the sizes of the two feature maps, making the sizes of the two feature maps to be added and fused match, so as to achieve element-wise addition and fusion of the two feature maps. Specifically, in the first addition and fusion, a 1×1 convolutional kernel is used to downsample the local feature map, so that the size and number of channels of the downsampled local feature map are the same as those of the first feature map corresponding to the local feature map. The downsampled local feature map and the first feature map are added element-wise to obtain a fused feature map. In the second addition and fusion, when the fused feature map passes through the deformable convolutional layer in the convolutional block, the size and number of channels of the fused feature map will not change, and there is no need to use a 1×1 convolutional kernel for downsampling, and element-wise addition can be directly performed. That is, the fused feature map and the second feature map are added element-wise to obtain the spatial feature map.

[0066] It can be understood that element-wise addition and fusion is a common feature fusion method. By adding the corresponding elements of the feature maps in different paths, features at different levels and of different types can be fused, enriching the feature representation, so as to obtain the target template image features respectively through multiple convolutional blocks in sequence.

[0067] S102: Generate the spatial feature corresponding to the current video frame according to the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image.

[0068] In this step, the spatial feature maps of the current video frame and the template image are comprehensively utilized. The template image usually contains the typical features of the known tracking target, while the current video frame reflects the actual situation in the current scene. By combining the two, the prior knowledge of each tracking target in the template image and the real-time information in the current video frame can be fully utilized to describe the target in the current video frame more comprehensively and accurately, improving the recognition ability of the tracking target features. In addition, the spatial feature map corresponding to the template image can be used as a benchmark to be compared and fused with the spatial feature map of the current video frame, so as to better adapt to the changes of the tracking target. For example, when the target gradually approaches or moves away from the camera in the video, its scale will change. By combining the features of the template image, this scale change information can be more accurately captured, generating spatial features that can adapt to the dynamic changes of the target, and improving the tracking and recognition effect of the target.

[0069] In this step, when generating the spatial feature corresponding to the current video frame according to the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image, the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image are connected to obtain a one-dimensional feature sequence, which specifically includes:

[0070] Perform linear projection and reshaping on the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image in sequence.

[0071] It is understandable that the essence of linear projection is to map input data from one vector space to another through matrix multiplication, which is usually used to change the dimension of data. Its core is to learn a weight matrix to re-express the input through linear transformation or linear transformation plus bias. By performing linear projection on the spatial feature map, the number of channels of the spatial feature map can be changed from the current C to the required D, so as to flexibly adjust the dimension of the spatial feature map according to the actual situation, making it better adapt to subsequent network layers or specific task requirements, thereby improving the accuracy and efficiency of tracking.

[0072] Furthermore, the two reshaped spatial feature maps are concatenated, and positional encoding is embedded in the one-dimensional spatial feature sequence obtained by the concatenation process to obtain a one-dimensional feature sequence. Specifically, the two reshaped spatial feature maps are flattened and sorted to form a one-dimensional sequence, and the one-dimensional sequences of the two spatial feature maps are directly concatenated. Positional encoding is added to the concatenated one-dimensional spatial feature sequence to obtain a one-dimensional feature sequence.

[0073] Furthermore, information interaction processing is performed on the one-dimensional feature sequence through cross-attention to obtain the spatial features corresponding to the current video frame. Specifically, the one-dimensional feature sequence is input into cross-attention, and spatial feature interaction is performed on the one-dimensional feature sequence through the cross-attention mechanism to achieve effective interaction of the spatial information between the current video frame and the template image, and the spatial features based on the current video frame are obtained.

[0074] It should be noted that cross-attention is a mechanism widely used in deep learning, especially in the Transformer architecture, to enhance the understanding of the relationships between different elements and plays a key role in many fields such as natural language processing and computer vision.

[0075] S103: Feature interaction enhancement processing is performed on the spatial features corresponding to the current video frame and the temporal features of the previous video frame based on a preset decoder to obtain the temporal features corresponding to the current video frame.

[0076] In this step, the spatial features corresponding to the current video frame and the temporal features of the previous video frame are input into the decoder, and the received spatial features and temporal features are interacted and enhanced by the preset decoder. This process is repeated N times, that is, the output of the decoder is re-input into the decoder for interaction and enhancement, and iterated N times. After iterating N times, the decoder outputs the temporal features of the current video frame based on spatial information and historical time information.

[0077] By combining the spatial features of the current video frame and the temporal features of the previous video frame, the spatio-temporal correlation in satellite video data can be fully exploited. The spatial features can capture the feature information such as the shape, texture, and position of the tracking target in the current frame, while the temporal features can reflect the feature information such as the motion changes and action continuation of the tracking target. Additionally, using the temporal features of the previous video frame can provide temporal coherence for the tracking process of the current video frame, and better handle problems such as motion blur, occlusion, and light changes in the video. Even if the feature information in the current video frame is not clear due to various reasons, through interaction with the temporal features of the previous frame, these feature information can be restored or supplemented to a certain extent, with stronger robustness. Compared with only using a single spatial feature or temporal feature, this fusion and enhancement method can give full play to the advantages of spatio-temporal information, thereby improving the tracking accuracy.

[0078] Specifically, referring to Figure 4 as shown, based on a preset decoder, feature interaction enhancement processing is performed on the spatial features corresponding to the current video frame and the temporal features of the previous video frame, including:

[0079] Determine the respective corresponding query vectors according to the spatial features corresponding to the current video frame and the temporal features of the previous video frame.

[0080] In the cross-attention layer, each position of the decoder generates a query vector (Q) to calculate the attention weights at all positions of the encoder. Specifically, for the query vector of the current frame, by learning a weight matrix and multiplying the weight matrix with the spatial features, the query vector of the current frame is obtained; for the query vector of the previous frame, it is a vector formed by concatenating the outputs of the temporal decoder when tracking the previous N video frames.

[0081] It should be noted that there are two types of queries in the decoder in this step. One is the query vector Q passed down from the previous video frame out , and the other is the query vector Q learned from the current video frame cur , where the query vector Q cur can describe the state of the target in the current video frame and can synthesize spatio-temporal information. The decoder input is composed of the query vector Q all and the query vector Q cur . The query vector Q all is obtained by the concatenation operation of the query vector Q pre and the query vector Q out . Among them, the query vector Q pre refers to the N query vectors Q output by the decoder when tracking N video frames before the previous video frame (excluding the previous video frame) out concatenated to obtain Q pre .

[0082] Input the query vector corresponding to the current video frame and the query vector corresponding to the previous video frame into the first multi-head attention of the decoder. The output vector of the first multi-head attention sequentially passes through the first normalization layer and the first skip connection of the decoder to generate the input vector of the second multi-head attention of the decoder.

[0083] Take the query vector Q all as the key vector (K) and value vector (V) of the first multi-head attention. Input the query vector Q cur as the query Q of the first multi-head attention into the first multi-head attention of the decoder together. The output vector of the first multi-head attention sequentially passes through the first normalization layer and the first skip connection to form the query Q of the next multi-head attention, that is, generate the input vector of the second multi-head attention of the decoder.

[0084] Input the input vector of the second multi-head attention and the spatial features corresponding to the current video frame into the second multi-head attention. The output vector of the second multi-head attention sequentially passes through the second normalization layer, the fully connected feed-forward network layer, the third normalization layer and the second skip connection of the decoder to obtain the temporal features corresponding to the current video frame.

[0085] Take the spatial features corresponding to the current video frame as the key vector (K) and value vector (V) of the second multi-head attention. Input the output of the first multi-head attention as the query Q of the second multi-head attention into the second multi-head attention of the decoder together. The output vector of the second multi-head attention sequentially passes through the second normalization layer, the fully connected feed-forward network layer (FFN layer), the third normalization layer and the second skip connection of the decoder. Execute the foregoing process N times to obtain the temporal features corresponding to the current video frame.

[0086] S104: Perform spatio-temporal information fusion processing on the spatial features and temporal features respectively corresponding to the current video frame to obtain the spatio-temporal fusion features corresponding to the current video frame.

[0087] In this step, individual spatial features or temporal features can only reflect partial information of the video frame, while spatio-temporal information fusion aims to combine these two features, make full use of their respective advantages, and more comprehensively and accurately describe the content of the video frame. Thereby improving the accuracy and robustness of target tracking.

[0088] Specifically, adopt the feature fusion method of dot product to fuse the spatial features and temporal features respectively corresponding to the current video frame to obtain the spatio-temporal fusion features corresponding to the current video frame.

[0089] After spatio-temporal information fusion processing, the spatio-temporal fusion feature corresponding to the current video frame is obtained. The spatio-temporal fusion feature contains the comprehensive information of the current video frame in both the spatial and temporal dimensions, and can represent the content of the current video frame more comprehensively. Thus, it can more accurately judge the action features of the tracked object in the current video frame.

[0090] S105: Use the prediction head of the deep learning model to perform prediction processing based on the spatio-temporal fusion feature, and generate the tracking result of the single object in the current video frame based on the prediction processing result.

[0091] In this step, the prediction result output by the prediction head represents the prediction of the state of the tracked object. That is, information such as the bounding box, offset, and multiple classification results of the tracked object are output. Based on these prediction results, further processing is performed to generate the tracking result of the single object. Thus, accurate tracking of the single object in the current video frame is achieved.

[0092] Specifically, input the spatio-temporal fusion feature into the prediction head of the learning model, and output the bounding box, offset, and multiple classification results of the single object in the current video frame through the prediction head of the deep learning model. Among them, the prediction head takes the spatio-temporal fusion feature as input, and uses its internal network structure and parameters to calculate and reason about the spatio-temporal fusion feature. During the calculation process, the prediction head will learn the relationship between the spatio-temporal fusion feature and the state of the tracked object. Through operations such as matrix operations and activation functions, the prediction head outputs the prediction result, and the prediction result reflects its estimation of the state of the tracked object in the current video frame.

[0093] Select the position corresponding to the classification result with the highest classification score from the multiple classification results as the tracking position of the single object, and perform correction processing on the tracking position through the offset and the bounding box to obtain the tracking result of the single object in the current video frame.

[0094] Among them, the classification result output by the prediction head is the probability distribution of the tracked object belonging to different categories. Select the position corresponding to the category with the highest classification score (i.e., the highest probability), and identify it as the position of the tracked object in the current frame. After determining the approximate position of the tracked object (determined by the position with the highest classification score), combined with the offset O and the bounding box B, the final position and range of the tracked object can be further accurately calculated. The offset O is used to fine-tune the position of the tracked object to make it more accurately correspond to the actual position of the tracked object in the current frame; the bounding box B determines the range of the tracked object. By integrating these two pieces of information, a more accurate bounding box can be generated, and the tracking object position and range determined by this bounding box are the final tracking result, which can accurately represent the position of the tracked object in the current video frame.

[0095] It should be noted that in a deep learning model, a prediction head is a component or module in the model architecture, usually located at the end of the model, and is used to make predictions for specific tasks based on the features extracted by the model, and output prediction results related to the tasks, such as the category, location, bounding box information of the target, or the category to which each pixel in the image belongs, etc.

[0096] As can be seen from the above description, in the present invention, a CNN feature extraction network is used to extract features of a single target. The CNN feature extraction network can better focus on the boundaries and detailed features of the single target, and adapt to changes in the appearance of the single target, so as to effectively extract the significant features of the single target in the satellite video, and then improve the accuracy of target tracking. The decoder interacts with the features extracted by the CNN feature extraction network to generate spatio-temporal fusion features. The spatio-temporal fusion features can represent the change trajectory of the single target in all directions and from multiple angles. Based on the spatio-temporal fusion features, the single target in the satellite video can be accurately identified, and then the accuracy of target tracking can be improved. In addition, using spatio-temporal information for tracking can effectively handle the problems of the changing appearance of a single target and the occlusion of a single target in complex scenes, and then improve the tracking accuracy.

[0097] In the above embodiments, although the various steps are described in the above order, those skilled in the art can understand that in order to achieve the effects of the present embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reversed order, and these simple changes are all within the protection scope of the present invention.

[0098] An embodiment of the present invention also provides an embodiment of the full process of a satellite video single-target tracking method containing the above combination of CNN and decoder, see Figure 5 as shown, including:

[0099] Step 1: Input the target template to be tracked and the satellite video into a CNN (Convolutional Neural Network) feature extraction network.

[0100] In this step, the target template to be tracked and the Nth frame of the satellite video are input into the CNN feature extraction network, and the CNN feature extraction network performs feature extraction.

[0101] Step 2: Based on the CNN feature extraction network for small targets, extract features from the target image and the Nth frame to obtain their spatial feature maps.

[0102] In this step, the weights of the CNN feature extraction network for extracting features from the target template and the Nth frame are shared. When extracting features from the Nth frame, in order to reduce the computational amount, a search area can be set on the Nth frame, and feature extraction is performed within the search area.

[0103] Step 3: Input the spatial feature map into the cross-attention mechanism.

[0104] In this step, the spatial feature map obtained in Step 2 is reshaped and linearly projected to obtain two one-dimensional feature sequences. The two one-dimensional feature sequences are concatenated, and the concatenated sequence is added with position encoding. The feature sequence after adding the position encoding is input into the cross-attention mechanism for spatial feature interaction to obtain the spatial feature of the Nth frame.

[0105] Step 4: Input the spatial features obtained in Step 3 into the temporal decoder and the spatio-temporal information fusion module respectively.

[0106] Step 5: The temporal decoder receives the spatial features from Step 3 and the temporal features of the previous video frame (the (N - 1)th frame). The temporal decoder interacts and enhances these two features, and this process is repeated N times to output the temporal features of the Nth frame based on spatial information and historical temporal information.

[0107] Step 6: The spatio-temporal information fusion module performs feature fusion on the spatial features from Step 3 and the temporal features from Step 5 and inputs the result into the prediction head.

[0108] Step 7: The prediction head classifies and regresses the bounding box of the fused features from Step 6 to determine the position of the target tracking, that is, to output the tracking result.

[0109] The second embodiment of the present invention provides a satellite video single-object tracking system combining a CNN and a decoder, including:

[0110] A feature extraction module, configured to perform feature extraction processing on the current video frame of the satellite video and the template image corresponding to the single object respectively through a preset CNN feature extraction network to obtain the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image;

[0111] A feature module, configured to generate the spatial features corresponding to the current video frame according to the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image;

[0112] An interaction module, configured to perform feature interaction and enhancement processing on the spatial features corresponding to the current video frame and the temporal features of the previous video frame based on a preset decoder to obtain the temporal features corresponding to the current video frame;

[0113] A fusion module, configured to perform spatio-temporal information fusion processing on the spatial features and temporal features corresponding to the current video frame respectively to obtain the spatio-temporal fusion features corresponding to the current video frame;

[0114] A prediction module, configured to perform prediction processing based on the spatio-temporal fusion features by using a prediction head of a deep learning model, and generate a tracking result of a single target in the current video frame based on the prediction processing result.

[0115] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and related descriptions of the above-described system can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0116] It should be noted that the satellite video single-target tracking system combining a CNN and a decoder provided in the above embodiments is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the modules or steps in the embodiments of the present invention can be further decomposed or combined. For example, the modules in the above embodiments can be combined into one module, or further split into multiple sub-modules to complete all or part of the functions described above. The names of the modules and steps involved in the embodiments of the present invention are only for distinguishing each module or step, and are not regarded as an improper limitation of the present invention.

[0117] An electronic device according to a third embodiment of the present invention includes:

[0118] At least one processor;

[0119] And a memory communicatively connected to at least one of the processors;

[0120] Wherein, the memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the above-mentioned satellite video single-target tracking method combining a CNN and a decoder.

[0121] A computer-readable storage medium according to a fourth embodiment of the present invention, the computer-readable storage medium stores computer instructions, and the computer instructions are used to be executed by the computer to implement the above-mentioned satellite video single-target tracking method combining a CNN and a decoder.

[0122] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and related descriptions of the above-described storage device and processing device can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0123] Those skilled in the art should be able to realize that the modules and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. The programs corresponding to the software modules and method steps can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the technical field. To clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in the form of electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0124] Reference is made below to Figure 6 , which shows a schematic structural diagram of a computer system of a server for implementing the method, system, and device embodiments of the present application. Figure 6 The server shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0125] As Figure 6 shown, the computer system includes a central processing unit (CPU, Central Processing Unit) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM, Read Only Memory) 602 or the program loaded from the storage section 608 into the random access memory (RAM, Random Access Memory) 603. In the RAM 603, various programs and data required for system operation are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other through a bus 604. The input / output (I / O, Input / Output) interface 605 is also connected to the bus 604.

[0126] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 610 as needed so that a computer program read therefrom is installed into the storage section 608 as needed.

[0127] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above functions defined in the method of the present application are executed. It should be noted that the computer-readable medium in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0128] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0129] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0130] The terms "first", "second", etc. are used to distinguish similar objects and not to describe or represent a specific order or sequence.

[0131] The term "comprising" or any other similar term is intended to cover non-exclusive inclusion, such that a process, method, article, or device / equipment that comprises a series of elements includes not only those elements but also other elements not expressly listed, or also includes elements inherent to those processes, methods, articles, or devices / equipment.

[0132] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easily understood by those skilled in the art that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.

Claims

1. A satellite video single-object tracking method combining CNN and a decoder, characterized in that, Including: Performing feature extraction processing on the current video frame of the satellite video and the template image corresponding to a single target respectively through a preset CNN feature extraction network to obtain the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image; Generating the spatial feature corresponding to the current video frame according to the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image; Performing feature interaction enhancement processing on the spatial feature corresponding to the current video frame and the temporal feature of the previous video frame based on a preset decoder to obtain the temporal feature corresponding to the current video frame; Performing spatio-temporal information fusion processing on the spatial feature and the temporal feature respectively corresponding to the current video frame to obtain the spatio-temporal fusion feature corresponding to the current video frame; Using the prediction head of the deep learning model to perform prediction processing based on the spatio-temporal fusion feature, and generating the tracking result of the single target in the current video frame based on the prediction processing result; The performing feature extraction processing on the current video frame of the satellite video and the template image corresponding to a single target respectively through a preset CNN feature extraction network includes: Performing local feature extraction on the current video frame and the template image respectively through the convolutional layer of the CNN feature extraction network to obtain the respective corresponding local feature maps; Performing dimensionality reduction processing on the two local feature maps respectively through the max pooling layer of the CNN feature extraction network; Performing target feature extraction on the two local feature maps after dimensionality reduction processing respectively through the convolutional blocks of the CNN feature extraction network to obtain the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image; The performing target feature extraction on the two local feature maps after dimensionality reduction processing respectively through the convolutional blocks of the CNN feature extraction network includes: Extracting the local feature map through the convolutional layer in the convolutional block to obtain the first feature map corresponding to the local feature map; Performing the first addition fusion on the local feature map and the first feature map corresponding to the local feature map to obtain a fusion feature map; Extracting the fusion feature map through the deformable convolutional layer in the convolutional block to obtain the second feature map corresponding to the local feature map; Performing the second addition fusion on the fusion feature map and the second feature map to obtain the spatial feature map corresponding to the local feature map; The performing feature interaction enhancement processing on the spatial feature corresponding to the current video frame and the temporal feature of the previous video frame based on a preset decoder includes: Determining the respective corresponding query vectors according to the spatial feature corresponding to the current video frame and the temporal feature of the previous video frame; Inputting the query vector corresponding to the current video frame and the query vector corresponding to the previous video frame into the first multi-head attention of the decoder, and the output vector of the first multi-head attention sequentially passes through the first normalization layer and the first skip connection of the decoder to generate the input vector of the second multi-head attention of the decoder; Input the input vector of the second multi-head attention and the spatial features corresponding to the current video frame into the second multi-head attention. The output vector of the second multi-head attention sequentially passes through the second normalization layer, the fully-connected feed-forward network layer, the third normalization layer, and the second skip connection of the decoder to obtain the temporal features corresponding to the current video frame.

2. The satellite video single-object tracking method combining CNN and decoder according to claim 1, wherein, The first addition fusion of the local feature map and the first feature map corresponding to the local feature map to obtain a fused feature map includes: Use a 1×1 convolutional kernel to downsample the local feature map, and perform element-wise addition on the downsampled local feature map and the first feature map to obtain the fused feature map.

3. The satellite video single-object tracking method combining CNN and a decoder according to claim 1, characterized in that, The second addition fusion of the fused feature map and the second feature map to obtain the spatial feature map corresponding to the local feature map includes: Perform element-wise addition on the fused feature map and the second feature map to obtain the spatial feature map.

4. The satellite video single-object tracking method combining CNN and decoder according to claim 1, characterized in that, The generation of the spatial features corresponding to the current video frame based on the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image includes: Perform connection processing on the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image to obtain a one-dimensional feature sequence; Perform information interaction processing on the one-dimensional feature sequence in a cross-attention manner to obtain the spatial features corresponding to the current video frame.

5. The satellite video single-object tracking method combining CNN and decoder according to claim 3, wherein The connection processing of the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image includes: Perform linear projection and reshaping on the spatial feature map corresponding to the current video frame and the spatial feature map corresponding to the template image in sequence; Perform splicing processing on the two reshaped spatial feature maps, and embed position encoding in the one-dimensional spatial feature sequence obtained by the splicing processing to obtain a one-dimensional feature sequence.

6. The satellite video single-object tracking method combining CNN and a decoder according to claim 1, characterized in that The spatio-temporal information fusion processing of the spatial features and temporal features respectively corresponding to the current video frame includes: Use the dot product feature fusion method to fuse the spatial features and temporal features respectively corresponding to the current video frame to obtain the spatio-temporal fusion features corresponding to the current video frame.

7. The satellite video single-object tracking method combining CNN and a decoder according to claim 1, characterized in that, The prediction processing using the prediction head of the deep learning model based on the spatio-temporal fusion features, and the generation of the tracking result of a single target in the current video frame based on the prediction processing result includes: Input the spatio-temporal fusion features into the prediction head of the deep learning model, and output the bounding box, offset, and multiple classification results of a single target in the current video frame through the prediction head of the deep learning model; Select the position corresponding to the classification result with the highest classification score from the multiple classification results as the tracking position of the single target; Perform correction processing on the tracking position through the offset and the bounding box to obtain the tracking result of a single target in the current video frame.

Citation Information

Patent Citations

  • Single-target long-time tracking method

    CN115187799A

  • Target tracking method of low-altitude aircraft visual angle

    CN115880332A

  • Video action recognition method and device, electronic equipment and storage medium

    CN119380415A