Two-stage feature fusion target tracking method fusing time information
Through the method of combining two-stage feature fusion and time information, the problem of weak target appearance deformation and similar interfering objects in drone target tracking is solved, and a higher tracking accuracy and success rate are achieved.
Patent Information
- Application Number
- CN202510593800.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-15
AI Technical Summary
There is a problem in the tracking of drone targets that have weak distinction ability between target appearance deformation and similar interferers. The existing methods are difficult to effectively distinguish targets from similar interferers, resulting in tracking failure.
The two-stage feature fusion method is adopted to extract features through the local information aggregation module and the residual feature fusion network, and combine time information branches to enhance feature interaction and target details capture capabilities, use self-attention and cross-attention mechanisms to perform feature fusion, and add residual structure and time information filters.
It improves the success rate and accuracy of drone target tracking, enhances the ability to discriminate similar interferers, reduces the risk of target loss, and improves the tracking accuracy in complex scenarios.
Smart Images

Figure CN120495956A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of unmanned aerial vehicle (UAV) target tracking, and in particular relates to a two-stage feature fusion target tracking method that fuses time information. Background Art
[0002] Vision-based drone target tracking is currently a hot research topic in the drone field and is gaining increasing application in various fields, including civil, military, and scientific research. Due to the complexity of application scenarios, drones can encounter interference from similar objects, target motion, and occlusion during target tracking. Currently, a backbone network combined with a Transformer approach is used for drone target tracking. This single-stream, single-stage target tracking method allows for early information interaction between the search area and the template, making it effective for tracking features that generate specific targets. For example, the application document with the number "CN202411563870.X" discloses "A Target-Aware Transformer Drone Tracking Method Focusing on Key Information." This method utilizes a single-stream tracking framework that integrates feature learning and target search to improve information interaction between tokens. Furthermore, an adaptive relationship modeling mechanism and a multi-layer feature aggregation module are used to further enhance the discriminative power of feature representation. However, the single-stream framework still suffers from incorrect feature interaction. In addition, another document with the application number "CN202410524415.2" discloses "a temporal multi-scale target tracking method based on Transformer", which attempts to reconstruct the target from different time scales by fusing features of multiple time scales and utilizing the global information capture capability of Transformer, thereby improving the tracking success rate in scenarios such as interference from similar objects, deformation, and occlusion. However, due to the global attention mechanism of Transformer, it is also easy to learn a large amount of redundant information during the re-learning process, which pollutes the features and affects subsequent tracking. It can be seen from the above document that: due to the incorrect feature interaction in the Transformer backbone network and the fact that the extracted features are not effective enough, it is difficult to distinguish the target from similar interferers. Therefore, when applied to UAV target tracking tasks, there is a problem that the target appearance is deformed, the target is easily lost, and the ability to distinguish between similar interferers is weak. Summary of the Invention
[0003] The present invention proposes a two-stage feature fusion target tracking method that integrates temporal information to overcome the problems in the prior art of target appearance deformation, easy loss of target, and weak ability to distinguish similar interferers.
[0004] To achieve the above object, the technical solution of the present invention is as follows: a two-stage feature fusion target tracking method integrating time information, comprising the following steps:
[0005] Step 1: Split the template and search area into image block sequences, map the one-dimensional position code, and input it into the backbone network integrated into the local information aggregation module to extract the first-stage features;
[0006] Step 2: The first-stage features extracted in step 1 are re-divided into template features and search area features, and input into the constructed residual feature fusion network for interactive feature fusion to obtain the second-stage features. At the same time, after feature fusion is performed in the CFA module of the second feature fusion layer in the residual feature fusion network, a copy of the search area features is saved; then the residual structure is introduced to realize the two-stage feature fusion and obtain the two-stage fusion features;
[0007] Step 3: Send the search area features saved in step 2 and the search area features re-divided by the first stage features in the next frame into the time information branch to capture time information;
[0008] Step 4: Adaptively fuse the time information with the two-stage fusion features obtained in step 2 and send them to the prediction head for tracking.
[0009] Furthermore, in the above step 2, the residual feature fusion network is constructed by designing two layers of feature fusion layers that are serially spliced, and then connecting a separate CFA module and a residual structure in series after feature fusion. Each feature fusion layer is constructed by arranging two groups of ECA modules containing a self-attention mechanism and CFA modules containing a cross-attention mechanism in parallel;
[0010] The specific calculations of the ECA module with self-attention mechanism and the CFA module with cross-attention mechanism are as follows:
[0011] ECA module with self-attention mechanism:
[0012] X ECA =Norm(X+MultiHead(X+P x ,X+P x ,X)) (1)
[0013] In formula (1), X represents the input of the ECA module, and the sine and cosine functions are used to generate the spatial position code P x , X ECA Represents the output of the ECA module;
[0014] CFA module with cross attention mechanism:
[0015]
[0016] In the above formula, X q and X kv Represents two different inputs, Pq and P k represents the spatial position encoding corresponding to two different inputs, Norm represents layer normalization, FFN represents feedforward network, X CFA Represents the output of the CFA module.
[0017] Furthermore, in the above step 2, the method of introducing the residual structure to realize the two-stage feature fusion is shown in formula (4);
[0018] F=(F1+F2) / 2 (4)
[0019] In the above formula, F is the two-stage fusion feature; F1 is the first-stage feature; F2 is the second-stage feature; and finally the two are fused by summing and averaging them.
[0020] Furthermore, the adaptive fusion method in the above step 4 is shown in the following formula (8):
[0021] f=wf1+(1-w)f2(8)
[0022] In the above formula, f is the feature after fusing the time information feature and the output of the feature fusion network, f1 is the feature output by the feature fusion network, f2 is the time information feature output by the constructed time information branch, and w is a learnable parameter.
[0023] Compared with the prior art, the present invention has the following beneficial effects:
[0024] 1. This paper constructs a Transformer module that combines global information and local information. By building a local information aggregation module in the conventional Transformer structure, it reduces the duplication and redundancy of certain features and enhances the traditional VIT's ability to capture local information.
[0025] 2. The residual feature fusion network constructed by the present invention combines the features extracted by the backbone network and the second-stage features output by the feature interaction fusion for learning, so that the model pays attention to the effective detail information of the target in the feature interaction fusion stage, thereby improving the algorithm's ability to discriminate similar interferers.
[0026] 3. The present invention adds an additional time information branch and captures the time information through the features between the previous and next frames, so that the target appearance is not easily deformed.
[0027] 4. This method effectively improves the success rate and accuracy of the target tracking method under challenges such as similar interference objects and target deformation through the interactive fusion of two-stage feature extraction. It is not easy to lose the target and can accurately distinguish similar interference objects with strong resolution ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is the overall flow chart of the method of the present invention;
[0029] Figure 2 This is the structural diagram of the local information aggregation module;
[0030] Figure 3 This is the residual feature fusion network structure diagram;
[0031] Figure 4 The diagram shows the visualization results of the method of the present invention and the comparative method. DETAILED DESCRIPTION
[0032] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0033] See also Figure 1 The present invention provides a two-stage feature fusion target tracking method that integrates time information. The method extracts features through a backbone network that incorporates local information and uses a residual feature fusion network to fuse features. The features are processed in two stages and then fed together with the time information into a prediction head to achieve tracking. The method specifically includes the following steps:
[0034] Step 1: Split the template and search area into image block sequences, map the one-dimensional position code, input it into the backbone network integrated into the local information aggregation module, and extract the features of the first stage. The specific operations are as follows:
[0035] 1.1: Split the template and search area into image blocks, reconstruct them into one-dimensional positional encoding, and use the concatenated result as input for subsequent steps.
[0036] 1.2: First, design a backbone network that incorporates a local information aggregation module: construct a local information aggregation module by using deep convolution, such as Figure 2 As shown in the figure, the local information aggregation module is inserted between the multi-head attention mechanism layer and the LayerNorm layer in the first layer of the Transformer module of the backbone network constructed by the conventional 12-layer Transformer module, and the other 11 layers remain unchanged, thereby constructing a Transformer backbone network that combines global information and local information. It is used as the backbone network of this method for feature extraction;
[0037] The concatenated results from step 1.1 are then fed into the backbone network. After the input encoded information passes through the first layer of multi-head attention and before entering the local information aggregation module, it is divided into a template and a search region patch embedding as input to the local information aggregation module. This is reshaped into a two-dimensional response map, and two separable convolutions with residual structures are used to aggregate local information for the template and search region, respectively. The template and search region 2D response maps, after aggregating local information, are then reshaped back into a one-dimensional template patch embedding sequence and a search region patch embedding sequence, respectively. After concatenation, these are fed into the subsequent LayerNorm layer and the remaining modules.
[0038] Finally, after the backbone network extraction, the extracted first-stage features are used as the input of step two.
[0039] Step 2: The first-stage features extracted in step 1 are re-divided into template features and search area features, and then input into the residual feature fusion network constructed by this method for interactive feature fusion to obtain the second-stage features. At the same time, after feature fusion is performed in the CFA module of the second feature fusion layer in the residual feature fusion network, a copy of the search area features is saved. The residual structure is then introduced to achieve two-stage feature fusion and obtain two-stage fused features. The specific operation is as follows:
[0040] 2.1: First, the first-stage features output in step 1 are divided again into template features and search area features, and then input into the residual feature fusion network constructed by this method for feature fusion.
[0041] like Figure 3 As shown, the residual feature fusion network constructed by the present invention has two layers of feature fusion layers serially spliced, and after feature fusion, a separate CFA module and a residual structure are connected in series. Each feature fusion layer is composed of two groups of ECA modules containing self-attention mechanism and CFA modules containing cross-attention mechanism arranged in parallel.
[0042] The self-attention mechanism ECA module uses multi-head self-attention with a residual structure to enhance the global context. Its specific calculation process is shown in formula (1).
[0043] X ECA =Norm(X+MultiHead(X+P x ,X+P x ,X)) (1)
[0044] In formula (6), X represents the input of the ECA module. In addition, since the attention mechanism is unable to distinguish the position information of the input feature sequence, it uses sine and cosine functions to generate the spatial position code P x .XECA Represents the output of the ECA module.
[0045] The Cross-Attention Mechanism (CFA) module uses a residual multi-head cross-attention to fuse feature vectors from two inputs. Similar to ECA, CFA also uses spatial position encoding and also uses the FFN module to enhance the model's fitting capabilities. The CFA calculation process is shown in Equations (2) and (3).
[0046]
[0047] In the above formula, X q and X kv Represents two different inputs, P q and P k Represents the spatial position encoding corresponding to two different inputs, Norm represents layer normalization, and FFN represents feedforward network. CFA Represents the output of the CFA module.
[0048] 2.2: At the same time, after the search area features pass through the CFA module of the second feature fusion layer, a separate copy of the search area features is saved.
[0049] 2.3: After the features pass through the two feature fusion layers, a CFA module is used to decode the template branch features and the search area branch features. At the same time, a residual structure is introduced to fuse the features of the feature extraction stage and the feature interaction stage to obtain the two-stage fusion features. The specific fusion method is shown in formula (4):
[0050] F=(F1+F2) / 2(4)
[0051] In the above formula, F is the two-stage fusion feature, F1 is the first-stage feature, and F2 is the second-stage feature. Finally, the two are fused by summing and averaging them to obtain more discriminative features.
[0052] Step 3: The search area features saved in step 2 and the search area features re-divided by the first stage features in the next frame are sent to the time information branch to capture the time information. The time information is adaptively fused with the output of the residual feature fusion network and sent to the prediction head for tracking. The specific operations are as follows:
[0053] 3.1: The temporal information branch designed in this invention consists of an adaptive temporal Transformer module and two ECA modules.
[0054] First, an adaptive temporal Transformer module is constructed. The search area features saved in step 2 are used to interact with the search area features output by the backbone network in step 1 in the next frame. The temporal information between the two frames is captured. At the same time, redundant context is filtered out through the temporal information filter. The specific calculation process is shown in formula (5):
[0055]
[0056] In the above formula, F t Represents the current search frame feature, F t-1 Represents the search frame features of the previous frame, Norm represents layer normalization, and MultiHead represents the multi-head attention mechanism.
[0057] By attaching a feed-forward network FFN to the GAP obtained by global average pooling The global descriptor of is used to generate a concise temporal information filter, namely Filtered time information Obtained by formula (6):
[0058]
[0059] In the above formula, f represents the convolution layer, Cat represents the concatenation operation, and * represents the dot product. The final obtained time information feature is shown in formula (7):
[0060]
[0061] In the above formula, F O is the temporal information feature finally output by the adaptive temporal Transformer module, Represents the filtered time information, Norm represents layer normalization, and MultiHead represents the multi-head attention mechanism.
[0062] 3.2: After the template and search area features pass through the adaptive temporal Transformer module, they are corrected by two ECA modules to obtain more reliable temporal information. Therefore, this method designs a multi-search frame training method to train the temporal information branch:
[0063] First, multiple search frame images are obtained by sampling at certain frame intervals in the dataset. The search frames collected in each dataset are arranged in chronological order as training samples and fed into the model for training along with the template: First, the template and search frame 1 are fed into the network constructed in the first two steps, and the search area branch features of the second layer in the residual feature fusion network are saved for the next frame calculation. Then, the above operation is repeated for the template and search frame 2. At the same time, the search frame 2 features extracted by the Transformer backbone network and the search area branch features in the saved residual feature fusion network of the previous frame are interacted through the proposed temporal information branch, and the captured temporal information is adaptively fused with the fusion features output by the residual feature fusion network. The process of template and search frame 2 is repeated for subsequent search frames to achieve the implementation of the temporal information branch.
[0064] Step 4: Adaptively fuse the time information obtained in step 3 with the two-stage fusion features obtained in step 2. The fusion method is shown in formula (8).
[0065] f=wf1+(1-w)f2(8)
[0066] In the above formula, f is the feature after fusing the temporal information feature and the output of the feature fusion network, f1 is the feature output by the feature fusion network, f2 is the temporal information feature output by the constructed temporal information branch, w is a learnable parameter, and finally the fused feature is sent to the prediction head for tracking.
[0067] like Figure 4 Figure 2 shows a comparison of experimental results. The OSTrack target tracking method and the present invention are used for visual comparison, using the DTB70 dataset. (A) represents the target tracking method designed in this paper, and (B) represents the OSTrack tracking method. It can be seen that in some frames, OSTrack tracking fails due to the target moving out of the frame or deforming. However, the present method uses two-stage feature fusion, allowing the model to focus on more detailed information. Furthermore, by capturing temporal information, it can better cope with changes in target appearance, further improving tracking accuracy.
[0068] In addition to comparing the proposed method with other methods, the proposed method also conducts comparative experiments with common target tracking algorithms: OSTrack, AbaViTrack, HIFT, TCTrack, and other high-performance target tracking methods. Table 1 shows the test results of the proposed method and OSTrack and other target tracking methods on the UAV123 and DTB70 datasets. Both tables show that the proposed method improves tracking accuracy and success rate in common drone scenarios.
[0069] Table 1 Comparison results on two datasets
[0070]
[0071] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention are still within the scope of protection of the technical solution of the present invention.
Claims
1. A two-stage feature fusion target tracking method integrating temporal information, characterized by: The following steps are involved: Step 1: Split the template and search area into image block sequences, map the one-dimensional position code, and input it into the backbone network integrated into the local information aggregation module to extract the first-stage features; Step 2: The first-stage features extracted in step 1 are re-divided into template features and search area features, and input into the constructed residual feature fusion network for interactive feature fusion to obtain the second-stage features. At the same time, after feature fusion is performed in the CFA module of the second feature fusion layer in the residual feature fusion network, a copy of the search area features is saved; Then the residual structure is introduced to realize the two-stage feature fusion and obtain the two-stage fusion features; Step 3: Send the search area features saved in step 2 and the search area features re-divided by the first stage features in the next frame into the time information branch to capture time information; Step 4: Adaptively fuse the time information with the two-stage fusion features obtained in step 2 and send them to the prediction head for tracking.
2. The two-stage feature fusion target tracking method integrating temporal information according to claim 1, characterized in that: In the step 2, the residual feature fusion network is constructed by serially splicing two feature fusion layers, and then serially connecting a separate CFA module and a residual structure after feature fusion. Each feature fusion layer is constructed by arranging two groups of ECA modules containing self-attention mechanisms and CFA modules containing cross-attention mechanisms in parallel. The specific calculations of the ECA module with self-attention mechanism and the CFA module with cross-attention mechanism are as follows: ECA module with self-attention mechanism: X ECA =Norm(X+MultiHead(X+P x ,X+P x ,X)) (1) In formula (1), X represents the input of the ECA module, and the sine and cosine functions are used to generate the spatial position code P x , X ECA Represents the output of the ECA module; CFA module with cross attention mechanism: In the above formula, X q and X kv Represents two different inputs, P q and P k represents the spatial position encoding corresponding to two different inputs, Norm represents layer normalization, FFN represents feedforward network, X CFA Represents the output of the CFA module.
3. The two-stage feature fusion target tracking method integrating temporal information according to claim 2, characterized in that: In the step 2, the two-stage feature fusion method of introducing the residual structure is shown in formula (4): F=(F1+F2) / 2 (4) In the above formula, F is the two-stage fusion feature; F1 is the first-stage feature; F2 is the second-stage feature; and finally the two are fused by summing and averaging them.
4. The two-stage feature fusion target tracking method integrating temporal information according to claim 3, characterized in that: The adaptive fusion method in step 4 is shown in the following formula (8): f=wf1+(1-w)f2 (8) In the above formula, f is the feature after fusing the time information feature and the output of the feature fusion network, f1 is the feature output by the feature fusion network, f2 is the time information feature output by the constructed time information branch, and w is a learnable parameter.