Unmanned aerial vehicle target tracking method and system based on mamba feature extraction
By adopting a mamba feature extraction method on the drone, using the synergy between VSS Block and SPPF modules, high-precision and real-time drone target tracking under limited computing power is achieved, and the tracking accuracy and real-time problems of traditional algorithms in complex scenarios are solved.
Patent Information
- Application Number
- CN202510615597.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-14
AI Technical Summary
With limited computing power, existing drone target tracking algorithms are difficult to achieve high-precision and real-time tracking in complex scenarios, and target loss or tracking drift is prone to occur.
The UAV target tracking method based on mamba feature extraction is adopted, and multi-scale feature extraction and global feature fusion are realized through the synergy between the VSS Block module and the SPPF module. The specific steps include real-time acquisition of video streams, extracting search areas and template image features, inputting features to the backbone network for feature extraction and fusion, and ultimately used to track target prediction.
Under the condition of limited computing power, efficient and accurate tracking of specific goals is achieved, tracking accuracy and real-time performance are improved, and target tracking tasks in complex scenarios can be better cope with.
Smart Images

Figure CN120125618A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) monitoring, and particularly to a UAV target tracking method and system based on mamba feature extraction. Background Art
[0002] The monitoring of UAVs can be achieved by carrying various sensors and cameras, enabling efficient, accurate, and real-time monitoring of environmental conditions. Compared with traditional monitoring methods, UAV monitoring has many advantages, such as a wide monitoring range, fast data acquisition speed, and low labor cost. UAV monitoring based on AI technology can achieve more intelligent and automated applications, which can solve problems such as efficiency, accuracy, monitoring and early warning, operation safety, and autonomous operation in traditional UAV operations, meeting the current requirements for intelligent and efficient UAV applications.
[0003] Target tracking from the perspective of UAVs is one of the basic tasks of UAV AI monitoring. In the existing technology, it is usually achieved by using a target tracking algorithm based on CNN (Convolutional Neural Network) feature extraction, that is, using a large amount of training data to train CNN to achieve the detection and tracking of targets. However, UAVs often have limited computing power. In the case of limited computing power, the receptive field of CNN has limitations, resulting in limited tracking accuracy of traditional target tracking algorithms based on CNN feature extraction. Especially in complex scenarios, such as when the target undergoes large deformations, similar interference objects appear, or experiences rapid movement, traditional target tracking algorithms based on CNN feature extraction are prone to target loss or tracking drift, and cannot guarantee the accuracy and stability of tracking.
[0004] Using the global feature fusion ability of Transformer for feature extraction to achieve target tracking can, to a certain extent, improve the accuracy of the target tracking algorithm and better handle target tracking tasks in complex scenarios. However, the self-attention mechanism of Transformer has quadratic complexity, which will greatly increase the computational burden of the model. If the Transformer model is directly applied to UAVs with limited computing power to achieve target tracking, due to the large amount of model calculations, it is actually difficult to achieve real-time tracking and cannot balance tracking accuracy and real-time performance. Summary of the Invention
[0005] The technical problem to be solved by the present invention is: aiming at the technical problems existing in the prior art, the present invention provides a UAV target tracking method and system based on mamba feature extraction with high accuracy and good real-time performance, which can achieve efficient and accurate tracking of specific targets under the condition of limited computing power to meet the high-efficiency and real-time requirements of UAVs in different application scenarios.
[0006] To solve the above technical problems, the technical solution proposed by the present invention is as follows: An unmanned aerial vehicle target tracking method based on mamba feature extraction, comprising: Obtain a video stream in real time, obtain a search area image from the real-time video stream, and obtain a template image after confirming the target to be tracked; Input the obtained search area image and template image into a mamba-based backbone network for feature extraction respectively. The mamba-based backbone network includes a VSS Block module for preliminary feature extraction and an SPPF module for global feature fusion. In the VSS Block module, multi-scale feature extraction is performed through a DWRFEM unit to enhance the context semantic features of pixel features. The DWRFEM unit includes two branches. One branch is used to perform feature extraction on the input data through convolution operations, and the other is used to further extract multi-scale global features using a DWRFE component with wavelet depth separable dilated convolution (WTDConv) after convolution operations on the input data. The features extracted by the two branches are concatenated and fused through convolution operations to obtain the finally extracted multi-scale global features.
[0007] Input the extracted template image features and search area image features into a bottleneck layer for further global feature extraction and fusion to obtain fused features; Input the fused features output by the bottleneck layer into a detection head, and use the fused features to predict the tracking target.
[0008] As a further improvement of the method of the present invention: the VSS Block module further includes a first normalization layer, a first linear transformation layer, a second linear transformation layer, a third linear transformation layer, an SS2D unit, and a second normalization layer. The result obtained after the input data passes through the first normalization layer, the first linear transformation layer, the DWRFEM unit, the SS2D unit, and the second normalization layer in sequence is multiplied by the result obtained after the input data passes through the first normalization layer and the second linear transformation layer in sequence. The result after multiplication passes through the third linear transformation layer and is added to the input data to obtain the extracted features. The SS2D unit is used to further perform global feature extraction and fusion on the input features using a state space model.
[0009] As a further improvement of the method of the present invention: the steps of the SS2D unit for global feature extraction and fusion include: Flatten the features of the input two-dimensional feature map into one-dimensional vectors along multiple different directions, so that each vector contains different direction information of the input features; Perform state space feature fusion operations on each one-dimensional vector to generate a final one-dimensional output vector; Fuse the generated final one-dimensional output vector into a two-dimensional feature output.
[0010] As a further improvement of the method of the present invention: the DWRFE component includes a plurality of parallel independent branches, each independent branch includes a 1×1 convolutional layer, a wavelet depthwise separable 3×3 dilated convolution, a 1×1 convolutional layer, and a GELU activation function connected in sequence. In each branch, each wavelet depthwise separable 3×3 dilated convolution performs a convolution operation based on wavelet transform at different dilation convolution rates. The result after passing through the 1×1 convolution, depthwise separable 3×3 dilated convolution, 1×1 convolutional layer, and GELU activation function for the input data in each branch is added to the input data to obtain the output of each branch. The outputs of each branch are concatenated in the channel dimension and then compressed in channels through a 1×1 convolutional layer to obtain the finally extracted global fusion feature.
[0011] As a further improvement of the method of the present invention: in the SPPF module, a hybrid pooling layer is used to randomly select the way of max pooling or average pooling for pooling operation. The calculation formula of the hybrid pooling layer is as follows: , where, represents the hybrid pooling output value corresponding to the th feature map in the rectangular region , represents the element at (p,q) in the rectangular region , represents the number of elements in the rectangular region , is a random value of 0 or 1 to correspondingly represent the selection of using average pooling or max pooling.
[0012] As a further improvement of the method of the present invention: the bottleneck layer is a multi-layer pyramid structure based on the self-attention mechanism. Input the extracted template image features and search region image features into the bottleneck layer for global feature extraction and fusion. The obtained fused features include: Perform feature extraction at different levels on the template image features and the search image region features in sequence through a multi-layer attention structure; Upsample and element-wise add the search region image features passing through each layer of the attention structure to fuse the features extracted from the high layer and the low layer of the search region image features to obtain a search region feature map; perform global average pooling on the template image features after passing through the multi-layer attention structure to convert them into a one-dimensional vector, and use it as a convolution kernel to perform a convolution operation on the search region feature map to obtain a region similarity map between the template image features and the search region image features; Multiply the region similarity map with the search region feature map to obtain the final fused feature.
[0013] As a further improvement of the method of the present invention: each self-attention structure in the multi-layer self-attention structure includes an SA attention module and a stage extraction and fusion module. In the stage extraction and fusion module, global feature extraction and fusion are performed through an MHA unit. In the MHA unit, the DiTAC activation function is used to perform a non-linear transformation on the input features. The calculation process of the MHA unit includes: , , , , , where, represents the attention mechanism function, , X represents the input features, , , respectively represent the learnable parameter matrices for calculating the Q value, K value, and V value. i represents the i-th self-attention detection head, represents the relative position encoding in the i-th self-attention detection head, represents the input feature dimension, represents the coordinates of two different pixel points on the input feature map, represents the calculation parameter matrix, represents the output of the i-th self-attention detection head, represents the output of the final MHA unit, represents the output of the N-th self-attention head, represents the concatenation operation, represents the DiTAC activation function, represents the cumulative distribution function of the standard normal distribution, represents the diffeomorphic transformation, and a, b represent the upper and lower bounds of the domain of, represents the calculated value of the diffeomorphic transformation of the input features on different domains, represents the input of the activation function.
[0014] As a further improvement of the method of the present invention: in the stage extraction and fusion module, KAN units are also provided at the input end and output end of the MHA unit. The characterization structure of the KAN unit is: , , where, Represents an activation function , represents a spline curve function, represents the output of the KAN module corresponding to the input feature x, L represents the number of network layers, represents a weight coefficient, represents the composite operation of functions.
[0015] As a further improvement of the method of the present invention: the fused feature output by the bottleneck layer is input to the detection head to use the fused feature for tracking target prediction, and the following loss function is used: , , , wherein, , respectively represent , weights, represents the bounding box regression loss function of the tracking target, represents the head object prediction loss function, , , respectively represent , , weights of, represents the detection box regression loss function based on GioU, represents the distribution focal loss function, represents the L2 regularization loss function, represents the tracking target prediction probability, represents the tracking target true label, represents the overlap degree between the predicted tracking detection box and the true box, represents the binary cross entropy loss function.
[0016] The present invention also provides a drone target tracking system based on mamba feature extraction, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the drone target tracking method based on mamba feature extraction.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Through the collaborative effect of the VSS Block module and the SPPF module, the present invention realizes multi-scale feature extraction and global feature fusion. The DWRFEM unit in the VSS Block module not only extracts features through convolutional operations but also further extracts global fusion features using dilated convolution of wavelet transform, thereby enhancing the context semantic features of pixel features. This multi-scale feature extraction method enables the unmanned aerial vehicle (UAV) to capture more comprehensively the local details and global information of a single target, improves the recognition and positioning accuracy of the target, and achieves an effective balance between the single-target tracking accuracy and real-time performance in the UAV scenario.
[0018] 2. The present invention further uses a hybrid pooling layer in the SPPF module for pooling operations. By randomly selecting max pooling or average pooling during training, when the SPPF module performs global feature fusion, it can more effectively extract multi-scale features, further enhancing the context semantic features of pixel features, enabling the tracking system to capture more comprehensively the local details and global information of the target, and improving the recognition and positioning accuracy of the target. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a structural diagram of a single-target tracking algorithm based on mamba feature extraction in an embodiment of the present invention.
[0020] Figure 2 It is a flowchart of a UAV target tracking method based on mamba feature extraction in an embodiment of the present invention.
[0021] Figure 3 It is a schematic diagram of the hardware architecture of an airborne single-target tracking system in an embodiment of the present invention.
[0022] Figure 4 It is a business flowchart of an airborne single-target tracking system in an embodiment of the present invention.
[0023] Figure 5 It is a schematic diagram of the mamba network structure in an embodiment of the present invention.
[0024] Figure 6 It is a schematic diagram of the VSS Block structure in an embodiment of the present invention.
[0025] Figure 7 It is a schematic diagram of the DWFREM structure in an embodiment of the present invention.
[0026] Figure 8 It is a schematic diagram of the DWREM structure in an embodiment of the present invention.
[0027] Figure 9 It is a schematic diagram of the wavelet depth separable dilated convolution WTDConv structure in an embodiment of the present invention.
[0028] Figure 10 This is a schematic diagram of the SS2D module structure in an embodiment of the present invention.
[0029] Figure 11 This is a schematic diagram of the SPPF layer structure in an embodiment of the present invention.
[0030] Figure 12 This is a schematic diagram of the module structure in the first stage of an embodiment of the present invention.
[0031] Figure 13 This is a schematic diagram of the module structures in the second and third stages of an embodiment of the present invention.
[0032] Figure 14 This is a schematic diagram of the MHA structure in an embodiment of the present invention.
[0033] Figure 15 This is a schematic diagram of the SA structure in an embodiment of the present invention. Detailed implementation manners
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0035] As Figure 1 shown, the model corresponding to the UAV target tracking method based on mamba feature extraction in this embodiment can be decomposed into three main parts: a backbone network based on mamba, a bottleneck layer, and a detection head. The template image and the search area image are used as input images and are respectively input into the backbone network based on mamba. After the backbone network based on mamba extracts features from the template image and the search area image respectively, it enters the bottleneck layer for global feature extraction and fusion. The fused features are transmitted to the detection head, and the detection head realizes the bounding box regression of the tracking target and the foreground target prediction through the fully convolutional neural network layer, and outputs the final tracking result.
[0036] As Figure 2 shown, the UAV target tracking method based on mamba feature extraction in this embodiment includes the following steps: Step 1: Obtain the video stream in real time, obtain the search area image from the real-time video stream, and obtain the template image after confirming the target to be tracked.
[0037] During the flight of the drone, a video stream can be collected in real time through a camera or other types of image acquisition devices carried on the drone. Then, the search area image is obtained from the collected real-time video stream. The search area image is the image corresponding to the area to be searched. At the same time, according to the target to be tracked, the corresponding template image, that is, the tracking template image, is obtained.
[0038] As Figure 3 , Figure 4 shown, in a specific application embodiment, visible light, infrared and other cameras and gas information acquisition devices in an airborne optoelectronic pod can be used to obtain video stream data in real time, and the video stream data is transmitted to an AI edge device. The above-mentioned drone target tracking model is stored in the AI edge device; after receiving the data, the AI edge device uses the drone target tracking model to perform target recognition and tracking processing, and transmits the processed video stream and detection results back to the ground control platform. The control platform confirms whether it is necessary to track a specific target through a human-computer interaction interface. If tracking is required, the tracking template or the target detection method is selected to automatically select the template image, and the template image and the search area image are transmitted into the drone target tracking model in real time to achieve real-time tracking of a specific target; if tracking is not required, the video stream is displayed. In this embodiment, by deploying and reasoning the model at the AI edge, the real-time performance of the system operation can be improved, the dependence of the system on the network can be reduced, the robustness of the system operation can be improved, and at the same time, the computing pressure on the server can be reduced, and the overall hardware deployment cost can be reduced.
[0039] Step 2: Input the obtained search area image and template image into the backbone network based on mamba for feature extraction. The backbone network based on mamba includes a VSS Block module for preliminary feature extraction and an SPPF module for global feature fusion. In the VSS Block module, multi-scale feature extraction is performed through the DWRFEM unit to enhance the context semantic features of pixel features. The DWRFEM unit includes two branches. One branch is used to extract features from the input data through convolution operations, and the other is used to further extract global fusion features using the DWRFE component with multi-branch wavelet depth separable dilated convolution (WTDConv) after the input data passes through convolution operations. The features extracted by the two branches are concatenated and fused through convolution operations to obtain the finally extracted multi-scale features.
[0040] As Figure 5As shown in the figure, in this embodiment, the backbone network based on Mamba includes a plurality of sequentially arranged VSS Block modules and an SPPF module. A downsampling module is also provided at the output end of each VSS Block module for performing downsampling operations. After the input images (search region image and template image) pass through a plurality of VSS Block modules in sequence for multi-scale feature extraction and downsampling operations, they are globally feature fused by the SPPF module to obtain the template image features and search region image features for output. As the last layer of the feature extraction network, the SPPF module enables the model to better adapt to scale changes by fusing multi-scale information, thereby improving the robustness of detection.
[0041] The VSS Block module is the core module of the backbone network. The core module of the VSS Block module includes a DWRFEM unit and an SS2D unit. The DWRFEM unit is used to perform preliminary feature extraction on the input feature layer, further abstract the features, and improve the feature expression ability of the feature layer. The SS2D unit further performs global feature fusion on the input feature layer, which can further improve the global semantic features of local pixels, thereby improving the feature expression ability of the model. As Figure 6 shown in the figure, in this embodiment, the VSS Block module includes a first normalization layer, a first linear transformation layer, a second linear transformation layer, a third linear transformation layer, a DWRFEM unit, an SS2D unit, and a second normalization layer. The result obtained after the input data passes through the first normalization layer, the first linear transformation layer, the DWRFEM unit, the SS2D unit, and the second normalization layer in sequence is multiplied by the result obtained after the input data passes through the first normalization layer and the second linear transformation layer in sequence. The result after multiplication is added to the input data after passing through the third linear transformation layer to obtain the extracted features. The SS2D unit is used to use the state space model to perform global feature extraction and fusion on the input features.
[0042] Specifically, as Figure 6 shown in the figure, the overall structure of the VSS Block module is similar to the residual structure. The input features are divided into two branches. After the 1 branch passes through the first normalization layer LN, the first linear transformation layer (linear), the DWRFEM unit, the SS2D unit, and the second normalization layer LN in sequence, it is multiplied by the result obtained after the input data passes through the first normalization layer and the second linear transformation layer in sequence. The result after multiplication then passes through the third linear transformation layer LN. After a series of the above-mentioned feature extraction transformations, more abstract features are obtained, and this feature layer is added to the input feature layer (original input data) of the 2 branch to obtain the final output layer to make up for the loss of data information during network transmission. Through the above structure, on the one hand, the feature fitting ability of the network can be improved, and on the other hand, the gradient can be propagated faster, alleviating the problems of gradient disappearance and gradient explosion, and facilitating the expansion to form a deeper and wider model.
[0043] In the above VSS Block module of this embodiment, the overall structure of Branch 1 is an attention structure. The input features are transformed through this branch. Branch 1 includes a feature extraction branch for extracting features from the data stream. The core modules of the feature extraction branch are the DWRFEM unit and the SS2D unit for core feature abstraction. Branch 1 also includes an attention branch for performing a linear transformation (the third linear transformation layer LN) on the original data stream and multiplying it with the result of the feature extraction branch, which can improve the feature expression ability of the model for the region of interest.
[0044] In this embodiment, the DWRFEM unit initially extracts features from the input feature layer, further abstracts the features, and improves the feature expression ability of the feature layer. Considering that in the single-object tracking process, on the one hand, due to the change in the flight altitude of the drone, the search and tracking image presents multi-scale changes. Therefore, by extracting and fusing multi-scale features of the image, the tracking accuracy can be improved; on the other hand, since it is crucial to match the template image features with the search and tracking image in single-object tracking, thus, extracting the global features of the image can enhance the feature expression ability of the template image features and the search and tracking image, and improve the robustness of tracking. This embodiment adopts the DWRFEM module based on wavelet transform-based depthwise separable dilated convolution to initially capture multi-scale and global semantic information using wavelet transform-based depthwise separable dilated convolution. As Figure 7 shown, the DWRFEM unit of this embodiment includes two branches. One branch performs feature extraction on the input features using 1×1 convolution, and the other branch performs 1×1 convolution on the input features and then further feature extraction by the DWRFE component. Then, the features of the two branches are concatenated (C) to enhance the semantic expression ability of the output features. Finally, the concatenated features are fused and dimension-reduced using 1×1 convolution. The DWRFE component performs convolution operations using wavelet depthwise separable dilated convolutions with different dilation rates. Compared with traditional dilated convolutions, its receptive field range is larger, enabling the model to capture multi-scale information while having a larger range of context semantic information, thereby improving the feature expression ability of the feature map.
[0045] In the DWRFEM unit, its core feature extraction module is the DWRFE component, as Figure 8As shown, in this embodiment, the DWRFE component includes multiple parallel independent branches. Each independent branch includes a 1×1 convolutional layer, a wavelet depthwise separable 3×3 dilated convolution WTDConv, a 1×1 convolutional layer, and a GELU activation function connected in sequence. In each branch, each wavelet depthwise separable 3×3 dilated convolution WTDConv performs convolution operations based on wavelet transform at different dilation rates. The input data is fed into each branch and passes through the 1×1 convolution, the wavelet depthwise separable 3×3 dilated convolution, the 1×1 convolutional layer, and the GELU activation function in sequence. Then the result is added to the input data to obtain the output of each branch. The outputs of all branches are concatenated in the channel dimension and then compressed in the channel dimension through a 1×1 convolutional layer to obtain the finally extracted global fusion features. The DWRFE component uses dilated convolution based on wavelet transform to initially capture multi-scale and global semantic information. Each branch uses wavelet depthwise separable dilated convolution with different dilation rates for convolution operations. Compared with traditional dilated convolution, a larger receptive field range can be obtained, enabling the model to capture multi-scale information while having a larger range of context semantic information, thereby improving the feature expression ability of the feature map.
[0046] The core structure of the DWRFE component is the wavelet depthwise separable dilated convolution (WTDConv) with different dilation rates. In the traditional WTDConv convolution module, an ordinary convolution module is used. In this embodiment, the traditional convolution is replaced by introducing depthwise separable dilated convolution. The depthwise separable convolution mode can accelerate the operation efficiency. At the same time, dilated convolution is used to replace the ordinary convolution in the depthwise separable convolution, which can make full use of the advantage of dilated convolution to expand the receptive field, so that both the convolution operation efficiency can be improved and the convolution receptive field can be increased, enhancing the context semantic information of the pixels in the feature layer. As Figure 9 shown, the calculation process of the WTDConv convolution module is as follows: Two branches are adopted for the input feature map a. One branch directly uses depthwise separable dilated convolution to perform convolution operations on the input feature map a to obtain the output feature map d. The second branch uses wavelet transform on the input feature map a to obtain different frequency maps b of the input feature map a. Then, on the one hand, depthwise separable dilated convolution is performed on the different frequency maps b to obtain the output feature 1 of the different frequency maps b. On the other hand, the low-frequency images (similar to the reduced images of the original images) in the different frequency maps b are transformed by wavelet transform to obtain different frequency maps c of the low-frequency images. Depthwise separable dilated convolution is used on the different frequency maps c to obtain the output feature map f, and the inverse wavelet transform is used on the output feature map f to obtain the reconstructed map g. The reconstructed map g is added to the low-frequency image in the output feature 1 to obtain the final convolution result map e of the different frequency maps b. The inverse wavelet transform is performed on the map e to obtain the reconstructed map h. The reconstructed map h is added to the map d to obtain the final output feature map i of the wavelet depthwise separable dilated convolution.
[0047] In this embodiment, the SS2D unit uses a state space model to perform global feature extraction and fusion on the input features, so as to enhance the context semantic information of the feature layer and improve the feature expression ability of the feature layer. The steps for the SS2D unit to perform global feature extraction and fusion include: Step 201: Flatten the features of the input two-dimensional (2D) feature map into one-dimensional vectors along multiple different directions (such as the upper left, lower right, lower left, upper right, etc.), so that each vector contains different direction information of the input features; Step 202: Perform state space feature fusion operations on each one-dimensional vector to generate a final one-dimensional output vector; Step 203: Fuse the generated final one-dimensional output vectors into a two-dimensional feature output.
[0048] As Figure 10 shown, in a specific application embodiment, when the SS2D unit performs global feature extraction and fusion, the scanning expansion module flattens a 2D feature into 1D vectors along 4 different directions (upper left, lower right, lower left, upper right); the S6 block module independently sends the 4 1D vectors obtained in the previous step into S6 operations; the scanning merging module fuses the obtained 4 1D vectors into a 2D feature output. Among them, the S6 operation is an efficient selective information retention and parallel scanning algorithm, which can independently process the flattened one-dimensional vectors to ensure that the information in each direction is thoroughly scanned, thereby capturing features in different directions, forming a global receptive field without increasing the linear computational complexity, and effectively extracting the global features of the image.
[0049] The traditional SPPF structure adopts a single maximum pooling strategy, which may limit the regularization ability of the model. As Figure 11 shown, in this embodiment, a mixed pooling layer (Mix Pool) is used in the SPPF module for pooling operations. The mixed pooling layer enhances the regularization ability of the model by randomly selecting maximum pooling or average pooling during training.
[0050] As a preferred implementation manner, the calculation formula of the mixed pooling layer is as follows: (1) where represents the mixed pooling output value corresponding to the th feature map in the rectangular region , represents the element at (p,q) in the rectangular region , represents the number of elements in the rectangular region , A random value of 0 or 1 is used to correspondingly represent the selection of average pooling or max pooling. Specifically, during model training, random pooling is adopted, that is is a random value of 0 or 1; during model inference, if it is 1, max pooling is adopted.
[0051] In the backbone network of this embodiment, by adopting the above structure, the SPPF module can implement a fast spatial pyramid pooling strategy, effectively reduce the computational amount, fuse information of different scales through an efficient pooling method, and enable the model to better adapt to scale changes, thereby improving the robustness of detection.
[0052] Step 3: Input the extracted template image features and search region image features into the bottleneck layer for global feature extraction and fusion to obtain the fused features; input the fused features output by the bottleneck layer into the detection head to use the fused features for tracking target prediction.
[0053] After the backbone network extracts and fuses the features of the template image and the search region image, the bottleneck layer extracts multi-scale global features layer by layer from the template image features and the search region image features and performs cross-level global feature fusion. To further reduce the impact of multi-scale changes in the search region pictures under the UAV perspective on target tracking, in the UAV target tracking model based on the mamba Siamese network in this embodiment, the overall structure of the bottleneck layer adopts a multi-layer pyramid structure based on the self-attention mechanism. The bottleneck layer module includes a bottom-up path and a top-down path. Through the bottom-up path, a multi-layer self-attention structure is adopted to perform feature extraction at different levels on the template image features and the search region image features, and global feature fusion is performed at each layer to capture information of different scales; through the top-down path, the search region features are upsampled and element-wise added to fuse the fine features of the high layer and the rough features of the low layer, so that the feature map of each layer contains information of different scales. The template features are globally averaged pooled to be converted into a one-dimensional vector and used as a convolutional kernel to perform a convolutional operation on the search region features to obtain a region similarity map; the region similarity map is multiplied by the search region feature map to obtain the final fused features.
[0054] In this embodiment, the step of inputting the extracted template image features and search region image features into the bottleneck layer for global feature extraction and fusion to obtain the fused features includes: Step S301. The template image features and the search region image features are successively passed through a multi-layer attention structure for feature extraction at different levels; Step S302. Upsample and element-wise add the search region image features that have passed through each layer of the attention structure to fuse the features extracted from the high-level and low-level of the search region image features to obtain a search region feature map; perform global average pooling on the template image features after passing through multiple layers of the attention structure to convert them into a one-dimensional vector, and use it as a convolution kernel to perform a convolution operation on the search region feature map to obtain a regional similarity map between the template image features and the search region image features; Step S303. Multiply the regional similarity map by the search region feature map to obtain the final fused feature.
[0055] See Figure 1 , the bottleneck layer module consists of a bottom-up path and a top-down path. For the template feature and the search region feature, the bottom-up path adopts a multi-layer self-attention structure to extract features at different levels and further perform global feature fusion; for the search region feature, the top-down path fuses the fine features of the high-level of the search region feature and the rough features of the low-level of the search region feature by upsampling and element-wise addition. In this way, the feature map of each layer contains information at different scales, so that objects at different scales can be better processed; for the template feature, the top-down path directly performs global average pooling on it to convert it into a one-dimensional vector. Finally, use the template feature as a convolution kernel to perform a convolution operation on the search region feature to obtain a regional similarity map between the template and the search region. Finally, multiply the regional similarity map by the search region feature map to further enhance the feature representation ability of the region to be matched in the search region feature and improve the tracking accuracy.
[0056] In this embodiment, each layer of the self-attention structure in the multi-layer self-attention structure of the bottleneck layer includes an SA attention module and a stage extraction and fusion module. In the stage extraction and fusion module, global feature extraction and fusion are performed through the MHA unit, and the DiTAC activation function is used in the MHA unit to perform a non-linear transformation on the input feature. See Figure 1 , the core feature module in the bottleneck layer includes a first stage extraction module (stage 1), a second stage extraction module (stage 2), a third stage extraction module (stage 3), and an SA attention module. As Figure 12 shown, the first stage extraction module contains an MHA unit and a KAN unit. The input data passes through the MHA unit for global feature extraction and fusion. After the output data of the MHA unit is added to the input data, the result after the first addition is obtained; the result after the first addition is input into the KAN unit for feature enhancement, and the output of the KAN unit is added to the result after the first addition again to obtain the result after the second addition. Through the above processing, the global features and detailed features of the data can be fully fused, and the feature expression ability is significantly enhanced. AsFigure 13 As shown in Figure 13 , both the second-stage extraction module (Stage 2) and the third-stage extraction module (Stage 3) include an MHA unit and two KAN units. By adding a KAN unit at the input end based on the structure of the first-stage extraction module, global feature extraction and fusion can be fully achieved.
[0057] The structure of the MHA unit is as Figure 14 shown. After the input data is linearly transformed, it is input into the first BatchNorm (batch normalization) layer for normalization processing. Subsequently, the data is divided into three parts and enters the Q, K, and V branches respectively. After the Q and K branches are multiplied and added to the output of the Attention Bias layer, the output result obtained by inputting the added result into the Sigmoid activation function layer is then multiplied by the V branch, and then passes through the DiTAC activation function, the linear transformation layer, and the second BatchNorm layer to obtain the final output. Through the above processing, further global feature extraction and fusion of the target and the search feature region can be achieved. Traditional fixed activation functions have limited non-linearity and thus limited expressive power, impose learning biases on the network, and are difficult to adapt to different problem types and data complexities, resulting in limited model feature extraction capabilities. In this embodiment, in the MHA unit, the sigmoid activation function is used to replace the softmax activation function. The sigmoid function enables the model to pay more evenly attention to all features, avoiding the competition effect in the Softmax function, thereby capturing more information and having a better regularization effect, which helps to improve the generalization ability of the model. The DiTAC activation function is used to replace the traditional hardswish activation function. The DiTAC activation function has the characteristics of high efficiency and strong expressive power, can learn a set of suitable parameters according to the problem type and data complexity, enhance the adaptability of the activation function to the data set, and further improve the expressive power of the model.
[0058] In this embodiment, the calculation expression of the MHA unit in the multi-layer self-attention structure can be expressed as: (2) (3) (4) (5) (6) Among them, represents the attention mechanism function, , X represents the input, , , respectively represent the learnable parameter matrices for calculating the Q value, the K value, and the V value. i represents the i-th self-attention detection head. Denote the relative position encoding in the $i$-th self-attention detection head, Denote the input feature dimension, Denote the coordinates of two different pixel points on the input feature map, Denote the calculation of A learnable parameter matrix, Denote the output of the $i$-th self-attention detection head, Denote the output of the final MHA unit, Denote the output of the $N$-th self-attention head, Denote the concatenation operation, Denote the trainable activation function of diffeomorphism, Denote the cumulative distribution function of the standard normal distribution, Denote a learnable diffeomorphic transformation, $(a, b)$ represents The domain of definition, Denote the calculated value of the input feature under the diffeomorphic transformation on different domains of definition, Denote The input of the activation function.
[0059] In this embodiment, the DiTAC activation function is used in the multi-layer self-attention structure. The DiTAC activation function is a trainable activation function based on diffeomorphism, and the calculation process of the activation function is defined as: (7) Where, Denote the trainable activation function of diffeomorphism, Denote the cumulative distribution function of the standard normal distribution, Denote a learnable diffeomorphic transformation, $(a, b)$ represents The domain of definition. Since the function fitting ability of the traditional MLP (Multi-Layer Perceptron) is limited, the present invention uses the KAN unit to replace the MLP module in the first to third stage modules. The KAN unit has the advantage of stronger function fitting ability, and can enhance the model's ability to describe the features of the template and the search image.
[0060] Specifically, KAN units are also provided at the input end and output end of the MHA unit in the stage extraction and fusion module. KAN is essentially a combination of a spline curve and an MLP. It absorbs the advantages of both, that is, the spline curve has high precision in low dimensions and can fully utilize the advantages of the spline curve with high precision and strong interpretability in low dimensions, as well as the convenience of MLP for dimension expansion, with strong interpretability and convenient dimension expansion. In this embodiment, by replacing the multi-layer perceptron MLP with the KAN unit in the multi-layer self-attention structure, the description ability of the model for the template image features and the search region image features can be enhanced, and the problem of limited fitting function ability of the traditional MLP (multi-layer perceptron) can be solved.
[0061] For example, the general representation structure of the KAN unit can be expressed as: (8) (9) represents the activation function , represents the spline curve function, represents the output of the KAN module corresponding to the input feature x, L represents the number of network layers, represents the weight coefficient, represents the composite operation of the function.
[0062] As Figure 15 shown, the SA module is responsible for connecting the feature layers of each stage module and downsampling the features. Its overall structure is similar to that of the MHA unit. Different from the MHA unit, in order to generate the query Q, the SA module splits the 2D input feature into the template image feature (denoted as T) and the search region image feature (denoted as S) according to the positions of the template feature and the search region feature, reshapes the template image feature and the search region image feature into 3D features, performs secondary sampling on them in multiples of 2 in each spatial direction, flattens the features again, and connects the flattened features in the spatial dimension. In this way, the size of Q can be reduced by up to 4 times. Therefore, the final output of SA is also downsampled. Doubling the number of V channels can also reduce the loss of information caused by downsampling and increase the number of channels of the output features.
[0063] In this embodiment, a single-layer fully convolutional neural network layer is used in the detection head to implement the bounding box regression of the tracking target and the foreground target prediction. The calculation formula of its overall loss function is as follows: (10) Among them, , respectively represent , weights, Represents the bounding box regression loss function for tracking targets, and its loss consists of three parts, including: (11) Among them, 、 、 respectively represent 、 、 weights, where represents the detection box regression loss function based on GioU, represents the DFL Loss (Distribution Focal Loss) distribution focal loss function, represents the L2 regularization loss.
[0064] Represents the head object prediction loss function: (12) Among them, is the loss function, represents the predicted probability of the tracking target, represents the true label of the tracking target, represents the overlap degree between the predicted tracking detection box and the true box.
[0065] This embodiment also provides a drone target tracking system based on mamba feature extraction, including a microprocessor and a memory connected to each other. The microprocessor is programmed or configured to execute the drone target tracking method based on mamba feature extraction.
[0066] It can be understood that the above method of this embodiment can be executed by a single device, such as a computer or a server, etc., or can also be applied to a distributed scenario where multiple devices cooperate with each other to complete. In the case of a distributed scenario, one of the multiple devices can only execute one or more steps of the above method of this embodiment, and the multiple devices interact with each other to complete the above method. The processor can be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., for executing relevant programs to implement the above method of this embodiment. The memory can be implemented in the form of a read-only memory ROM, a random access memory RAM, a static storage device, and a dynamic storage device, etc. The memory can store an operating system and other application programs. When implementing the above method of this embodiment through software or firmware, the relevant program codes are stored in the memory and called by the processor for execution.
[0067] The above are only the preferred embodiments of the present invention and do not impose any form of limitation on the present invention. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention shall fall within the scope of protection of the technical solution of the present invention.
Claims
1. A UAV target tracking method based on mamba feature extraction, characterized in that: include: Acquire video stream in real time, obtain search area image from the real-time video stream and obtain template image after confirming the target to be tracked; The acquired search area image and template image are respectively input into a mamba-based backbone network for feature extraction, wherein the mamba-based backbone network includes a VSS Block module for preliminary feature extraction and an SPPF module for global feature fusion, wherein the VSS Block module uses a DWRFEM unit to perform multi-scale global feature extraction to enhance the contextual semantic features of pixel features, wherein the DWRFEM unit includes two branches, one branch is used to perform feature extraction on input data through a convolution operation, and the other branch is used to perform a multi-branch wavelet deep separable hole convolution WTDConv on the input data after the convolution operation, and the features extracted by the two branches are connected and fused through a convolution operation to obtain the final extracted multi-scale global features; The extracted template image features and search area image features are input into the bottleneck layer for further global feature extraction and fusion to obtain fused features; The fused features output by the bottleneck layer are input to the detection head to use the fused features for tracking target prediction.
2. The unmanned aerial vehicle target tracking method based on mamba feature extraction according to claim 1 is characterized in that: The VSS Block module also includes a first normalization layer, a first linear transformation layer, a second linear transformation layer, a third linear transformation layer, an SS2D unit and a second normalization layer. The result obtained by sequentially passing the input data through the first normalization layer, the first linear transformation layer, the DWRFEM unit, the SS2D unit and the second normalization layer is multiplied with the result of sequentially passing the input data through the first normalization layer and the second linear transformation layer. The result after multiplication is passed through the third linear transformation layer and added to the input data to obtain the extracted features. The SS2D unit is used to realize global feature extraction and fusion of the input features using the state space model.
3. The unmanned aerial vehicle target tracking method based on mamba feature extraction according to claim 2 is characterized in that: The steps of extracting and fusing global features by the SS2D unit include: Flatten the input two-dimensional feature map features into one-dimensional vectors along multiple different directions, so that each vector contains different directional information of the input features; Perform state space feature fusion operation on each one-dimensional vector to generate the final one-dimensional output vector; The resulting final one-dimensional output vector is fused into a two-dimensional feature output.
4. The unmanned aerial vehicle target tracking method based on mamba feature extraction according to claim 1 is characterized in that: The DWRFE component includes multiple parallel independent branches, each independent branch includes a 1×1 convolution layer, a wavelet depth-separable 3×3 dilated convolution, a 1×1 convolution layer and a GELU activation function connected in sequence, each wavelet depth-separable 3×3 dilated convolution in each branch uses a different extended convolution rate to perform a convolution operation based on wavelet transform, each branch accesses input data and sequentially passes through 1×1 convolution, wavelet depth-separable 3×3 dilated convolution, 1×1 convolution layer and GELU activation function, and the result is added to the input data to obtain the output of each branch, and the output of each branch is connected in the channel dimension and then compressed through a 1×1 convolution layer to obtain the multi-scale global fusion feature finally extracted.
5. The unmanned aerial vehicle target tracking method based on mamba feature extraction according to claim 1 is characterized in that: The SPPF module uses a hybrid pooling layer to randomly select the maximum pooling or average pooling method to perform pooling operations. The calculation formula of the hybrid pooling layer is as follows: , in, Indicates The feature map corresponds to the rectangular area The mixed pooling output value, Represents a rectangular area The element at (p,q) in Represents a rectangular area The number of elements in is a random value of 0 or 1 to indicate whether average pooling or maximum pooling is used.
6. The unmanned aerial vehicle target tracking method based on mamba feature extraction according to any one of claims 1 to 5, characterized in that: The bottleneck layer is a multi-layer pyramid structure based on the self-attention mechanism. The extracted template image features and search area image features are input into the bottleneck layer for global feature extraction and fusion. The fused features include: The template image features and the search area image features are sequentially passed through a multi-layer attention structure to extract features at different levels; The search area image features that have passed through each layer of attention structure are upsampled and element-wise added to fuse the features extracted from the high-level features of the search area image features with the features extracted from the low-level features to obtain a search area feature map; the template image features that have passed through the multi-layer attention structure are globally mean pooled to be converted into a one-dimensional vector, and used as a convolution kernel to perform a convolution operation with the search area feature map to obtain a regional similarity map between the template image features and the search area image features; The region similarity map is multiplied with the search region feature map to obtain a final fusion feature.
7. The unmanned aerial vehicle target tracking method based on mamba feature extraction according to claim 6 is characterized in that: Each layer of the self-attention structure in the multi-layer self-attention structure includes an SA attention module and a stage extraction and fusion module. The stage extraction and fusion module uses an MHA unit to extract and fuse global features. The MHA unit uses a DiTAC activation function to perform a nonlinear transformation on the input features. The calculation process of the MHA unit includes: , , , , , in, represents the attention mechanism function, , X represents the input features, , , They represent the learnable parameter matrices for calculating Q value, K value, and V value respectively, i represents the i-th self-attention detection head, represents the relative position encoding in the i-th self-attention detection head, represents the input feature dimension, Represents the coordinates of two different pixels on the input feature map, Representation calculation The parameter matrix of represents the output of the i-th self-attention detection head, represents the final output of the MHA unit, represents the output of the Nth self-attention head, Represents a splicing operation, is the activation function, represents the cumulative distribution function of the standard normal distribution, represents differential homeomorphism, a, b represent The upper and lower bounds of the domain of Represents the calculated value of differential homeomorphic transformation of input features on different definition domains, express Activation function input.
8. The unmanned aerial vehicle target tracking method based on mamba feature extraction according to claim 7 is characterized in that: In the stage extraction and fusion module, a KAN unit is further provided at the input and output ends of the MHA unit, and the representation structure of the KAN unit is: , , in, Represents the activation function , represents the spline function, represents the KAN module output corresponding to the input feature x, L represents the number of network layers, represents the weight coefficient, Represents a composite operation of a function.
9. The unmanned aerial vehicle target tracking method based on mamba feature extraction according to any one of claims 1 to 5, characterized in that: The fused features output by the bottleneck layer are input to the detection head to use the fused features for tracking target prediction, using the following loss function: , , , in, , Respectively , Weight, represents the bounding box regression loss function of the tracked target, represents the head object prediction loss function, , , Respectively , , The weight of Represents the regression loss function based on the GioU detection box, represents the distribution focal loss function, represents the L2 regularization loss function, represents the predicted probability of tracking target, represents the real label of the tracking target, Indicates the overlap between the predicted tracking detection box and the real box. represents the binary cross entropy loss function.
10. A UAV target tracking system based on mamba feature extraction, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the UAV target tracking method based on mamba feature extraction as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Single-target tracking network based on combination of Vmamba and Transform
CN118537369A
Target tracking method, system and device based on selective state space, and medium
CN119228853A
Medical image segmentation method and device, electronic equipment and readable storage medium
CN119559197A
Improved infrared small target detection method and system based on visual state space model
CN119810681A
Multi-target tracking method in complex scene based on adaptive association
CN119887850A
Cited By
Target tracking method and system based on Mama visual hybrid module
CN120298458A
Unmanned aerial vehicle multi-target tracking method based on CNN-Transform-Mama network and space-time Mama motion model
CN121600022A
Unmanned aerial vehicle multi-target tracking method based on CNN-Transformer-Mamba network and space-time Mamba motion model
CN121600022B