UAV Target Tracking Method and System Based on Mamba Feature Extraction
Through a mamba feature extraction method, combined with multi-scale feature extraction and global feature fusion of VSS Block and SPPF modules, the accuracy and real-time problems of the drone target tracking algorithm in complex scenarios are solved, and efficient and accurate target tracking is achieved.
Patent Information
- Application Number
- CN202510615597.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-14
AI Technical Summary
With limited computing power, existing drone target tracking algorithms are difficult to achieve a balance of high precision and real-time in complex scenarios, especially when target deformation, similar interference objects or rapid movement, target loss or tracking drift is prone to occur.
Using a method based on mamba feature extraction, multi-scale feature extraction and global feature fusion are performed through the VSS Block module and the SPPF module, and convolution operations are performed by combining DWRFEM units and DWRFE components. The wavelet depth can be used to separate the context semantic features of the pixel features, and the multi-layer pyramid structure with self-attention mechanism is used to perform feature extraction and fusion at the bottleneck layer, and finally target prediction is performed in the detection head.
Under the condition of limited computing power, the high accuracy and real-time performance of drone target tracking is achieved, which can better capture local details and global information of a single target, improve the target recognition and positioning accuracy, and adapt to tracking tasks in complex scenarios.
Smart Images

Figure CN120125618B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unmanned aerial vehicle (UAV) monitoring, and particularly to a UAV target tracking method and system based on mamba feature extraction. Background Art
[0002] The monitoring of UAVs can be achieved by carrying various sensors and cameras, enabling efficient, accurate, and real-time monitoring of environmental conditions. Compared with traditional monitoring methods, UAV monitoring has many advantages, such as a wide monitoring range, fast data acquisition speed, and low labor cost. UAV monitoring based on AI technology can achieve more intelligent and automated applications, which can solve problems such as efficiency, accuracy, monitoring and early warning, operation safety, and autonomous operation in traditional UAV operations, meeting the current requirements for intelligent and efficient UAV applications.
[0003] Target tracking from the perspective of UAVs is one of the basic tasks of UAV AI monitoring. In the prior art, it is usually achieved by using a target tracking algorithm based on CNN (Convolutional Neural Network) feature extraction, that is, using a large amount of training data to train CNN to achieve the detection and tracking of targets. However, UAVs often have limited computing power. In the case of limited computing power, the receptive field of CNN has limitations, resulting in limited tracking accuracy of traditional target tracking algorithms based on CNN feature extraction. Especially in complex scenarios, such as when the target undergoes large deformations, similar interference objects appear, or experiences rapid movement, traditional target tracking algorithms based on CNN feature extraction are prone to target loss or tracking drift, and cannot guarantee the accuracy and stability of tracking.
[0004] Using the global feature fusion ability of Transformer for feature extraction to achieve target tracking can improve the accuracy of the target tracking algorithm to a certain extent and can better handle target tracking tasks in complex scenarios. However, the self-attention mechanism of Transformer has quadratic complexity, which will greatly increase the computational burden of the model. If the Transformer model is directly applied to UAVs with limited computing power to achieve target tracking, due to the large amount of model calculations, it is actually difficult to achieve real-time tracking and cannot balance tracking accuracy and real-time performance. Summary of the Invention
[0005] The technical problem to be solved by the present invention is: aiming at the technical problems existing in the prior art, the present invention provides a UAV target tracking method and system based on mamba feature extraction with high accuracy and good real-time performance, which can achieve efficient and accurate tracking of specific targets under the condition of limited computing power to meet the high-efficiency and real-time requirements of UAVs in different application scenarios.
[0006] To solve the above technical problems, the technical solution proposed by the present invention is as follows:
[0007] A drone target tracking method based on mamba feature extraction, comprising:
[0008] Obtain the video stream in real time, obtain the search area image from the real-time video stream, and obtain the template image after confirming the target to be tracked;
[0009] Input the obtained search area image and template image into the mamba-based backbone network for feature extraction respectively. The mamba-based backbone network includes a VSS Block module for preliminary feature extraction and an SPPF module for global feature fusion. In the VSS Block module, multi-scale feature extraction is performed through the DWRFEM unit to enhance the context semantic features of pixel features. The DWRFEM unit includes two branches. One branch is used to perform feature extraction on the input data through convolution operations, and the other is used to further extract multi-scale global features using the DWRFE component with wavelet depth separable dilated convolution (WTDConv) after convolution operations on the input data. The features extracted by the two branches are concatenated and fused through convolution operations to obtain the finally extracted multi-scale global features.
[0010] Input the extracted template image features and search area image features into the bottleneck layer for further global feature extraction and fusion to obtain the fused features;
[0011] Input the fused features output by the bottleneck layer into the detection head, and use the fused features for tracking target prediction.
[0012] As a further improvement of the method of the present invention: the VSS Block module further includes a first normalization layer, a first linear transformation layer, a second linear transformation layer, a third linear transformation layer, an SS2D unit, and a second normalization layer. The result obtained after the input data passes through the first normalization layer, the first linear transformation layer, the DWRFEM unit, the SS2D unit, and the second normalization layer in sequence is multiplied by the result obtained after the input data passes through the first normalization layer and the second linear transformation layer in sequence. The result after multiplication passes through the third linear transformation layer and is added to the input data to obtain the extracted features. The SS2D unit is used to further perform global feature extraction and fusion on the input features using the state space model.
[0013] As a further improvement of the method of the present invention: the steps of the SS2D unit for global feature extraction and fusion include:
[0014] Flatten the features of the input two-dimensional feature map into one-dimensional vectors along multiple different directions, so that each vector contains different direction information of the input features;
[0015] Perform a state - space feature fusion operation on each one - dimensional vector to generate a final one - dimensional output vector;
[0016] Fuse the generated final one - dimensional output vectors into a two - dimensional feature output.
[0017] As a further improvement of the method of the present invention: The DWRFE component includes multiple parallel independent branches. Each independent branch includes a 1×1 convolutional layer, a wavelet depth - separable 3×3 dilated convolutional layer, a 1×1 convolutional layer, and a GELU activation function connected in sequence. Each wavelet depth - separable 3×3 dilated convolutional layer in each branch performs a convolution operation based on wavelet transform with different dilation convolution rates. The result after passing through the 1×1 convolution, depth - separable 3×3 dilated convolution, 1×1 convolutional layer, and GELU activation function for the input data in each branch is added to the input data to obtain the output of each branch. The outputs of each branch are concatenated in the channel dimension and then compressed in channels through a 1×1 convolutional layer to obtain the finally extracted global fusion features.
[0018] As a further improvement of the method of the present invention: In the SPPF module, a hybrid pooling layer is used to perform pooling operations by randomly selecting either max - pooling or average - pooling. The calculation formula of the hybrid pooling layer is as follows:
[0019] ,
[0020] where, represents the hybrid pooling output value corresponding to the th feature map in the rectangular region , represents the element at (p,q) in the rectangular region , represents the number of elements in the rectangular region , is a random value of 0 or 1 to correspondingly represent the selection of using average - pooling or max - pooling.
[0021] As a further improvement of the method of the present invention: The bottleneck layer is a multi - layer pyramid structure based on the self - attention mechanism. Input the extracted template image features and search region image features into the bottleneck layer for global feature extraction and fusion. The obtained fused features include:
[0022] Perform feature extraction at different levels on the template image features and search image region features in sequence through a multi - layer attention structure;
[0023] Upsample the image features of the search area passing through each layer of the attention structure and perform element-wise addition to fuse the features extracted from the high layer and the low layer of the search area image features to obtain the search area feature map; perform global average pooling on the template image features after passing through multiple layers of the attention structure to convert them into one-dimensional vectors, and use them as convolutional kernels to perform convolutional operations on the search area feature map to obtain the regional similarity map between the template image features and the search area image features;
[0024] Multiply the regional similarity map with the search area feature map to obtain the final fused feature.
[0025] As a further improvement of the method of the present invention: each self-attention structure in the multi-layer self-attention structure includes an SA attention module and a stage extraction and fusion module. In the stage extraction and fusion module, global feature extraction and fusion are performed through an MHA unit. In the MHA unit, the DiTAC activation function is used to perform non-linear transformation on the input features. The calculation process of the MHA unit includes:
[0026] ,
[0027] ,
[0028] ,
[0029] ,
[0030] ,
[0031] Among them, represents the attention mechanism function, , X represents the input feature, 、 、 respectively represent the learnable parameter matrices for calculating the Q value, K value, and V value. i represents the i-th self-attention detection head, represents the relative position encoding in the i-th self-attention detection head, represents the input feature dimension, represents the coordinates of two different pixel points on the input feature map, represents the calculation of the parameter matrix, represents the output of the i-th self-attention detection head, represents the output of the final MHA unit, represents the output of the N-th self-attention head, represents the concatenation operation, represents the DiTAC activation function, represents the cumulative distribution function of the standard normal distribution, represents a diffeomorphic transformation, and a, b represent the upper and lower bounds of the domain of, represents the calculated value of the diffeomorphic transformation of the input feature on different domains, represents the input of the activation function.
[0032] As a further improvement of the method of the present invention: a KAN unit is further provided at the input end and the output end of the MHA unit in the stage extraction and fusion module, and the characterization structure of the KAN unit is:
[0033] ,
[0034] ,
[0035] wherein, represents the activation function , represents the spline curve function, represents the output of the KAN module corresponding to the input feature x, L represents the number of network layers, represents the weight coefficient, represents the composite operation of the function.
[0036] As a further improvement of the method of the present invention: the fused feature output by the bottleneck layer is input to the detection head to use the fused feature for tracking target prediction, and the following loss function is used:
[0037] ,
[0038] ,
[0039] ,
[0040] wherein, , respectively represent , weights, represents the bounding box regression loss function of the tracking target, represents the head object prediction loss function, , , respectively represent , , weights of, represents the detection box regression loss function based on GioU, represents the distribution focal loss function, represents the L2 regularization loss function, represents the predicted probability of the tracking target, represents the true label of the tracking target, represents the overlap degree between the predicted tracking detection box and the true box, represents the binary cross - entropy loss function.
[0041] The present invention also provides a drone target tracking system based on mamba feature extraction, including a microprocessor and a memory connected to each other. The microprocessor is programmed or configured to execute the drone target tracking method based on mamba feature extraction.
[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0043] 1. Through the collaborative action of the VSS Block module and the SPPF module, the present invention realizes multi - scale feature extraction and global feature fusion. The DWRFEM unit in the VSS Block module not only extracts features through convolution operations, but also further extracts global fusion features using dilated convolution of wavelet transform, thereby enhancing the context semantic features of pixel features. This multi - scale feature extraction method enables the drone to capture the local details and global information of a single target more comprehensively, improves the recognition and positioning accuracy of the target, and realizes an effective balance between the single - target tracking accuracy and real - time performance in the drone scenario.
[0044] 2. The present invention further uses a hybrid pooling layer for pooling operations in the SPPF module. By randomly selecting max - pooling or average - pooling during training, when the SPPF module performs global feature fusion, it can more effectively extract multi - scale features, further enhancing the context semantic features of pixel features, enabling the tracking system to capture the local details and global information of the target more comprehensively, and improving the recognition and positioning accuracy of the target. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is the structural diagram of the single - target tracking algorithm based on mamba feature extraction in the embodiment of the present invention.
[0046] Figure 2 It is the flowchart of the drone target tracking method based on mamba feature extraction in the embodiment of the present invention.
[0047] Figure 3 It is the schematic diagram of the hardware architecture of the airborne single - target tracking system in the embodiment of the present invention.
[0048] Figure 4 It is the business flowchart of the airborne single - target tracking system in the embodiment of the present invention.
[0049] Figure 5Schematic diagram of the mamba network structure in the embodiments of the present invention.
[0050] Figure 6 Schematic diagram of the VSS Block structure in the embodiments of the present invention.
[0051] Figure 7 Schematic diagram of the DWFREM structure in the embodiments of the present invention.
[0052] Figure 8 Schematic diagram of the DWREM structure in the embodiments of the present invention.
[0053] Figure 9 Schematic diagram of the wavelet depth separable dilated convolution WTDConv structure in the embodiments of the present invention.
[0054] Figure 10 Schematic diagram of the SS2D module structure in the embodiments of the present invention.
[0055] Figure 11 Schematic diagram of the SPPF layer structure in the embodiments of the present invention.
[0056] Figure 12 Schematic diagram of the first stage module structure in the embodiments of the present invention.
[0057] Figure 13 Schematic diagram of the second and third stage module structures in the embodiments of the present invention.
[0058] Figure 14 Schematic diagram of the MHA structure in the embodiments of the present invention.
[0059] Figure 15 Schematic diagram of the SA structure in the embodiments of the present invention. Detailed implementation manners
[0060] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0061] Such as Figure 1As shown in the figure, the model corresponding to the UAV target tracking method based on mamba feature extraction in this embodiment can be decomposed into three main parts: the mamba-based backbone network, the bottleneck layer, and the detection head. The template image and the search area image are used as input images and input into the mamba-based backbone network respectively. After the mamba-based backbone network extracts features from the template image and the search area image respectively, it enters the bottleneck layer for global feature extraction and fusion. The fused features are transmitted to the detection head, and the detection head realizes the bounding box regression of the tracking target and the foreground target prediction through the fully convolutional neural network layer, and outputs the final tracking result.
[0062] As Figure 2 shown in the figure, the UAV target tracking method based on mamba feature extraction in this embodiment includes the following steps:
[0063] Step 1: Obtain the video stream in real time, obtain the search area image from the real-time video stream, and obtain the template image after confirming the target to be tracked.
[0064] During the flight of the UAV, the video stream can be collected in real time through the camera or other types of image acquisition devices carried on the UAV, and then the search area image is obtained from the collected real-time video stream. The search area image is the image corresponding to the area to be searched. At the same time, the corresponding template image, that is, the tracking template image, is obtained according to the target to be tracked.
[0065] As Figure 3 、 Figure 4 shown in the figure, in a specific application embodiment, visible light, infrared and other cameras and gas information acquisition devices in the airborne optoelectronic pod can be used to obtain video stream data in real time, and the video stream data is transmitted to the AI edge device. The above UAV target tracking model is stored in the AI edge device; after receiving the data, the AI edge device uses the UAV target tracking model for target recognition and tracking processing, and transmits the processed video stream and detection results back to the ground control platform. The control platform confirms whether a specific target needs to be tracked through the human-computer interaction interface. If tracking is required, the tracking template or the target detection method is selected to automatically select the template image, and the template image and the search area image are transmitted into the UAV target tracking model in real time to realize the real-time tracking of the specific target; if tracking is not required, the video stream is displayed. By deploying and reasoning the model at the AI edge end in this embodiment, the real-time performance of the system operation can be improved, the dependence of the system on the network can be reduced, the robustness of the system operation can be improved, and at the same time, the computing pressure on the server can be reduced, and the overall hardware deployment cost can be reduced.
[0066] Step 2: Input the obtained search area image and the template image into the mamba-based backbone network for feature extraction. The mamba-based backbone network includes a VSS Block module for preliminary feature extraction and an SPPF module for global feature fusion. In the VSS Block module, multi-scale feature extraction is performed through the DWRFEM unit to enhance the context semantic features of pixel features. The DWRFEM unit includes two branches. One branch is used to extract features from the input data through convolution operations, and the other is used to further extract global fusion features using the DWRFE component with multi-branch wavelet depth separable dilated convolution (WTDConv) after convolution operations on the input data. The features extracted from the two branches are concatenated and fused through convolution operations to obtain the finally extracted multi-scale features.
[0067] As Figure 5 shown, in this embodiment, the mamba-based backbone network includes multiple sequentially arranged VSS Block modules and an SPPF module. A downsampling module is also provided at the output end of each VSS Block module for downsampling operations. After the input images (search area image and template image) pass through multiple VSS Block modules for multi-scale feature extraction and downsampling operations in sequence, they are globally feature-fused by the SPPF module to output the template image features and search area image features. As the last layer of the feature extraction network, the SPPF module can better adapt the model to scale changes by fusing multi-scale information, thereby improving the robustness of detection.
[0068] The VSS Block module is the core module of the backbone network. The core module of the VSS Block module includes a DWRFEM unit and an SS2D unit. The DWRFEM unit is used to perform preliminary feature extraction on the input feature layer, further abstract the features, and improve the feature expression ability of the feature layer. The SS2D unit further performs global feature fusion on the input feature layer, which can further improve the global semantic features of local pixels, thereby improving the feature expression ability of the model. As Figure 6 shown, in this embodiment, the VSS Block module includes a first normalization layer, a first linear transformation layer, a second linear transformation layer, a third linear transformation layer, a DWRFEM unit, an SS2D unit, and a second normalization layer. The result obtained after the input data passes through the first normalization layer, the first linear transformation layer, the DWRFEM unit, the SS2D unit, and the second normalization layer in sequence is multiplied by the result obtained after the input data passes through the first normalization layer and the second linear transformation layer in sequence. The result after multiplication is added to the input data through the result of the third linear transformation layer to obtain the extracted features. The SS2D unit is used to use the state space model to achieve global feature extraction and fusion of the input features.
[0069] Specifically, as Figure 6 shown, the overall structure of the VSS Block module is similar to the residual structure. The input features are divided into two branches. After the first branch passes through the first normalization layer LN, the first linear transformation layer (linear), the DWRFEM unit, the SS2D unit, and the second normalization layer LN in sequence, it is multiplied by the result of the input data passing through the first normalization layer and the second linear transformation layer in sequence. The result after multiplication then passes through the third linear transformation layer LN. Through the above series of feature extraction transformations, more abstract features are obtained, and this feature layer is added to the input feature layer (original input data) of the second branch to obtain the final output layer, so as to make up for the loss of data information during network transmission. Through the above structure, on the one hand, the feature fitting ability of the network can be improved, and on the other hand, the gradient can be propagated faster, alleviating the problems of gradient disappearance and gradient explosion, and facilitating the expansion to form a deeper and wider model.
[0070] In the above VSS Block module of this embodiment, the overall structure of the first branch is an attention structure. The input features are subjected to feature transformation through this branch. The first branch includes a feature extraction branch for extracting features from the data stream. The core modules of the feature extraction branch are the DWRFEM unit and the SS2D unit for core feature abstraction. The first branch also includes an attention branch for performing a linear transformation (the third linear transformation layer LN) on the original data stream and multiplying it by the result of the feature extraction branch, which can improve the feature expression ability of the model for the region of interest.
[0071] In this embodiment, the DWRFEM unit is used to perform preliminary feature extraction on the input feature layer, further abstract the features, and improve the feature expression ability of the feature layer. Considering that in the process of single-object tracking, on the one hand, due to the change in the flight altitude of the drone, the search and tracking image presents multi-scale changes, so by extracting and fusing the multi-scale features of the image, the tracking accuracy can be improved; on the other hand, since it is crucial to match the template image features with the search and tracking image in single-object tracking, therefore, extracting the global features of the image can enhance the feature expression ability of the template image features and the search and tracking image, and improve the robustness of the tracking. This embodiment adopts the DWRFEM module based on wavelet transform-based depthwise separable dilated convolution to initially capture multi-scale and global semantic information. As Figure 7As shown, the DWRFEM unit of this embodiment includes two branches. One branch performs feature extraction on the input features using 1×1 convolution, and the other branch performs 1×1 convolution on the input features and then further feature extraction is performed by the DWRFE component. Then, the features of the two branches are concatenated (C) to enhance the semantic expression ability of the output features. Finally, the concatenated features are fused and dimension-reduced using 1×1 convolution. The DWRFE component performs convolution operations using wavelet depth-separable dilated convolutions with different dilation rates. Compared with traditional dilated convolutions, its receptive field range is larger, enabling the model to capture multi-scale information while having a larger range of context semantic information, thereby improving the feature expression ability of the feature map.
[0072] In the DWRFEM unit, its core feature extraction module is the DWRFE component, as Figure 8 shown. In this embodiment, the DWRFE component includes multiple parallel independent branches. Each independent branch includes a 1×1 convolutional layer, a wavelet depth-separable 3×3 dilated convolution WTDConv, a 1×1 convolutional layer, and a GELU activation function connected in sequence. Each wavelet depth-separable 3×3 dilated convolution WTDConv in each branch performs convolution operations based on wavelet transform with different dilation rates. The result after each branch accesses the input data and passes through the 1×1 convolution, the wavelet depth-separable 3×3 dilated convolution, the 1×1 convolutional layer, and the GELU activation function is added to the input data to obtain the output of each branch. The outputs of each branch are concatenated in the channel dimension and then compressed in the channel by a 1×1 convolutional layer to obtain the finally extracted global fusion features. The DWRFE component uses dilated convolution based on wavelet transform to initially capture multi-scale and global semantic information. Each branch uses wavelet depth-separable dilated convolutions with different dilation rates for convolution operations. Compared with traditional dilated convolutions, a larger receptive field range can be obtained, enabling the model to capture multi-scale information while having a larger range of context semantic information, thereby improving the feature expression ability of the feature map.
[0073] The core structure of the DWRFE component is the wavelet depth-separable dilated convolution (WTDConv) with different dilation rates. In the traditional WTDConv convolution module, an ordinary convolution module is used. In this embodiment, the traditional convolution is replaced by a depth-separable dilated convolution. The depth-separable convolution mode can accelerate the operation efficiency. At the same time, the dilated convolution is used to replace the ordinary convolution in the depth-separable convolution, which can make full use of the advantage of the dilated convolution to expand the receptive field, so that both the convolution operation efficiency can be improved and the convolution receptive field can be increased, enhancing the context semantic information of the pixels in the feature layer. As Figure 9As shown in the figure, the calculation process of the WTDConv convolution module is as follows: For the input feature map a, two branches are adopted. One branch directly uses depthwise separable dilated convolution to perform convolution operation on the input feature map a to obtain the output feature map d. The second branch performs wavelet transform on the input feature map a to obtain different frequency maps b of the input feature map a. Then, on the one hand, depthwise separable dilated convolution is performed on different frequency maps b to obtain output feature 1 of different frequency maps b. On the other hand, the low-frequency images (similar to the reduced images of the original images) in different frequency maps b are subjected to wavelet transform to obtain different frequency maps c of the low-frequency images. Depthwise separable dilated convolution is used for different frequency maps c to obtain the output feature map f, and the inverse wavelet transform is used for the output feature map f to obtain the reconstructed map g. The reconstructed map g is added to the low-frequency image in output feature 1 to obtain the final convolution result map e of different frequency maps b. The inverse wavelet transform is performed on the map e to obtain the reconstructed map h, and the reconstructed map h is added to the map d to obtain the final output feature map i of the wavelet depthwise separable dilated convolution.
[0074] In this embodiment, the SS2D unit uses a state space model to achieve global feature extraction and fusion of the input features, so as to enhance the context semantic information of the feature layer and improve the feature expression ability of the feature layer. The steps for the SS2D unit to perform global feature extraction and fusion include:
[0075] Step 201: Flatten the features of the input two-dimensional (2D) feature map into one-dimensional vectors along multiple different directions (such as the upper left, lower right, lower left, upper right, etc. directions), so that each vector contains different direction information of the input features;
[0076] Step 202: Perform state space feature fusion operation on each one-dimensional vector to generate the final one-dimensional output vector;
[0077] Step 203: Fuse the generated final one-dimensional output vectors into a two-dimensional feature output.
[0078] As Figure 10 shown, in a specific application embodiment, when the SS2D unit performs global feature extraction and fusion, the scanning expansion module flattens a 2D feature into 1D vectors along 4 different directions (upper left, lower right, lower left, upper right); the S6 block module independently sends the 4 1D vectors obtained in the previous step to perform S6 operation; the scanning merging module fuses the 4 obtained 1D vectors into a two-dimensional feature output. Among them, the S6 operation is an efficient selective information retention and parallel scanning algorithm, which can independently process the flattened one-dimensional vectors to ensure that the information in each direction is thoroughly scanned, so as to capture the features in different directions, form a global receptive field without increasing the linear computational complexity, and effectively extract the global features of the image.
[0079] The traditional SPPF structure adopts a single max pooling strategy, which may limit the regularization ability of the model. As Figure 11 shown, in this embodiment, a Mix Pool (mixed pooling layer) is used in the SPPF module for pooling operations. The mixed pooling layer enhances the regularization ability of the model by randomly selecting max pooling or average pooling during training.
[0080] As a preferred embodiment, the calculation formula of the mixed pooling layer is as follows:
[0081] (1)
[0082] where, represents the output value of the mixed pooling in the rectangular region corresponding to the th feature map, represents the element at (p, q) in the rectangular region , represents the number of elements in the rectangular region , is a random value of 0 or 1 to correspondingly represent the selection of using average pooling or max pooling. Specifically, during model training, random pooling is adopted, that is is a random value of 0 or 1; during model inference, is 1, and max pooling is adopted.
[0083] By adopting the above structure in the backbone network of this embodiment, a fast spatial pyramid pooling strategy can be realized, effectively reducing the computational amount, and fusing information of different scales through an efficient pooling method. It can also enable the model to better adapt to scale changes, thereby improving the robustness of detection.
[0084] Step 3: Input the extracted template image features and search region image features into the bottleneck layer for global feature extraction and fusion to obtain the fused features; input the fused features output by the bottleneck layer into the detection head to use the fused features for tracking target prediction.
[0085] After the backbone network extracts and fuses the features of the template image and the search area image, the bottleneck layer extracts multi-scale global features layer by layer from the template image features and the search area image features and performs cross-layer global feature fusion. To further reduce the impact of the multi-scale changes in the search area pictures under the UAV perspective on target tracking, in this embodiment, the overall structure of the bottleneck layer in the UAV target tracking model based on the mamba siamese network adopts a multi-layer pyramid structure based on the self-attention mechanism. The bottleneck layer module includes a bottom-up path and a top-down path. Through the bottom-up path, a multi-layer self-attention structure is adopted to extract features of different levels from the template image features and the search area image features, and global feature fusion is performed at each layer to capture information of different scales. Through the top-down path, the search area features are upsampled and element-wise added to fuse the fine features of the high layer and the rough features of the low layer, so that the feature map of each layer contains information of different scales. The template features are globally average pooled and converted into a one-dimensional vector, which is used as a convolution kernel to perform a convolution operation on the search area features to obtain a regional similarity map. The regional similarity map is multiplied by the search area feature map to obtain the final fused features.
[0086] In this embodiment, the steps of inputting the extracted template image features and search area image features into the bottleneck layer for global feature extraction and fusion to obtain the fused features include:
[0087] Step S301. Sequentially pass the template image features and the search area image features through a multi-layer attention structure for feature extraction of different levels;
[0088] Step S302. Upsample and element-wise add the search area image features that have passed through each layer of the attention structure to fuse the features extracted from the high layer and the low layer of the search area image features to obtain a search area feature map. Globally average pool the template image features after passing through the multi-layer attention structure to convert them into a one-dimensional vector, and use it as a convolution kernel to perform a convolution operation on the search area feature map to obtain a regional similarity map between the template image features and the search area image features;
[0089] Step S303. Multiply the regional similarity map by the search area feature map to obtain the final fused features.
[0090] See Figure 1, the bottleneck layer module consists of a bottom-up path and a top-down path. For the template feature and the search region feature, in the bottom-up path, a multi-layer self-attention structure is adopted to extract features at different levels and further perform global feature fusion. For the search region feature, in the top-down path, the fine features of the high-level search region feature and the rough features of the low-level search region feature are fused through upsampling and element-wise addition. In this way, the feature map of each layer contains information at different scales, enabling better processing of objects at different scales. For the template feature, the top-down path directly performs global average pooling on it to convert it into a one-dimensional vector. Finally, the template feature is used as a convolution kernel to perform a convolution operation on the search region feature to obtain the region similarity map between the template and the search region. Finally, the region similarity map is multiplied by the search region feature map to further enhance the feature representation ability of the region to be matched in the search region feature and improve the tracking accuracy.
[0091] In this embodiment, each self-attention structure in the multi-layer self-attention structure of the bottleneck layer includes an SA attention module and a stage extraction and fusion module. In the stage extraction and fusion module, global feature extraction and fusion are performed through an MHA unit, and the DiTAC activation function is used in the MHA unit to perform a non-linear transformation on the input feature. See Figure 1 , the core feature module in the bottleneck layer includes a first stage extraction module (stage 1), a second stage extraction module (stage 2), a third stage extraction module (stage 3), and an SA attention module. As Figure 12 shown, the first stage extraction module contains an MHA unit and a KAN unit. The input data undergoes global feature extraction and fusion through the MHA unit. After the output data of the MHA unit is added to the input data, the result after the first addition is obtained; the result after the first addition is input into the KAN unit for feature enhancement, and the output of the KAN unit is added to the result after the first addition again to obtain the result after the second addition. Through the above processing, the global features and detailed features of the data can be fully fused, significantly enhancing the feature expression ability. As Figure 13 shown, the second stage extraction module (stage 2) and the third stage extraction module (stage 3) both include an MHA unit and two KAN units, and a KAN unit is added at the input end on the basis of the structure of the first stage extraction module to fully perform global feature extraction and fusion.
[0092] The structure of the MHA unit is as Figure 14As shown, the input data is linearly transformed and then input into the first BatchNorm (batch normalization) layer for normalization processing. Subsequently, the data is divided into three parts, which enter the Q, K, and V branches respectively. After the Q and K branches are multiplied and added to the output of the Attention Bias layer, the result of the addition is input into the Sigmoid activation function layer. The output obtained is then multiplied by the V branch and successively passes through the DiTAC activation function, the linear transformation layer, and the second BatchNorm layer to obtain the final output. Through the above processing, it is possible to further perform global feature extraction and fusion on the target and the search feature region. Traditional fixed activation functions have limited non-linearity, so their expression ability is limited, and they impose learning biases on the network, making it difficult to adapt to different problem types and data complexities, resulting in limited feature extraction ability of the model. In this embodiment, in the MHA unit, the sigmoid activation function is used to replace the softmax activation function. The sigmoid function enables the model to more evenly focus on all features, avoiding the competition effect in the Softmax function, thereby capturing more information and having a better regularization effect, which helps to improve the generalization ability of the model. The DiTAC activation function is used to replace the traditional hardswish activation function. The DiTAC activation function has the characteristics of high efficiency and strong expression ability, and can learn a set of suitable parameters according to the problem type and data complexity, enhancing the adaptability of the activation function to the data set, and further improving the expression ability of the model.
[0093] In this embodiment, the calculation expression of the MHA unit in the multi-layer self-attention structure can be expressed as:
[0094] (2)
[0095] (3)
[0096] (4)
[0097] (5)
[0098] (6)
[0099] Among them, represents the attention mechanism function, , X represents the input, 、 、 respectively represent the learnable parameter matrices for calculating the Q value, the K value, and the V value. i represents the i-th self-attention detection head, represents the relative position encoding in the i-th self-attention detection head, represents the input feature dimension, Denote the coordinates of two different pixel points on the input feature map, Denote the calculation A learnable parameter matrix, Denote the output of the i-th self-attention detection head, Denote the output of the final MHA unit, Denote the output of the N-th self-attention head, Denote the concatenation operation, Denote the trainable activation function of diffeomorphism, Denote the cumulative distribution function of the standard normal distribution, Denote a learnable diffeomorphic transformation, (a, b) denote The domain of definition, Denote the calculated value of the input feature under the diffeomorphic transformation on different domains of definition, Denote The input of the activation function.
[0100] In this embodiment, the DiTAC activation function is used in the multi-layer self-attention structure. The DiTAC activation function is a trainable activation function based on diffeomorphism. The calculation process of the activation function is defined as:
[0101] (7)
[0102] Among them, Denote the trainable activation function of diffeomorphism, Denote the cumulative distribution function of the standard normal distribution, Denote a learnable diffeomorphic transformation, (a, b) denote The domain of definition.
[0103] Due to the limited function fitting ability of the traditional MLP (Multi-Layer Perceptron), the present invention uses the KAN unit to replace the multi-layer perceptron MLP in the first to third stage modules. The KAN unit has the advantage of stronger function fitting ability, and can enhance the model's ability to describe the features of the template and the search image.
[0104] Specifically, KAN units are also provided at the input end and the output end of the MHA unit in the stage extraction and fusion module. KAN is essentially a combination of a spline curve and an MLP. It absorbs the advantages of both, that is, the spline curve has high precision in low dimensions, can give full play to the advantages of the spline curve with high precision and strong interpretability in low dimensions, and the MLP is convenient for dimension expansion, with strong interpretability and convenient dimension expansion. In this embodiment, by using the KAN unit to replace the multi-layer perceptron MLP in the multi-layer self-attention structure, the model's ability to describe the features of the template image and the search area image can be enhanced, and the problem of the limited function fitting ability of the traditional MLP (Multi-Layer Perceptron) can be solved.
[0105] For example, the general representation structure of the KAN unit can be expressed as:
[0106] (8)
[0107] (9)
[0108] Denotes the activation function , Denotes the spline curve function, Denotes the output of the KAN module corresponding to the input feature x, L denotes the number of network layers, Denotes the weight coefficient, Denotes the composite operation of the function.
[0109] Such as Figure 15 As shown, the SA module is responsible for connecting the feature layers of each stage module and downsampling the features. Its overall structure is similar to that of the MHA unit. Different from the MHA unit, in order to generate the query Q, the SA module splits the 2D input feature into the template image feature (denoted as T) and the search region image feature (denoted as S) according to the positions of the template feature and the search region feature, reshapes the template image feature and the search region image feature into 3D features, subsamples them by a factor of 2 in each spatial direction, then flattens the features again and connects the flattened features in the spatial dimension. By the above method, the size of Q can be reduced by up to 4 times. Therefore, the final output of SA is also downsampled. Doubling the number of V channels can also reduce the loss of information caused by downsampling and increase the number of channels of the output features.
[0110] In this embodiment, a single-layer fully convolutional neural network layer is used in the detection head to implement the bounding box regression of the tracking target and the foreground target prediction. The calculation formula of its overall loss function is as follows:
[0111] (10)
[0112] Among them, , Respectively denote , Weights, Denotes the bounding box regression loss function of the tracking target, and its loss consists of three parts including:
[0113] (11)
[0114] Among them, , , Respectively denote , , weights, where represents the regression loss function based on the GioU detection box, represents the DFL Loss (Distribution Focal Loss) distribution focal loss function, represents the L2 regularization loss.
[0115] represents the head object prediction loss function:
[0116] (12)
[0117] where is the loss function, represents the tracking target prediction probability, represents the true label of the tracking target, represents the overlap degree between the predicted tracking detection box and the true box.
[0118] This embodiment also provides a drone target tracking system based on mamba feature extraction, including a microprocessor and a memory connected to each other. The microprocessor is programmed or configured to execute the drone target tracking method based on mamba feature extraction.
[0119] It can be understood that the above method of this embodiment can be executed by a single device, such as a computer or a server, etc., or can also be applied to a distributed scenario where multiple devices cooperate with each other to complete. In the case of a distributed scenario, one of the multiple devices can only execute one or more steps of the above method of this embodiment, and the multiple devices interact with each other to complete the above method. The processor can be implemented in ways such as a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., for executing relevant programs to implement the above method of this embodiment. The memory can be implemented in forms such as a read-only memory ROM, a random access memory RAM, a static storage device, and a dynamic storage device, etc. The memory can store an operating system and other application programs. When implementing the above method of this embodiment through software or firmware, the relevant program codes are stored in the memory and called by the processor for execution.
[0120] The above is only a preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Although the present invention has been disclosed above with a preferred embodiment, it is not intended to limit the present invention. Therefore, any simple modification, equivalent change, and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention should fall within the scope of the technical solution of the present invention.
Claims
1. A method for unmanned aerial vehicle target tracking based on mamba feature extraction, characterized in that Including: Obtain the video stream in real time, obtain the search area image from the real-time video stream, and obtain the template image after confirming the target to be tracked; Input the obtained search area image and template image into the backbone network based on mamba for feature extraction. The backbone network based on mamba includes a VSS Block module for preliminary feature extraction and an SPPF module for global feature fusion. In the VSS Block module, multi-scale global feature extraction is performed through the DWRFEM unit to enhance the context semantic features of pixel features. The DWRFEM unit includes two branches. One branch is used to extract features from the input data through convolution operations, and the other is used to further extract multi-scale global features using the DWRFE component with multi-branch wavelet depth separable dilated convolution (WTDConv) after convolution operations on the input data. The features extracted by the two branches are concatenated and fused through convolution operations to obtain the finally extracted multi-scale global features; Input the extracted template image features and search area image features into the bottleneck layer for further global feature extraction and fusion to obtain the fused features; Input the fused features output by the bottleneck layer into the detection head to use the fused features for tracking target prediction.
2. The method for tracking an unmanned aerial vehicle target based on mamba feature extraction according to claim 1, wherein The VSS Block module further includes a first normalization layer, a first linear transformation layer, a second linear transformation layer, a third linear transformation layer, an SS2D unit, and a second normalization layer. The result obtained after the input data passes through the first normalization layer, the first linear transformation layer, the DWRFEM unit, the SS2D unit, and the second normalization layer in sequence is multiplied by the result obtained after the input data passes through the first normalization layer and the second linear transformation layer in sequence. The result after multiplication is added to the input data after passing through the third linear transformation layer to obtain the extracted features. The SS2D unit is used to implement global feature extraction and fusion of the input features using the state space model.
3. The drone target tracking method based on mamba feature extraction according to claim 2, wherein, The steps for the SS2D unit to perform global feature extraction and fusion include: Flatten the features of the input two-dimensional feature map into one-dimensional vectors along multiple different directions, so that each vector contains different direction information of the input features; Perform state space feature fusion operations on each one-dimensional vector to generate the final one-dimensional output vector; Fuse the generated final one-dimensional output vectors into a two-dimensional feature output.
4. The method for unmanned aerial vehicle target tracking based on mamba feature extraction according to claim 1, wherein, The DWRFE component includes multiple parallel independent branches. Each independent branch includes a 1×1 convolutional layer, a wavelet depthwise separable 3×3 dilated convolution, a 1×1 convolutional layer, and a GELU activation function connected in sequence. In each branch, different dilation rates are used for the wavelet depthwise separable 3×3 dilated convolution to perform convolution operations based on wavelet transform. The result after the input data passes through the 1×1 convolution, the wavelet depthwise separable 3×3 dilated convolution, the 1×1 convolutional layer, and the GELU activation function in each branch is added to the input data to obtain the output of each branch. The outputs of each branch are concatenated in the channel dimension and then compressed in channels through a 1×1 convolutional layer to obtain the finally extracted multi-scale global fusion features.
5. The drone target tracking method based on mamba feature extraction according to claim 1, wherein, In the SPPF module, a hybrid pooling layer randomly selects the way of max pooling or average pooling for pooling operations. The calculation formula of the hybrid pooling layer is as follows: , Among them, represents the output value of hybrid pooling corresponding to the th feature map in the rectangular region . represents the element at (p, q) in the rectangular region . represents the number of elements in the rectangular region . is a random value of 0 or 1 to correspondingly represent the selection of using average pooling or max pooling.
6. The method for tracking an unmanned aerial vehicle target based on mamba feature extraction according to any one of claims 1 to 5, characterized in that The bottleneck layer is a multi-layer pyramid structure based on the self-attention mechanism. The extracted template image features and search region image features are input into the bottleneck layer for global feature extraction and fusion. The fused features obtained include: The template image features and search region image features are successively passed through a multi-layer attention structure for feature extraction at different levels; The search region image features passing through each layer of the attention structure are upsampled and element-wise added to fuse the features extracted from the high layer and the low layer of the search region image features to obtain a search region feature map; the template image features passing through the multi-layer attention structure are globally average pooled to be converted into a one-dimensional vector, which is used as a convolution kernel to perform a convolution operation on the search region feature map to obtain a region similarity map between the template image features and the search region image features; The region similarity map and the search region feature map are multiplied to obtain the final fused features.
7. The method for UAV target tracking based on mamba feature extraction according to claim 6, characterized in that, Each layer of the multi-layer self-attention structure includes an SA attention module and a stage extraction and fusion module. In the stage extraction and fusion module, global feature extraction and fusion are performed through an MHA unit. In the MHA unit, the DiTAC activation function is used to perform a non-linear transformation on the input features. The calculation process of the MHA unit includes: , , , , , Among them, represents the attention mechanism function, , X represents the input feature, , , respectively represent the learnable parameter matrices for calculating Q value, K value, and V value. i represents the i-th self-attention detection head, represents the relative position encoding in the i-th self-attention detection head, represents the input feature dimension, represents the coordinates of two different pixel points on the input feature map, represents the calculation of parameter matrix, represents the output of the i-th self-attention detection head, represents the output of the final MHA unit, represents the output of the N-th self-attention head, represents the concatenation operation, is the activation function, represents the cumulative distribution function of the standard normal distribution, represents the diffeomorphic transformation, and a, b represent the upper and lower bounds of the domain of, represents the calculated value of the input feature under the diffeomorphic transformation on different domains, represents the input of the activation function.
8. The method for tracking an unmanned aerial vehicle target based on mamba feature extraction according to claim 7, wherein In the stage extraction and fusion module, KAN units are also set at the input end and output end of the MHA unit. The representation structure of the KAN unit is: , , Among them, represents the activation function , represents the spline curve function, represents the output of the KAN module corresponding to the input feature x, L represents the number of network layers, represents the weight coefficient, represents the composite operation of the function.
9. The method for tracking an unmanned aerial vehicle target based on mamba feature extraction according to any one of claims 1 to 5, characterized in that, The fused features output by the bottleneck layer are input into the detection head to use the fused features for tracking target prediction. The following loss function is used: , , , Among them, and respectively represent and weights, represents the bounding box regression loss function of the tracking target, represents the head object prediction loss function, and and respectively represent and and weights, represents the regression loss function of the detection box based on GioU, represents the distribution focal loss function, represents the L2 regularization loss function, represents the prediction probability of the tracking target, represents the true label of the tracking target, represents the coincidence degree between the predicted tracking detection box and the true box, represents the binary cross-entropy loss function.
10. A drone target tracking system based on mamba feature extraction, comprising a microprocessor and a memory connected to each other, characterized in that, The microprocessor is programmed or configured to execute the drone target tracking method based on mamba feature extraction as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Single-target tracking network based on combination of Vmamba and Transform
CN118537369A
Improved infrared small target detection method and system based on visual state space model
CN119810681A