Lightweight unmanned aerial vehicle target tracking method based on separable convolution

Through a lightweight UAV target tracking method based on separable convolution, using deep separable convolution and feature fusion modules, the problems of large number of parameters and high computational complexity in UAV remote sensing target tracking are solved, efficient real-time tracking is achieved on limited resource equipment, and tracking accuracy and efficiency are improved.

CN120747796APending Publication Date: 2025-10-03CHANGCHUN UNIV OF SCI & TECH

Patent Information

Application Number
CN202511021072.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing UAV remote sensing target tracking methods have large parameters and high computational complexity, resulting in high memory usage and long processing time. They are difficult to run in real time on devices with limited computing resources, affecting tracking accuracy and efficiency.

Method used

A lightweight UAV target tracking method based on separable convolution is adopted. The number of parameters and computational complexity are reduced through deep separable convolution and feature fusion modules. The Infini-Attention and CFA modules are combined to enhance feature expression and long sequence processing capabilities. The mean absolute error evaluation framework is used to optimize model performance.

Benefits of technology

The number of parameters and calculation amount are significantly reduced, and the real-time performance and tracking accuracy of the model are improved. The number of parameters is reduced by 24%, and the calculation amount is reduced by 45%, maintaining efficient and accurate tracking performance in lightweight UAV target tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747796A_ABST
    Figure CN120747796A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image processing, and particularly relates to a lightweight unmanned aerial vehicle target tracking method based on separable convolution. The method comprises the following steps: step 1, preparing two remote sensing image data sets which are respectively used for training and testing; 2, inputting a frame into a separable convolution block for feature extraction to obtain a feature sequence; 3, inputting the feature sequence into an inverted bottleneck block and a forward feedback network, carrying out feature transformation and information mixing, and enhancing the feature expression capability; 4, the extracted category number is input into a fusion module to be processed, and a fusion sequence is obtained; and step 5, obtaining a classification regression vector through loss calculation, and then outputting a result graph. According to the method, a lightweight remote sensing target tracking network architecture is innovated, an improved feature extraction module UIB-P and a fusion module ICA are added, and an innovative loss function is adopted, so that the model greatly improves the global representation capability and precision of target tracking while reducing the calculation amount.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a lightweight UAV target tracking method based on separable convolution. Background Art

[0002] Remote sensing tracking is an important method for monitoring and analyzing the Earth's surface and its changes using remote sensing technology. With advances in science and technology, particularly the development of satellite and drone technology, remote sensing tracking has shown broad application potential in various fields, such as drone surveillance, military reconnaissance, and environmental monitoring. Transformer models based on the self-attention mechanism have achieved significant success in object detection and tracking. However, the high parameter count and computational complexity of object detection and tracking in video sequences make many complex models and algorithms difficult to meet real-time requirements with limited computing resources, resulting in tracking difficulties. Therefore, for object tracking in such large-scale models, a feature extraction network with a universal inverted bottleneck (UIB-P) module is used to increase the effective input of feature vectors, reduce the number of parameters and computation, and thus improve the network's computational efficiency. In the feature fusion network, a cross-attention architecture based on ICA and CFA modules is used to combine the feature extraction sequences to obtain richer representations from different network layers, improving model performance. The resulting fused feature vector is then passed through a series of fully connected layers to output the final object bounding box coordinates and object category.

[0003] In the existing technology, China's invention patent publication number "CN110807795B" is titled "A UAV remote sensing target tracking method and device based on MDnet". This method performs tracking tasks by constructing an MDnet-based tracking model and a preset update strategy that incorporates an adaptive context-aware correlation filter. This speeds up the update speed and efficiency of the tracking model and better improves the robustness and adaptability of tracking.

[0004] However, during the tracking process, the large number of parameters causes the model to occupy a large amount of memory resources, resulting in high video memory overhead during the calculation process, making the algorithm difficult to deploy on embedded devices or mobile terminals with limited computing resources. At the same time, the high computational complexity significantly increases processing time, causing the algorithm to fail to meet real-time requirements and experience severe delays. The combined effect of these factors not only reduces the real-time performance of the tracking algorithm, but also affects the overall tracking accuracy and system efficiency. Therefore, we propose a lightweight UAV remote sensing target tracking method to address the above issues. The overall effect is that while maintaining tracking accuracy, the number of parameters is reduced by 24% and the computational complexity is reduced by 45%, achieving the goal of lightweightness. Summary of the Invention

[0005] (1) Technical problems solved

[0006] In view of the shortcomings of the existing technology, the present invention provides a lightweight UAV target tracking method based on separable convolution to solve the problems of low tracking efficiency and accuracy and complex calculation.

[0007] (2) Technical solution

[0008] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:

[0009] A lightweight UAV target tracking method based on separable convolution includes the following steps:

[0010] Step 1: Prepare two remote sensing image datasets. Dataset 1 is used for network training, and dataset 2 is used for model testing.

[0011] Step 2: The template frame and the search frame are input into an additional depth-wise separable convolution block for feature extraction, and spatial mixing and feature extraction are performed to obtain a feature sequence;

[0012] Step 3: Input the inverted bottleneck block and perform feature transformation on the extracted feature sequence; use the forward feedback network to mix the channel information and perform pooling operations to enhance the expressiveness of the features;

[0013] Step 4: Input the output category number of the extraction network into the feature fusion module and perform image processing to obtain a fusion sequence;

[0014] Step 5: Calculate the classification vector and regression vector obtained by the fusion vector through loss, and then output the result graph.

[0015] Furthermore, in step 1, dataset 1 is the UAV123 dataset, and dataset 2 is the GOT-10k dataset, which are used for training and testing the model, respectively.

[0016] Furthermore, the fusion design in step 2 combines standard convolution and point-by-point convolution to optimize the computational process. 3×3 convolutions are first used to achieve channel expansion and spatial mixing, followed by 1×1 convolutions to compress the channels and reduce computational steps. This fusion design sacrifices some theoretical efficiency to improve actual inference speed, making it particularly suitable for hardware acceleration scenarios. Additional depthwise convolution blocks can optionally be equipped with 5×5 depthwise convolutions to further extract spatial features. Depthwise separable convolution (DW) is the core operation, decomposed into depthwise convolution for channel-by-channel processing, significantly reducing computational effort; 1×1 point-by-point convolutions fuse channel information and compress the dimensionality. This structure significantly reduces the number of parameters, especially with high-resolution inputs, balancing efficiency and feature expression.

[0017] Furthermore, the inverted bottleneck block (IB) in step 3 optimizes feature transformation through a three-step process of "expansion-depthwise convolution-compression." 1×1 point-by-point convolution expands the channel, 3×3 depthwise convolution performs spatial mixing, and 1×1 convolution compresses the channel, balancing computational efficiency and feature expression, making it suitable for lightweight mobile scenarios. The ConvNext-Like module is further optimized: 7×7 large-kernel depthwise convolution is used to expand the receptive field and enhance feature capture; the inverted bottleneck adopts a "wide-narrow-wide" channel design to reduce computational complexity; the GELU activation function and LayerNorm improve training stability, and residual connections mitigate gradient vanishing; efficient downsampling is achieved, and when the step size is greater than 1, the residual path is resized using 2×2 average pooling. This design significantly improves performance while maintaining lightweightness, balancing mobile efficiency and the requirements of visual tasks.

[0018] Furthermore, the FFN module in step 3 is improved by combining ConvNeXt and adopting a 1×1 point-by-point convolution stacking structure to enhance nonlinear expression capabilities. Feature fusion and information extraction are optimized on a lightweight basis, taking into account both efficiency and performance. The input is first processed and normalized using LayerNorm. 1×1 convolution then expands the channel and increases the feature dimension. 7×7 depthwise convolution is then used to replace traditional MLP local modeling to expand the receptive field and enhance spatial interaction. After GELU activation, the 1×1 convolution is compressed back to the original channel to form a lightweight bottleneck structure. The output and input are connected using residuals to stabilize training and alleviate gradient vanishing. The pooling layer is designed to highlight significant features such as edges and textures, smooth noise, retain global information, and improve training stability.

[0019] Furthermore, in step 4, an efficient feature fusion architecture is proposed. This architecture achieves multi-scale feature interaction by alternating four layers of fusion modules (each layer contains two ICA and two CFA modules). The ICA module utilizes an innovative Infini-Attention mechanism, combining local dot-product attention with a global memory network. This mechanism achieves memory compression through a trainable association matrix, and introduces a gated scalar β to dynamically fuse local and global context. The CFA module utilizes a channel-wise attention mechanism, using 1×1 convolutions to learn channel importance weights. This design achieves three breakthroughs within the Transformer framework: supporting infinite-length context modeling through recurrent memory states; maintaining processing efficiency through multi-head parallel computing; and achieving precise feature selection through a dual channel-spatial attention mechanism. Experiments demonstrate that this architecture significantly improves performance on long sequence tasks while maintaining low computational overhead.

[0020] Furthermore, step 5 utilizes a model error assessment framework that decomposes the mean absolute error into three component error assessments, resulting in improved robustness, purity, and interpretability. This framework provides precise diagnostic evidence for model optimization and can provide targeted guidance for improvements such as bias correction, scaling, and random error suppression.

[0021] (3) Beneficial effects

[0022] Compared with the existing technology, the present invention provides a lightweight UAV target tracking method based on separable convolution, which has the following beneficial effects:

[0023] 1. This paper designs a feature extraction module based on a universal inverted bottleneck block. The standard convolution is decomposed into depthwise convolution and pointwise convolution, significantly reducing the amount of computation. Each channel information is independently spatially filtered and blended, significantly reducing the number of parameters and computational complexity. High-dimensional features are used to enhance expressiveness, and the number of channels is compressed to reduce the amount of computation. This design increases model capacity while avoiding excessive growth in computational complexity. The parameter computation of the entire feature extraction network architecture is significantly reduced while maintaining efficient feature extraction capabilities.

[0024] 2. This paper utilizes a feature fusion algorithm with a self-attention contextual enhancement module (ICA) and a cross-feature enhancement module (CFA). This allows the fusion module to capture long-range dependencies in long-sequence tasks, making it suitable for processing global contextual information. This effectively selects and retains important information while suppressing unnecessary interference, helping to improve the efficiency of tracking tasks.

[0025] 3. This paper uses an innovative mean absolute error loss function for regression problems, decomposing the mean square error and root mean square error into three components: bias error, scale error, and unsystematic error, to more accurately assess model error. This method balances the evaluation metrics of MobileNetV3 and MobileNetV4, achieving a Top-1 accuracy of 72.4% with a parameter count of 2.9M and a computational load of 0.11 G FLOPs. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 Schematic diagram of the remote sensing target tracking method of the present invention;

[0027] Figure 2 This is a network architecture diagram of the remote sensing target tracking method of the present invention;

[0028] Figure 3 Schematic diagram of the structure of the feature extraction module UIB-P in the remote sensing target tracking method of the present invention;

[0029] Figure 4Schematic diagram showing the structure comparison between the feature extraction module UIB-P in the remote sensing target tracking method of the present invention and other existing methods;

[0030] Figure 5 This is a schematic diagram of the ICA structure of the attention mechanism module in the remote sensing target tracking method of the present invention;

[0031] Figure 6 Schematic diagram of the CFA structure of the attention mechanism module in the remote sensing target tracking method of the present invention;

[0032] Figure 7 A comparison chart of the parameters, computational complexity, and accuracy of MobileNetV3, MobileNetV4, and the method proposed in this invention; DETAILED DESCRIPTION

[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0034] Example

[0035] like Figure 1 As shown, an embodiment of the present invention proposes a lightweight UAV target tracking method based on separable convolution, including feature extraction, feature transformation, information fusion and loss calculation processes. The specific process is as follows Figure 2 shown.

[0036] As described in Step 1, prepare two remote sensing image datasets: Dataset 1: the UAV123 dataset, and Dataset 2: the GOT-10k dataset. These datasets feature diverse scenes and large amounts of annotated data. They are designed for practical challenges and support multi-task research. Dataset 1 will be used for network training, while Dataset 2 will be used for model testing.

[0037] like Figure 3 As shown in Figure 1, the feature extraction network of this method includes a fused inverted bottleneck block, an additional depth convolution block, an inverted bottleneck block, a ConvNext-Like block, a FFN module, and a pooling layer.

[0038] Furthermore, the Fused Inverted Bottleneck Block (FusedIB) primarily consists of two steps: a 3×3 standard convolution for channel expansion and spatial blending, and a 1×1 pointwise convolution for channel compression. This module combines the expansion and depthwise convolution into a single standard convolution, reducing the number of computational steps. In hardware-acceleration-friendly scenarios, this fused operation simplifies the process, sacrificing some theoretical computational efficiency in exchange for improved actual inference speed.

[0039] As described in step 2, the additional depthwise convolution block further includes two optional depthwise convolution operations with a convolution kernel of 5×5 for further spatial mixing and feature extraction. Figure 4 As shown in the figure, referring to the feature extraction module of the MobileNetV4 algorithm, an inverted bottleneck architecture with separable depthwise convolutional blocks is designed. The dynamic selection of the location and number of depthwise convolutions reduces computational effort while enhancing feature extraction flexibility. This algorithmic module is more suitable for real-time long-sequence visual tracking. The two optional dependency convolutions in the UIB block have four possible instantiations, resulting in different trade-offs. UIB extends the MobileNet inverted bottleneck (IB) block, becoming a standard building block for efficient networks. It introduces optional depthwise convolutions before the expansion layer and also enables optional depthwise convolutions between the expansion layer and the projection layer. The UIB-P block of this method not only contains optional depthwise convolution blocks but also includes a pooling module, resulting in a simple and efficient structure for feature extraction while maintaining good feature accuracy. This building block unifies important existing blocks: the original IB block, the ConvNext block, and the FFN block in ViT.

[0040] Depthwise separable convolution DW is one of the key operations in the UIB module, which decomposes the standard convolution into two steps: depthwise convolution and pointwise convolution. Depthwise Convolution: Perform convolution operations on each channel of the input separately, instead of convolving all channels like standard convolution. This can significantly reduce the amount of computation, especially when the number of channels is large. Pointwise Convolution: After the depthwise convolution, a 1×1 convolution kernel is used to mix the information of different channels. The purpose of this step is to compress the number of channels back to a smaller number. This step not only reduces the number of parameters and computation of the model, but also helps the model to fuse information between channels in a lower dimension. In the UIB module, depthwise separable convolution is often used to reduce computational complexity, especially in the early stages of the network. When the input resolution is high, the use of depthwise separable convolution can effectively reduce the amount of computation.

[0041] As described in step 3, the inverted bottleneck block (IB) is further used for feature transformation. It consists of three steps: 1×1 point-by-point convolution in the expansion layer to expand the number of channels; 3×3 depthwise separable convolution for spatial blending; and 1×1 point-by-point convolution in the projection layer to compress the number of channels. This three-step separation of "expansion-depthwise convolution-compression" balances computational efficiency and feature representation. Prioritizing theoretical computational efficiency, depthwise separable convolution minimizes FLOPs, making it suitable for lightweight networks and low-power mobile applications.

[0042] Furthermore, the ConvNext-Like block performs spatial mixing before expansion, allowing for cheaper spatial mixing with larger kernel sizes. This module draws on modern design principles from ConvNeXt, achieving efficient feature extraction through large-kernel depthwise convolutions, an inverted bottleneck architecture, optimized activation functions, and residual connections. The core of the large-kernel depthwise convolution module replaces traditional small convolution kernels with 7×7 large-kernel depthwise separable convolutions. This large receptive field enhances local feature capture while maintaining a low parameter count. The inverted bottleneck architecture employs a "wide-narrow-wide" channel design: 1×1 pointwise convolutions in the first layer expand the number of channels, followed by depthwise convolutions, and a final 1×1 convolution compresses the channels. This architecture balances computational effort and feature representation. GELU replaces ReLU in the activation function, providing smoother nonlinear characteristics. LayerNorm replaces BatchNorm for improved training stability. Residual connections introduce identity shortcuts, directly adding the input to the output, alleviating the vanishing gradient problem in deep networks. When the stride is greater than 1, the residual path adds 2×2 average pooling to match the size. The ConvNext-Like module significantly improves model performance while maintaining the lightweight nature of MobileNet by combining large kernel convolutions with modern components. Its design balances mobile computing constraints with the requirements of visual tasks.

[0043] Furthermore, the FFN module introduces ConvNeXt-style improvements based on the traditional MobileNet. It consists of a stack of two 1×1 point-by-point convolutions, separated by an activation and normalization layer. This not only enhances the model's nonlinear expressiveness but also maintains its lightweight nature. The steps include input normalization and channel expansion, depthwise convolution to enhance local modeling, activation function and feature compression, and residual connections. The input feature map is first normalized using LayerNorm to adapt to dynamic input sizes. Subsequently, 1×1 point-by-point convolutions are used to expand the number of channels, increasing feature dimensionality and enhancing model capacity. Deep convolutions are then used to enhance local modeling. The expanded features are then spatially fused using 7×7 depthwise convolutions. Large convolution kernels provide a wider receptive field, enhance local feature interaction, and replace the fully connected operations in traditional MLPs, reducing computational complexity. After the depthwise convolutions, the GELU activation function is used, followed by 1×1 point-by-point convolutions to compress the original number of channels, creating a bottleneck structure and reducing the number of parameters. The final output is residually summed with the original input, alleviating the degradation problem of deep networks.

[0044] Furthermore, the model uses two pooling layers: max pooling and average pooling. Max pooling slides a fixed-size window over the input feature map. For the covered area, the maximum value of the pixels is taken as the output. This preserves salient features while being computationally efficient, making it suitable for detecting edges, corners, and other key information and highlighting them. Average pooling, on the other hand, takes the average value of the covered area as the output. Compared to max pooling, average pooling has a stronger noise smoothing capability, is more stable for training, and is also suitable for extracting global information.

[0045] As described in step 4, feature fusion part: f x and f z The input feature fusion network includes two ICA modules, which enhance the model's expressiveness by introducing infinite-scale attention computation. Two CFA modules receive feature maps and fuse them using multi-head cross-attention. Thus, the two ICA modules and two CFA modules form a fusion layer. This fusion layer is repeated four times, and the result is then input into a CFA to obtain the decoded image f.

[0046] like Figure 5As shown in the figure, the infinite-attention in the ICA module is a cyclic attention mechanism that calculates local and global context states and combines them as output. Unlike traditional ones, infinite-attention emphasizes capturing dependencies in a wider range, which enables the model to better handle long-range dependencies, especially at different scales and channel relationships of feature maps. Similar to multi-head attention (MHA), it maintains the same number of parallel compressed memories as the attention heads on each attention layer. In addition to dot product attention, like RNN and MNM, it maintains a cyclic memory state to effectively track long sequence contexts:

[0047]

[0048] in, For attention output, For memory state, is a fragment sequence.

[0049] Furthermore, multi-head extended dot product attention is the main building block of LLM. MHA's powerful ability to model context-dependent dynamic computations and its convenience of temporal masking have been widely used in autoregressive generative models. We compute H attention context vectors in parallel for each sequence element, concatenate them along the second dimension, and then finally project the concatenated vectors into the model space to obtain the attention output. The attention context is calculated as a weighted average of all other values ​​as follows:

[0050]

[0051]

[0052] in, is a trainable projection matrix. K, V, and Q are the key, value, and query states.

[0053] Furthermore, in Infini-Attention, instead of computing new memory entries for compressed memory, we reuse the states (Q, K, and V) from the dot-product attention computation. The state sharing and reuse between dot-product attention and compressed memory not only enables efficient plug-and-play long-context adaptation, but also speeds up training and inference. For simplicity and computational efficiency, the memory is parameterized using an association matrix. This approach further treats the memory update and retrieval process as a linear attention mechanism and leverages the stable training techniques of related methods. About Memory Retrieval It can be expressed as:

[0054]

[0055] in, ,and and are the nonlinear activation function and the normalization term, respectively. Since the choice of nonlinearity and norm method is crucial to training stability, the sum of all keys is recorded as the normalization term z s , and uses element ELU+1 as the activation function. The storage update calculation formula is:

[0056]

[0057]

[0058] in, is called the association binding operator. If the KV binding already exists in memory, the update rule keeps the association matrix unchanged while still tracking the same normalization term as the previous one to obtain numerical stability. Long-term context binding requires aggregating the local attention state A through the learned gating scalar β dot and memory retrieval content A mem :

[0059]

[0060] Here, only a single scalar value is added as a training parameter for each head, while allowing a learnable trade-off between long-term and local information flow in the model. For multi-head infinite attention, H context states are computed in parallel and concatenated and projected into the final attention output O:

[0061]

[0062] in, are trainable weights.

[0063] Furthermore, Infini-Transformer can realize infinite context windows with limited memory usage. The memory complexity of storing compressed context in each head of a single layer is d key ×d value +d key , while for other models, the complexity increases with the sequence dimension - the memory complexity depends on the cache size or the soft hint size of RTM and AutoCompressors.

[0064] like Figure 6As shown in the figure, the CFA module process first performs global pooling on the input feature map, compressing the spatial dimension information into a 1×1 channel vector. A 1×1 convolution operation is then performed on the compressed channel vector to learn the relative importance of different channels. Finally, the generated channel attention weights are multiplied channel by channel with the original input feature map to adjust the feature response of each channel, thereby obtaining a feature map weighted by channel attention. Channel feature compression extracts the global information of each channel and compresses it into a smaller representation. Channel feature learning aims to strengthen attention on important channels while suppressing unimportant channels.

[0065] As described in Step 5, a common approach to evaluating the performance of quantitative models has traditionally been to decompose the sum of squared errors into systematic and unsystematic errors, representing distinct components of model error. However, these sum-of-squares-based metrics have been shown to be imprecise and sometimes misleading indicators of the mean error and its components. Therefore, the evaluation of model estimates and forecasts should increasingly utilize error metrics based on absolute values. To address the need to decompose the MAE into its components, a new metric, formed as a weighted average of the absolute errors, has been proposed. Consequently, the MAE can now be decomposed into three components, representing bias (MAEb), proportionality (MAEp), and unsystematic (MSEu) errors. These metrics provide a more interpretable metric for assessing model error while also more specifically identifying the types of errors that may be mitigated.

[0066] Furthermore, MAE is decomposed into three components: offset error, scale error, and non-systematic error. Both offset error and scale error are systematic errors, with one representing the amount of bias in the model and the other representing the degree to which the model systematically under- or over-predicts.

[0067] Among them, the mean deviation error (MBE) formula is as follows:

[0068]

[0069] In addition to being able to indicate overall over- or underestimation of predicted values, the mean deviation (MBE) can also be used to generate a set of i ) corresponds to the unbiased predicted value.

[0070]

[0071] The magnitude (absolute value) of the MBE also serves as a weight to determine the relative importance of the bias to the overall MAE:

[0072]

[0073] Among them, the proportional error is a systematic error related to under- or over-prediction. This error is reflected in To O i If the slope of the relationship is not 1, there is a proportional error in the model predictions. If the slope is less than 1, it indicates that the model is underestimating the observed value. There is a systematic overestimation of the part that is higher than the observed value, and a systematic underestimation of the part that is higher than the observed value. On the contrary, when the slope is greater than 1, the model is The value of is systematically underestimated. To estimate the proportional error, we use the unbiased predictive regressor:

[0074]

[0075] because The ordinary least squares (OLS) solution of is constrained to pass , and the straight line pass , so the estimate is unbiased. The proportional error (for each O i ) is weighted using the difference between the unbiased forecast and the observed value:

[0076]

[0077] Among them, non-systematic error: After correcting the bias error and proportional error, the remaining error is Similar to the composition of each component of MSEu, the relative importance weight of each prediction non-systematic error is determined by the difference between the unbiased prediction value and the unbiased regression value:

[0078]

[0079] Similarly, if OLS regression is used for , then the biased prediction and regression values ​​produce the same weight:

[0080]

[0081] Furthermore, the three weights of offset error, scale error, and non-systematic error can now be used to weigh the individual components of the absolute error:

[0082]

[0083]

[0084]

[0085] The MAE is the sum of the components:

[0086]

[0087] A clear advantage of this weighted average error decomposition is that it uses MAE instead of MSE as a baseline. Another advantage is that predictions without errors do not affect the components. This is not the case with the MSE-based decomposition, where predictions without errors can have a significant impact on the values ​​of MSE and MSEu.

[0088] This invention provides a remote sensing target tracking method that employs a feature extraction module and a feature fusion module to track feature information in remote sensing images. During training, the method utilizes a target tracking loss function to explicitly measure the deviation between predicted and true values, optimizing the target outcome. Under the same conditions, the feasibility and superiority of this method were further verified by calculating relevant indicators of target tracking categories compared to existing methods.

[0089] like Figure 7 As shown in the figure, the method proposed in the present invention is close to MobileNetV3 in terms of parameter quantity, which is 2.9M, and the computational complexity is 0.11 G FLOPs. In terms of Top-1 accuracy, it is even closer to MobileNetV4, which is 72.4%. The method proposed in the present invention has a good balance between performance and efficiency, which further illustrates that the present invention exhibits strong robustness in lightweight network models and can still maintain high performance and generalization capabilities.

[0090] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A lightweight UAV target tracking method based on separable convolution, comprising the following steps, characterized in that: Step 1: Prepare data sets. Prepare two remote sensing image data sets. Data set 1 is used for network training, and data set 2 is used for model testing. Step 2: The template frame and the search frame are input into an additional depth-wise separable convolution block for feature extraction, and spatial mixing and feature extraction are performed to obtain a feature sequence; Step 3: Input the inverted bottleneck block and perform feature transformation on the extracted feature sequence; Using the forward feedback network, channel information is mixed and pooled to enhance the expressiveness of features; Step 4: Input the output category number of the extraction network into the feature fusion module and perform image processing to obtain a fusion sequence; Step 5: Calculate the classification vector and regression vector obtained by the fusion vector through loss, and then output the result graph.

2. The lightweight UAV target tracking method based on separable convolution according to claim 1 is characterized in that: In step 1, dataset 1 is the UAV123 dataset, and dataset 2 is the GOT-10k dataset, which are used for training and testing the model respectively.

3. The lightweight UAV target tracking method based on separable convolution according to claim 1 is characterized in that: In step 2, a convolution operation is performed on the template frame and the search frame to capture local features, enhance the expressiveness of the features, and obtain a feature extraction sequence that realizes spatial mixing.

4. The lightweight UAV target tracking method based on separable convolution according to claim 1 is characterized in that: In step 3, the extracted feature sequence is fed into the inverted bottleneck block, where it is expanded through convolution and spatial filtering, and finally compressed back to its original dimensions. This design enhances the expressive power of features while maintaining efficient computation, providing a richer feature representation.

5. The lightweight UAV target tracking method based on separable convolution according to claim 1 is characterized in that: Step 3 utilizes a feedforward network to mix channel information through two convolutions, expanding and compressing the number of channels to their original dimensions. This increases the expressiveness of features while also integrating information between channels. Subsequently, global pooling is performed on the feature map, compressing the spatial dimensions of each channel to a single scalar value, preserving channel information while reducing computational complexity.

6. The lightweight UAV target tracking method based on separable convolution according to claim 1, characterized in that: Step 4 extracts the number of output categories from the network and inputs it into the feature fusion module to generate a fused sequence. This fusion module utilizes the global context modeling capabilities of the Infini-attention mechanism and the multi-head attention mechanism for cross-operation, capturing long-range dependencies and multi-scale information. This enhances the model's adaptability to complex scenarios and improves its robustness.

7. The lightweight UAV target tracking method based on separable convolution according to claim 1, characterized in that: In step 5, the classification vector and regression vector are obtained by fusing the vectors, and the classification vector outputs the probability distribution of each category for prediction of the target category; The regression vector directly outputs the bounding box coordinates of the target for precise positioning. Finally, the classification and regression results are combined to generate an output result image, which contains the target's category label and precise bounding box information.

Citation Information

Patent Citations

  • A method and apparatus for UAV remote sensing target tracking based on MDnet

    CN110807795B

Cited By

  • Road surface category recognition model construction method based on sub-channel deep convolution and application

    CN121982435A

  • Road surface category recognition model construction method based on split-channel deep convolution and application

    CN121982435B

  • A multi-stage feature fusion-based cross-modal identity authentication method

    CN122471418A