Lightweight multi-task small target detection algorithm based on adaptive pyramid and multi-stage path aggregation

By introducing an adaptive feature pyramid network and multi-stage path aggregation module in the small object detection technology, combining the lightweight Transformer module and an optimized channel attention mechanism, the problem of lack of adaptability and high computational complexity of feature fusion weights in the existing technology is solved, and efficient small object detection and multi-task learning are achieved.

CN120070857AActive Publication Date: 2025-05-30CENT SOUTH UNIV

Patent Information

Application Number
CN202510136324.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-30
Estimated Expiration
2045-02-07

AI Technical Summary

Technical Problem

In the existing small-objective detection technology, there are problems such as lack of adaptability in feature fusion weights, insufficient modeling capabilities of small-objective feature, high computational complexity, and insufficient sharing of multi-task learning features.

Method used

Adaptive feature pyramid network (AdaptiveFPN) and multi-stage path aggregation module (MPAM) are adopted, combined with the lightweight Transformer module and the optimized channel attention mechanism, dynamically adjust the fusion weight of multi-scale features to realize global feature modeling, and object detection, bounding box regression and semantic segmentation are performed through multi-task learning heads to share multi-scale features.

Benefits of technology

It improves the accuracy and efficiency of small object detection, reduces computing overhead, enhances the feature sharing ability of multi-task learning, and significantly improves the model's detection ability of small objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070857A_ABST
    Figure CN120070857A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision, and particularly discloses an adaptive pyramid and multi-stage path aggregation-based lightweight multi-task small target detection algorithm, which comprises an adaptive feature pyramid network, a multi-stage path aggregation module, a lightweight Transform module, an optimized channel attention mechanism, a multi-task learning head and an algorithm training method. According to the lightweight multi-task small target detection algorithm based on adaptive pyramid and multi-stage path aggregation, dynamic weight adjustment of different levels of features is realized by combining an adaptive FPN, and the expression ability of multi-scale features is enhanced. The MPAM module introduces a lightweight Transform module on the original basis, global feature modeling is performed by using an axial attention mechanism, the perception ability of the model to a small target is improved, and a channel attention mechanism is optimized to reduce calculation overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly to a lightweight multi-task small target detection algorithm with adaptive pyramid and multi-stage path aggregation. Background Art

[0002] With the rapid development of computer vision technology, object detection based on deep learning has made remarkable progress in many application fields such as intelligent monitoring, remote sensing image analysis, and medical image processing. However, small target detection remains a challenging task, mainly because small targets occupy limited pixels in the image and are unevenly distributed, which easily leads to high false alarm rates and missed detection rates of existing algorithms. Therefore, researchers have been continuously exploring new model architectures and algorithm optimization methods to improve the accuracy and efficiency of small target detection.

[0003] The existing Feature Pyramid Network (FPN) improves the performance of object detection through multi-scale feature fusion, but there is still a problem that the feature fusion weights are not adaptive when dealing with small targets. In addition, the Transformer module performs well in global feature modeling, but its computational cost is relatively high, which is not conducive to real-time detection requirements. Multi-task learning improves the generalization ability of the model by sharing feature extraction modules. However, in existing methods, the feature sharing mechanism between multi-tasks is not yet perfect, and it is difficult to make full use of multi-scale feature information. Summary of the Invention

[0004] The purpose of the present invention is to provide a lightweight multi-task small target detection algorithm with adaptive pyramid and multi-stage path aggregation and its training method, aiming to solve several key problems existing in the existing small target detection technology, including the lack of adaptability of feature fusion weights, insufficient small target feature modeling ability, high computational complexity, and insufficient multi-task learning feature sharing.

[0005] To achieve the above purpose, the present invention provides a lightweight multi-task small target detection algorithm with adaptive pyramid and multi-stage path aggregation, including an Adaptive Feature Pyramid Network (AdaptiveFPN), a Multi-Stage Path Aggregation Module (MPAM), a lightweight Transformer module, an optimized channel attention mechanism, a multi-task learning head, and an algorithm training method;

[0006] The Adaptive Feature Pyramid Network uses learnable weight parameters to automatically optimize the fusion ratio of features at different levels;

[0007] The Multi-Stage Path Aggregation Module (MPAM) is connected after the Adaptive Feature Pyramid Network, and realizes the fusion of multi-scale information through path aggregation in multiple stages. Each stage is embedded with a lightweight Transformer module for global feature modeling;

[0008] The optimized channel attention mechanism adopts axial attention and calculates attention in the height and width directions respectively;

[0009] The multi-task learning head includes an object detection head, a bounding box regression head, and a semantic segmentation head. While performing object detection, it adds the tasks of bounding box regression and semantic segmentation and shares the multi-scale features extracted by MPAM.

[0010] Preferably, the architecture of the adaptive feature pyramid network includes bottom-layer feature extraction, a feature fusion layer, and multi-scale feature output. The bottom-layer feature extraction uses the backbone network to extract multi-level feature maps, including high-resolution features at low levels and low-resolution features at high levels;

[0011] In the feature fusion layer, learnable weight parameters are introduced between each feature layer, and features at different levels are fused by weighted summation. For each feature layer, AdaptiveFPN uses a small fully connected network to generate fusion weights, and these weights are dynamically adjusted according to the characteristics of the input features. The feature fusion process is expressed as:

[0012]

[0013] where, F 1 , F 2 , …, F n are the feature maps extracted by the backbone network, F′ i represents the feature map of the i-th layer, w i,j are learnable weight parameters dynamically generated by a small fully connected network, satisfying

[0014] The multi-scale feature output is the feature map after adaptive fusion, and the multi-scale feature output is subsequently passed to the MPAM module as the input of the multi-scale features;

[0015] The adaptive feature pyramid network adaptively optimizes the fusion weights of features at different levels through a dynamic weight adjustment mechanism and in combination with a self-supervised learning method.

[0016] Preferably, the MPAM module mainly consists of multi-stage path aggregation, a lightweight Transformer module, and feature merging. The feature fusion F agg of the multi-stage path aggregation part is expressed as:

[0017] F agg = Concat(F stage1 , F stage2 , F stagr3 );

[0018] where, F stage1 , F stage2 , Fstage3 respectively represent the feature maps after path aggregation in each stage;

[0019] The global feature modeling F of the lightweight Transformer module global is expressed as:

[0020] F global = AxialAttention(F agg );

[0021] The final feature merging F final is expressed as:

[0022] F final = Merge(F agg , F global );

[0023] The MPAM module adopts model compression methods in the design, and the model compression methods include pruning, quantization, and knowledge distillation;

[0024] Among them, the pruning technique removes redundant neural network connections or channels to reduce the model complexity; the quantization technique converts the model parameters and activations from 32-bit floating-point numbers to lower-bit representations; the knowledge distillation technique trains a lightweight student model to imitate the behavior of a larger and more performant teacher model;

[0025] The MPAM module is in the pyramid residual feature mapping sub-module PRFM and the dilated convolution path aggregation sub-module DCPAM, and combines convolution kernels and pooling operations of different sizes. Among them, the PRFM module is a sub-module for multi-scale feature fusion; the DCPAM module uses dilated convolution to aggregate feature information from different scales;

[0026] The MPAM module adopts a multi-branch structure and cross-layer feature fusion technology to aggregate multi-scale information at different levels.

[0027] Preferably, the lightweight Transformer module is embedded in the MPAM module and adopts a decomposed axial attention mechanism. The axial attention expression is:

[0028]

[0029] where Q, K, and V are the query, key, and value matrices respectively, and d k is the dimension of the key. Attention is calculated in the height axis direction to capture vertical context information; attention is calculated in the width axis direction to capture horizontal context information, and attention calculations are performed in the height and width directions respectively.

[0030] Preferably, the multi-task learning head includes an object detection head, a bounding box regression head, and a semantic segmentation head, which share the multi-scale features extracted by MPAM. Among them, the object detection head is responsible for predicting the category and location of the object and outputting the detection result; the bounding box regression head accurately locates the bounding box of the object and optimizes the regression of the bounding box; the semantic segmentation head performs pixel-level classification on the image;

[0031] The multi-task learning head adopts an ensemble learning method. While combining detection models optimized by multiple MPAMs with different structures or different training strategies, it performs weighted averaging or uses a voting mechanism.

[0032] Preferably, the loss function of the multi-task learning head is composed of the loss functions of each task, which are weighted by weight coefficients to form a joint loss function. The joint loss function includes Focal Loss and GIoU Loss. The expression of the joint loss function is as follows:

[0033]

[0034] Among them, is the object detection loss, which adopts Focal Loss, is the bounding box regression loss, which adopts GIoU Loss, is the semantic segmentation loss, which adopts the cross-entropy loss function, and λ and μ are weight coefficients;

[0035] The calculation formula of Focal Loss is as follows:

[0036] FL(p t )=-α t (1-p t )γlog(p t );

[0037] Among them, p t is the prediction probability, α t is the class weight, and γ is the adjustment factor;

[0038] The calculation formula of GIoU Loss is as follows:

[0039]

[0040] Among them, C is the smallest closed matrix containing the predicted box A and the ground truth box B; |(A∩B)| represents the intersection area of the predicted box A and the ground truth box B, |A∪B| represents the union area of the predicted box A and the ground truth box B, IoU represents the overlapping degree of the predicted box and the ground truth box, and is a key indicator to measure the accuracy of object detection; \ represents the area of the region in the smallest closed rectangle that is not occupied by the predicted box and the ground truth box.

[0041] Preferably, the algorithm training method adopts a method including a cosine annealing learning rate scheduling strategy, data augmentation technology, and mixed-precision training technology.

[0042] Preferably, for the cosine annealing learning rate scheduling strategy, in combination with the learning rate warm-up stage in the first three epochs, the training process is optimized. Here, an epoch refers to the process in which the entire training dataset is processed and learned once in the neural network; cosine annealing means that during the training process, the learning rate gradually decreases according to a cosine curve as the number of training rounds increases; learning rate warm-up means that within the first three epochs of training, the learning rate linearly increases to the initial learning rate. The learning rate lr(t) at any time is:

[0043]

[0044] where t is the current training round, T is the total number of training rounds, T warmup is the number of rounds in the warm-up stage, lr 0 is the initial learning rate, lr final is the final learning rate.

[0045] Preferably, the data augmentation technology includes random cropping, rotation, and scaling;

[0046] Among them, random cropping is to randomly crop different regions of the image to ensure that the model can pay attention to various parts of the image, especially the locations of small targets; rotation is to randomly rotate the image by a certain angle; scaling is to randomly scale the size of the image to simulate the object detection scenarios at different resolutions.

[0047] Preferably, for the mixed-precision training technology, torch.cuda.amp in PyTorch is used to perform calculations using both 16-bit and 32-bit floating-point numbers simultaneously.

[0048] Therefore, the present invention adopts the above lightweight multi-task small target detection algorithm with adaptive pyramid and multi-stage path aggregation and its training method, and the beneficial effects are as follows:

[0049] (1) By combining Adaptive FPN, the present invention realizes the dynamic fusion weight adjustment of features at different levels, enhancing the expression ability of multi-scale features.

[0050] (2) In the present invention, the MPAM module introduces a lightweight Transformer module on the original basis, uses the axial attention mechanism for global feature modeling, improves the model's perception ability for small targets, and optimizes the channel attention mechanism to reduce the computational overhead.

[0051] (3) The present invention adopts a multi-task learning strategy to simultaneously perform object detection, bounding box regression, and semantic segmentation, fully sharing multi-scale features and improving the overall detection performance.

[0052] (4) The present invention combines the cosine annealing learning rate scheduling strategy with a warm-up stage, applies data augmentation techniques such as random cropping, rotation, and scaling, and adopts mixed-precision training techniques to further improve the training efficiency and model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is the overall model architecture diagram of the embodiment of the lightweight multi-task small object detection algorithm and its training method of the adaptive pyramid and multi-stage path aggregation of the present invention;

[0054] Figure 2 is the detailed structure diagram of the MPAM module of the embodiment of the lightweight multi-task small object detection algorithm and its training method of the adaptive pyramid and multi-stage path aggregation of the present invention;

[0055] Figure 3 is the learning rate scheduling curve during the training process of the embodiment of the lightweight multi-task small object detection algorithm and its training method of the adaptive pyramid and multi-stage path aggregation of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The technical solutions of the present invention will be further described below with reference to the drawings and embodiments.

[0057] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the field to which the present invention belongs.

[0058] Embodiment

[0059] As Figure 1 shown, for the lightweight multi-task small object detection algorithm and its training method of the adaptive pyramid and multi-stage path aggregation, the model significantly improves the accuracy and efficiency of small object detection by dynamically adjusting the fusion weights of features at different levels, introducing a lightweight Transformer module for global feature modeling, adopting an optimized axial attention mechanism (Axial Attention), and realizing feature sharing for multi-task learning (object detection, bounding box regression, and semantic segmentation).

[0060] As Figure 2 shown, the detailed structure of the multi-stage path aggregation module (MPAM) includes the implementation details of multi-stage path aggregation (Multi-stage Path Aggregation) and the axial attention mechanism (Axial Attention). Figure 2 Intuitively describes the flow of data inside the MPAM and the interaction between components. The processing flow based on the multi-stage path aggregation module (MPAM) is as follows:

[0061] S1. Feature Maps (Input Feature Maps), which are multi-scale feature maps from an Adaptive Feature Pyramid Network (Adaptive FPN), usually including feature layers of different resolutions (e.g., low-level high-resolution and high-level low-resolution feature maps), are input into a Multi-stage Path Aggregation module. It contains multiple parallel paths, and each path uses convolutional kernels of different sizes (such as 3×3, 5×5) and pooling operations to extract multi-scale features. This step indicates that the input multi-scale feature maps enter the multi-stage path aggregation module for processing and fusion.

[0062] S2. Similar to S1, but may use different convolutional kernel sizes or pooling strategies to further enrich the feature representation.

[0063] S3. Further extract higher-level multi-scale features to enhance the diversity and representation ability of the feature maps. The outputs of each path point to a Feature Fusion Node. The multi-scale features from each stage are fused through concatenation or weighted sum to form a comprehensive multi-scale feature representation. This step indicates that the multi-scale features extracted by different stage paths point to the feature fusion node through arrows for comprehensive fusion.

[0064] S4. The fused feature maps are output to the convolutional layer of a Lightweight Transformer Module (Lightweight TransformerModule). Using the decomposed axial attention mechanism (Axial Attention Mechanism), attention calculations are performed in the height and width directions respectively to achieve global feature modeling. This step indicates that the fused feature maps enter the lightweight Transformer module for global feature modeling.

[0065] S5. The feature maps are output to the Query (Q), Key (K), and Value Convolutional Layers (Value Convolutions). A 1×1 convolutional layer is used to convert the fused feature maps into Query (Q), Key (K), and Value (V) matrices.

[0066] S6. Point the Q, K, and V convolutional layers to the Height Axis Attention calculation module to calculate attention in the height axis direction and capture vertical context information. Point the Q, K, and V convolutional layers to the Width Axis Attention calculation module to calculate attention in the width axis direction and capture horizontal context information. This step means that after the Q, K, and V matrices are generated through the convolutional layer, they enter the attention calculation modules in the height and width directions respectively.

[0067] S7. Perform Feature Reconstruction based on the results of the height and width attention calculation modules. Superimpose the attention calculation results in the height and width directions to reconstruct the enhanced feature map. The global feature map processed by the axial attention mechanism further enhances the feature expression ability. This step means that the feature maps after height and width attention calculations are reconstructed and superimposed.

[0068] S8. Feed the Output Feature Map to the Transformer module. From the feature reconstruction node to the output feature map, this step means that the feature map processed by the axial attention mechanism is output for feature merging.

[0069] S9. Input the multi-stage path aggregation module and the lightweight Transformer module into the feature merging node. This step means that the global feature map output by the Transformer module is fused with the feature map output by the multi-stage path aggregation module.

[0070] S10. From the feature merging node to the Multi-task Learning Heads, including the Detection Head, Bounding Box Regression Head, and Semantic Segmentation Head, perform feature map fusion. Through addition or concatenation operations, form the final high-quality feature representation. The final enhanced feature map after fusion is used for subsequent object detection, bounding box regression, and semantic segmentation tasks. This step means that the final high-quality feature map after fusion is passed to the multi-task learning head to complete tasks such as object detection, bounding box regression, and semantic segmentation.

[0071] The Adaptive Feature Pyramid Network (Adaptive FPN) dynamically adjusts the fusion weights of features at different levels, enhancing the expressive power of multi-scale features; through learnable weight parameters, it adaptively optimizes the fusion ratio of features at different levels. Through Multi-stage Path Aggregation, it realizes the deep fusion of multi-scale information; each stage contains multiple parallel paths, using convolutional kernels and pooling operations of different sizes to extract rich multi-scale features. Through the Lightweight Transformer Module, it conducts global feature modeling, improving the model's perception ability for small targets; it embeds the Axial Attention mechanism, decomposes the attention calculation in the height and width directions, and reduces the computational complexity. While maintaining the global feature modeling ability, it significantly reduces the computational overhead through the Axial Attention Mechanism; it generates query, key, and value matrices through 1×1 convolution; calculates attention along the height axis to capture context information in the vertical direction; calculates attention along the width axis to capture context information in the horizontal direction; superimposes the attention results in the height and width directions to reconstruct the enhanced feature map. Through the Feature Merge Node, it fuses the feature maps output by the multi-stage path aggregation and the Transformer module to form the final high-quality feature representation; it fuses the feature maps from two sources through addition or concatenation operations. Through the Multi-task Learning Heads, it simultaneously completes object detection, bounding box regression, and semantic segmentation tasks; implementation details: each task has an independent exclusive branch, and the multi-scale features extracted by the shared MPAM are used for task-specific outputs.

[0072] Combined with the Adaptive Feature Pyramid Network and the Axial Attention Mechanism, the MPAM module effectively improves the model's detection ability for small targets while maintaining the advantage of computational efficiency. Figure 2 It shows the processing process of data inside the module, clearly demonstrating the core position and innovative design of the multi-stage path aggregation module (MPAM) in the present invention, and reflecting the technical advantages of the present invention in multi-scale feature fusion, global feature modeling, and multi-task learning.

[0073] 1. Adaptive Network Pyramid Network (Adaptive FPN)

[0074] The Adaptive Feature Pyramid Network (Adaptive FPN) aims to enhance the expressive power of multi-scale features by dynamically adjusting the fusion weights of features at different levels. Traditional FPN performs multi-scale feature fusion through a fixed topological structure and cannot adaptively adjust the feature fusion strategy according to the specific tasks and dataset requirements. Adaptive FPN introduces learnable fusion weights, enabling the model to adaptively optimize the fusion ratio of features at different levels according to the characteristics of the input data, thereby improving the expression effect of multi-scale features, especially when dealing with small objects.

[0075] The architecture of Adaptive FPN is based on traditional FPN and mainly consists of the following parts:

[0076] Low-level feature extraction: Use the backbone network to extract multi-level feature maps, usually including high-resolution features at low levels and low-resolution features at high levels.

[0077] Feature fusion layer: Introduce learnable weight parameters between each feature layer and fuse features at different levels through weighted summation. Specifically, for each feature layer, Adaptive FPN uses a small fully connected network (such as 1×1 convolution) to generate fusion weights, and these weights are dynamically adjusted according to the characteristics of the input features.

[0078] Multi-scale feature output: The feature maps after adaptive fusion are passed to the subsequent MPAM module as the input of multi-scale features.

[0079] The feature fusion process of Adaptive FPN can be expressed as:

[0080]

[0081] Among them, F 1 , F 2 , …, F n are the feature maps extracted by the backbone network, F′ i represents the feature map of the i-th layer, w i,j are learnable weight parameters, dynamically generated by a small fully connected network, satisfying

[0082] 2. Multi-stage Path Aggregation Module (MPAM)

[0083] The Multi-stage Path Aggregation Module (MPAM) enhances the expressive power of the feature map through multi-level and multi-scale information fusion. In the present invention, the MPAM further introduces a lightweight Transformer module for global feature modeling on the original basis to improve the model's perception ability for small targets. In addition, the MPAM optimizes the cross-layer fusion and global modeling capabilities of features through multi-stage path aggregation and axial attention mechanism.

[0084] The MPAM module mainly consists of multi-stage path aggregation, a lightweight transformer module, and feature merging;

[0085] Multi-stage Path Aggregation:

[0086] (1) Stage 1, Stage 2, Stage 3: Each stage contains multiple parallel paths, and each path extracts multi-scale features through convolution kernels and pooling operations of different sizes.

[0087] (2) Feature fusion: The outputs of each path are fused by concatenation or weighted summation to form a comprehensive expression of multi-scale features.

[0088] The MPAM module is in the Pyramid Residual Feature Mapping Sub-module (PRFM) and the Dilated Convolution Path Aggregation Sub-module (DCPAM), and combines convolution kernels and pooling operations of different sizes.

[0089] The PRFM module is a sub-module for multi-scale feature fusion, especially used in pyramid networks (such as FPN). The working principle of PRFM is to use a pyramid structure to process feature maps of different resolutions, then fuse these feature maps of different scales with residual connections, and finally optimize the flow of cross-scale information, reduce information loss, and enhance the feature expression ability. Therefore, the main task of this sub-module is to process features from different scales, enhance the flow of multi-scale information through residual connections (Residual Connections), prevent information loss, and improve the stability and convergence speed of the model;

[0090] The DCPAM module first makes the convolution kernel "jump" and sample features in the original image through dilated convolution, thereby expanding the receptive field; then through a multi-scale aggregation strategy, combines features from different resolutions to further improve the model's detection ability for multi-scale targets, enhance the diversity of the feature map, and enable the model to better handle small targets of different sizes and complex backgrounds. Therefore, the design of the DCPAM sub-module may utilize dilated convolution to aggregate feature information from different scales and enhance the model's perception ability for small targets.

[0091] Define the pseudo-code of the MPAM module:

[0092]

[0093]

[0094] 3. Lightweight Transformer module:

[0095] (1) Axial Attention mechanism: The decomposed axial attention mechanism is adopted to calculate the attention in the height and width directions respectively, significantly reducing the computational complexity.

[0096] (2) Global feature modeling: Through the lightweight Transformer module, global modeling is performed on the fused multi-scale features to capture the long-distance context information in the image and enhance the perception ability of small targets.

[0097] (3) Feature Merge:

[0098] The feature maps after multi-stage path aggregation and global feature modeling are further fused to form the final high-quality feature representation for the multi-task learning head to use.

[0099] The feature fusion F of the multi-stage path aggregation part agg can be expressed as:

[0100] F agg = Concat(F stage1 , F stage2 , F stage3 );

[0101] Among them, F stage1 , F stage2 , F stage3 respectively represent the feature maps after each stage of path aggregation. The global feature modeling F of the lightweight Transformer module global can be expressed as:

[0102] F global = AxialAttention(F agg );

[0103] The final feature merge F final is expressed as:

[0104] F final = Merge(F agg , F global ).

[0105] Define the pseudo-code of the lightweight Transformer module:

[0106] classdef LightweightTransformer<handle

[0107] properties

[0108] axialAttention

[0109] end

[0110] methods

[0111] function obj = LightweightTransformer(channels, heads)

[0112] obj.axialAttention = AxialAttention(channels, heads);

[0113] end

[0114] function out = forward(obj, x)

[0115] out = obj.axialAttention(x);

[0116] end

[0117] end

[0118] end

[0119] 4. Optimized Channel Attention Mechanism (Axial Attention)

[0120] The channel attention mechanism aims to enhance the expression ability of important features by assigning different weights to different channels. By adopting the axial attention mechanism (Axial Attention), while maintaining the global feature modeling ability, the computational overhead is significantly reduced, and the overall performance of the model is improved.

[0121] The Axial Attention mechanism decomposes global attention into axial attention and calculates attention in the height and width directions respectively, thereby reducing the computational complexity. The specific design is as follows:

[0122] Query, key, and value calculation: Use a 1×1 convolutional layer to convert the input feature map into query, key, and value.

[0123] Axial Attention Calculation: Calculate attention in the height axis direction to capture vertical context information; calculate attention in the width axis direction to capture horizontal context information, and superimpose the attention calculation results in the two directions to obtain the final attention-enhanced feature map. Let the input feature map be X ∈ R H×W×C , and the calculation process is as follows:

[0124] (1) Query, Key, and Value Calculation:

[0125] Q = Conv1×1(X), K = Conv1×1(X), V = Conv1×1(X);

[0126] (2) Height Direction Attention Calculation:

[0127]

[0128] (3) Width Direction Attention Calculation:

[0129]

[0130] (4) Feature Reconstruction:

[0131] X out = Attention H (Q, K, V) + Attention W (Q, K, V).

[0132] Define the pseudo-code of the Axial Attention module:

[0133]

[0134]

[0135] 5. Multi-task Learning

[0136] Multi-task learning can improve the generalization ability and robustness of the model by optimizing multiple related tasks simultaneously. In the present invention, in addition to the object detection task, the bounding box regression and semantic segmentation tasks are also introduced, and the multi-scale features extracted by MPAM are shared to improve the overall detection performance. Specifically, multi-task learning can promote information sharing between different tasks and enhance the expressive ability of features. Especially when dealing with small objects, fine-grained feature information can be better captured through shared features.

[0137] The multi-task learning part includes the following components:

[0138] Shared Feature Extraction: The multi-scale features extracted by the MPAM module are shared for multiple tasks, including object detection, bounding box regression, and semantic segmentation.

[0139] Task-Specific Branches:

[0140] Object Detection Head: Responsible for predicting the class and location of objects and outputting detection results.

[0141] Bounding Box Regression Head: Precisely locates the bounding boxes of objects and optimizes the regression of bounding boxes.

[0142] Semantic Segmentation Head: Classifies images at the pixel level to enhance the model's fine-grained understanding of target regions.

[0143] Let the shared feature be F final , and the outputs of each task can be expressed as:

[0144] Object Detection Output: Detection = DetectionHead(F final );

[0145] Bounding Box Regression Output: BBoxRegression = BBoxRegressionHead(F final );

[0146] Semantic Segmentation Output: Segmentation = SegmentationHead(F final ).

[0147] The loss function of multi-task learning consists of the loss functions of each task, which are weighted by weight coefficients to form a joint loss function. This joint loss function includes Focal Loss and GIoU Loss, and the expression of the joint loss function is as follows:

[0148]

[0149] Among them, is the object detection loss, using Focal Loss, is the bounding box regression loss, using GIoU Loss, is the semantic segmentation loss, using the cross-entropy loss function, and λ and μ are weight coefficients used to balance the influence of each loss;

[0150] The calculation formula of Focal Loss is as follows:

[0151] FL(p t) = -α t (1 - p t )γlog(p t );

[0152] Where p t is the predicted probability, α t is the class weight, and γ is the adjustment factor;

[0153] The calculation formula of GIoU Loss is as follows:

[0154]

[0155] Where C is the smallest closed matrix containing the predicted box A and the ground truth box B; |(A ∩ B)| represents the intersection area of the predicted box A and the ground truth box B, |A ∪ B| represents the union area of the predicted box A and the ground truth box B, IoU represents the overlapping degree of the predicted box and the ground truth box, which is a key indicator to measure the accuracy of object detection; \ represents the area of the region in the smallest closed rectangle that is not occupied by the predicted box and the ground truth box, which helps GIoU consider the spatial relationship between the object boxes, thereby providing a more accurate loss function optimization.

[0156] The multi - task learning head is trained using a joint loss function to optimize class balance and bounding box regression, improving the precision and recall of small object detection.

[0157] Define the pseudo - code of the multi - task learning module:

[0158]

[0159]

[0160] Define the pseudo - code of the joint loss function:

[0161]

[0162]

[0163] 6. Algorithm training method

[0164] To fully exploit the performance of the model, the present invention adopts a series of optimized training methods, including learning rate scheduling strategies, data augmentation techniques, mixed - precision training, and model compression methods.

[0165] Adopt the cosine annealing learning scheduling strategy, combined with the learning rate warm-up stage in the first three epochs, to optimize the training process and improve the final performance of the model. Here, an epoch refers to the process in which the entire training dataset is processed and learned once in the neural network. That is, in one epoch, the model will perform one forward propagation and one backward propagation on all training samples, and then update the model's parameters according to the feedback of the loss function. Therefore, the number of epochs represents the number of times the dataset is reused during the training process. Cosine annealing means that during the training process, the learning rate gradually decreases according to the cosine curve as the number of training rounds increases, which helps the model achieve better convergence in the later stage. Learning rate warm-up means that within the first three epochs of training, the learning rate warm-up stage in the first 3 epochs means that in the initial stage of training, the learning rate will gradually increase from a small value until it reaches the set initial learning rate. This can avoid the learning rate being too large in the initial stage of training, resulting in overly drastic gradient updates and affecting the convergence of the model. The cosine annealing learning rate scheduling strategy means that in the later stage of training, the learning rate will gradually decrease to help the model find a more accurate optimal solution. The learning rate lr(t) at any time:

[0166]

[0167] where t is the current training round, T is the total number of training rounds, and T warmup is the number of rounds in the warm-up stage, which is T in the embodiments of the present invention warmup = 3, lr 0 is the initial learning rate, and lr final is the final learning rate.

[0168] Data augmentation techniques include random cropping, rotation, and scaling. Among them, random cropping randomly crops different regions of the image to ensure that the model can pay attention to all parts of the image, especially the locations of small targets; rotation randomly rotates the image by a certain angle to enhance the model's robustness to target rotation changes; scaling randomly scales the size of the image to simulate target detection scenarios at different resolutions. The digital augmentation technique specifically targets small targets, enhances the diversity of training data, and improves the generalization ability of the model.

[0169] Mixed-precision training technology uses both 16-bit and 32-bit floating-point numbers for calculation, improves training efficiency, reduces video memory occupancy, and is suitable for the training of large models and high-resolution images.

[0170] Model compression methods include pruning, quantization, and knowledge distillation, which optimize the deployment efficiency of the model on edge devices and are suitable for application scenarios with real-time detection requirements;

[0171] Among them, pruning techniques reduce model complexity by removing redundant neural network connections or channels to meet the deployment requirements of resource-constrained edge devices; quantization techniques convert model parameters and activations from 32-bit floating-point numbers to lower-bit representations, reducing the storage requirements and computational overhead of the model; knowledge distillation techniques train a lightweight student model to mimic the behavior of a larger and more performant teacher model, reducing model complexity while maintaining model performance.

[0172] Training results

[0173] Experimental results on the VisDrone2019 and MS COCO2017 datasets show that the model proposed in the present invention exhibits significant performance improvement in small object detection tasks.

[0174] VisDrone2019

[0175] The mAP50 increased by 3.5%, from 42.3% to 45.8%.

[0176] The APS (accuracy of small object detection) increased by 2.9%, from 14.8% to 17.7%.

[0177] The inference speed remained at the real-time detection standard (>30 FPS), suitable for edge device deployment.

[0178] MS COCO2017

[0179] The mAP increased by 3.2%, from 0.646 to 0.678.

[0180] The accuracy of small object detection increased by 3.0%, from 0.288 to 0.318.

[0181] The number of model parameters only increased by about 15%, maintaining a low computational overhead while improving detection performance.

[0182] These results indicate that by introducing Adaptive FPN, optimized Axial Attention mechanism, multi-task learning, and a series of training method optimizations, the model of the present invention has significant advantages in small object detection tasks and is suitable for various practical application scenarios.

[0183] Such as Figure 3As shown, parameter settings: total_epochs is the total number of training epochs (default is 50). armup_epochs is the number of epochs in the warm-up stage (default is 3). initial_lr is the initial learning rate (default is 1e-4). final_lr is the final learning rate (default is 1e-6). Learning rate calculation, warm-up stage (the first 3 epochs): The learning rate linearly increases from 0 to initial_lr. Cosine annealing stage (subsequent epochs): The learning rate gradually decreases according to a cosine curve from initial_lr to final_lr. Comparison curve: The red dashed line represents the change in the learning rate when only the cosine annealing strategy is adopted. Figure 3 Intuitively demonstrates the advantages of the learning rate scheduling strategy that combines cosine annealing and the warm-up stage during training. Through the linear increase in the learning rate in the warm-up stage, the model can quickly adapt at the beginning of training, avoiding instability caused by too high a learning rate; subsequently, the cosine annealing strategy causes the learning rate to gradually decrease, which helps the model achieve more stable convergence in the later stage and improve the final detection performance.

[0184] Therefore, the present invention adopts the above lightweight multi-task small target detection algorithm and its training method of adaptive pyramid and multi-stage path aggregation. By combining the adaptive feature pyramid network and the multi-stage path aggregation module, the fusion weights of multi-scale features are dynamically adjusted, a lightweight Transformer module is introduced for global feature modeling, the optimized Axial Attention mechanism is adopted, and multi-task learning is achieved, effectively improving the accuracy and efficiency of small target detection. Combining the cosine annealing learning rate scheduling strategy, data augmentation technology, mixed precision training, and model compression method further optimizes the training process and deployment efficiency of the model. The experimental results of the method of the present invention on multiple public datasets verify its superiority, and it has broad application prospects and commercial value.

[0185] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A lightweight multi-task small target detection algorithm based on adaptive pyramid and multi-stage path aggregation, characterized by: Including Adaptive FPN, Multi-stage Path Aggregation Module MPAM, Lightweight Transformer Module, Optimized Channel Attention Mechanism, Multi-task Learning Head and Algorithm Training Method; The adaptive feature pyramid network uses learnable weight parameters to automatically optimize the fusion ratio of features at different levels; The multi-stage path aggregation module MPAM is connected to the adaptive feature pyramid network to achieve the fusion of multi-scale information through multi-stage path aggregation. Each stage is embedded with a lightweight Transformer module for global feature modeling. The optimized channel attention mechanism uses axial attention to perform attention calculations in the height and width directions respectively; The multi-task learning head includes the target detection head, the bounding box regression head and the semantic segmentation head. While performing target detection, the multi-task learning head adds the bounding box regression and semantic segmentation tasks and shares the multi-scale features extracted by MPAM.

2. The lightweight multi-task small target detection algorithm based on adaptive pyramid and multi-stage path aggregation according to claim 1 is characterized in that: The architecture of the adaptive feature pyramid includes bottom-level feature extraction, feature fusion layer, and multi-scale feature output. The bottom-level feature extraction uses the backbone network to extract multi-level feature maps, including low-level high-resolution features and high-level low-resolution features. In the feature fusion layer, learnable weight parameters are introduced between each feature layer. Features at different levels are fused by weighted summation. For each feature layer, AdaptiveFPN uses a small fully connected network to generate fusion weights. These weights are dynamically adjusted according to the characteristics of the input features. The feature fusion process is expressed as: Among them, F1, F2, ..., F n is the feature map extracted by the backbone network, F′ i represents the feature map of the i-th layer, w i,j are learnable weight parameters, dynamically generated through a small fully connected network, satisfying The multi-scale feature output is the feature map after adaptive fusion, and the multi-scale feature output is subsequently passed to the MPAM module as the input of the multi-scale feature; The adaptive feature pyramid network uses a dynamic weight adjustment mechanism combined with a self-supervised learning method to adaptively optimize the fusion weights of features at different levels.

3. The lightweight multi-task small target detection algorithm of adaptive pyramid and multi-stage path aggregation according to claim 2 is characterized in that: The MPAM module is mainly composed of multi-stage path aggregation, lightweight Transformer module, feature merging, and feature fusion of the multi-stage path aggregation part. agg It is expressed as: F agg =Concat(F stage1 ,F stage2 ,F stage3 ); Among them, F stage1 , F stage2 , F stage3 They represent the feature graphs after path aggregation at each stage; Global feature modeling of lightweight Transformer module F global It is expressed as: F global =AxialAttention(F agg ); The final feature merging F final It is expressed as: F final =Merge(F agg ,F global ); The MPAM module adopts model compression methods in its design, including pruning, quantization and knowledge distillation; Among them, pruning technology removes redundant neural network connections or channels to reduce model complexity; quantization technology converts model parameters and activation 32-bit floating-point numbers into lower-bit representations; knowledge distillation technology trains a lightweight student model to imitate the behavior of a larger and superior teacher model; The MPAM module is in the pyramid residual feature mapping submodule PRFM and the dilated convolution path aggregation submodule DCPAM, and combines convolution kernels of different sizes and pooling operations. Among them, the PRFM module is a submodule for multi-scale feature fusion; the DCPAM module uses dilated convolution to aggregate feature information from different scales; The MPAM module adopts a multi-branch structure and cross-layer feature fusion technology to aggregate multi-scale information at different levels.

4. The lightweight multi-task small target detection algorithm of adaptive pyramid and multi-stage path aggregation according to claim 3 is characterized in that: The lightweight Transformer module is embedded in the MPAM module and adopts a decomposed axial attention mechanism. The axial attention expression is: Among them, Q, K, and V are query, key, and value matrices respectively, and d k The dimension of the key is used to calculate the attention in the direction of the height axis to capture the vertical context information; the attention is calculated in the direction of the width axis to capture the horizontal context information, and the attention is calculated in the height and width directions respectively.

5. The lightweight multi-task small target detection algorithm of adaptive pyramid and multi-stage path aggregation according to claim 4 is characterized in that: The multi-task learning head includes a target detection head, a bounding box regression head and a semantic segmentation head, which share the multi-scale features extracted by MPAM. The target detection head is responsible for predicting the category and position of the target and outputting the detection result; the bounding box regression head accurately locates the bounding box of the target and optimizes the regression of the bounding box; the semantic segmentation head classifies the image at the pixel level; The multi-task learning head adopts an ensemble learning method, combining multiple detection models optimized by MPAM with different structures or different training strategies, and performing weighted averaging or using a voting mechanism.

6. The lightweight multi-task small target detection algorithm of adaptive pyramid and multi-stage path aggregation according to claim 5 is characterized in that: The loss function of the multi-task learning head is composed of the loss functions of each task, which are weighted by the weight coefficient to form a joint loss function. The joint loss function includes Focal Loss and GIoU Loss. The expression of the joint loss function is as follows: in, For target detection loss, FocalLoss is used. For the bounding box regression loss, GIoU Loss is used. is the semantic segmentation loss, using the cross entropy loss function, with λ and μ as weight coefficients; The FocalLoss calculation formula is as follows: FL(p t )D-α t (1-p t )γlog(p t )4 Among them, p t is the predicted probability, α t is the category weight, γ is the adjustment factor; The GIoU Loss calculation formula is as follows: Among them, C is the minimum closed matrix containing the predicted box A and the real box B; |(A∩B)| represents the intersection area of ​​the predicted box A and the true box B, |A∪B| represents the union area of ​​the predicted box A and the true box B, IoU represents the overlap between the predicted box and the true box, and is a key indicator for measuring the accuracy of target detection; \ represents the area of ​​the minimum closed rectangle that is not occupied by the predicted box and the true box.

7. The lightweight multi-task small target detection algorithm based on adaptive pyramid and multi-stage path aggregation according to claim 6 is characterized in that: The algorithm training method adopts cosine annealing learning rate scheduling strategy, data enhancement technology and mixed precision training technology.

8. The lightweight multi-task small target detection algorithm of adaptive pyramid and multi-stage path aggregation according to claim 7 is characterized in that: The cosine annealing learning rate scheduling strategy optimizes the training process in combination with the learning rate warm-up phase of the first three epochs, where epoch refers to the process in which the entire training data set is processed and learned once in the neural network; cosine annealing is that during the training process, the learning rate gradually decreases according to the cosine curve with the number of training rounds; learning rate warm-up is that within the first three epochs of training, the learning rate increases linearly to the initial learning rate, and the learning rate lr(t) at any time is: Among them, t is the current training round number, T is the total training round number, T warmup is the number of rounds in the warm-up phase, lr0 is the initial learning rate, lr final is the final learning rate.

9. The lightweight multi-task small target detection algorithm based on adaptive pyramid and multi-stage path aggregation according to claim 8, characterized in that: Data augmentation techniques include random cropping, rotation, and scaling; Among them, random cropping is to randomly crop different areas of the image to ensure that the model can focus on various parts of the image, especially the location of small targets; rotation is to randomly rotate the image by a certain angle; scaling is to randomly scale the image size to simulate target detection scenarios at different resolutions.

10. The lightweight multi-task small target detection algorithm and training method thereof of adaptive pyramid and multi-stage path aggregation according to claim 9, characterized in that: Mixed precision training technology uses PyTorch's torch.cuda.amp to calculate using both 16-bit and 32-bit floating point numbers.

Citation Information

Patent Citations

  • Remote-sensing image building change detection method

    CN110705457A

  • High-resolution remote sensing image target detection method of M-F-Y type lightweight convolutional neural network

    CN111666836A

  • Small target detection algorithm fusing attention and multi-scale double pyramids

    CN115311524A

  • YOLOv5-based lightweight ultraviolet image target detection method and system

    CN117671235A

  • Medical small target segmentation method and system based on multi-scale feature fusion and two-stage joint learning

    CN117876677A

Cited By

  • Lightweight target detection model integrating attention mechanism, optimization feature fusion method and self-supervised learning

    CN120070858A

  • A lightweight target detection system integrating attention mechanism, optimizing feature fusion method and self-supervised learning

    CN120070858B

  • Multi-task AI video analysis and detection method and system of shared network mechanism

    CN120499415A

  • A multi-task AI video analysis detection method and system sharing a network mechanism

    CN120499415B

  • Lightweight AI-based distribution line unmanned aerial vehicle edge end real-time visual identification and target detection method and system

    CN121459227A