A Small Target Detection Method from the Perspective of UAV Based on Self-Attention Mechanism
Through the drone perspective small object detection method with self-attention mechanism, combined with feature extraction module, local self-attention and cross-scale self-attention mechanism, the adaptability and real-time problems of small object detection in the drone perspective are solved, and efficient and accurate small object detection is achieved.
Patent Information
- Application Number
- CN202510478698.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The existing small-objective detection technology from the perspective of drones has problems such as poor adaptability, high computational complexity and insufficient real-time performance in complex environments, lighting changes and target scale differences, especially in extreme environments, there is room for further optimization of detection accuracy and multi-scale target recognition.
A drone perspective small object detection method based on self-attention mechanism is adopted. Through the feature extraction module, local self-attention module and cross-scale self-attention mechanism, combined with the decoder and adaptive loss function, a new feature extraction and loss function is designed to enhance the robustness and real-time nature of the network.
It effectively improves the recognition ability of small targets, improves detection accuracy and positioning accuracy, solves the detection problems of small targets in complex backgrounds, and achieves efficient and accurate small target detection.
Smart Images

Figure CN119992393B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target detection, and particularly to a small target detection method from the perspective of an unmanned aerial vehicle (UAV) based on a self-attention mechanism. Background Art
[0002] In recent years, UAV technology has developed rapidly and is widely used in fields such as agricultural monitoring, environmental protection, and disaster management. Among them, small target detection is an important link to improve the application efficiency and accuracy of UAVs. Existing small target detection technologies from the perspective of UAVs have certain limitations in complex environments, light changes, and target scale differences.
[0003] As an important technology in the field of deep learning in recent years, the self-attention mechanism can effectively capture long-distance feature information by modeling global dependencies, and particularly shows strong advantages in target detection tasks. Different from traditional convolutional neural networks that focus on local receptive fields, the self-attention mechanism can focus on key information in the image through weight assignment to achieve more accurate feature extraction and information fusion.
[0004] In the prior art, Chinese Patent with publication number CN119048730A discloses "A Target Detection Method for UAV Perspective". This method belongs to the field of target recognition and detection technology, aiming to improve the target detection ability of UAVs in complex backgrounds and high-dynamic environments, especially for the recognition of small aerial targets. The core innovation is to construct a target detection network CT-RODN, which includes multiple modules: a backbone network, including a frequency-domain based block composite attention module FBAM and a frequency-domain based CNN-Transformer module FCTB, a feature fusion module, and an attention prediction head. This method mainly lies in being able to improve the detection accuracy in complex backgrounds, especially for small targets and multi-scale targets, and ensuring the reliability and stability of UAV target detection. However, the above method also has some deficiencies: First, its adaptability to complex weather and light changes is poor, and the detection accuracy will decrease in extreme environments; second, the computational complexity is relatively high and cannot meet the real-time requirement; finally, although this method can effectively detect multi-scale targets, there is still room for further optimization in the recognition accuracy of extremely small or extremely large targets.
[0005] Therefore, there is an urgent need to design a small target detection method from the perspective of UAV based on the self-attention mechanism to solve the problems existing in the above prior art. Summary of the Invention
[0006] The object of the present invention is to provide a small target detection method from the perspective of an unmanned aerial vehicle based on the self-attention mechanism. Through the feature extraction module for feature extraction, combined with the local self-attention module and the cross-scale self-attention mechanism, the recognition ability of the model for small targets can be effectively improved. The decoder and the adaptive loss function are used to further enhance the robustness and real-time performance of the network, and finally efficient and accurate small target detection is achieved.
[0007] To achieve the above object, the present invention adopts the following technical solutions:
[0008] In the first aspect of the present invention, a small target detection method from the perspective of an unmanned aerial vehicle based on the self-attention mechanism is provided. The method includes the following steps:
[0009] S1: Obtain an image data set suitable for the perspective of an unmanned aerial vehicle and preprocess it;
[0010] S2: Construct a small target detection model based on the self-attention mechanism. The small target detection model includes a feature extraction module, a local self-attention mechanism, a cross-scale self-attention mechanism, a decoder, and an adaptive loss function;
[0011] The specific steps of S2 include:
[0012] S21: Design a new feature extraction module to extract features from the input data set;
[0013] S22: Further strengthen the information modeling of the small target area, and design a new type of local self-attention module in the encoder part;
[0014] S23: In the fusion process of the multi-scale feature maps of the encoder, the cross-scale self-attention mechanism is used to combine the context information between different scales to enhance the detection effect of small targets;
[0015] S24: Introduce a decoder to decode the candidate target boxes through a query mechanism. To further strengthen the detection ability of small targets, position encoding is introduced to retain the spatial information of small targets in the image;
[0016] S25: Design a new UGWD adaptive loss function in network training. The UGWD adaptive loss function includes a positioning loss function and a classification loss function;
[0017] S3: Feed the preprocessed image data set into the small target detection model for training;
[0018] S4: Apply the trained small target detection model to detect the target to be detected and output the result.
[0019] As an embodiment of the present application, the specific steps of S1 include:
[0020] S11: Obtain the perspective image dataset captured by the drone, clean, annotate, and format the data, and standardize the original image dataset to ensure that the image size, resolution, and color channels are consistent to meet the network input requirements;
[0021] S12: Enhance the image dataset using GPU-accelerated image enhancement techniques, including rotation, scaling, and cropping;
[0022] S13: Denoise the enhanced image dataset by applying the Gaussian filtering method to remove the noise components in the images;
[0023] S14: Finally, perform format conversion and storage optimization on the image dataset.
[0024] As an embodiment of this application, the step S21 specifically includes:
[0025] S211: Input feature map , where is the batch size, is the number of channels, , are the height and width of the feature map;
[0026] S212: Use depthwise separable convolution to extract the local information of the input feature map and fuse the spatial position information through position encoding. The depthwise separable convolution includes depth convolution and pointwise convolution;
[0027] S213: Add a new adaptive attention mechanism to the result output by the depthwise separable convolution. Its calculation formula is as follows:
[0028]
[0029] Where, is the weighted feature map, is the feature map input to the current layer, represents element-wise multiplication, and are the learning parameters of the attention network, which are the weight matrix and the bias term respectively, is the feature map obtained through the convolution operation, representing the high-order features extracted from the input feature map, is the calculated attention coefficient matrix, indicating the degree of attention at each spatial position;
[0030] S214: Then introduce a cross-layer feature importance adjustment strategy. This strategy inserts specific adjustment units between different network modules to perform dynamic weighting and importance evaluation on multi-layer features. Its weighting formula is as follows:
[0031]
[0032]
[0033] Among them, is the attention weight, is the attention network, is the output feature map of the th layer, and are respectively the height, width, and number of channels of the layer feature map, is the weighted feature map, represents element-wise multiplication;
[0034] S215: Finally, use the structural reparameterization technique to merge multiple convolutional layers in the training phase into a standard 3×3 convolutional operation during the inference phase.
[0035] As an embodiment of the present application, the step S22 specifically includes:
[0036] S221: Use a 1×1 convolutional kernel to obtain the region containing small targets, and obtain a feature map focusing on the small target region ;
[0037] S222: Divide the feature map processed through step S21 into multiple small regions , and for each local region, calculate the relationship between pixels within the region through the local self-attention mechanism. The calculation formula is as follows:
[0038]
[0039]
[0040] Among them, and and represent the query, key, and value matrices respectively, and and are the learned parameter matrices, with a dimension of , is the number of channels after convolution, is the dimension of the query and key vectors, is transpose matrix, is the attention score matrix between pixels within each region;
[0041] S223: Concatenate the output features of each small region to obtain the enhanced region features.
[0042] As an embodiment of the present application, step S24 specifically includes:
[0043] S241: First, initialize the input of the decoder by adding the positional encoding and the query embedding;
[0044] S242: The query vector of each decoder and the image features output by the encoder are weighted and summed through the cross-attention mechanism;
[0045] S243: Inside the decoder, each query will interact with other queries through self-attention mechanism to capture the spatial relationship between targets;
[0046] S244: After passing through the self-attention mechanism, the output of the decoder is fed into a feed-forward network, which contains two fully connected layers, and a ReLU activation function is added in the middle of the two fully connected layers for non-linear transformation;
[0047] S245: After being processed by the feed-forward network, the final output is mapped to the category and position of the target through a fully connected layer, and finally the probability distribution of the target category corresponding to each query and the exact position of the target are generated and output.
[0048] As an embodiment of the present application, step S25 specifically includes:
[0049] S251: The UGWD adaptive loss function is a positioning loss function based on the improved Gaussian Wasserstein distance. For a single target, the positioning loss function of the Gaussian Wasserstein distance The calculation formula is as follows:
[0050]
[0051]
[0052] where is the Wassertein distance between two multivariate Gaussian distributions, and are the center points of the bounding boxes, and are the covariance matrices of the bounding boxes, used to describe the target shape, is a hyperparameter, used to adjust the behavior of the function, is a non-linear transformation function;
[0053] S252: It is set that the UGWD adaptive loss function includes a positioning loss function and a classification loss function, and its calculation formula is as follows:
[0054]
[0055] Among them, is the localization loss, is the classification loss, is the weight coefficient of the classification loss;
[0056] S253: A dynamic weighting mechanism based on the Gaussian function is introduced in the localization loss function such that the weight dynamically depends on the context features of the target and the uncertainty of the network prediction. Apply this loss function to the network outputs at different scales to optimize the localization of small and large targets. The localization loss function has the following calculation formula:
[0057]
[0058]
[0059] Among them, is the set of multi-scale feature layers, represents the set of targets in the th layer, is the loss weight based on Gaussian weighting, is the scale of the target, is the standard deviation, used to control the weighting range, is a dynamic adjustment function, depending on the target features and the prediction probability ; when approaches 1, indicating that no additional attention is required;
[0060] S254: The Gaussian weighting mechanism is also used in the classification loss function and its calculation formula is as follows:
[0061]
[0062] Among them, is the prediction probability of the th target, is the same weighting coefficient as in the localization loss to ensure consistency;
[0063] S255: The UGWD adaptive loss function finally combines the localization loss function and the classification loss function , and balances the two through the weight . Its calculation formula is as follows:
[0064] .
[0065] As an embodiment of the present application, step S3 specifically includes:
[0066] S31: Divide the preprocessed image dataset into a training set and a validation set and feed them into the small target detection model;
[0067] S32: During the training process, use the AdamW optimization algorithm to adjust the model parameters, and use the early stopping method to avoid overfitting and accelerate the training process;
[0068] S33: Set the weight decay value to 0.0001, gradually increase the learning rate at the beginning of training until the set maximum value, and then start to decay, ensuring that the model can converge better and avoid overfitting or underfitting.
[0069] As an embodiment of the present application, step S4 specifically includes:
[0070] S41: Transmit the real-time image or video stream collected from the UAV perspective to the small target detection model for scaling and normalization processing;
[0071] S42: Use the trained small target detection model to process the input image dataset;
[0072] S43: Finally, the small target detection model outputs the class labels, bounding box coordinates, and corresponding confidence scores of the detected targets;
[0073] S44: Use the PySide6 framework to visually present the target detection results to the user and the system, facilitating the intuitive display of the bounding box, class name, and confidence level, thus supporting manual intervention and subsequent processing.
[0074] The beneficial effects of the present invention are as follows:
[0075] (1) By designing a brand-new feature extraction module, including depthwise separable convolution, position encoding, adaptive attention mechanism, cross-layer feature importance adjustment strategy, and hierarchical reparameterization technology, the present invention can extract effective target features from images of different resolutions and scales, effectively improving the multi-scale feature extraction ability of the input data, especially having a finer representation ability for the detailed features of small target regions; through the adaptive attention mechanism and cross-layer feature adjustment, the feature extraction module can dynamically adjust the attention to different regions and different scale targets and avoid information loss, effectively solving the problem that small targets are easily overlooked in complex backgrounds.
[0076] (2) Through the novel local self-attention module, the present invention further strengthens the modeling ability for small target regions. Through feature aggregation within a local range, the network's perception and capture ability for small target features are improved, effectively reducing the phenomena of missed detection and false detection.
[0077] (3) During the multi-scale feature fusion process, the present invention innovatively introduces a cross-scale self-attention mechanism, which can dynamically combine the context information of different-scale features, significantly enhancing the detection effect of small targets and solving the problem that small targets are prone to loss during multi-scale feature fusion.
[0078] (4) The present invention introduces a decoder based on a query mechanism to decode candidate target boxes, retaining the spatial information of small targets in the image and further improving the localization accuracy and feature decoding ability of the targets. This design enables the network to achieve more accurate localization and classification when processing multi-target tasks.
[0079] (5) Through the new UGWD adaptive loss function, the present invention comprehensively considers the classification loss and the bounding box regression loss. Especially through the dynamic weighting mechanism, it can adaptively adjust the loss weights during the training process, enabling the network to obtain better performance in small target detection and solving the problem that small targets are easily ignored or misdetected. This loss function not only improves the detection accuracy but also enhances the training stability and reduces the risk of overfitting. Brief Description of the Drawings
[0080] Figure 1 It is a schematic flowchart of a small target detection method from the perspective of an unmanned aerial vehicle based on a self-attention mechanism provided in an embodiment of the present invention;
[0081] Figure 2 It is a schematic diagram of a small target detection model of a small target detection method from the perspective of an unmanned aerial vehicle based on a self-attention mechanism provided in an embodiment of the present invention;
[0082] Figure 3 It is a schematic diagram of the structure of a feature extraction module of a small target detection method from the perspective of an unmanned aerial vehicle based on a self-attention mechanism provided in an embodiment of the present invention;
[0083] Figure 4 It is a schematic diagram of the structure of a decoder of a small target detection method from the perspective of an unmanned aerial vehicle based on a self-attention mechanism provided in an embodiment of the present invention. Detailed Embodiment
[0084] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0085] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.
[0086] In the present invention, unless otherwise clearly defined and limited, terms such as "connection" and "fixation" shall be understood in a broad sense. For example, "fixation" can be a fixed connection, a detachable connection, or an integral body; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the internal connection of two components or the interaction relationship between two components, unless otherwise clearly limited. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0087] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the meaning of "and / or" appearing throughout the text includes three parallel solutions. Taking "A and / or B" as an example, it includes solution A, solution B, or the solution where A and B are satisfied simultaneously. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0088] Referring to Figures 1 to 4 , the first aspect of the present invention provides a small target detection method from the perspective of an unmanned aerial vehicle based on a self-attention mechanism. The method includes the following steps:
[0089] S1: Obtain an image data set suitable for the perspective of an unmanned aerial vehicle and preprocess it;
[0090] S2: Construct a small object detection model based on the self-attention mechanism. The small object detection model includes a feature extraction module, a local self-attention mechanism, a cross-scale self-attention mechanism, a decoder, and an adaptive loss function;
[0091] The specific steps of step S2 include:
[0092] S21: Design a brand-new feature extraction module to extract features from the input data set;
[0093] S22: Further strengthen the information modeling of small object regions, and design a new type of local self-attention module in the encoder part;
[0094] S23: During the fusion process of the multi-scale feature maps of the encoder, through the cross-scale self-attention mechanism, the context information between different scales can be dynamically combined, and the image information from different feature layers and perspectives can be fused, further improving the detection accuracy and the discrimination ability of target features, enhancing the detection effect of small objects, and solving the problem that small objects are easily lost in multi-scale feature fusion;
[0095] S24: Introduce a decoder to decode the candidate target boxes through a query mechanism. To further enhance the detection ability of small objects, position encoding is introduced to retain the spatial information of small objects in the image;
[0096] S25: Design a new UGWD adaptive loss function in network training. The UGWD adaptive loss function includes a localization loss function and a classification loss function;
[0097] Specifically, by introducing the self-attention mechanism, the present invention can effectively improve the performance of small object detection; the self-attention mechanism enables the small object detection model to focus on important regions in the image, especially those target regions with small sizes and less information, thereby enhancing the feature extraction ability of the model in these regions.
[0098] S3: Send the preprocessed image data set into the small object detection model for training;
[0099] S4: Apply the trained small object detection model to detect the target to be detected and output the result.
[0100] Specifically, the present invention optimizes the object detection network by designing a brand-new feature extraction module. Combining the local self-attention module and the cross-scale self-attention mechanism, it can effectively improve the recognition ability of small objects, especially in complex backgrounds. By performing refined feature extraction on small objects in the drone-view images, the detection accuracy and positioning accuracy can be improved. At the same time, designing the UGWD adaptive loss function and decoder further enhances the robustness and real-time performance of the network, thereby improving the overall performance of small object detection. Especially in dynamic environments, it shows higher stability and adaptability, and finally realizes efficient and accurate small object detection.
[0101] As an embodiment of the present application, step S1 specifically includes:
[0102] S11: Obtain the dataset of perspective images captured by the drone, and clean, annotate, and format the data to ensure that the dataset is suitable for subsequent processing. Perform standardization processing on the original image dataset to ensure that the image size, resolution, and color channels are consistent to meet the network input requirements.
[0103] S12: Use GPU-accelerated image enhancement techniques to enhance the image dataset, including rotation, scaling, and cropping. Accelerate these operations through GPU parallel computing to reduce the performance bottlenecks that may be encountered under traditional CPU processing, and significantly improve the data processing speed. This not only accelerates the preparation process of the dataset but also improves the robustness of the model by enhancing image diversity.
[0104] S13: Denoise the enhanced image dataset. Apply the Gaussian filtering method to remove the noise components in the images to improve the accuracy of object detection. The GPU-accelerated denoising algorithm can process a large amount of image data and quickly complete the denoising operation to ensure the quality of the dataset.
[0105] S14: Finally, perform format conversion and storage optimization on the image dataset to ensure that the data can be quickly loaded into the training process and is compatible with the deep learning framework. Adopt the GPU-accelerated data loading technology, NVIDIA's DALI library, to accelerate the batch processing and parallel loading of data and avoid the impact on the training speed during the data loading process.
[0106] Specifically, by selecting a representative and diverse dataset of drone images, the present invention can provide rich image information for subsequent tasks; these image datasets cover different geographical environments, weather conditions, and time variations, ensuring the generalization ability of the trained model; during the preprocessing process, the images are first cropped and scaled to ensure that the size and quality of the input data are adapted to the requirements of the model; secondly, through data augmentation techniques such as rotation, mirroring, and random cropping, the diversity of the data is increased, and the robustness of the model is improved; for the noise in the images, filtering and denoising algorithms are used to reduce the impact of the external environment on the image quality; at the same time, the images are normalized so that the pixel values of the images are within a unified range, facilitating efficient learning of the model; these preprocessing steps not only improve the quality of the data but also lay a solid foundation for subsequent feature extraction and classification tasks.
[0107] As an embodiment of the present application, the step S21 specifically includes:
[0108] S211: Input feature map , where is the batch size, is the number of channels, , are the height and width of the feature map;
[0109] S212: Use depthwise separable convolution to extract local information of the input feature map and fuse spatial position information through position encoding. The depthwise separable convolution includes depth convolution and pointwise convolution;
[0110] S213: Add a new adaptive attention mechanism to the result output by the depthwise separable convolution. Its calculation formula is as follows:
[0111]
[0112] where, is the weighted feature map, is the feature map input at the current layer, represents element-wise multiplication, and are the learning parameters of the attention network, which are the weight matrix and the bias term respectively, is the feature map obtained through convolution operations, representing the high-order features extracted from the input feature map, is the calculated attention coefficient matrix, indicating the degree of attention at each spatial position;
[0113] S214: Next, a cross-layer feature importance adjustment strategy is introduced. This strategy inserts specific adjustment units between different network modules to perform dynamic weighting and importance evaluation on multi-layer features, enabling the model to flexibly adjust the feature flow and contribution between different layers, thereby enhancing the ability to capture small target details. The weighting formula is as follows:
[0114]
[0115]
[0116] Among them, is the attention weight, is the attention network, is the output feature map of the -th layer, , , are respectively the height, width, and number of channels of the -th layer feature map, is the weighted feature map, represents element-wise multiplication;
[0117] S215: Finally, the structural reparameterization technology is used to merge multiple convolutional layers in the training stage into a standard 3×3 convolution operation during the inference stage.
[0118] Specifically, through the multi-level feature extraction module of the present invention, including depthwise separable convolution, position encoding, adaptive attention mechanism, cross-layer feature importance adjustment strategy, and hierarchical reparameterization technology, effective target features can be extracted from images of different resolutions and scales, effectively improving the multi-scale feature extraction ability of the input data, especially having a more refined representation ability for the detailed features of small target regions; by adopting position encoding and self-attention mechanism, the network can perform adaptive feature selection and weighting in both spatial and channel dimensions, thereby significantly improving the detection accuracy of small targets and further improving the accuracy and efficiency of the model; through the adaptive attention mechanism and cross-layer feature adjustment, the feature extraction module can dynamically adjust the attention to different regions and targets of different scales and avoid information loss, effectively solving the problem that small targets are easily ignored in complex backgrounds.
[0119] As an embodiment of the present application, the step S22 specifically includes:
[0120] S221: Use a 1×1 convolution kernel to obtain the region containing the small target, and obtain a feature map focusing on the small target region ;
[0121] S222: Divide the feature map processed through step S21 into multiple small regions , for each local region, the relationship between pixels within the region is calculated through a local self-attention mechanism, and its calculation formula is as follows:
[0122]
[0123]
[0124] Among them, , , respectively represent the query, key, and value matrices, , , are learned parameter matrices, with a dimension of , is the number of channels after convolution, is the dimension of the query and key vectors, is 's transpose matrix, is the attention score matrix between pixels within each region;
[0125] S223: Concatenate the output features of each small region to obtain enhanced region features.
[0126] Specifically, through the local self-attention module, the present invention further strengthens the modeling ability for small target regions, effectively reduces the phenomena of missed detection and false detection, and improves the network's perception and capture ability of small target features through feature aggregation within a local range.
[0127] As Figure 4 shown, as an embodiment of the present application, the step S24 specifically includes:
[0128] S241: Initialize the input of the decoder by adding position encoding and query embedding. The position encoding is used to provide spatial position information for each query in the decoder. Since the self-attention mechanism itself has no sequential or spatial information, except for adding to the query embedding, the position encoding usually adds a fixed or learnable vector to each position of the input sequence;
[0129] S242: The query vector of each decoder and the image features output by the encoder are weighted and summed through the cross-attention mechanism;
[0130] S243: Inside the decoder, each query will interact with other queries through self-attention to capture the spatial relationship between targets. The calculation method of the self-attention mechanism is that each query vector calculates the similarity with all query vectors in the decoder and adjusts its representation according to these similarities; in this way, the decoder can capture the spatial relationship and context information between targets;
[0131] S244: After the decoder passes through the self-attention mechanism, the output is sent to a feed-forward network. The feed-forward network includes two fully connected layers, and a ReLU activation function is added in the middle of the two fully connected layers for non-linear transformation to extract higher-level features; after the feed-forward network processes the self-attention mechanism, it will independently process the features at each position;
[0132] S245: After being processed by the feed-forward network, the final output is mapped to the target category and location (bounding box) through a fully connected layer, normalized through the Softmax function, and finally the probability distribution of the target category corresponding to each query and the exact location of the target are generated and output.
[0133] Specifically, the present invention introduces a decoder based on the query mechanism to decode the candidate target boxes, retains the spatial information of small targets in the image, and further improves the target positioning accuracy and feature decoding ability; this design enables the network to achieve more accurate positioning and classification when processing multi-target tasks.
[0134] As an embodiment of the present application, the step S25 specifically includes:
[0135] S251: The UGWD adaptive loss function is a positioning loss function based on the improved Gaussian Wasserstein distance, which is applicable to the regression task of bounding boxes and has good smoothness and the ability to model the target shape; for a single target, the positioning loss function of the Gaussian Wasserstein distance The calculation formula is as follows:
[0136]
[0137]
[0138] Among them, is the Wassertein distance between two multivariate Gaussian distributions, and are the center points of the bounding box, and are the covariance matrices of the bounding box, which are used to describe the target shape, is a hyperparameter used to adjust the behavior of the function, is a non-linear transformation function;
[0139] S252: It is set that the UGWD adaptive loss function includes a positioning loss function and a classification loss function, and its calculation formula is as follows:
[0140]
[0141] Among them, is the localization loss, is the classification loss, is the weight coefficient of the classification loss;
[0142] S253: A dynamic weighting mechanism based on the Gaussian function is introduced in the localization loss function . This mechanism particularly focuses on the accuracy performance of small objects. By improving the traditional Gaussian weighting method, the weight dynamically depends on the context features of the object and the uncertainty of the network prediction. Applying this loss function to the network outputs at different scales is used to significantly optimize the localization of small and large objects. The localization loss function has the following calculation formula:
[0143]
[0144]
[0145] Among them, is the set of multi-scale feature layers, represents the object set of the th layer, is the loss weight based on Gaussian weighting, is the scale of the object, is the standard deviation, used to control the weighting range, is a dynamic adjustment function, depending on the object features and the prediction probability ; when approaches 1, , indicating that no additional attention is required;
[0146] S254: The Gaussian weighting mechanism is also used in the classification loss function so that the classification errors in small object regions are more punished. Its calculation formula is as follows:
[0147]
[0148] Among them, is the prediction probability of the th object, is the same weighting coefficient as in the localization loss to ensure consistency;
[0149] S255: The UGWD adaptive loss function finally combines the localization loss function and the classification loss function , and balances the two through the weight . Its calculation formula is as follows:
[0150] 。
[0151] Specifically, through the designed UGWD adaptive loss function, the present invention comprehensively considers the classification loss and the bounding box regression loss. In particular, through the dynamic weighting mechanism, it can adaptively adjust the loss weights during the training process, enabling the network to achieve better performance in small target detection and solving the problem that small targets are easily ignored or misdetected. This loss function not only improves the detection accuracy but also enhances the training stability and reduces the risk of overfitting.
[0152] As an embodiment of the present application, step S3 specifically includes:
[0153] S31: Divide the preprocessed image dataset into a training set and a validation set and send them into the small target detection model;
[0154] S32: During the training process, use the AdamW optimization algorithm to adjust the model parameters and use early stopping to avoid overfitting and accelerate the training process;
[0155] S33: Set the weight decay value to 0.0001, gradually increase the learning rate at the beginning of the training until the set maximum value, and then start to decay to ensure that the model can converge better and avoid overfitting or underfitting.
[0156] Specifically, first divide the preprocessed image dataset into a training set and a validation set to ensure that the model performance can be effectively evaluated during the training process; during the training process, input the preprocessed images into the small target detection model, and the model will calculate the loss function through forward propagation and adjust the weights in the network according to the loss value; since small targets usually occupy a small pixel area in the image, the UGWD loss function is designed to more strictly optimize the positioning and classification accuracy of small target areas; after the training is completed, use the validation set to evaluate the model performance, including evaluation metrics such as detection accuracy, recall rate, and F1 value; if the performance of the model on the validation set is satisfactory, further test and deploy it; if there are performance bottlenecks, further optimize the training process by adjusting the model architecture or increasing the diversity of the dataset.
[0157] As an embodiment of the present application, step S4 specifically includes:
[0158] S41: Transmit the real-time images or video streams collected from the UAV perspective to the small target detection model for scaling and normalization processing;
[0159] S42: Use the trained small target detection model to process the input image dataset;
[0160] S43: Finally, the small object detection model outputs the class labels, bounding box coordinates, and corresponding confidence scores of the detected objects;
[0161] S44: Using the PySide6 framework, the object detection results are presented to the user and the system in a visual way, facilitating the intuitive display of the bounding boxes, class names, and confidence levels, thereby supporting manual intervention and subsequent processing.
[0162] Specifically, the trained and verified model can be applied to actual detection tasks; in this process, first, the image or video frame to be detected is input into the trained small object detection model after undergoing the same preprocessing process as the training dataset; the model will, through its forward propagation process, utilize the self-attention mechanism and the ability of multi-scale feature fusion to identify small object regions in the image; finally, the results output by the model include the positions (i.e., the coordinates of the bounding boxes), class labels, and corresponding confidence scores of the detected small objects, and the visualization of the detection results is achieved through the Pyside6 framework.
[0163] Through the small object detection model constructed based on the self-attention mechanism, the present invention will automatically adjust the detection strategy and parameters according to the target features in different scenarios, and improve the detection accuracy and robustness through continuous optimized training and verification processes. Each type of small object has its unique visual features, and the small object detection model can effectively distinguish according to these features, thereby improving the accuracy and classification performance of object detection; through this method, not only can small-sized objects be accurately identified, but also the classification can be further refined to provide more refined detection results; for example, in the traffic monitoring scenario, not only can vehicles and pedestrians be identified, but also their specific types (such as sedans, trucks, electric vehicles, etc.) can be further identified, providing important support for subsequent monitoring, tracking, and early warning systems; in addition, this module can also handle problems such as light changes and noise interference in complex environments, ensuring the stability and reliability of the detection results, and meeting the high-precision requirements for object detection in different application scenarios.
[0164] The present invention has high-precision small object detection capabilities, excellent feature extraction and enhancement capabilities, and strong information fusion advantages. It can effectively handle diverse and dynamically changing small objects in the drone shooting perspective, and has great practical application potential. It is especially suitable for small object detection tasks of drones in the fields of security, monitoring, agricultural monitoring, etc., and can provide effective support for drones in a variety of application scenarios; its application fields include military reconnaissance, disaster monitoring, environmental protection, agricultural monitoring, etc., and it has broad application prospects and market value.
[0165] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the embodiments of the present disclosure that have similar functions.
Claims
1. A small target detection method from the perspective of an unmanned aerial vehicle based on the self-attention mechanism, characterized in that, The method includes the following steps: S1: Obtain an image dataset suitable for the perspective of the drone and preprocess it; S2: Build a small target detection model based on the self-attention mechanism. The small target detection model includes a feature extraction module, a local self-attention module, a cross-scale self-attention mechanism, a decoder, and a UGWD adaptive loss function; Step S2 specifically includes: S21: Design a new feature extraction module to extract features from the input dataset; S22: Further strengthen the information modeling of small target regions and design a new type of local self-attention module in the encoder part; S23: In the process of multi-scale feature map fusion of the encoder, use the cross-scale self-attention mechanism to combine the context information between different scales and enhance the detection effect of small targets; S24: Introduce a decoder to decode the candidate target boxes through a query mechanism. To further strengthen the detection ability of small targets, introduce position encoding to retain the spatial information of small targets in the image; S25: Design a new UGWD adaptive loss function in network training. The UGWD adaptive loss function includes a localization loss function and a classification loss function; S3: Feed the preprocessed image dataset into the small target detection model for training; S4: Apply the trained small target detection model to detect the target to be detected and output the results; Step S21 specifically includes: S211: Input feature map , where is the batch size, is the number of channels, , are the height and width of the feature map; S212: Use depthwise separable convolution to extract the local information of the input feature map and fuse the spatial position information through position encoding. The depthwise separable convolution includes depth convolution and pointwise convolution; S213: Add a new adaptive attention mechanism to the result output by the depthwise separable convolution. The calculation formula is as follows: Among them, is the weighted feature map, is the feature map input to the current layer, represents element-wise multiplication, and are the learning parameters of the attention network, which are the weight matrix and the bias term respectively, is the feature map obtained through the convolution operation, representing the high-order features extracted from the input feature map, is the calculated attention coefficient matrix, indicating the degree of attention at each spatial position; S214: Then introduce a cross-layer feature importance adjustment strategy. This strategy dynamically weights and evaluates the importance of multi-layer features by inserting specific adjustment units between different network modules. The weighting formula is as follows: Among them, is the attention weight, is the attention network, is the output feature map of the -th layer, and are respectively the height, width and number of channels of the -th layer feature map, is the weighted feature map, represents element-wise multiplication; S215: Finally, use the structural reparameterization technique to merge multiple convolutional layers in the training stage into a standard 3×3 convolution operation in the inference stage; Introduce a dynamic weighting mechanism based on the Gaussian function and use weights to improve the localization loss function and the classification loss function. The weights are calculated as follows: Among them, is the loss weight based on Gaussian weighting, is the scale of the target, is the standard deviation, which is used to control the weighting range, is a dynamic adjustment function that depends on the target features and the predicted probability ; when approaches 1, , indicating that no additional attention is required.
2. The small target detection method from the perspective of an unmanned aerial vehicle based on the self-attention mechanism according to claim 1, characterized in that Step S1 specifically includes: S11: Obtain the perspective image dataset captured by the drone, clean, annotate, and format the data, and standardize the original image dataset to ensure that the image size, resolution, and color channels are consistent to meet the network input requirements; S12: Use GPU-accelerated image enhancement techniques to enhance the image dataset, including rotation, scaling, and cropping; S13: Denoise the enhanced image dataset and apply the Gaussian filtering method to remove the noise components in the image; S14: Finally, perform format conversion and storage optimization on the image dataset.
3. A small target detection method from the perspective of an unmanned aerial vehicle based on the self-attention mechanism according to claim 1, characterized in that Step S22 specifically includes: S221: Obtain the region containing small targets using a 1×1 convolution kernel, and obtain a feature map focusing on the small target region ; S222: Divide the feature map into multiple small regions , and for each local region, calculate the relationship between each pixel in the region through the local self-attention mechanism. The calculation formula is as follows: Among them, , , represent query, key, and value matrices respectively, , , are learned parameter matrices with a dimension of , is the number of channels after convolution, is the dimension of the query and key vectors, is 's transpose matrix, is the attention score matrix between pixels within each region; S223: Concatenate the output features of each small region to obtain the enhanced region features.
4. A small target detection method from the perspective of an unmanned aerial vehicle based on the self-attention mechanism according to claim 1, characterized in that Step S24 specifically includes: S241: Initialize the input of the decoder by adding position encoding and query embedding; S242: The query vector of each decoder and the image features output by the encoder are weighted and summed through the cross-attention mechanism; S243: Inside the decoder, each query interacts with other queries through self-attention mechanism to capture the spatial relationship between targets; S244: After passing through the self-attention mechanism, the output of the decoder is fed into a feed-forward network which contains two fully-connected layers, and a ReLU activation function is added between the two fully-connected layers for non-linear transformation; S245: After being processed by the feed-forward network, the final output is mapped to the categories and positions of the targets through a fully-connected layer, and finally the probability distribution of the target category corresponding to each query and the precise position of the target are generated and output.
5. A small target detection method from the perspective of an unmanned aerial vehicle based on the self-attention mechanism according to claim 1, characterized in that Step S25 specifically includes: S251: The UGWD adaptive loss function is a positioning loss function based on the improved Gaussian Wasserstein distance. For a single target, the positioning loss function of the Gaussian Wasserstein distance The calculation formula is as follows: where, is the Wasserstein distance between two multivariate Gaussian distributions, and are the center points of the bounding boxes, and are the covariance matrices of the bounding boxes, which are used to describe the target shape, is a hyperparameter used to adjust the behavior of the function, is a non - linear transformation function; S252: It is set that the UGWD adaptive loss function includes a localization loss function and a classification loss function, and its calculation formula is as follows: Among them, is the localization loss, is the classification loss, is the weight coefficient of the classification loss; S253: A dynamic weighting mechanism based on the Gaussian function is introduced into the localization loss function, so that the weights dynamically depend on the context features of the target and the uncertainty of the network prediction. Apply this loss function to the network outputs at different scales to optimize the localization of small and large targets. The calculation formula of the localization loss function is as follows: Among them, is a set of multi-scale feature layers, represents the set of targets of the th layer; S254: In the classification loss function the Gaussian weighting mechanism is also used, and its calculation formula is as follows: Among them, is the predicted probability of the th target, and is the same weighting coefficient as in the localization loss to ensure consistency; S255: The UGWD adaptive loss function finally combines the positioning loss function and the classification loss function , and balances the two through weights . Its calculation formula is as follows: 。 6. The method for small target detection from the perspective of an unmanned aerial vehicle based on the self-attention mechanism according to claim 1, characterized in that Step S3 specifically includes: S31: The preprocessed image dataset is divided into a training set and a validation set and fed into the small target detection model; S32: During the training process, the AdamW optimization algorithm is used to adjust the model parameters, and the early stopping method is used to avoid overfitting and accelerate the training process; S33: Set the weight decay value to 0.0001, gradually increase the learning rate at the beginning of the training until the set maximum value, and then start to decay.
7. A small target detection method from the perspective of an unmanned aerial vehicle based on the self-attention mechanism according to claim 1, characterized in that, Step S4 specifically includes: S41: The real-time image or video stream collected from the perspective of the drone is transmitted to the small target detection model for scaling and normalization processing; S42: The trained small target detection model is used to process the input image dataset; S43: Finally, the small target detection model outputs the category labels, bounding box coordinates and corresponding confidence scores of the detected targets; S44: Using the PySide6 framework, the target detection results are presented to the user and the system in a visual way.
Citation Information
Patent Citations
Target detection method for view angle of unmanned aerial vehicle
CN119048730A
Small target detection method based on multi-scale feature fusion
CN116503641A
Load prediction method and system based on improved Transform
CN118295888A
Structured lightweight fire smoke detection method for petrochemical plant area
CN119107595A