Small target detection method under view angle of unmanned aerial vehicle based on self-attention mechanism

By introducing self-attention mechanism, local self-attention module and cross-scale self-attention mechanism into the small-object detection technology from the perspective of the drone, combined with the decoder and adaptive loss function, the limitations of the existing technology under the differences in complex environments and target scales are solved, and efficient and accurate small-object detection is achieved.

CN119992393AActive Publication Date: 2025-05-13WUHAN TEXTILE UNIV

Patent Information

Application Number
CN202510478698.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-13
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing small-target detection technology from the perspective of drones has limitations in complex environments, lighting changes and target scale differences, especially in the recognition accuracy of extremely small targets or extremely large targets.

Method used

The small object detection method based on the self-attention mechanism is adopted, and feature extraction is performed through the feature extraction module, combining the local self-attention module and the cross-scale self-attention mechanism to enhance the model's ability to identify small objects, and the robustness and real-timeness of the network are improved through the decoder and adaptive loss function.

Benefits of technology

It effectively improves the ability to identify small targets, enhances the detection accuracy in complex backgrounds, reduces the computational complexity, meets the real-time requirements, and improves the recognition accuracy of extremely small or large targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992393A_ABST
    Figure CN119992393A_ABST
Patent Text Reader

Abstract

The invention discloses a self-attention mechanism-based small target detection method under a visual angle of an unmanned aerial vehicle, and the method comprises the following steps: S1, obtaining an image data set suitable for the visual angle of the unmanned aerial vehicle, and carrying out the preprocessing of the image data set; s2, constructing a small target detection model based on a self-attention mechanism; s3, sending the preprocessed image data set into a small target detection model for training; and S4, detecting a to-be-detected target by using the trained small target detection model and outputting a result. According to the invention, feature extraction is carried out through the feature extraction module, the local self-attention module and the cross-scale self-attention mechanism are combined, the small target recognition capability of the model can be effectively improved, the robustness and real-time performance of the network are further enhanced by adopting the decoder and the adaptive loss function, and finally efficient and accurate small target detection is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection, and in particular to a small target detection method from the perspective of a drone based on a self-attention mechanism. Background Art

[0002] In recent years, drone technology has developed rapidly and has been widely used in agricultural monitoring, environmental protection, disaster management and other fields. Among them, small target detection is an important link to improve the efficiency and accuracy of drone applications. The existing small target detection technology from the perspective of drones has certain limitations in complex environments, lighting changes, and target scale differences.

[0003] As an important technology in the field of deep learning in recent years, the self-attention mechanism can effectively capture long-distance feature information by modeling global dependencies, especially showing strong advantages in target detection tasks. Unlike traditional convolutional neural networks that focus on local receptive fields, the self-attention mechanism can focus on key information in the image through weight distribution, achieving more accurate feature extraction and information fusion.

[0004] In the prior art, the Chinese patent with publication number CN119048730A discloses "a target detection method for drone perspective", which belongs to the field of target recognition and detection technology and aims to improve the target detection capability of drones in complex backgrounds and high dynamic environments, especially the recognition of small targets in the air. The core innovation is to build a target detection network CT-RODN, which includes multiple modules: a backbone network, including a block composite attention module FBAM based on the frequency domain and a CNN-Transformer module FCTB based on the frequency domain, a feature fusion module and an attention prediction head. The method is mainly capable of improving the detection accuracy of small targets and multi-scale targets in complex backgrounds, and ensuring the reliability and stability of drone target detection. However, the above method also has some shortcomings: first, it has poor adaptability to complex weather and lighting changes, and the detection accuracy will be reduced in extreme environments; second, the computational complexity is high and cannot meet the real-time requirements; finally, although the method can effectively detect multi-scale targets, there is still room for further optimization in the recognition accuracy of extremely small or extremely large targets.

[0005] Therefore, it is urgent to design a small target detection method from the perspective of a drone based on the self-attention mechanism to solve the problems existing in the above-mentioned existing technologies. Summary of the invention

[0006] The purpose of the present invention is to provide a small target detection method from the perspective of a drone based on a self-attention mechanism. Feature extraction is performed through a feature extraction module. Combined with a local self-attention module and a cross-scale self-attention mechanism, the model's ability to recognize small targets can be effectively improved. The decoder and adaptive loss function are used to further enhance the robustness and real-time performance of the network, ultimately achieving efficient and accurate small target detection.

[0007] In order to achieve the above object, the present invention adopts the following technical solutions: A first aspect of the present invention provides a small target detection method from the perspective of a drone based on a self-attention mechanism, the method comprising the following steps: S1: Obtain an image dataset suitable for the drone’s perspective and preprocess it; S2: construct a small target detection model based on the self-attention mechanism, wherein the small target detection model includes a feature extraction module, a local self-attention mechanism, a cross-scale self-attention mechanism, a decoder, and an adaptive loss function; The step S2 specifically includes: S21: Design a new feature extraction module to extract features from the input data set; S22: Further strengthen the information modeling of small target areas, and design a new local self-attention module in the encoder part; S23: In the fusion process of the multi-scale feature maps of the encoder, the context information between different scales is combined through the cross-scale self-attention mechanism to enhance the detection effect of small objects; S24: A decoder is introduced to decode the candidate target box through a query mechanism. In order to further enhance the detection capability of small targets, position encoding is introduced to retain the spatial information of small targets in the image. S25: designing a new UGWD adaptive loss function in network training, wherein the UGWD adaptive loss function includes a positioning loss function and a classification loss function; S3: Send the preprocessed image dataset to the small object detection model for training; S4: Apply the trained small target detection model to detect the target to be detected and output the result.

[0008] As an embodiment of the present application, the step S1 specifically includes: S11: Obtain the perspective image dataset taken by the drone, clean, annotate and format the data, and standardize the original image dataset to ensure that the image size, resolution and color channel are consistent to meet the network input requirements; S12: Use GPU-accelerated image enhancement techniques to enhance image datasets, including rotation, scaling, and cropping; S13: performing denoising processing on the enhanced image data set, applying a Gaussian filtering method to remove noise components in the image; S14: Finally, the image dataset is format converted and storage optimized.

[0009] As an embodiment of the present application, step S21 specifically includes: S211: Input feature map ,in is the batch size, is the number of channels, , is the height and width of the feature map; S212: extracting local information of the input feature map using depthwise separable convolution, and fusing spatial position information through position encoding, wherein the depthwise separable convolution includes depthwise convolution and pointwise convolution; S213: Add a new adaptive attention mechanism to the output of the depthwise separable convolution. The calculation formula is as follows:

[0010] in, is the weighted feature map, is the feature map of the current layer input, represents element-wise multiplication, and are the learning parameters of the attention network, which are the weight matrix and bias term, It is a feature map obtained through convolution operation, representing the high-order features extracted from the input feature map. is the calculated attention coefficient matrix, which indicates the degree of attention at each spatial position; S214: Next, a cross-layer feature importance adjustment strategy is introduced. This strategy dynamically weights and evaluates the importance of multi-layer features by inserting specific adjustment units between different network modules. The weighting formula is as follows:

[0011]

[0012] in, is the attention weight, is the attention network, For the The output feature map of the layer, , , They are The height, width and number of channels of the layer feature map, is the weighted feature map, Represents element-wise multiplication; S215: Finally, the structure reparameterization technique is used to merge multiple convolutional layers in the training phase into a standard 3×3 convolution operation in the inference phase.

[0013] As an embodiment of the present application, step S22 specifically includes: S221: Use a 1×1 convolution kernel to obtain the area containing the small target and obtain a feature map focusing on the small target area. ; S222: The feature map processed by step S21 Divide into multiple small areas , for each local area, the relationship between pixels in the area is calculated through the local self-attention mechanism, and the calculation formula is as follows:

[0014]

[0015] in, , , Represent query, key, and value matrices respectively, , , is the learned parameter matrix with dimension , is the number of channels after convolution, is the dimension of the query,key vector, yes The transposed matrix of is the attention score matrix between pixels in each region; S223: splicing the output features of each small area to obtain enhanced regional features.

[0016] As an embodiment of the present application, step S24 specifically includes: S241: First, the decoder input is initialized by adding the position encoding and the query embedding; S242: The query vector of each decoder and the image features output by the encoder are weighted summed through a cross-attention mechanism; S243: Inside the decoder, each query interacts with other queries to capture the spatial relationship between targets through the self-attention mechanism; S244: After the decoder passes through the self-attention mechanism, the output is sent to a feed-forward network, wherein the feed-forward network includes two fully connected layers, and a ReLU activation function is added between the two fully connected layers for nonlinear transformation; S245: After being processed by the feedforward network, the final output is mapped to the category and position of the target through a fully connected layer, and finally the probability distribution of the target category corresponding to each query and the precise position of the target are generated and output.

[0017] As an embodiment of the present application, step S25 specifically includes: S251: The UGWD adaptive loss function is a positioning loss function based on the improved Gaussian Wasserstein distance. For a single target, the positioning loss function of the Gaussian Wasserstein distance is The calculation formula is as follows:

[0018]

[0019] in, is the Wassertein distance between two multivariate Gaussian distributions, and is the center point of the bounding box, and is the covariance matrix of the bounding box, which describes the object shape, is a hyperparameter used to adjust the behavior of the function. is a nonlinear transformation function; S252: The UGWD adaptive loss function is set to include a positioning loss function and a classification loss function, and its calculation formula is as follows:

[0020] in, is the positioning loss, is the classification loss, is the weight coefficient of classification loss; S253: In the positioning loss function A dynamic weighting mechanism based on Gaussian function is introduced in Dynamically depends on the contextual features of the target and the uncertainty of network prediction. The loss function is applied to network outputs of different scales to optimize the positioning of small and large targets. The positioning loss function The calculation formula is as follows:

[0021]

[0022] in, is a collection of multi-scale feature layers, Indicates The target set of the layer, is the loss weight based on Gaussian weighting, is the scale of the goal, is the standard deviation, used to control the weighting range, is a dynamically adjusted function that depends on the target features and predicted probability ;when Approaching 1, , indicating that no additional attention is needed; S254: In the classification loss function The Gaussian weighting mechanism is also used in , and its calculation formula is as follows:

[0023] in, It is The predicted probability of a target, is the same weighting coefficient as in the localization loss to ensure consistency; S255: The UGWD adaptive loss function is finally combined with the positioning loss function And the classification loss function , and by weight Balancing the two, the calculation formula is as follows: .

[0024] As an embodiment of the present application, step S3 specifically includes: S31: Divide the preprocessed image dataset into a training set and a validation set and feed them into a small object detection model; S32: During the training process, the AdamW optimization algorithm is used to adjust the model parameters, and the early stopping method is used to avoid overfitting and speed up the training process; S33: Set the weight decay value to 0.0001, gradually increase the learning rate at the beginning of training until it reaches the set maximum value, and then start decaying to ensure that the model can converge better and avoid overfitting or underfitting.

[0025] As an embodiment of the present application, step S4 specifically includes: S41: The real-time image or video stream collected from the drone's perspective is transmitted to the small target detection model for scaling and normalization; S42: Processing the input image dataset using the trained small object detection model; S43: Finally, the small object detection model outputs the category label, bounding box coordinates and corresponding confidence score of the detected object; S44: Use the PySide6 framework to present the target detection results to users and systems in a visual way, so as to facilitate the intuitive display of bounding boxes, category names and confidence levels, thereby supporting manual intervention and subsequent processing.

[0026] The beneficial effects of the present invention are: (1) The present invention designs a new feature extraction module, including deep separable convolution, position encoding, adaptive attention mechanism, cross-layer feature importance adjustment strategy and hierarchical reparameterization technology, which can extract effective target features from images of different resolutions and scales, effectively improving the multi-scale feature extraction capability of input data, especially having a more refined characterization capability for the detailed features of small target areas; through the adaptive attention mechanism and cross-layer feature adjustment, the feature extraction module can dynamically adjust the focus on targets of different areas and different scales and avoid information loss, effectively solving the problem that small targets are easily ignored in complex backgrounds.

[0027] (2) The present invention further enhances the modeling capability of small target areas through a novel local self-attention module. By aggregating features within a local range, the network’s ability to perceive and capture small target features is enhanced, effectively reducing missed detections and false detections.

[0028] (3) The present invention innovatively introduces a cross-scale self-attention mechanism in the process of multi-scale feature fusion, which can dynamically combine the contextual information of features of different scales, significantly enhance the detection effect of small targets, and solve the problem that small targets are easily lost in multi-scale feature fusion.

[0029] (4) The present invention introduces a decoder based on a query mechanism to decode the candidate target frame, which retains the spatial information of small targets in the image and further improves the positioning accuracy and feature decoding capability of the target. This design enables the network to achieve more accurate positioning and classification when processing multi-target tasks.

[0030] (5) The present invention adopts a new UGWD adaptive loss function, which comprehensively considers the classification loss and bounding box regression loss. In particular, through the dynamic weighting mechanism, the loss weight can be adaptively adjusted during the training process, so that the network can achieve better performance in small target detection and solve the problem that small targets are easily ignored or misdetected. This loss function not only improves the detection accuracy, but also improves the stability of training and reduces the risk of overfitting. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 A schematic diagram of a flow chart of a small target detection method from the perspective of a drone based on a self-attention mechanism provided in an embodiment of the present invention; Figure 2A schematic diagram of a small target detection model of a small target detection method from the perspective of a drone based on a self-attention mechanism provided in an embodiment of the present invention; Figure 3 A schematic diagram of the structure of a feature extraction module of a small target detection method from the perspective of a drone based on a self-attention mechanism provided in an embodiment of the present invention; Figure 4 A schematic diagram of the decoder structure of a small target detection method from the perspective of a drone based on a self-attention mechanism provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0033] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative position relationship, movement status, etc. between the components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indication will also change accordingly.

[0034] In the present invention, unless otherwise clearly specified and limited, the terms "connection", "fixation", etc. should be understood in a broad sense. For example, "fixation" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements, unless otherwise clearly defined. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0035] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the meaning of "and / or" appearing in the full text includes three parallel schemes. Taking "A and / or B" as an example, it includes scheme A, or scheme B, or a scheme that satisfies both A and B. In addition, the technical solutions between the various embodiments can be combined with each other, but it must be based on the ability of ordinary technicians in the field to implement. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0036] Reference Figures 1 to 4 In a first aspect, the present invention provides a method for detecting small targets from the perspective of a drone based on a self-attention mechanism, the method comprising the following steps: S1: Obtain an image dataset suitable for the drone’s perspective and preprocess it; S2: construct a small target detection model based on the self-attention mechanism, wherein the small target detection model includes a feature extraction module, a local self-attention mechanism, a cross-scale self-attention mechanism, a decoder, and an adaptive loss function; The step S2 specifically includes: S21: Design a new feature extraction module to extract features from the input data set; S22: Further strengthen the information modeling of small target areas, and design a new local self-attention module in the encoder part; S23: In the fusion process of the encoder's multi-scale feature maps, the cross-scale self-attention mechanism can dynamically combine context information between different scales, and fuse image information from different feature layers and perspectives, further improving the detection accuracy and the ability to distinguish target features, enhancing the detection effect of small targets, and solving the problem that small targets are easily lost in multi-scale feature fusion; S24: A decoder is introduced to decode the candidate target box through a query mechanism. In order to further enhance the detection capability of small targets, position encoding is introduced to retain the spatial information of small targets in the image. S25: designing a new UGWD adaptive loss function in network training, wherein the UGWD adaptive loss function includes a positioning loss function and a classification loss function; Specifically, by introducing the self-attention mechanism, the present invention can effectively improve the performance of small target detection; the self-attention mechanism can enable the small target detection model to focus on important areas in the image, especially those target areas with smaller size and less information, thereby enhancing the model's feature extraction ability in these areas.

[0037] S3: Send the preprocessed image dataset to the small object detection model for training; S4: Apply the trained small target detection model to detect the target to be detected and output the result.

[0038] Specifically, the present invention optimizes the target detection network by designing a new feature extraction module, and combines the local self-attention module and the cross-scale self-attention mechanism to effectively improve the recognition ability of small targets, especially in complex backgrounds; by performing refined feature extraction of small targets in drone-view images, the detection accuracy and positioning accuracy can be improved. At the same time, the design of the UGWD adaptive loss function and decoder further enhances the robustness and real-time performance of the network, thereby improving the overall performance of small target detection, especially showing higher stability and adaptability in dynamic environments, and ultimately achieving efficient and accurate small target detection.

[0039] As an embodiment of the present application, the step S1 specifically includes: S11: Obtain the perspective image dataset taken by the drone, and clean, annotate and format the data to ensure that the dataset is suitable for subsequent processing; standardize the original image dataset to ensure that the image size, resolution and color channel are consistent to meet the network input requirements; S12: Use GPU-accelerated image enhancement technology to enhance image datasets, including rotation, scaling, and cropping; accelerate these operations through GPU parallel computing, reduce performance bottlenecks that may be encountered under traditional CPU processing, and significantly improve data processing speed. This not only speeds up the dataset preparation process, but also improves the robustness of the model by enhancing image diversity; S13: De-noising the enhanced image dataset, applying Gaussian filtering to remove noise components in the image to improve the accuracy of target detection; the GPU-accelerated denoising algorithm can process a large amount of image data and quickly complete the denoising operation to ensure the quality of the dataset; S14: Finally, the image dataset is converted and optimized for storage to ensure that the data can be quickly loaded into the training process and is compatible with the deep learning framework. The GPU-accelerated data loading technology NVIDIA's DALI library is used to accelerate batch processing and parallel loading of data to avoid the impact of data loading on training speed.

[0040] Specifically, the present invention can provide rich image information for subsequent tasks by selecting drone image datasets with representative and diverse scenes; these image datasets cover different geographical environments, weather conditions and time changes, ensuring the generalization ability of the training model; in the preprocessing process, the image is first cropped and scaled to ensure that the size and quality of the input data are adapted to the model requirements; secondly, through data enhancement techniques such as rotation, mirroring, random cropping, etc., the diversity of the data is increased and the robustness of the model is improved; for the noise in the image, filtering and denoising algorithms are used to process it to reduce the impact of the external environment on the image quality; at the same time, the image is normalized so that the pixel value of the image is within a uniform range, which is convenient for efficient model learning; these preprocessing steps not only improve the quality of the data, but also lay a solid foundation for subsequent feature extraction and classification tasks.

[0041] As an embodiment of the present application, step S21 specifically includes: S211: Input feature map ,in is the batch size, is the number of channels, , is the height and width of the feature map; S212: extracting local information of the input feature map using depthwise separable convolution, and fusing spatial position information through position encoding, wherein the depthwise separable convolution includes depthwise convolution and pointwise convolution; S213: Add a new adaptive attention mechanism to the result of the depthwise separable convolution output, and its calculation formula is as follows:

[0042] in, is the weighted feature map, is the feature map of the current layer input, represents element-wise multiplication, and are the learning parameters of the attention network, which are the weight matrix and bias term, It is a feature map obtained through convolution operation, representing the high-order features extracted from the input feature map. is the calculated attention coefficient matrix, which indicates the degree of attention at each spatial position; S214: Next, a cross-layer feature importance adjustment strategy is introduced. This strategy dynamically weights and evaluates the importance of multi-layer features by inserting specific adjustment units between different network modules, so that the model can flexibly adjust the feature flow and contribution between different layers, thereby improving the ability to capture small target details. The weighted formula is as follows:

[0043]

[0044] in, is the attention weight, is the attention network, For the The output feature map of the layer, , , They are The height, width and number of channels of the layer feature map, is the weighted feature map, Represents element-wise multiplication; S215: Finally, the structure reparameterization technique is used to merge multiple convolutional layers in the training phase into a standard 3×3 convolution operation in the inference phase.

[0045] Specifically, the present invention can extract effective target features from images of different resolutions and scales through a multi-level feature extraction module, including deep separable convolution, position encoding, adaptive attention mechanism, cross-layer feature importance adjustment strategy and hierarchical reparameterization technology, effectively improving the multi-scale feature extraction capability of input data, especially having a more refined characterization capability for detail features of small target areas; by adopting position encoding and self-attention mechanism, the network can perform adaptive feature selection and weighting in spatial and channel dimensions, thereby significantly improving the detection accuracy of small targets, and further improving the accuracy and efficiency of the model; through adaptive attention mechanism and cross-layer feature adjustment, the feature extraction module can dynamically adjust the focus on targets of different areas and different scales, and avoid information loss, effectively solving the problem that small targets are easily ignored in complex backgrounds.

[0046] As an embodiment of the present application, step S22 specifically includes: S221: Use a 1×1 convolution kernel to obtain the area containing the small target and obtain a feature map focusing on the small target area. ; S222: The feature map processed by step S21 Divide into multiple small areas , for each local area, the relationship between pixels in the area is calculated through the local self-attention mechanism, and the calculation formula is as follows:

[0047]

[0048] in, , , Represent query, key, and value matrices respectively, , , is the learned parameter matrix with dimension , is the number of channels after convolution, is the dimension of the query,key vector, yes The transposed matrix of is the attention score matrix between pixels in each region; S223: splicing the output features of each small area to obtain enhanced regional features.

[0049] Specifically, the present invention further enhances the modeling capability of small target areas through the local self-attention module, effectively reduces missed detection and false detection, and improves the network's ability to perceive and capture small target features through feature aggregation in a local range.

[0050] like Figure 4 As shown, as an embodiment of the present application, step S24 specifically includes: S241: Initialize the decoder input by adding the position encoding and the query embedding. The position encoding is used to provide spatial position information for each query in the decoder. Since the self-attention mechanism itself has no sequential or spatial information, in addition to adding to the query embedding, the position encoding usually adds a fixed or learnable vector to each position of the input sequence. S242: The query vector of each decoder and the image features output by the encoder are weighted summed through a cross-attention mechanism; S243: Inside the decoder, each query interacts with other queries to capture the spatial relationships between objects through a self-attention mechanism, which is calculated by calculating the similarity of each query vector with all query vectors in the decoder and adjusting its representation based on these similarities; in this way, the decoder is able to capture the spatial relationships between objects and contextual information; S244: After the decoder passes through the self-attention mechanism, the output is sent to a feedforward network, which includes two fully connected layers, and a ReLU activation function is added between the two fully connected layers for nonlinear transformation to extract higher-level features; after processing the self-attention mechanism, the feedforward network will independently process the features of each position; S245: After being processed by the feedforward network, the final output is mapped to the category and location (bounding box) of the target through a fully connected layer, normalized by the Softmax function, and finally generates and outputs the probability distribution of the target category corresponding to each query and the precise location of the target.

[0051] Specifically, the present invention introduces a decoder based on a query mechanism to decode the candidate target frame, retains the spatial information of small targets in the image, and further improves the positioning accuracy and feature decoding capability of the target; this design enables the network to achieve more accurate positioning and classification when processing multi-target tasks.

[0052] As an embodiment of the present application, step S25 specifically includes: S251: The UGWD adaptive loss function is a positioning loss function based on the improved Gaussian Wasserstein distance, which is suitable for bounding box regression tasks and has good smoothness and modeling capabilities for target shapes. For a single target, the positioning loss function of the Gaussian Wasserstein distance is The calculation formula is as follows:

[0053]

[0054] in, is the Wassertein distance between two multivariate Gaussian distributions, and is the center point of the bounding box, and is the covariance matrix of the bounding box, which describes the object shape, is a hyperparameter used to adjust the behavior of the function. is a nonlinear transformation function; S252: The UGWD adaptive loss function is set to include a positioning loss function and a classification loss function, and its calculation formula is as follows:

[0055] in, is the positioning loss, is the classification loss, is the weight coefficient of classification loss; S253: In the positioning loss function A dynamic weighting mechanism based on Gaussian function is introduced in this paper. This mechanism pays special attention to the accuracy of small targets. By improving the traditional Gaussian weighting method, the weight Dynamically depends on the contextual features of the target and the uncertainty of the network prediction. The loss function is applied to the network output at different scales to significantly optimize the positioning of small and large targets. The positioning loss function The calculation formula is as follows:

[0056]

[0057] in, is a collection of multi-scale feature layers, Indicates The target set of the layer, is the loss weight based on Gaussian weighting, is the scale of the goal, is the standard deviation, used to control the weighting range, is a dynamically adjusted function that depends on the target features and predicted probability ;when Approaching 1, , indicating that no additional attention is needed; S254: In the classification loss function The Gaussian weighting mechanism is also used in , so that the classification errors in small target areas are more penalized. The calculation formula is as follows:

[0058] in, It is The predicted probability of a target, is the same weighting coefficient as in the localization loss to ensure consistency; S255: The UGWD adaptive loss function is finally combined with the positioning loss function And the classification loss function , and by weight Balancing the two, the calculation formula is as follows: .

[0059] Specifically, the present invention comprehensively considers the classification loss and bounding box regression loss through the designed UGWD adaptive loss function, and in particular, can adaptively adjust the loss weight during the training process through the dynamic weighting mechanism, so that the network can achieve better performance in small target detection, solving the problem that small targets are easily ignored or misdetected. This loss function not only improves the detection accuracy, but also improves the stability of training and reduces the risk of overfitting.

[0060] As an embodiment of the present application, step S3 specifically includes: S31: Divide the preprocessed image dataset into a training set and a validation set and feed them into a small object detection model; S32: During the training process, the AdamW optimization algorithm is used to adjust the model parameters, and the early stopping method is used to avoid overfitting and speed up the training process; S33: Set the weight decay value to 0.0001, gradually increase the learning rate at the beginning of training until it reaches the set maximum value, and then start decaying to ensure that the model can converge better and avoid overfitting or underfitting.

[0061] Specifically, the preprocessed image dataset is first divided into a training set and a validation set to ensure that the model performance can be effectively evaluated during the training process; during the training process, the preprocessed images are input into the small target detection model, and the model will calculate the loss function through forward propagation, and adjust the weights in the network through back propagation according to the loss value; since small targets usually occupy less pixel area in the image, the UGWD loss function is designed to perform more stringent optimization on the positioning and classification accuracy of the small target area; after the training is completed, the validation set is used to evaluate the performance of the model, including evaluation indicators such as detection accuracy, recall rate, and F1 value; if the performance of the model on the validation set is satisfactory, it is further tested and deployed; if there is a performance bottleneck, the training process is further optimized by adjusting the model architecture or increasing the diversity of the dataset.

[0062] As an embodiment of the present application, step S4 specifically includes: S41: The real-time image or video stream collected from the drone's perspective is transmitted to the small target detection model for scaling and normalization; S42: Processing the input image dataset using the trained small object detection model; S43: Finally, the small object detection model outputs the category label, bounding box coordinates and corresponding confidence score of the detected object; S44: Use the PySide6 framework to present the target detection results to users and systems in a visual way, so as to facilitate the intuitive display of bounding boxes, category names and confidence levels, thereby supporting manual intervention and subsequent processing.

[0063] Specifically, the trained and verified model can be applied to actual detection tasks; in this process, the image or video frame to be detected is first preprocessed in the same way as the training data set and then input into the trained small target detection model; the model will use its forward propagation process, the self-attention mechanism and the ability of multi-scale feature fusion to identify the small target area in the image; finally, the model outputs the results including the position of the detected small target (i.e. the coordinates of the bounding box), the category label and the corresponding confidence score, and the detection results are visualized through the Pyside6 framework.

[0064] The present invention constructs a small target detection model based on the self-attention mechanism, which automatically adjusts the detection strategy and parameters according to the target features in different scenarios, and improves the detection accuracy and robustness through continuously optimized training and verification processes. Each type of small target has its own unique visual features, and the small target detection model can effectively distinguish them based on these features, thereby improving the accuracy and classification performance of target detection; through this method, not only can small-sized targets be accurately identified, but also the classification can be further refined to provide more detailed detection results; for example, in traffic monitoring scenarios, not only can vehicles and pedestrians be identified, but also their specific types (such as cars, trucks, electric vehicles, etc.) can be further identified, providing important support for subsequent monitoring, tracking and early warning systems; in addition, the module can also cope with problems such as lighting changes and noise interference in complex environments, ensure the stability and reliability of the detection results, and meet the high-precision requirements of target detection in different application scenarios.

[0065] The present invention has high-precision small target detection capability, excellent feature extraction and enhancement capability, and powerful information fusion advantages. It can effectively deal with diverse and dynamically changing small targets in the shooting angle of drones. It has strong practical application potential, and is particularly suitable for small target detection tasks of drones in the fields of security, monitoring, agricultural monitoring, etc. It can provide effective support for drones in various application scenarios; its application fields include military reconnaissance, disaster monitoring, environmental protection, agricultural monitoring, etc., and it has broad application prospects and market value.

[0066] The above descriptions are only some preferred embodiments of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the above features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A small target detection method from the perspective of a drone based on a self-attention mechanism, characterized in that: The method comprises the following steps: S1: Obtain an image dataset suitable for the drone’s perspective and preprocess it; S2: Building a small target detection model based on the self-attention mechanism, the small target detection model includes a feature extraction module, a local self-attention module, a cross-scale self-attention mechanism, a decoder and a UGWD adaptive loss function; The step S2 specifically includes: S21: Design a new feature extraction module to extract features from the input data set; S22: Further strengthen the information modeling of small target areas, and design a new local self-attention module in the encoder part; S23: In the multi-scale feature map fusion process of the encoder, the context information between different scales is combined through the cross-scale self-attention mechanism to enhance the detection effect of small objects; S24: A decoder is introduced to decode the candidate target box through a query mechanism. In order to further enhance the detection capability of small targets, position encoding is introduced to retain the spatial information of small targets in the image. S25: designing a new UGWD adaptive loss function in network training, wherein the UGWD adaptive loss function includes a positioning loss function and a classification loss function; S3: Send the preprocessed image dataset to the small object detection model for training; S4: Apply the trained small target detection model to detect the target to be detected and output the result.

2. According to claim 1, a small target detection method based on self-attention mechanism from the perspective of a drone is characterized in that: The step S1 specifically includes: S11: Obtain the perspective image dataset taken by the drone, clean, annotate and format the data, and standardize the original image dataset to ensure that the image size, resolution and color channel are consistent to meet the network input requirements; S12: Use GPU-accelerated image enhancement techniques to enhance image datasets, including rotation, scaling, and cropping; S13: performing denoising processing on the enhanced image data set, applying a Gaussian filtering method to remove noise components in the image; S14: Finally, the image dataset is format converted and storage optimized.

3. According to the method of small target detection from the perspective of a drone based on the self-attention mechanism in claim 1, it is characterized in that: The step S21 specifically includes: S211: Input feature map ,in is the batch size, is the number of channels, , is the height and width of the feature map; S212: extracting local information of the input feature map using depthwise separable convolution, and fusing spatial position information through position encoding, wherein the depthwise separable convolution includes depthwise convolution and pointwise convolution; S213: Add a new adaptive attention mechanism to the output of the depthwise separable convolution. The calculation formula is as follows: in, is the weighted feature map, is the feature map of the current layer input, represents element-wise multiplication, and are the learning parameters of the attention network, which are the weight matrix and bias term, It is a feature map obtained through convolution operation, representing the high-order features extracted from the input feature map. is the calculated attention coefficient matrix, which indicates the degree of attention at each spatial position; S214: Next, a cross-layer feature importance adjustment strategy is introduced. This strategy dynamically weights and evaluates the importance of multi-layer features by inserting specific adjustment units between different network modules. The weighting formula is as follows: in, is the attention weight, is the attention network, For the The output feature map of the layer, , , They are The height, width and number of channels of the layer feature map, is the weighted feature map, Represents element-wise multiplication; S215: Finally, the structure reparameterization technique is used to merge multiple convolutional layers in the training phase into a standard 3×3 convolution operation in the inference phase.

4. According to claim 1, a method for detecting small targets from the perspective of a drone based on a self-attention mechanism is characterized in that: The step S22 specifically includes: S221: Use a 1×1 convolution kernel to obtain the area containing the small target and obtain a feature map focusing on the small target area. ; S222: Feature map Divide into multiple small areas , for each local area, the relationship between pixels in the area is calculated through the local self-attention mechanism, and the calculation formula is as follows: in, , , Represent query, key, and value matrices respectively, , , is the learned parameter matrix with dimension , is the number of channels after convolution, is the dimension of the query,key vector, yes The transposed matrix of is the attention score matrix between pixels in each region; S223: splicing the output features of each small area to obtain enhanced regional features.

5. According to claim 1, a method for detecting small targets from the perspective of a drone based on a self-attention mechanism is characterized in that: The step S24 specifically includes: S241: Initialize the decoder input by adding the position encoding and the query embedding; S242: The query vector of each decoder and the image features output by the encoder are weighted summed through a cross-attention mechanism; S243: Inside the decoder, each query interacts with other queries to capture the spatial relationship between targets through the self-attention mechanism; S244: After the decoder passes through the self-attention mechanism, the output is sent to a feed-forward network, wherein the feed-forward network includes two fully connected layers, and a ReLU activation function is added between the two fully connected layers for nonlinear transformation; S245: After being processed by the feedforward network, the final output is mapped to the category and position of the target through a fully connected layer, and finally the probability distribution of the target category corresponding to each query and the precise position of the target are generated and output.

6. The small target detection method based on self-attention mechanism from the perspective of a drone according to claim 1, characterized in that: The step S25 specifically includes: S251: The UGWD adaptive loss function is a positioning loss function based on the improved Gaussian Wasserstein distance. For a single target, the positioning loss function of the Gaussian Wasserstein distance is The calculation formula is as follows: in, is the Wassertein distance between two multivariate Gaussian distributions, and is the center point of the bounding box, and is the covariance matrix of the bounding box, which describes the object shape, is a hyperparameter used to adjust the behavior of the function. is a nonlinear transformation function; S252: Setting the UGWD adaptive loss function to include a positioning loss function and a classification loss function, and the calculation formula thereof is as follows: in, is the positioning loss, is the classification loss, is the weight coefficient of classification loss; S253: In the positioning loss function A dynamic weighting mechanism based on Gaussian function is introduced in Dynamically depends on the contextual features of the target and the uncertainty of network prediction. The loss function is applied to network outputs of different scales to optimize the positioning of small and large targets. The positioning loss function The calculation formula is as follows: in, is a collection of multi-scale feature layers, Indicates The target set of the layer, is the loss weight based on Gaussian weighting, is the scale of the goal, is the standard deviation, used to control the weighting range, is a dynamically adjusted function that depends on the target features and predicted probability ;when Approaching 1, , indicating that no additional attention is needed; S254: In the classification loss function The Gaussian weighting mechanism is also used in , and its calculation formula is as follows: in, It is The predicted probability of a target, is the same weighting coefficient as in the localization loss to ensure consistency; S255: The UGWD adaptive loss function is finally combined with the positioning loss function And the classification loss function , and by weight Balancing the two, the calculation formula is as follows: 。 7. The method for detecting small targets from the perspective of a drone based on a self-attention mechanism according to claim 1, characterized in that: The step S3 specifically includes: S31: Divide the preprocessed image dataset into a training set and a validation set and feed them into a small object detection model; S32: During the training process, the AdamW optimization algorithm is used to adjust the model parameters, and the early stopping method is used to avoid overfitting and speed up the training process; S33: Set the weight decay value to 0.0001, gradually increase the learning rate at the beginning of training until it reaches the set maximum value, and then start decaying.

8. The method for detecting small targets from the perspective of a drone based on a self-attention mechanism according to claim 1, characterized in that: The step S4 specifically includes: S41: The real-time image or video stream collected from the drone's perspective is transmitted to the small target detection model for scaling and normalization; S42: Processing the input image dataset using the trained small object detection model; S43: Finally, the small object detection model outputs the category label, bounding box coordinates and corresponding confidence score of the detected object; S44: Use the PySide6 framework to present the target detection results to users and systems in a visual way.

Citation Information

Patent Citations

  • Target detection method for view angle of unmanned aerial vehicle

    CN119048730A

  • Lightweight video behavior recognition method based on structure reparameterization

    CN115661939A

  • Small target detection method based on multi-scale feature fusion

    CN116503641A

  • Dense pest image detection method for enhancing self-attention based on Gaussian receptive field

    CN118053074A

  • Load prediction method and system based on improved Transform

    CN118295888A

Cited By

  • Subway water leakage identification method based on infrared image

    CN120163972A

  • Subway water leakage identification method based on infrared image

    CN120163972B

  • Unmanned aerial vehicle target detection method based on cross-spatial frequency domain and electronic equipment

    CN120318499A

  • Steel bar binding point detection method based on improved YOLOv8

    CN120726296A

  • A steel bar binding point detection method based on improved YOLOv8

    CN120726296B