Light-weight high-performance anti-unmanned aerial vehicle detection method based on AE-YOLO11
By integrating the AKConv module and ELA module into the YOLO11 model, the AE-YOLO11 model is constructed, and the existing drone detection methods have solved the problems of slow detection speed, low accuracy and large model size, achieving efficient and lightweight drone detection effects.
Patent Information
- Application Number
- CN202510058516.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-13
AI Technical Summary
The existing drone detection methods have problems such as slow detection speed, low accuracy, and large model size, which is difficult to meet the requirements of real-time and lightweight in practical applications.
A lightweight high-performance anti-UAV detection method based on AE-YOLO11 is proposed. By fusing the AKConv module and the ELA module into the original YOLO11 object detection model, the AE-YOLO11 model is constructed, and the model structure and training process are optimized to improve detection accuracy and speed.
The model is lightweight, the detection accuracy and speed are improved, and it can operate efficiently on resource-constrained devices, meeting the requirements for device compatibility and portability in practical applications.
Smart Images

Figure CN119992379A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of low-altitude UAV intelligent detection, and specifically, relates to a lightweight and high-performance anti-UAV detection method based on AE-YOLO11. Background Art
[0002] Drones, which occupy a key position in the low-altitude field, have been widely used in many fields such as aerial photography, logistics, and agriculture. However, the widespread use of drones has also brought a series of safety hazards. Therefore, it is crucial to quickly and accurately detect and identify low-altitude drones.
[0003] At present, among the existing drone detection methods, vision-based detection methods have attracted much attention due to their advantages such as low cost, high accuracy, and easy deployment. Among them, deep learning target detection algorithms have achieved remarkable results. However, most traditional deep learning target detection algorithms face problems such as complex models, large computational complexity, and slow detection speed, making it difficult to meet the strict requirements of real-time and lightweight in practical applications. For example, on resource-constrained devices (such as embedded devices and mobile devices), these algorithms may not operate normally or achieve the desired detection effect. Summary of the invention
[0004] In view of the problems of slow detection speed, low precision, large model size, etc. in the prior art anti-UAV detection methods, a lightweight and high-performance anti-UAV detection method based on AE-YOLO11 is proposed to achieve fast and accurate detection of UAVs and run efficiently on resource-constrained devices. The method aims to improve the detection accuracy and speed while reducing the computational complexity and memory usage of the model to meet the anti-UAV detection needs in different scenarios.
[0005] The present invention is achieved through the following technical solutions:
[0006] A lightweight and high-performance anti-UAV detection method based on AE-YOLO11:
[0007] The method specifically comprises the following steps:
[0008] Step 1: Integrate the AKConv module and the ELA module into the original YOLO11 target detection model to construct the AE-YOLO11 model.
[0009] Step 2: Collect drone image data from different background environments and preprocess them;
[0010] Step 3: According to the characteristics of the anti-UAV detection task, design the loss function, perform weighted processing on the detection loss of small targets, select the optimizer, and train the AE-YOLO11 model;
[0011] Step 4: Input the image to be detected into the trained AE-YOLO11 model and output the detection result including the drone target location, category and confidence.
[0012] Furthermore, in step 1, it includes:
[0013] Step 1.1, build the foundation of the YOLO11 model: according to the original structure and parameter settings of the YOLO11 model, use the deep learning framework to build a basic neural network model;
[0014] Step 1.2, insert AKCon v Module: Insert AKCon into the backbone network v module; during the insertion process, ensure that AKCon v The module is correctly connected to the surrounding network layer, and parameter transmission is normal;
[0015] Step 1.3, integrating the ELA module: Insert the ELA module at the end of the feature extraction network to ensure that the ELA module can correctly learn the importance of different regions in the feature map and effectively weight the feature map.
[0016] Furthermore, in step 1.2,
[0017] AKCon v The module fixes the features at the corresponding positions through the sampling grid. Let R represent the regular sampling grid, then R is expressed as follows:
[0018] R={(-1,-1),(-1,0),…,(0,1),(1,1)}(1)
[0019] For irregular convolution of any size, set the upper left corner (0, 0) as the sampling origin, and set the initial coordinates of the irregular convolution to P n , then position P 0 The corresponding convolution operation at is as follows, where w represents the convolution parameter:
[0020] Conv(P 0 )=∑w×(P 0 +P n )(2)
[0021] Furthermore, in step 1.3,
[0022] For the ELA module, consider the output of a convolutional block, expressed as denote the height, width and channel dimensions respectively,
[0023] Average pooling is performed on each channel in two spatial ranges: along the horizontal direction (H, 1) and along the vertical direction (1, W). Given an input x, the output of the cth channel at height h and width w is represented as:
[0024]
[0025] z h and z w It not only captures the full-field perception, but also captures the precise location information. In order to effectively utilize these features, this module applies one-dimensional convolution to the location information in both horizontal and vertical directions to enhance its information; then, group normalization G is used n To process the enhanced position information, we can get the representation of position attention in the vertical and horizontal directions:
[0026] y h =σ(G n (F h (z h ))) (5)
[0027] y w =σ(G n (F w (z w ))) (6)
[0028] Among them, σ is a nonlinear activation function, F h and F w represents a one-dimensional convolution, and the convolution kernel is set to 5 or 7. Thus, the final output of the ELA module is obtained, denoted as Y:
[0029] Y = x c ×y h ×y W (7)
[0030] Furthermore, in step 2, it includes:
[0031] Step 2.1: The collected images cover different types of drones, different flight postures, different scenes, different weather conditions, and different time periods. The collected images are accurately annotated, and 80% of them are used as training data sets and 20% as test data sets;
[0032] Step 2.2, enhancing the collected image, including randomly rotating the image, flipping the image horizontally and vertically, scaling the image, and adding Gaussian noise;
[0033] Step 2.3, normalize the enhanced image data and map the pixel values of the image from the range of [0, 255] to the range of [0, 1];
[0034] The specific formula is: Among them, is the pixel value of the original image, I norm is the normalized pixel value.
[0035] Furthermore, in step 3, it includes:
[0036] Step 3.1, initialization parameters: set the initial learning rate, number of iterations, batch size of the model, and select an optimizer to adaptively adjust the learning rate;
[0037] Step 3.2, forward propagation: input the training data into the AE-YOLO11 model according to the batch size, and the data passes through the input layer, feature extraction network and AKCon in turn. v Module, ELA module, detection head and output layer to get the prediction results;
[0038] During feature extraction, AKCon v The module works together with the ELA module to extract features from the input image and extract feature maps containing the features of the drone;
[0039] Step 3.3, calculate the loss: use the mean square error loss function (MSE) to calculate the loss between the predicted result and the true label;
[0040] For the target position prediction, the coordinate error between the predicted box and the real box is calculated; for the target category prediction, the probability error between the predicted category and the real category is calculated;
[0041] Step 3.4, back propagation: Calculate the gradient of the loss function to the model parameters through the back propagation algorithm; use the automatic derivation tool provided by the deep learning framework to automatically calculate the gradient, and then update the model parameters according to the gradient;
[0042] Step 3.5, repeat training: Repeat steps 3.1 to 3.4, and verify the performance of the model on the validation set after a certain number of iterations;
[0043] The model parameters are adjusted according to the validation results, and training is stopped if the loss on the validation set no longer decreases or reaches the preset number of iterations.
[0044] Furthermore, in step 4,
[0045] Step 4.1, feature extraction: Input the image data to be detected into the trained AE-YOLO11 model, and the model first extracts features from the image. The AKConv module extracts a feature map containing the features of the drone; the feature map is enhanced by the feature fusion of the neck and the weighted processing of the ELA mechanism to express the features of the drone.
[0046] Step 4.2, object detection: Based on the extracted feature map, the model predicts the location, category, and confidence of the drone in the image through the convolutional layer and fully connected layer in the detection head; the detection head converts the feature map into a prediction result, and obtains the bounding box representing the drone's location, its category, and confidence information;
[0047] Step 4.3, result output: Perform non-maximum suppression processing on the detection results output by the model, and only retain the detection boxes with the highest confidence and accurate position, so as to obtain the final accurate detection results;
[0048] The non-maximum suppression process is used to remove duplicate detection frames: set a suitable NMS threshold, and for detection frames with higher confidence, calculate the intersection over union (IoU) between them and other detection frames; if the IoU is greater than the threshold, these detection frames are considered to be duplicates, retain the detection frame with the highest confidence, and remove other detection frames; finally, output the accurate position, category, confidence and other information of the drone in the image.
[0049] A lightweight and high-performance anti-UAV detection system based on AE-YOLO11:
[0050] The detection system includes: a construction module, a collection and processing module, a training module and a detection output module;
[0051] The construction module integrates the AKConv module and the ELA module into the original YOLO11 target detection model to construct an AE-YOLO11 model;
[0052] The acquisition and processing module collects drone image data from different background environments and performs preprocessing;
[0053] The training module designs a loss function based on the characteristics of the anti-UAV detection task, performs weighted processing on the detection loss of small targets, selects an optimizer, and trains the AE-YOLO11 model;
[0054] The detection output module inputs the image to be detected into the trained AE-YOLO11 model and outputs the detection result including the drone target position, category and confidence.
[0055] An electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0056] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the steps of the above method are implemented.
[0057] Beneficial effects of the present invention
[0058] The present invention realizes the lightweight model: by introducing the variable kernel convolution (AKConv) module and the efficient local attention (ELA) mechanism, the number of model parameters and the amount of calculation are effectively reduced without reducing the detection accuracy, making the AE-YOLO11 model more efficient and lightweight. This enables the model to run efficiently on resource-constrained devices (such as embedded devices, mobile terminals, etc.), meeting the requirements for device compatibility and portability in practical applications.
[0059] The present invention has high performance in detecting drones. The variable kernel convolution mechanism of the AKConv module can better adapt to drone targets of different scales and shapes, and the ELA module can enhance the model's ability to perceive targets. The combination of the two greatly improves the model's detection accuracy and robustness for drone targets. In various complex environments and different scenarios, the method of the present invention can accurately and stably detect drones.
[0060] The present invention realizes the function of real-time tracking of drones. The lightweight model structure and optimized calculation method enable the method of the present invention to have a high detection speed and meet the real-time requirements. For example, in real-time video stream monitoring, the position and status of drones can be detected in real time, providing strong support for taking timely countermeasures.
[0061] The method of the present invention can achieve efficient and fast reasoning under limited hardware resources, and can track the collected low-altitude drone videos in real time, taking into account both lightweight and high speed. It can be applied to places such as airports, urban areas, military restricted areas, and areas around important infrastructure where real-time prevention of drones is required. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 It is a flow chart of the anti-UAV system of the present invention.
[0063] Figure 2 This is the AE-YOLO11 network structure diagram of the present invention. DETAILED DESCRIPTION
[0064] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0065] The experimental methods used in the following examples are conventional methods unless otherwise specified. The materials, reagents, methods and instruments used are conventional materials, reagents, methods and instruments in the art unless otherwise specified, and can be obtained through commercial channels by those skilled in the art.
[0066] A lightweight and high-performance anti-UAV detection method based on AE-YOLO11: Figure 1 As shown,
[0067] The method specifically comprises the following steps:
[0068] Step 1: Build the model: Convolution with variable kernel (AKCon v ) and the lightweight efficient local attention (ELA) are integrated into the original YOLO11 target detection model to construct an efficient and lightweight AE-YOLO11 (AKCon v &ELAYOLO11) Model: Figure 2 As shown,
[0069] Step 1.1, build the foundation of the YOLO11 model: According to the original structure and parameter settings of the YOLO11 model, use a deep learning framework (such as PyTorch) to build a basic neural network model. This includes building basic components such as convolutional layers, pooling layers, and fully connected layers to ensure that the basic architecture of the model is correct.
[0070] That is, the YOLO11 target detection model is used as the basic framework, and two lightweight modules are integrated to construct the improved AE-YOLO11. One of them is the variable kernel convolution module (AKCon v ), which will replace the convolutional module of the backbone network. v The module adopts an innovative variable kernel design, which allows the convolution kernel to have any number of parameters and sampling shapes. When processing different images and targets, AKCon v The convolution kernel can automatically adjust its sampling shape. And it allows the number of convolution parameters to increase or decrease linearly. This feature is very beneficial to the hardware environment because it can dynamically adjust the number of parameters according to actual needs, thereby achieving more efficient resource utilization. In the case of limited resources, AKCon v It can help the network model reduce the number of parameters and computational overhead while maintaining performance, which is especially important for the design of lightweight models.
[0071] Step 1.2, insert AKCon v Module: In the backbone network e ) v Module. Based on experiments and analysis, we choose to insert AKCon after the convolutional layer which is critical for extracting features of targets of different scales. v For example, insert AKCon after the convolutional layer responsible for extracting small-scale features. vmodule to enhance the model's ability to capture features of different types of drones. During the insertion process, ensure that AKCon v The module is correctly connected to the surrounding network layer and parameter passing is normal.
[0072] AKCon v First, the features are positioned at the corresponding positions through the sampling grid. Let R represent the sampling grid, then R is expressed as follows:
[0073] R={(-1,-1),(-1,0),…,(0,1),(1,1)}(1)
[0074] For irregular convolution of any size, set the upper left corner (0, 0) as the sampling origin, and set the initial coordinates of the irregular convolution to P n , then position P 0 The corresponding convolution operation at is as follows, where w represents the convolution parameter:
[0075] Conv(P 0 )=∑w×(P 0 +P n ) (2)
[0076] Step 1.3, integrating ELA module: At the end of the feature extraction network, such as after several convolutional layers and AKCon v After the module, insert the ELA module. The insertion position of the ELA module is adjusted according to the performance of the model and experimental results to achieve the best feature enhancement effect. When implementing the ELA module, ensure that it can correctly learn the importance of different regions in the feature map and perform effective weighting on the feature map.
[0077] The efficient local attention (ELA) mechanism will be inserted into the detection head. This module proposes a feature enhancement technique that combines one-dimensional convolution and Group Normalization. This method can accurately locate the region of interest without dimensionality reduction by effectively encoding two one-dimensional position feature maps, while allowing lightweight and fast implementation. It obtains feature vectors in the horizontal and vertical directions in the spatial dimension, maintains a narrow kernel shape to capture long-range dependencies, and prevents irrelevant areas from affecting label predictions, thereby generating rich target position features in their respective directions. And the above feature vectors are processed independently for each direction to obtain attention predictions, and then combined using a dot multiplication operation to ensure accurate location information of the region of interest, thereby enhancing the model's perception and extraction capabilities of drone features and improving detection accuracy.
[0078] For ELA, consider the output of a convolutional block, expressed as Denote the height, width, and channel dimensions, respectively. Average pooling is performed on each channel in two spatial ranges: along the horizontal direction (H, 1) and along the vertical direction (1, W). Given an input x, the output of the cth channel at height h and width w is represented as:
[0079]
[0080] z h and z w In order to effectively utilize these features, the module not only captures the full-field perception but also the precise location information. The module applies one-dimensional convolution to the location information in two directions (horizontally and vertically) to enhance its information. Subsequently, the group normalization G is used. n To process the enhanced position information, we can get the representation of position attention in the vertical and horizontal directions:
[0081] y h =σ(G n (F h (z h ))) (5)
[0082] y w =σ(G n (F w (z w ))) (6)
[0083] Among them, σ is a nonlinear activation function, F h and F w represents a one-dimensional convolution, and the convolution kernel is set to 5 or 7. Thus, the final output of the ELA module can be obtained, denoted as Y:
[0084] Y = x c ×y h ×y W (7)
[0085] Step 2: Data preprocessing: Collect drone image data from different background environments and preprocess them;
[0086] Step 2.1, Image acquisition: Use a variety of devices, such as cameras, drone-mounted cameras, etc., to collect a large amount of drone image data in different scenes and conditions. The collected images should cover different types of drones, different flight postures and background environments to ensure the comprehensiveness and representativeness of the data. If enough images are collected, 80% of them can be used as training data sets and 20% as test data sets.
[0087] Collect drone image data from different scenes, different weather conditions, and different time periods. At the same time, ensure that the collected data covers drones of various types, sizes, and postures to build a rich and diverse anti-drone detection dataset. Accurately annotate the collected images and mark the location of the drone (represented by a bounding box), category (such as fixed-wing drone, rotary-wing drone, etc.), and other information.
[0088] Step 2.2, image enhancement: Use image enhancement libraries (such as OpenCV, Augmentor, etc.) to enhance the collected images. Specific operations include randomly rotating images (rotation angle range is [-30°, 30°]), flipping images horizontally and vertically, scaling images (scaling ratio range is [0.8, 1.2]), adding Gaussian noise (noise intensity is adjusted according to actual conditions), etc. Through these operations, the diversity of data is increased and the generalization ability of the model is improved.
[0089] Perform a series of data enhancement operations on the collected image data, such as rotation, flipping, scaling, translation, adding noise, blurring, etc. Through data enhancement, the diversity and scale of the data set are increased, so that the model can learn richer feature patterns, improve the generalization ability of the model, and reduce the occurrence of overfitting.
[0090] Step 2.2, normalization: normalize the enhanced image data and map the pixel values of the image from the range of [0, 255] to the range of [0, 1]. The specific formula is: Among them, is the pixel value of the original image, I norm is the normalized pixel value. Normalization helps improve the training efficiency and stability of the model.
[0091] Normalize the image data and map the pixel values of the image from the original range (usually 0-255) to the interval [0, 1]. Normalization helps to speed up the training convergence of the model and improve the stability and training efficiency of the model.
[0092] Step 3: Model training: Based on the characteristics of the anti-UAV detection task, a comprehensive loss function is designed. Considering that small UAVs account for a small proportion of the image and are difficult to detect, the detection loss of small targets is weighted, and more attention is paid to small target detection, so as to improve the model's detection accuracy for small-sized UAVs. Then, a suitable optimizer is selected. Finally, GPU is used for training.
[0093] Step 3.1, Initialize parameters: Set the initial learning rate, number of iterations, and batch size of the model. Select a suitable optimizer (such as SGD, Adam) as the optimizer of the model to adaptively adjust the learning rate and speed up the convergence of the model.
[0094] Step 3.2, forward propagation: input the training data into the AE-YOLO11 model according to the batch size, and the data passes through the input layer, feature extraction network and AKConv module, ELA module, detection head and output layer in turn to obtain the prediction result. During the feature extraction process, the AKConv module and the ELA module work together to extract features from the input image and extract feature maps containing drone features.
[0095] Step 3.3, calculate the loss: use the mean square error loss function (MSE) to calculate the loss between the predicted result and the true label. For the location prediction of the target, calculate the coordinate error between the predicted box and the true box; for the category prediction of the target, calculate the probability error between the predicted category and the true category.
[0096] Step 3.4, back propagation: Calculate the gradient of the loss function to the model parameters through the back propagation algorithm. Use the automatic differentiation tool provided by the deep learning framework (such as PyTorch's autograd module) to automatically calculate the gradient, and then update the model parameters according to the gradient.
[0097] Step 3.5, repeat training: Repeat the above steps, and verify the performance of the model on the validation set after a certain number of iterations. Adjust model parameters such as learning rate based on the verification results. If the loss on the validation set no longer decreases or reaches the preset number of iterations, stop training.
[0098] Step 4: Model detection: Input the image to be detected into the trained AE-YOLO11 model.
[0099] Step 4.1, feature extraction: Input the image data to be detected into the trained AE-YOLO11 model, and the model first extracts features from the image. The AKConv module and the ELA module play an important role in the feature extraction process. Through a series of convolution operations, feature maps containing drone features are extracted.
[0100] The image first goes through a preprocessing step and then enters the backbone network of the model for feature extraction. After multi-scale feature extraction by the AKConv module and layer-by-layer processing by the backbone network, a feature map with rich semantic information is obtained. Then, the feature map is further enhanced through feature fusion of the neck and weighted processing of the ELA mechanism to express the characteristics of the drone.
[0101] Step 4.2, object detection: Based on the extracted feature map, the model predicts the location, category, and confidence of the drone in the image through the convolutional layer and the fully connected layer in the detection head. The detection head converts the feature map into a prediction result, and obtains the bounding box representing the drone's location, the category it belongs to, and the confidence information.
[0102] Detection prediction is performed in the head part, and the detection results containing information such as the drone target location, category and confidence level are output.
[0103] Step 4.3, result output: Post-process the prediction results and use the non-maximum suppression (NMS) algorithm to remove duplicate detection frames. Set a suitable NMS threshold (such as 0.5) and calculate the intersection over union (IoU) of detection frames with higher confidence and other detection frames. If the IoU is greater than the threshold, these detection frames are considered to be duplicates, and the detection frames with the highest confidence are retained, while other detection frames are removed. Finally, the accurate location, category, and confidence of the drone in the image are output.
[0104] The detection results output by the model are processed by non-maximum suppression, and only the detection boxes with the highest confidence and accurate position are retained, so as to obtain the final accurate detection results.
[0105] After the above steps, this lightweight model can also be used to quickly detect and track drones in the airspace in real time on resource-constrained platforms.
[0106] A lightweight and high-performance anti-UAV detection system based on AE-YOLO11:
[0107] The detection system includes: a construction module, a collection and processing module, a training module and a detection output module;
[0108] The construction module integrates the AKConv module and the ELA module into the original YOLO11 target detection model to construct an AE-YOLO11 model;
[0109] The acquisition and processing module collects drone image data from different background environments and performs preprocessing;
[0110] The training module designs a loss function based on the characteristics of the anti-UAV detection task, performs weighted processing on the detection loss of small targets, selects an optimizer, and trains the AE-YOLO11 model;
[0111] The detection output module inputs the image to be detected into the trained AE-YOLO11 model and outputs the detection result including the drone target position, category and confidence.
[0112] An electronic device comprises a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0113] A computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the steps of the above method are implemented.
[0114] The memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory, ROM, a programmable read-only memory, PROM, an erasable programmable read-only memory, EPROM, an electrically erasable programmable read-only memory, EEPROM, or a flash memory. The volatile memory may be a random access memory, RAM, which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory static RAM, SRAM, dynamic random access memory dynamic RAM, DRAM, synchronous dynamic random access memory synchronous DRAM, SDRAM, double data rate synchronous dynamic random access memory double data rate SDRAM, DDR SDRAM, enhanced synchronous dynamic random access memory enhanced SDRAM, ESDRAM, synchronous link dynamic random access memory synchlink DRAM, SLDRAM, and direct memory bus random access memory direct rambus RAM, DR RAM. It should be noted that memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0115] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented by software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website site, computer, server or data center through a wired method such as coaxial cable, optical fiber, digital subscriber line digital subscriber line, DSL or wireless such as infrared, wireless, microwave, etc. to another website site, computer, server or data center. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium such as a floppy disk, a hard disk, a tape, an optical medium such as a high-density digital video disc digital video disc, DVD, or a semiconductor medium such as a solid state hard disk solid state disc, SSD, etc.
[0116] In the implementation process, each step of the above method can be completed by an integrated logic circuit of hardware in a processor or an instruction in the form of software. The steps of the method disclosed in conjunction with the embodiment of the present application can be directly embodied as a hardware processor for execution, or a combination of hardware and software modules in a processor for execution. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it is not described in detail here.
[0117] It should be noted that the processor in the embodiment of the present application can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method embodiment can be completed by an integrated logic circuit of hardware in the processor or an instruction in the form of software. The above processor can be a general-purpose processor, a digital signal processor DSP, an application-specific integrated circuit ASIC, a field programmable gate array FPGA or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as a hardware decoding processor to perform, or the hardware and software modules in the decoding processor can be combined to perform. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in a memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0118] The above is a detailed introduction to the lightweight and high-performance anti-UAV detection method based on AE-YOLO11 proposed in the present invention, and the principles and implementation methods of the present invention are explained. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. A lightweight and high-performance anti-UAV detection method based on AE-YOLO11, characterized by: The method specifically comprises the following steps: Step 1: Integrate the AKConv module and the ELA module into the original YOLO11 target detection model to construct the AE-YOLO11 model. Step 2: Collect drone image data from different background environments and preprocess them; Step 3: According to the characteristics of the anti-UAV detection task, design a loss function, perform weighted processing on the detection loss of small targets, select an optimizer, and train the AE-YOLO11 model; Step 4: Input the image to be detected into the trained AE-YOLO11 model, and output the detection result including the drone target location, category and confidence.
2. The detection method according to claim 1, characterized in that: In step 1, include: Step 1.1, build the foundation of the YOLO11 model: according to the original structure and parameter settings of the YOLO11 model, use the deep learning framework to build a basic neural network model; Step 1.2, insert AKConv module: insert AKConv module into the backbone network; during the insertion process, ensure that the connection between AKConv module and surrounding network layers is correct and the parameter transfer is normal; Step 1.3, integrating the ELA module: Insert the ELA module at the end of the feature extraction network to ensure that the ELA module can correctly learn the importance of different regions in the feature map and effectively weight the feature map.
3. The detection method according to claim 2, characterized in that: In step 1.2, The AKConv module fixes the features at the corresponding positions through the sampling grid. Let R represent the regular sampling grid, then R is expressed as follows: R={(-1,-1),(-1,0),…,(0,1),(1,1)} (1) For irregular convolution of any size, set the upper left corner (0, 0) as the sampling origin, and set the initial coordinates of the irregular convolution to P n , then the corresponding convolution operation at position P0 is as follows, where w represents the convolution parameter: Conv(P0)=∑w×(P0+P n ) (2)。 4. The detection method according to claim 3, characterized in that: In step 1.3, For the ELA module, consider the output of a convolutional block, expressed as denote the height, width and channel dimensions respectively, Average pooling is performed on each channel in two spatial ranges: along the horizontal direction (H, 1) and along the vertical direction (1, W). Given an input x, the output of the cth channel at height h and width w is represented as: z h and z w It not only captures the full-field perception, but also the precise location information. In order to effectively utilize these features, this module applies one-dimensional convolution to the location information in both horizontal and vertical directions to enhance its information. Then, using group normalization G n To process the enhanced position information, we can get the representation of position attention in the vertical and horizontal directions: y h =σ(G n (F h (z h ))) (5) y w =σ(G n (F w (z w ))) (6) Among them, σ is a nonlinear activation function, F h and F w represents a one-dimensional convolution, and the convolution kernel is set to 5 or 7; thus, the final output of the ELA module is obtained, denoted as Y: Y=x c ×y h ×y w (7)。 5. The detection method according to claim 4, characterized in that: In step 2, include: Step 2.1: The collected images cover different types of drones, different flight postures, different scenes, different weather conditions, and different time periods. The collected images are accurately annotated, and 80% of them are used as training data sets and 20% as test data sets; Step 2.2, enhancing the collected image, including randomly rotating the image, flipping the image horizontally and vertically, scaling the image, and adding Gaussian noise; Step 2.3, normalize the enhanced image data and map the pixel values of the image from the range of [0,255] to the range of [0,1]; The specific formula is: Where I is the pixel value of the original image, I norm is the normalized pixel value.
6. The detection method according to claim 5, characterized in that: In step 3, include: Step 3.1, initialization parameters: set the initial learning rate, number of iterations, batch size of the model, and select an optimizer to adaptively adjust the learning rate; Step 3.2, forward propagation: input the training data into the AE-YOLO11 model according to the batch size. The data passes through the input layer, feature extraction network and AKConv module, ELA module, detection head and output layer in turn to obtain the prediction result; In the feature extraction process, the AKConv module and the ELA module work together to extract features from the input image and extract a feature map containing the features of the drone; Step 3.3, calculate the loss: use the mean square error loss function (MSE) to calculate the loss between the predicted result and the true label; For the target position prediction, the coordinate error between the predicted box and the real box is calculated; for the target category prediction, the probability error between the predicted category and the real category is calculated; Step 3.4, back propagation: Calculate the gradient of the loss function to the model parameters through the back propagation algorithm; use the automatic derivation tool provided by the deep learning framework to automatically calculate the gradient, and then update the model parameters according to the gradient; Step 3.5, repeat training: Repeat steps 3.1 to 3.4, and verify the performance of the model on the validation set after a certain number of iterations; The model parameters are adjusted according to the validation results, and training is stopped if the loss on the validation set no longer decreases or reaches the preset number of iterations.
7. The detection method according to claim 6, characterized in that: In step 4, Step 4.1, feature extraction: The image data to be detected is input into the trained AE-YOLO11 model. The model first extracts features from the image. The AKConv module extracts a feature map containing the features of the drone. The feature map is enhanced by the feature fusion of the neck and the weighted processing of the ELA mechanism to express the features of the drone. Step 4.2, object detection: Based on the extracted feature map, the model predicts the location, category, and confidence of the drone in the image through the convolutional layer and fully connected layer in the detection head; the detection head converts the feature map into a prediction result, and obtains the bounding box representing the drone's location, its category, and confidence information; Step 4.3, result output: Perform non-maximum suppression processing on the detection results output by the model, and only retain the detection boxes with the highest confidence and accurate position, so as to obtain the final accurate detection results; The non-maximum suppression process is used to remove duplicate detection frames: set a suitable NMS threshold, and for detection frames with higher confidence, calculate the intersection over union (IoU) between them and other detection frames; if the IoU is greater than the threshold, these detection frames are considered to be duplicates, retain the detection frame with the highest confidence, and remove other detection frames; finally, output the accurate position, category, confidence and other information of the drone in the image.
8. A detection system for executing the lightweight and high-performance anti-UAV detection method based on AE-YOLO11 according to any one of claims 1 to 7, characterized in that: The detection system includes: a construction module, a collection and processing module, a training module and a detection output module; The construction module integrates the AKConv module and the ELA module into the original YOLO11 target detection model to construct an AE-YOLO11 model; The acquisition and processing module collects drone image data from different background environments and performs preprocessing; The training module designs a loss function based on the characteristics of the anti-UAV detection task, performs weighted processing on the detection loss of small targets, selects an optimizer, and trains the AE-YOLO11 model; The detection output module inputs the image to be detected into the trained AE-YOLO11 model and outputs the detection result including the drone target position, category and confidence.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium for storing computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Low-altitude unmanned aerial vehicle detection method based on super-resolution and multi-dimensional attention fusion
CN120783029A
Method and system for detecting low-altitude unmanned aerial vehicle based on improved model of Yolov11 and BRA mechanisms
CN121095864A