Unmanned aerial vehicle small target detection method inspired by eagle eye visual mechanism

By adopting the neural network design inspired by the Hawkeye vision mechanism in the detection of small targets in UAV, including DFB, SFB and NRt modules, the accuracy and resource balance problems of small target detection from the perspective of UAV are solved, and efficient small target detection and lightweight model deployment are achieved.

CN120032271AActive Publication Date: 2025-05-23GUANGXI UNIVERSITY OF TECHNOLOGY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411818072.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-11
Publication Date
2025-05-23
Estimated Expiration
2044-12-11

AI Technical Summary

Technical Problem

From the perspective of drone, small-object detection tasks face challenges of too small target size, complex background and excessive model parameters. The traditional method lacks balance between detection accuracy and computing resource requirements.

Method used

A neural network design inspired by the Hawkeye visual mechanism, including the DFB module and the SFB module, was used to simulate the deep and shallow fovea structures of the Hawkeye. Combined with the NRt module to simulate the large receptive field and global search characteristics of the thalamic nucleus, multi-scale feature fusion and detection were performed.

Benefits of technology

It significantly improves the detection accuracy of small targets under the drone image. The model has lightweight characteristics and is suitable for the deployment needs of edge devices. Compared with YOLOv9t on the VisDrone2019 dataset, the detection accuracy is 12.5%, and the parameter volume is only 1.46M.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention aims to provide an unmanned aerial vehicle small target detection method inspired by an eagle eye vision mechanism, and the method comprises the following steps: A, constructing a neural network which comprises a Backbone network, a Neck network module and a Head network module; in the Backbone network module, a DFB module is applied to simulate a deep fovea structure of an eagle eye, and an SFB module is applied to simulate a shallow fovea structure of the eagle eye; b, the original image is input into a Backbone network module, and obtained processing results are input into a Neck network module; and C, performing multi-scale feature fusion on the input features by the Neck network module to obtain two feature layers with different resolutions, and respectively inputting the two feature layers into the Head network module for detection to obtain a series of detection frames and category data, namely a final result. According to the method, the feature information of the tiny target can be efficiently extracted, and the target can be accurately positioned in a complex scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and image processing, and in particular to a method for detecting small targets of unmanned aerial vehicles inspired by an eagle-eye vision mechanism. Background Art

[0002] Object detection technology plays an important role in the field of computer vision and is widely used in tasks such as image classification, instance segmentation, and target tracking. However, the task of detecting small targets from the perspective of drones faces many challenges, including small target size, complex background, and excessive model parameters. Traditional object detection methods mostly use the "Backbone-Neck-Head" structure built by convolutional neural networks (CNNs) to improve detection performance by improving feature extraction and multi-scale feature fusion. In recent years, the YOLO series of methods and various improved models have shown good results in drone image detection, but the balance between detection accuracy and computing resource requirements is still insufficient.

[0003] Inspired by the bionic visual system, researchers have tried to simulate the characteristics of biological visual systems to improve detection performance. For example, the eagle's visual system has a double fovea structure, which enables it to have excellent small target capture capabilities, providing design inspiration for the bionic visual network. However, most existing bionic designs only consider the basic structure of the visual system, and do not fully utilize the characteristics of the eagle's tectal pathway, double fovea structure, and large receptive field of the thalamic nucleus, which leads to limited optimization between model performance and lightweight. Summary of the invention

[0004] The present invention aims to provide a UAV small target detection method inspired by the eagle-eye vision mechanism, which can efficiently extract the feature information of tiny targets and accurately locate targets in complex scenes.

[0005] The technical solution of the present invention is as follows:

[0006] The method for detecting small targets of unmanned aerial vehicles inspired by the eagle eye vision mechanism comprises the following steps:

[0007] A. Constructing a neural network, wherein the neural network includes a Backbone network, a Neck network module, and a Head network module;

[0008] The Backbone network module uses a DFB module and a SFB module;

[0009] B. The original image is input into the Backbone network module and processed by the Silence module. The processing results of the Silence module are divided into two paths:

[0010] The first route is processed by the first CBS module, the second CBS module, the first DFB module, the second DFB module, the first Adown module, the third DFB module, the second Adown module, and the fourth DFB module in sequence; the second route is processed by the third CBS module, the fourth CBS module, and the first SFB module in sequence;

[0011] The processing result of the second DFB module is processed by the first CBlinear module to obtain the processing result of the first CBlinear module; the processing result of the third DFB module is processed by the second CBlinear module to obtain the processing result of the second CBlinear module;

[0012] The processing result of the first SFB module, the processing result of the first CBlinear module, the processing result of the second CBlinear module, the processing result of the second DFB module, the processing result of the third DFB module, and the processing result of the fourth DFB module are respectively input into the Neck network module;

[0013] C. The Neck network module performs multi-scale feature fusion on the input features to obtain two feature layers with different resolutions, which are input into the Head network module for detection respectively to obtain a series of detection boxes and category data, which is the final result.

[0014] The first DFB module, the second DFB module, the third DFB module, and the fourth DFB module have the same structure, and the processing process is as follows:

[0015] After the input result is processed by the GELAN module, it is divided into four paths; the first path is processed by 1*1 convolution and 3*3 convolution in sequence to obtain the first path result; the second path is processed by 1*1 convolution, 1*3 convolution, 3*1 convolution, and 3*3Atrous modules in sequence to obtain the second path result; the third path is processed by 1*1 convolution, 3*1 convolution, 1*3 convolution, and 3*3Atrous modules in sequence to obtain the third path result; the fourth path is processed by 1*1 convolution to obtain the fourth path result;

[0016] The first result, the second result, the third result, and the fourth result are concatenated by the Concat function to obtain the output result.

[0017] The processing process in the Neck network module is as follows:

[0018] The processing result of the fifth DFB module and the processing result of the first CBlinear module are respectively input into the first CBFuse module for processing, and the processing result of the first CBFuse module is processed by the second SFB module and the first NRT module in turn to obtain the processing result of the first NRT module;

[0019] The processing result of the second SFB module is processed by the third Adown module to obtain the processing result of the third Adown module. The processing result of the third Adown module, the processing result of the first CBlinear module, and the processing result of the second CBlinear module are respectively input into the second module for processing. The processing result of the second CBFuse module is processed by the third SFB module and the second NRT module to obtain the processing result of the second NRT module.

[0020] The processing result of the fourth DFB module is processed by the SPPELAN module and then upsampled to obtain a first upsampling result. The first upsampling result and the processing result of the third DFB module are concatenated by a Concat function. The concatenated result is processed by the first GELAN module to obtain a first GELAN module processing result.

[0021] After the processing result of the first GELAN module is upsampled, it is concatenated with the processing result of the second DFB module through the Concat function, and the obtained concatenated result is processed by the second GELAN module and the third NRT module in turn to obtain the processing result of the third NRT module;

[0022] The processing result of the second GELAN module is processed by the third Adown module to obtain the processing result of the third Adown module; the processing result of the first GELAN module and the processing result of the third Adown module are concatenated by the Concat function, and the concatenated result is sequentially processed by the third GELAN module and the fourth NRT module to obtain the processing result of the fourth NRT module;

[0023] The processing results of the first NRT module, the second NRT module, the third NRT module, and the fourth NRT module are input into the Head network module.

[0024] The first SFB module, the second SFB module, and the third SFB module have the same structure, and the processing process is as follows:

[0025] After the input result is processed by the GELAN module, it is divided into four branches; the first branch is processed by 1*1 convolution and 7*7 convolution in sequence to obtain the first branch result; the second branch is processed by 1*1 convolution, 1*7 convolution, 7*1 convolution, and 7*7Atrous modules in sequence to obtain the second branch result; the third branch is processed by 1*1 convolution, 7*1 convolution, 1*7 convolution, and 7*7Atrous modules in sequence to obtain the third branch result; the fourth branch is processed by 1*1 convolution to obtain the fourth branch result;

[0026] The first branch result, the second branch result, the third branch result, and the fourth branch result are concatenated through the Concat function to obtain the output result.

[0027] The first NRT module, the second NRT module, the third NRT module, and the fourth NRT module have the same structure, and the processing process is as follows:

[0028] The input results are divided into three paths. The first and second paths are processed by 1*1 convolution respectively. The two convolution processing results are multiplied and fused, and then processed by the Softmax module to obtain the Softmax module processing result.

[0029] The third path is processed by the Rotundal cell module, and the result is multiplied and fused with the result of the Softmax module. Then, it is processed by 1*1 convolution, and the convolution result is added and fused with the input result to obtain the output result.

[0030] The processing process in the Rotundal cell module is as follows:

[0031] The input result is processed by 1*3 convolution, 3*1 convolution, 1*5 depth-wise separable convolution, 5*1 depth-wise separable convolution, and 1*1 convolution in sequence, and the processed result is multiplied and fused with the input result to obtain the output result.

[0032] The dilation coefficients of the 1*5 depthwise separable convolution and the 5*1 depthwise separable convolution are both 3.

[0033] In the Head network module, four Head detection heads are provided, and the processing result of each NRT module is respectively input into a Head detection head for processing.

[0034] The four Head detection heads are all built-in detection heads of the YOLO network.

[0035] The present invention adopts the DFB module to simulate the deep fovea structure of the eagle eye, and applies the SFB module to simulate the shallow fovea structure of the eagle eye.

[0036] DFB module simulation By introducing multi-branch atlas convolution, we can capture local contextual information of different scales, especially enhance the details of small objects, so as to have higher visual acuity and accurately perceive tiny objects.

[0037] In addition to introducing multi-path Atlas convolution, the SFB module uses larger convolution kernels (such as 1x7, 7x1, and 7x7) to enhance the model's feature extraction capabilities over a larger range. This design helps the model better match and fuse global semantic information when processing a wide range of background information, thereby improving the recognition of small objects in complex backgrounds.

[0038] The combination of these two structures enables Hawkeye to have an incomparable visual advantage in dynamic environments. Through the synergistic effect of the two, the EVMNet network of the present invention can better extract rich global-local semantic information in multi-scale feature fusion, effectively improving the detection ability of small objects, especially in scenes with complex backgrounds and dense objects.

[0039] The present invention also designs a bionic circular nucleus cell module, namely the NRt module, in the Neck network part. The NRt module simulates the large receptive field and global search characteristics of the thalamic circular nucleus by using deep separable convolution and global attention architecture. First, the equivalent large kernel convolution is used to simulate the circular nucleus cell. The equivalent large kernel convolution is composed of two different deep separable convolutions, which not only maintains lightweight, but also achieves a receptive field equivalent to the large kernel convolution effect with a convolution kernel of 11. Then, by combining the circular nucleus cell module into the global attention architecture, the visual processing function of the thalamic circular nucleus is further simulated. The global attention mechanism aggregates global contextual information by calculating the correlation between each pair of pixels, so that the model can effectively capture the relationship between different positions and improve the ability to capture long-distance dependencies, similar to the role of the thalamic circular nucleus in visual processing. In this way, the NRt module can not only focus on local features, but also capture long-distance dependency information on a global scale, thereby significantly improving the performance of small target detection.

[0040] The beneficial effects of the present invention are as follows:

[0041] The method of the present invention uses the eagle eye visual neural structure as a prototype, combines the structural characteristics of the double fovea and the tectal pathway, and improves the accuracy of small target detection in drone images through innovatively designed feature extraction modules and information fusion modules. At the same time, the model has lightweight characteristics and is suitable for the deployment requirements of edge devices.

[0042] The lightweight bionic UAV small target detection network EVMNet proposed in this invention effectively improves the accuracy of small target detection and shows superior performance on multiple test data sets. On the VisDrone2019 data set, compared with YOLOv9t, EVMNet's detection accuracy is improved by 12.5%, reaching 49.2% mAP0.5, while the number of parameters is only 1.46M, which is suitable for resource-constrained UAV platform applications. In addition, the bionic dual fovea module and NRt module have both verified their advantages in multi-scale feature extraction and background interference suppression in experiments, allowing EVMNet to achieve a good balance between performance and lightweight in all indicators. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 This is a schematic diagram of the structure of a neural network according to Embodiment 1 of the present invention;

[0044] Figure 2 It is a structural schematic diagram of the DFB module of Example 1;

[0045] Figure 3 It is a structural schematic diagram of the SFB module of Example 1;

[0046] Figure 4 This is a schematic diagram of the structure of the NRt module of Example 1;

[0047] Figure 5 This is a schematic diagram of the structure of the Rotundal cell module of Example 1;

[0048] Figure 6 This is a comparison chart of the detection results of the method in Example 1 and MHA-YOLO. DETAILED DESCRIPTION

[0049] The present invention is described in detail below with reference to the accompanying drawings and embodiments.

[0050] Example 1

[0051] The method for detecting small targets of unmanned aerial vehicles inspired by the eagle eye vision mechanism comprises the following steps:

[0052] A. Construct a neural network, such as Figure 1 As shown, the neural network includes a Backbone network, a Neck network module, and a Head network module;

[0053] The Backbone network module uses the DFB module and the SFB module; the Head network module is provided with four Head detection heads, and the four Head detection heads are all detection heads that come with the YOLO network.

[0054] B. The original image is input into the Backbone network module and processed by the Silence module. The processing results of the Silence module are divided into two paths:

[0055] The first path is processed by the first CBS module, the second CBS module, the first DFB module, the second DFB module, the first Adown module, the third DFB module, the second Adown module, and the fourth DFB module in sequence; the second path is processed by the third CBS module, the fourth CBS module, and the first SFB module in sequence;

[0056] The processing result of the second DFB module is processed by the first CBlinear module to obtain the processing result of the first CBlinear module; the processing result of the third DFB module is processed by the second CBlinear module to obtain the processing result of the second CBlinear module;

[0057] The processing result of the first SFB module, the processing result of the first CBlinear module, the processing result of the second CBlinear module, the processing result of the second DFB module, the processing result of the third DFB module, and the processing result of the fourth DFB module are respectively input into the Neck network module;

[0058] The first DFB module, the second DFB module, the third DFB module, and the fourth DFB module have the same structure. Figure 2 As shown, the processing process is as follows:

[0059] After the input result is processed by the GELAN module, it is divided into four paths; the first path is processed by 1*1 convolution and 3*3 convolution in sequence to obtain the first path result; the second path is processed by 1*1 convolution, 1*3 convolution, 3*1 convolution, and 3*3Atrous modules in sequence to obtain the second path result; the third path is processed by 1*1 convolution, 3*1 convolution, 1*3 convolution, and 3*3Atrous modules in sequence to obtain the third path result; the fourth path is processed by 1*1 convolution to obtain the fourth path result;

[0060] The first result, the second result, the third result, and the fourth result are concatenated by the Concat function to obtain the output result.

[0061] C. The Neck network module performs multi-scale feature fusion on the input features to obtain two feature layers with different resolutions, which are input into the Head network module for detection respectively to obtain a series of detection boxes and category data, which is the final result.

[0062] The processing process in the Neck network module is as follows:

[0063] The processing result of the fifth DFB module and the processing result of the first CBlinear module are respectively input into the first CBFuse module for processing, and the processing result of the first CBFuse module is processed by the second SFB module and the first NRT module in turn to obtain the processing result of the first NRT module;

[0064] The processing result of the second SFB module is processed by the third Adown module to obtain the processing result of the third Adown module. The processing result of the third Adown module, the processing result of the first CBlinear module, and the processing result of the second CBlinear module are respectively input into the second module for processing. The processing result of the second CBFuse module is processed by the third SFB module and the second NRT module to obtain the processing result of the second NRT module.

[0065] The processing result of the fourth DFB module is processed by the SPPELAN module and then upsampled to obtain a first upsampling result. The first upsampling result and the processing result of the third DFB module are concatenated by a Concat function. The concatenated result is processed by the first GELAN module to obtain a first GELAN module processing result.

[0066] After the processing result of the first GELAN module is upsampled, it is concatenated with the processing result of the second DFB module through the Concat function, and the obtained concatenated result is processed by the second GELAN module and the third NRT module in turn to obtain the processing result of the third NRT module;

[0067] The processing result of the second GELAN module is processed by the third Adown module to obtain the processing result of the third Adown module; the processing result of the first GELAN module and the processing result of the third Adown module are concatenated by the Concat function, and the concatenated result is sequentially processed by the third GELAN module and the fourth NRT module to obtain the processing result of the fourth NRT module;

[0068] The processing results of the first NRT module, the second NRT module, the third NRT module, and the fourth NRT module are input into the Head network module, and the processing results of each NRT module are respectively input into a Head detection head for processing.

[0069] The first SFB module, the second SFB module, and the third SFB module have the same structure. Figure 3 As shown, the processing process is as follows:

[0070] After the input result is processed by the GELAN module, it is divided into four branches; the first branch is processed by 1*1 convolution and 7*7 convolution in sequence to obtain the first branch result; the second branch is processed by 1*1 convolution, 1*7 convolution, 7*1 convolution, and 7*7Atrous modules in sequence to obtain the second branch result; the third branch is processed by 1*1 convolution, 7*1 convolution, 1*7 convolution, and 7*7Atrous modules in sequence to obtain the third branch result; the fourth branch is processed by 1*1 convolution to obtain the fourth branch result;

[0071] The first branch result, the second branch result, the third branch result, and the fourth branch result are concatenated through the Concat function to obtain the output result.

[0072] The first NRT module, the second NRT module, the third NRT module, and the fourth NRT module have the same structure. Figure 4 As shown, the processing process is as follows:

[0073] The input results are divided into three paths. The first and second paths are processed by 1*1 convolution respectively. The two convolution processing results are multiplied and fused, and then processed by the Softmax module to obtain the Softmax module processing result.

[0074] The third path is processed by the Rotundal cell module, and the result is multiplied and fused with the result of the Softmax module. Then, it is processed by 1*1 convolution, and the convolution result is added and fused with the input result to obtain the output result.

[0075] like Figure 5 As shown, the processing process in the Rotundal cell module is as follows:

[0076] The input result is processed by 1*3 convolution, 3*1 convolution, 1*5 depth-wise separable convolution, 5*1 depth-wise separable convolution, and 1*1 convolution in sequence, and the processed result is multiplied and fused with the input result to obtain the output result.

[0077] The dilation coefficients of the 1*5 depthwise separable convolution and the 5*1 depthwise separable convolution are both 3.

[0078] Example 2

[0079] For the quantitative performance evaluation of the final small target image from the drone's perspective, we adopt the performance measurement standard widely used in the field of target detection. The specific evaluation is shown in formulas (1) and (2).

[0080]

[0081] Where AP represents the integral of R (Recall) on P (precision), the confidence threshold ranges from 0 to 1, mAP represents the average AP value of all categories in the dataset, and N is the number of categories. The higher the mAP value, the stronger the detection performance of the model.

[0082] Table 1 summarizes the detection comparison experiments of Example 1 and the existing YOLOv9 on the AI ​​UAV perspective small target detection dataset (VisDrone2019). The higher the scores of the four indicators Precision, Recall, mAP0.5, and mAP0.5:0.95, the stronger the detection performance of the model. The lower the parameter amount (Param), the lighter the model. From the experimental results, compared with the YOLOv9 framework, Example 1 of the present invention has made significant progress in detection accuracy and parameter amount.

[0083] Table 1 Detection comparison with YOLOv9

[0084]

[0085] Table 2 summarizes the comparative data with the existing small target detection models on the small target detection dataset from the perspective of an artificial intelligence drone (VisDrone2019). It can be seen that the neural network model of Example 1 of the present invention has achieved significant results in detection performance, parameter quantity, and computational complexity (GFLOPS).

[0086] Table 2: Comparison of results with existing networks on the VisDrone dataset

[0087]

[0088] Example 3

[0089] In order to more intuitively understand the performance of the model on the visdrone dataset, this example shows the visualization of the detection results in various scenarios. The first column of pictures is the original picture with the Ground Truth (real label), the second column is the detection picture of the comparison network MHA-YOLO, and the third column is the detection picture of Example 1. Figure 6 As shown, EVMNet shows excellent detection capabilities, especially in the recognition of small objects in complex backgrounds. The first row of pictures is a parking lot, including detection targets of different scales. It can be seen that the model of Example 1 in the yellow box accurately recognizes targets of different scales, which is almost the same as the annotation of Ground Truth, while MHA-YOLO does not recognize objects of different scales very well. The second row is a scene containing dense small targets. MHA-YOLO does not detect enough small targets in the yellow box, while EVMNet successfully recognizes small targets of different categories. The third row also shows the accuracy of the model of Example 1.

Claims

1. A method for detecting small targets of unmanned aerial vehicles inspired by the eagle-eye vision mechanism, characterized in that: The following steps are involved: A. Constructing a neural network, wherein the neural network includes a Backbone network, a Neck network module, and a Head network module; The Backbone network module uses a DFB module and a SFB module; B. The original image is input into the Backbone network module and processed by the Silence module. The processing results of the Silence module are divided into two paths: The first route is processed in sequence by the first CBS module, the second CBS module, the first DFB module, the second DFB module, the first Adown module, the third DFB module, the second Adown module, and the fourth DFB module; The second route is processed by the third CBS module, the fourth CBS module, and the first SFB module in sequence; The processing result of the second DFB module is processed by the first CBlinear module to obtain the processing result of the first CBlinear module; The processing result of the third DFB module is processed by the second CBlinear module to obtain the processing result of the second CBlinear module; The processing result of the first SFB module, the processing result of the first CBlinear module, the processing result of the second CBlinear module, the processing result of the second DFB module, the processing result of the third DFB module, and the processing result of the fourth DFB module are respectively input into the Neck network module; C. The Neck network module performs multi-scale feature fusion on the input features to obtain two feature layers with different resolutions, which are input into the Head network module for detection respectively to obtain a series of detection boxes and category data, which is the final result.

2. The method for detecting small targets of unmanned aerial vehicles inspired by the eagle eye vision mechanism as claimed in claim 1, characterized in that: The first DFB module, the second DFB module, the third DFB module, and the fourth DFB module have the same structure, and the processing process is as follows: After the input result is processed by the GELAN module, it is divided into four paths; the first path is processed by 1*1 convolution and 3*3 convolution in sequence to obtain the first path result; the second path is processed by 1*1 convolution, 1*3 convolution, 3*1 convolution, and 3*3Atrous modules in sequence to obtain the second path result; the third path is processed by 1*1 convolution, 3*1 convolution, 1*3 convolution, and 3*3Atrous modules in sequence to obtain the third path result; the fourth path is processed by 1*1 convolution to obtain the fourth path result; The first result, the second result, the third result, and the fourth result are concatenated by the Concat function to obtain the output result.

3. The method for detecting small targets of unmanned aerial vehicles inspired by the eagle eye vision mechanism as claimed in claim 1, characterized in that: The processing process in the Neck network module is as follows: The processing result of the fifth DFB module and the processing result of the first CBlinear module are respectively input into the first CBFuse module for processing, and the processing result of the first CBFuse module is processed by the second SFB module and the first NRT module in turn to obtain the processing result of the first NRT module; The processing result of the second SFB module is processed by the third Adown module to obtain the processing result of the third Adown module. The processing result of the third Adown module, the processing result of the first CBlinear module, and the processing result of the second CBlinear module are respectively input into the second module for processing. The processing result of the second CBFuse module is processed by the third SFB module and the second NRT module to obtain the processing result of the second NRT module. The processing result of the fourth DFB module is processed by the SPPELAN module and then upsampled to obtain a first upsampling result. The first upsampling result and the processing result of the third DFB module are concatenated by a Concat function. The concatenated result is processed by the first GELAN module to obtain a first GELAN module processing result. After the processing result of the first GELAN module is upsampled, it is concatenated with the processing result of the second DFB module through the Concat function, and the obtained concatenated result is processed by the second GELAN module and the third NRT module in turn to obtain the processing result of the third NRT module; The processing result of the second GELAN module is processed by the third Adown module to obtain the processing result of the third Adown module; The processing result of the first GELAN module and the processing result of the third Adown module are concatenated by the Concat function, and the concatenated result is sequentially processed by the third GELAN module and the fourth NRT module to obtain the processing result of the fourth NRT module; The processing results of the first NRT module, the second NRT module, the third NRT module, and the fourth NRT module are input into the Head network module.

4. The method for detecting small targets of unmanned aerial vehicles inspired by the eagle eye vision mechanism as claimed in claim 3, characterized in that: The first SFB module, the second SFB module, and the third SFB module have the same structure, and the processing process is as follows: After the input result is processed by the GELAN module, it is divided into four branches; the first branch is processed by 1*1 convolution and 7*7 convolution in sequence to obtain the first branch result; the second branch is processed by 1*1 convolution, 1*7 convolution, 7*1 convolution, and 7*7Atrous modules in sequence to obtain the second branch result; the third branch is processed by 1*1 convolution, 7*1 convolution, 1*7 convolution, and 7*7Atrous modules in sequence to obtain the third branch result; the fourth branch is processed by 1*1 convolution to obtain the fourth branch result; The first branch result, the second branch result, the third branch result, and the fourth branch result are concatenated through the Concat function to obtain the output result.

5. The method for detecting small targets of unmanned aerial vehicles inspired by the eagle eye vision mechanism as claimed in claim 3, characterized in that: The first NRT module, the second NRT module, the third NRT module, and the fourth NRT module have the same structure, and the processing process is as follows: The input results are divided into three paths. The first and second paths are processed by 1*1 convolution respectively. The two convolution processing results are multiplied and fused, and then processed by the Softmax module to obtain the Softmax module processing result. The third path is processed by the Rotundal cell module, and the result is multiplied and fused with the result of the Softmax module. Then, it is processed by 1*1 convolution, and the convolution result is added and fused with the input result to obtain the output result.

6. The method for detecting small targets of unmanned aerial vehicles inspired by the eagle eye vision mechanism as claimed in claim 5, characterized in that: The processing process in the Rotundal cell module is as follows: The input result is processed by 1*3 convolution, 3*1 convolution, 1*5 depth-wise separable convolution, 5*1 depth-wise separable convolution, and 1*1 convolution in sequence, and the processed result is multiplied and fused with the input result to obtain the output result.

7. The method for detecting small targets of unmanned aerial vehicles inspired by the eagle eye vision mechanism as claimed in claim 6, characterized in that: The dilation coefficients of the 1*5 depthwise separable convolution and the 5*1 depthwise separable convolution are both 3.

8. The method for detecting small targets of unmanned aerial vehicles inspired by the eagle eye vision mechanism as claimed in claim 1, characterized in that: In the Head network module, four Head detection heads are provided, and the processing result of each NRT module is respectively input into a Head detection head for processing.

9. The method for detecting small targets of unmanned aerial vehicles inspired by the eagle eye vision mechanism as claimed in claim 8, characterized in that: The four Head detection heads are all built-in detection heads of the YOLO network.

Citation Information

Patent Citations

  • Eagle-eye-imitating adaptive mechanism-based unmanned aerial vehicle sea surface small target identification method

    CN112101099A

  • Bionic eagle eye multi-scale fusion super-resolution reconstruction model, method and device for single image and storage medium

    CN114549325A

  • Eagle-eye-vision-imitating double-fovea-center air target saliency detection method

    CN116863354A

  • Bionic visual navigation control system and method thereof for autonomous aerial refueling docking

    US20190031347A1