Unmanned Aerial Vehicle Target Detection Method and Device
By improving the convolutional block attention unit of the YOLOv7-tiny model and introducing the SCFM module, the problem of insufficient detection accuracy of small targets and performance degradation in complex contexts in UAV target detection is solved, and stronger recognition ability and robustness of small targets are achieved.
Patent Information
- Application Number
- CN202510520251.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The YOLOv7-tiny algorithm has problems with insufficient small target detection accuracy and performance degradation in complex contexts in UAV target detection.
By improving the convolutional block attention unit in the YOLOv7-tiny model, the weight allocation of spatial dimensions and channel dimensions is performed to improve the network's attention to drone feature information. At the same time, an SCFM module is introduced to filter out interference information of the feature map from both the channel and the space, and enhance the small-target feature information.
It improves the recognition ability of small targets and its robustness in dense contexts, and enhances the accuracy and performance of drone target detection.
Smart Images

Figure CN120047677B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of target detection, and particularly to a method and device for unmanned aerial vehicle (UAV) target detection. Background Art
[0002] The application of UAVs in target detection is becoming increasingly widespread, especially in the fields of monitoring, search and rescue, and military. Target detection algorithms play a core role in the real-time data processing and decision-making support of UAVs.
[0003] Traditional UAV target detection methods have problems such as window redundancy and poor feature robustness. Since the emergence of deep learning, great breakthroughs have been made in target detection, including two-stage algorithms represented by Faster R-CNN and single-stage target detection algorithms represented by the YOLO and SSD series. Among them, the YOLO (You Only Look Once) series of algorithms has become the mainstream choice for target detection due to its high efficiency and accuracy. YOLOv7-tiny is a simplified version of the YOLOv7 model. While maintaining the detection performance, it improves the processing speed and computational efficiency by reducing the complexity and the number of parameters of the network, meeting the requirements of embedded systems and edge devices. However, in the application of UAV target detection, the YOLOv7-tiny algorithm faces specific challenges. UAVs usually fly in complex environments, and the targets to be detected are often small or highly occluded. Although YOLOv7-tiny has efficient real-time processing capabilities, it has problems such as insufficient detection accuracy for small targets and performance degradation in complex backgrounds. Summary of the Invention
[0004] In view of this, this application provides a method and device for UAV target detection. By improving the convolutional block attention unit in the YOLOv7-tiny model, weight distribution can be performed on the spatial dimension and channel dimension of the input features, increasing the network's attention to the feature information of UAVs, thereby reducing the impact of complex backgrounds on UAV recognition to a certain extent; through the SCFM module, interference information in the feature map can be effectively filtered from both the channel and spatial aspects, enhancing the feature information of small targets. Therefore, the embodiments of this application can enhance the recognition ability of small targets and the robustness in dense backgrounds.
[0005] According to one aspect of this application, a method for UAV target detection is provided, including:
[0006] Obtain a data set;
[0007] Design an improved YOLOv7-tiny model;
[0008] The improved YOLOv7-tiny model is trained using the said dataset to obtain a trained improved YOLOv7-tiny model. Among them, the improved YOLOv7-tiny model includes an input layer, a backbone network layer, a neck network layer, and a head network layer. The backbone network layer includes a C-ELAN module, and the C-ELAN module is generated based on a convolutional block attention unit and an ELAN sub-module. The neck network layer includes an SCFM module and the C-ELAN module, and the SCFM module includes a spatial filtering sub-module and a channel filtering sub-module;
[0009] The drone image to be detected is input into the input layer of the trained improved YOLOv7-tiny model for image preprocessing, and the preprocessed image is output;
[0010] The preprocessed image is input into the backbone network layer, and feature extraction is performed based on the C-ELAN module, and the feature extraction image is output;
[0011] The feature extraction image is input into the neck network layer, and feature fusion is performed based on the SCFM module and the C-ELAN module, and the feature fusion image is output;
[0012] The feature fusion image is input into the head network layer for detection, and the target detection image is output.
[0013] According to another aspect of the present application, a drone target detection device is provided. The device includes:
[0014] An acquisition module, used to acquire a dataset;
[0015] A design module, used to design an improved YOLOv7-tiny model;
[0016] A training module, used to train the improved YOLOv7-tiny model using the said dataset to obtain a trained improved YOLOv7-tiny model. Among them, the improved YOLOv7-tiny model includes an input layer, a backbone network layer, a neck network layer, and a head network layer. The backbone network layer includes a C-ELAN module, and the C-ELAN module is generated based on a convolutional block attention unit and an ELAN sub-module. The neck network layer includes an SCFM module and the C-ELAN module, and the SCFM module includes a spatial filtering sub-module and a channel filtering sub-module;
[0017] The detection module is used to input the drone image to be detected into the input layer of the trained improved YOLOv7-tiny model for image preprocessing, and output the preprocessed image; input the preprocessed image into the backbone network layer, perform feature extraction based on the C-ELAN module, and output the feature extraction image; input the feature extraction image into the neck network layer, perform feature fusion based on the SCFM module and the C-ELAN module, and output the feature fusion image; input the feature fusion image into the head network layer for detection, and output the target detection image.
[0018] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned drone target detection method is implemented.
[0019] According to still another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the program, the above-mentioned drone target detection method is implemented.
[0020] With the above technical solution, a method and device for detecting an unmanned aerial vehicle (UAV) target, a storage medium, and a computer device provided by this application can, first, obtain a data set. Next, an improved YOLOv7-tiny model can be designed. Among them, the improved YOLOv7-tiny model can include an input layer, a backbone network layer, a neck network layer, and a head network layer. Specifically, the backbone network layer contains a C-ELAN module formed by combining a convolutional block attention unit and an ELAN sub-module, and the neck network layer includes an SCFM module and the above C-ELAN module. After that, the obtained data set can be used to train the improved YOLOv7-tiny model to obtain the trained improved YOLOv7-tiny model. In the model application stage, a UAV image to be detected can be obtained and input into the trained improved YOLOv7-tiny model. After being input into the trained improved YOLOv7-tiny model, first, the input layer of the model can perform data preprocessing on the UAV image to be detected. Subsequently, the backbone network layer of the model can use multiple C-ELAN modules therein to extract image features, thereby obtaining a feature extraction image. Further, the neck network layer of the model can use the SCFM module, C-ELAN module, etc. to fuse image features, thereby obtaining a feature fusion image. Finally, the head network layer of the model performs detection based on the received feature fusion image and outputs a target detection image. In the embodiment of this application, by improving the convolutional block attention unit in the YOLOv7-tiny model, weight distribution can be performed on the spatial dimension and channel dimension of the input features, improving the network's attention to UAV feature information, thereby reducing the impact of complex backgrounds on UAV recognition to a certain extent; through the SCFM module, interference information of the feature map can be effectively filtered from both the channel and spatial aspects, enhancing small target feature information. Therefore, the embodiment of this application can enhance the recognition ability of small targets and the robustness in a dense background.
[0021] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the following specifically illustrates the specific embodiments of this application. Brief Description of the Drawings
[0022] The drawings described herein are used to provide a further understanding of this application and constitute a part of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0023] Figure 1 A flowchart showing a method for detecting a UAV target provided by an embodiment of this application is shown;
[0024] Figure 2 Shows a schematic structural diagram of an improved YOLOv7-tiny model provided by an embodiment of the present application;
[0025] Figure 3 Shows a schematic structural diagram of a CBL provided by an embodiment of the present application;
[0026] Figure 4 Shows a schematic structural diagram of a C-ELAN module provided by an embodiment of the present application;
[0027] Figure 5 Shows a schematic structural diagram of an SPP module provided by an embodiment of the present application;
[0028] Figure 6 Shows a schematic structural diagram of an SCFM module provided by an embodiment of the present application;
[0029] Figure 7 Shows a schematic structural diagram of a channel filtering sub-module provided by an embodiment of the present application;
[0030] Figure 8 Shows a schematic structural diagram of a drone target detection device provided by an embodiment of the present application;
[0031] Figure 9 Shows a schematic structural diagram of a device of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0032] The present application will be described in detail below with reference to the drawings and in combination with embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0033] In the application of drone target detection, the YOLOv7-tiny algorithm faces specific challenges, such as insufficient detection accuracy for small targets and performance degradation in complex backgrounds. Drones usually fly in complex environments, and the targets to be detected are often small or highly occluded. Although YOLOv7-tiny has efficient real-time processing capabilities, its standard configuration may perform poorly in these situations. Therefore, it is particularly important to improve the YOLOv7-tiny algorithm to enhance its recognition ability for small targets and robustness in dense backgrounds.
[0034] In this embodiment, a drone target detection method is provided, as Figure 1 shown, the method includes:
[0035] Step 101, obtain a data set.
[0036] Step 102, design an improved YOLOv7-tiny model.
[0037] Step 103: Use the dataset to train the improved YOLOv7-tiny model to obtain a trained improved YOLOv7-tiny model. The improved YOLOv7-tiny model includes an input layer, a backbone network layer, a neck network layer, and a head network layer. The backbone network layer includes a C-ELAN module, which is generated based on a convolutional block attention unit and an ELAN sub-module. The neck network layer includes an SCFM module and the C-ELAN module. The SCFM module includes a spatial filtering sub-module and a channel filtering sub-module.
[0038] A drone target detection method provided by an embodiment of this application is implemented based on an improved YOLOv7-tiny model. First, a dataset can be obtained. Here, the training samples and test samples in the dataset can be collected based on an image acquisition device, and specifically can include labeled drone images in different scenarios (cities, mountains, seas) and different lighting conditions (daytime, dusk, cloudy days). For example, drone images that have been annotated with bounding boxes and carry annotation information such as target categories and coordinates.
[0039] In one embodiment, the dataset can also be augmented in the following ways to increase the diversity of training samples and test samples. The methods for dataset augmentation can include: (1) Geometric transformation. Geometric transformation can include rotation transformation, translation and scaling, mirror flipping, etc. Rotation transformation: Rotate at an arbitrary angle based on the image center to increase the diversity of target angles. Translation and scaling: Simulate the distance change of the drone. For example, increase the diversity of target positions by random cropping (such as randomly intercepting an 800x800 area from a 1024x1024 image). Mirror flipping: Use horizontal mirroring to increase the diversity of shooting directions. (2) Color enhancement. Color enhancement can include brightness adjustment, noise injection, dynamic combination strategies, etc. Brightness adjustment: Increase the diversity of light intensity changes at different times (such as the light differences between dawn and dusk) through Gamma transformation or histogram equalization. Noise injection: Add Gaussian noise to simulate sensor interference and improve the robustness of the model in haze or electromagnetic interference environments. Dynamic combination strategy: Adopt an online enhancement method to randomly apply 3-5 transformation combinations to each input image, such as "rotate 15 degrees + increase brightness by 20% + add salt and pepper noise", to ensure that the model continuously learns the features of different scenarios.
[0040] Next, an improved YOLOv7-tiny model can be designed. Among them, the improved YOLOv7-tiny model can include an input layer, a backbone network layer, a neck network layer, and a head network layer. Specifically, the input layer can receive a 3-channel RGB image and preprocess the RGB image; the backbone network layer is a feature extraction layer. Among them, in order to solve the problem that the background of the constructed UAV dataset is complex and interferes with UAV recognition, the convolutional block attention unit (CBAM, Convolutional Block Attention Module) is combined with the ELAN sub-module in the backbone network layer of the original YOLOv7-Tiny model to form a C-ELAN module. The neck network layer is a feature fusion layer, and the neck network layer includes an SCFM module and the above-mentioned C-ELAN module. Among them, the SCFM module includes a spatial filtering sub-module and a channel filtering sub-module. This structure can effectively filter out the interference information of the feature map from both the channel and spatial aspects and enhance the small target feature information.
[0041] After that, the obtained dataset can be used to train the improved YOLOv7-tiny model to obtain the trained improved YOLOv7-tiny model.
[0042] Step 104, input the UAV image to be detected into the input layer of the trained improved YOLOv7-tiny model for image preprocessing, and output the preprocessed image.
[0043] Step 105, input the preprocessed image into the backbone network layer, perform feature extraction based on the C-ELAN module, and output the feature extraction image.
[0044] Step 106, input the feature extraction image into the neck network layer, perform feature fusion based on the SCFM module and the C-ELAN module, and output the feature fusion image.
[0045] Step 107, input the feature fusion image into the head network layer for detection, and output the target detection image.
[0046] In this embodiment, during the model application phase, an unmanned aerial vehicle (UAV) image to be detected can be obtained and input into the trained improved YOLOv7-tiny model. After being input into the trained improved YOLOv7-tiny model, first, the input layer of the model can perform data preprocessing on the UAV image to be detected. For example, the input size of the UAV image to be detected can be normalized, data augmentation can be performed, etc., so as to obtain a preprocessed image. Subsequently, the backbone network layer of the model can extract image features by using multiple C-ELAN modules therein, etc., so as to obtain a feature extraction image. Further, the neck network layer of the model can fuse image features by using the SCFM module, the C-ELAN module, etc., so as to obtain a feature fusion image. Finally, the head network layer of the model performs detection based on the received feature fusion image and outputs a target detection image. Among them, the target detection image can include information such as bounding box coordinates, target confidence, class probability, etc.
[0047] By applying the technical solution of this embodiment, first, a data set can be obtained. Next, an improved YOLOv7-tiny model can be designed. The improved YOLOv7-tiny model can include an input layer, a backbone network layer, a neck network layer, and a head network layer. Specifically, the backbone network layer contains a C-ELAN module formed by combining a convolutional block attention unit and an ELAN sub-module, and the neck network layer includes an SCFM module and the above C-ELAN module. After that, the obtained data set can be used to train the improved YOLOv7-tiny model to obtain the trained improved YOLOv7-tiny model. In the model application stage, an unmanned aerial vehicle (UAV) image to be detected can be obtained and input into the trained improved YOLOv7-tiny model. After being input into the trained improved YOLOv7-tiny model, first, the input layer of the model can perform data preprocessing on the UAV image to be detected. Subsequently, the backbone network layer of the model can use multiple C-ELAN modules therein to extract image features, thereby obtaining a feature extraction image. Further, the neck network layer of the model can use the SCFM module, C-ELAN module, etc. to perform image feature fusion, thereby obtaining a feature fusion image. Finally, the head network layer of the model performs detection based on the received feature fusion image and outputs a target detection image. In the embodiment of the present application, by improving the convolutional block attention unit in the YOLOv7-tiny model, weight distribution can be performed on the spatial dimension and channel dimension of the input features, improving the network's attention to UAV feature information, thereby reducing the impact of complex backgrounds on UAV recognition to a certain extent; through the SCFM module, interference information of the feature map can be effectively filtered from both the channel and spatial aspects, enhancing small target feature information. Therefore, the embodiment of the present application can enhance the recognition ability of small targets and robustness in a dense background.
[0048] In the embodiment of the present application, optionally, step 105 specifically includes: sequentially inputting the preprocessed image into a first CBL module, a second CBL module, a first C-ELAN module, a first MP module, and a second C-ELAN module to output a first feature extraction image; sequentially inputting the first feature extraction image into a second MP module and a third C-ELAN module to output a second feature extraction image; sequentially inputting the second feature extraction image into a third MP module, a fourth C-ELAN module, a third CBL module, an SPP module, and a fourth CBL module to output a third feature extraction image.
[0049] In this embodiment, as Figure 2As shown, the backbone network layer may include a first CBL module, a second CBL module, a first C-ELAN module, a first MP module, a second C-ELAN module, a second MP module, a third C-ELAN module, a third MP module, a fourth C-ELAN module, a third CBL module, an SPP module, and a fourth CBL module. Among them, the structures of the first CBL module, the second CBL module, the third CBL module, and the fourth CBL module may be the same; the structures of the first C-ELAN module, the second C-ELAN module, the third C-ELAN module, and the fourth C-ELAN module may be the same; the structures of the first MP module, the second MP module, and the third MP module may be the same.
[0050] In a specific embodiment, as Figure 3 shown, each CBL module may include Conv (Convolutional Layer), BN (Batch Normalization), and a LeakyReLU activation function. Among them, Conv is a convolutional layer used to extract local features; BN is a batch normalization layer used to normalize the output of the convolutional layer; the LeakyReLU activation function is used to introduce non-linearity and enhance the expression ability of the model. Among them, the convolution kernel of Conv can be determined according to actual needs.
[0051] In the embodiment of the present application, optionally, the ELAN sub-module includes a splicing unit and multiple CBL units, and the C-ELAN module is obtained by combining the ELAN sub-module with the convolutional block attention unit; after the C-ELAN module receives the first input image, the method further includes: respectively inputting the first input image into the first CBL unit and the second CBL unit in the C-ELAN module, and outputting a first convolution result corresponding to the first CBL unit and a second convolution result corresponding to the second CBL unit; sequentially inputting the first convolution result into the convolutional block attention unit and the third CBL unit in the C-ELAN module, and outputting a third convolution result; inputting the third convolution result into the fourth CBL unit in the C-ELAN module, and outputting a fourth convolution result; splicing the first convolution result, the second convolution result, the third convolution result, and the fourth convolution result to generate a target splicing result; inputting the target splicing result into the fifth CBL unit in the C-ELAN module to obtain the output result corresponding to the C-ELAN module.
[0052] As Figure 4As shown, the C-ELAN module may include multiple CBL units, a convolutional block attention unit, and a splicing unit. Among them, the convolutional block attention unit improves the network's attention to the feature information of the drone by assigning weights to the spatial dimension and channel dimension of the input features, thereby reducing the impact of complex backgrounds on drone recognition to a certain extent. The structure of the CBL unit may be the same as that of the CBL module. "Unit" and "module" are only used to distinguish the levels to which the CBL belongs. The splicing unit may specifically be a CONCAT unit. It should be noted that, regardless of the input of which C-ELAN module, it can be uniformly referred to as the first input image. In the C-ELAN module, the first CBL unit and the second CBL unit are two parallel CBL units. In the C-ELAN module, the first convolution result is processed by the convolutional block attention unit. On the one hand, the channel attention in the convolutional block attention unit can learn the channel importance weights, and on the other hand, the spatial attention therein can learn the spatial importance weights. Thus, it can not only enhance the contribution of important feature channels but also highlight the spatial position of the target area, improve the model's sensitivity to key features, and at the same time suppress background noise interference, so that while maintaining a high inference speed, it can effectively improve the model's detection ability for multi-scale targets.
[0053] In a specific embodiment, the MP module may be a max pooling layer.
[0054] In a specific embodiment, as Figure 5 shown, the SPP (Spatial Pyramid Pooling) module may include a first CBL sub-module, three max pooling sub-modules, a first splicing sub-module, a second CBL sub-module, and a second splicing sub-module. Among them, the structures of the first CBL sub-module and the second CBL sub-module may be the same, and the size of the convolutional kernel can be determined according to requirements; the structures of the first splicing sub-module and the second splicing sub-module may be the same. It should be noted that the structure of the CBL sub-module here may be the same as that of the aforementioned CBL module, and "sub-module" and "module" are only used to distinguish the levels to which the CBL belongs. The max pooling sub-module here may be a Maxpool layer, and the splicing sub-module may be a CONCAT. SPP fuses the information of feature maps of different scales through pooling operations with convolutional kernels of different sizes to complete feature fusion.
[0055] In an embodiment of the present application, optionally, step 106 specifically includes: sequentially inputting the first feature extraction image into a first SCFM module and a first Conv module to generate a first image; splicing the first image with the upsampled image corresponding to the first feature extraction image to obtain a first spliced image; inputting the first spliced image into a fifth C-ELAN module to output a first feature fusion image; sequentially inputting the second feature extraction image into a second SCFM module and a second Conv module to generate a second image; splicing the second image with the upsampled image corresponding to the third feature extraction image to obtain a second spliced image; inputting the second spliced image into a sixth C-ELAN module to generate a third image; inputting the first feature fusion image into a fifth CBL module to generate a fourth image; splicing the third image and the fourth image to obtain a third spliced image; inputting the third spliced image into a seventh C-ELAN module to output a second feature fusion image; inputting the second feature fusion image into a sixth CBL module to generate a fifth image; splicing the third feature extraction image and the fifth image to obtain a fourth spliced image; inputting the fourth spliced image into an eighth C-ELAN module to output a third feature fusion image.
[0056] In this embodiment, as Figure 2 shown, the neck network layer is a feature fusion stage. The neck network layer includes multiple SCFM modules, multiple Conv modules, multiple CONCAT modules, multiple C-ELAN modules, multiple CBL modules, and multiple upsampling modules. The neck network layer uses a Path Aggregation Network (PANet) as the feature fusion part of the network. By adding a bottom-up path enhancement on the basis of a Feature Pyramid Network (FPN), the accurate low-level localization signals are used to enhance the entire feature hierarchy, thereby shortening the information path between the low-level and top-level features. In an embodiment of the present application, a spatial filtering sub-module and a channel filtering sub-module are introduced into the SCFM module, and this structure can effectively filter out the interference information of the feature map from the channel and space and enhance the small target feature information.
[0057] It should be noted that before upsampling the first feature extraction image and the third feature extraction image, they can also be processed by a CBL module first and then upsampled. The structure of the CBL module here is the same as that of the aforementioned CBL module, and the size of the convolution kernel can be specifically determined according to requirements; the splicing processing operation here is implemented by each CONCAT module.
[0058] In an embodiment of the present application, optionally, the SCFM module further includes a Conv sub-module, a first multiplication sub-module, a second multiplication sub-module, and a summation sub-module; after the SCFM module receives the second input image, the method further includes: inputting the second input image into the Conv sub-module to output a fifth convolution result; inputting the fifth convolution result into the spatial filtering sub-module and the channel filtering sub-module respectively to output a first filtering result corresponding to the spatial filtering sub-module and a second filtering result corresponding to the channel filtering sub-module; inputting the first filtering result and the fifth convolution result into the first multiplication sub-module together to obtain a first multiplication result, and inputting the second filtering result and the fifth convolution result into the second multiplication sub-module together to obtain a second multiplication result; inputting the first multiplication result and the second multiplication result into the summation sub-module together to obtain the output result corresponding to the SCFM module.
[0059] In this embodiment, as Figure 6 shown, the SCFM module includes a Conv sub-module, a first multiplication sub-module, a second multiplication sub-module, a summation sub-module, a spatial filtering sub-module, and a channel filtering sub-module. First, the second input image is compressed in the channel direction by the Conv sub-module, and then it is input into the spatial filtering sub-module and the channel filtering sub-module respectively. Among them, the spatial filtering sub-module is used to capture local spatial patterns (such as edges, textures) and enhance the feature expression in the spatial dimension; the channel filtering sub-module is used to learn the relationships between channels and enhance the feature expression in the channel dimension. Through these two sub-modules, small target features can be effectively enhanced. Here, the spatial filtering sub-module processes the compressed feature map using the Log_Softmax activation function to filter and enhance the feature information in the space. The first multiplication sub-module is used to fuse the spatial filtering features (i.e., the first filtering result) and the compressed features (i.e., the fifth convolution result) to obtain spatially enhanced features (i.e., the first multiplication result); the second multiplication sub-module is used to fuse the channel filtering features (i.e., the second filtering result) and the compressed features to obtain channel-enhanced features (i.e., the second multiplication result). The summation sub-module is used to integrate the spatially enhanced features and the channel-enhanced features. The embodiment of the present application enables the improved YOLOv7-tiny model to more effectively capture the multi-scale features and complex context information of the target through feature enhancement in the spatial and channel dimensions, and is particularly suitable for target detection tasks with large target scale changes and complex backgrounds.
[0060] It should be noted that, regardless of the input of which SCFM module, it can be uniformly referred to as the second input image. The structure of the Conv sub-module can be the same as that of the aforementioned Conv module. The "sub-module" and "module" are only used to distinguish the levels to which Conv belongs. The first product module and the second product module can be Multipy layers.
[0061] In a specific embodiment, the spatial filtering sub-module can capture local context information through a large convolution kernel (5x5), and the channel filtering sub-module can learn the non-linear relationship between channels through 1x1 convolution.
[0062] In the embodiment of the present application, optionally, the spatial filtering sub-module includes a Log_Softmax activation function, and the channel filtering sub-module includes an average pooling unit, a max pooling unit, a plurality of convolution units, a Hardswish activation function, a plurality of upsampling units, and a summation unit; after the channel filtering sub-module receives the third input image, the method further includes: sequentially inputting the third input image into the average pooling unit, the first convolution unit, the Hardswish activation function, the second convolution unit, and the first upsampling unit to output a first upsampling result, and, sequentially inputting the third input image into the max pooling unit, the third convolution unit, the Hardswish activation function, the fourth convolution unit, and the second upsampling unit to output a second upsampling result; jointly inputting the first upsampling result and the second upsampling result into the summation unit to obtain the output result corresponding to the channel filtering sub-module.
[0063] In this embodiment, the spatial filtering sub-module processes the compressed feature map (i.e., the fifth convolution result) using the Log_Softmax activation function to filter and enhance the feature information of the feature map spatially.
[0064] The spatial filtering sub-module includes a Log_Softmax activation function, which performs a Log operation on the result based on the Softmax function to generate the relative weights of all positions with respect to the channels.
[0065] The formula of the Softmax function is:
[0066] ;
[0067] The formula of the Log_Softmax function is:
[0068] ;
[0069] In the formula: represents the Softmax function, Z represents the input vector, Denotes the j-th element, Denotes the i-th element, e Denotes the exponential function, Denotes the summation function, Denotes the Log_Softmax function, ln Denotes the logarithmic function.
[0070] From the algorithm formulas of the two, although both Log_Softmax and Softmax are monotonic, their effects on the relative values of the loss function are different. Using the Log_Softmax function can make the algorithm converge faster.
[0071] Such as Figure 7 As shown, a channel filtering sub-module provided by an embodiment of the present application is shown. The channel filtering sub-module includes an average pooling unit, a max pooling unit, a plurality of convolutional units, a Hardswish activation function, and a plurality of upsampling units. After the third input image is input into the channel filtering sub-module, it passes through the branches corresponding to the average pooling unit and the max pooling unit respectively. The results of the subsequent two branches will be combined to obtain more detailed global features; 2 convolutional units and the Hardswish activation function are introduced in each branch to enhance the small target features. Finally, the size of the feature map is restored through the nearest neighbor upsampling operation, and the results of the two branches are added to obtain the final output of the channel filtering sub-module.
[0072] Hardswish is introduced here as the activation function because Hardswish not only has the advantages of the Swish function, that is, it helps to prevent the gradient from gradually approaching 0 and causing saturation during slow training, smoothness plays an important role in optimization and generalization, and the derivative is always greater than 0, but also is composed of common operator combinations on the basis of the Swish function, greatly reducing the computational amount of the algorithm model while achieving similar effects.
[0073] It should be noted that the average pooling unit here can be an AvgPool layer, the max pooling unit can be a MaxPool layer, the first convolutional unit, the second convolutional unit, the third convolutional unit, and the fourth convolutional unit can be Conv2d with the same structure, and the parameter settings can be determined according to actual needs. The first upsampling unit and the second upsampling unit can be Upsample layers.
[0074] In an embodiment of the present application, optionally, the head network layer includes a first head detection module, a second head detection module, and a third head detection module, and the first head detection module, the second head detection module, and the third head detection module are respectively used to detect targets of different scales; step 107 specifically includes: inputting the first feature fusion image into the first head detection module to output a first target detection image; inputting the second feature fusion image into the second head detection module to output a second target detection image; inputting the third feature fusion image into the third head detection module to output a third target detection image.
[0075] In this embodiment, as Figure 2 shown, the head network layer is in the detection stage. The head network layer may include a first head detection module, a second head detection module, and a third head detection module. Each head detection module may specifically be a YOLO HEAD module. Different head detection modules are used to detect targets of different scales. Among them, the first head detection module is used to process shallow-layer high-resolution features, the second head detection module is used to process middle-layer medium-resolution features, and the third head detection module is used to process deep-layer low-resolution features. Using three head detection modules to respectively output target detection images of three scales can detect large, medium, and small objects.
[0076] In an embodiment of the present application, optionally, the positioning loss function adopted by the improved YOLOv7-tiny model is the Wise-IoU loss function, and the Wise-IoU loss function is specifically:
[0077] ;
[0078] ;
[0079] ;
[0080] In the formula, represents the Wise-IoU loss function, represents the first hyperparameter, represents the second hyperparameter, represents the third hyperparameter, represents the penalty term, represents the IoU loss function, exp represents the exponential function, represents the coordinates of the center point of the predicted bounding box, represents the coordinates of the center point of the ground truth bounding box, represents the width of the smallest enclosing box of the predicted bounding box and the ground truth bounding box, represents the height of the smallest enclosing box of the predicted bounding box and the ground truth bounding box, represents the intersection over union.
[0081] In this embodiment, the loss function of the model is further optimized. The loss function in the traditional YOLOv7-tiny model consists of confidence loss, classification loss, and localization loss. Among them, the CIoU loss function is used for the localization loss, and its disadvantage is that it does not consider the problem of only detection mismatch between the predicted box and the ground truth box, resulting in a slow convergence speed. Therefore, the Wise-IoU is introduced as a new localization loss function in the embodiment of this application. In the Wise-IoU loss function, represents the intersection over union (IoU), which is a common measure of the overlap between the predicted box and the ground truth box. and are usually set to 1.9 and 3.0. When it is the case, it makes .
[0082] By introducing the outlier degree to describe the quality of the anchor box, is negatively correlated with the quality of the anchor box. The formula for calculating the outlier degree is:
[0083] ;
[0084] In the formula, is the gradient gain of the monotonic focusing coefficient, which has the same definition as . Here, means that it will be continuously calculated and changed according to the situation of each object detection during the training process; is the moving average of the momentum m. Introducing means that the highest gradient gain can be dynamically adjusted according to the training process. The formula for calculating the momentum m is:
[0085] ;
[0086] In the formula: t is the value of the epoch, and n is the value of the batch size. The significance of introducing the momentum m is that after t rounds of training, the WIoU assigns small gradient gains to low-quality anchor boxes to reduce harmful gradients.
[0087] In the embodiment of this application, optionally, step 104 specifically includes: inputting the drone image to be detected into the input layer of the trained improved YOLOv7-tiny model, and performing data augmentation and adaptive anchor frame calculation on the drone image to be detected through the input layer, and outputting the preprocessed image.
[0088] In this embodiment, the input layer of the improved YOLOv7-tiny model can preprocess the drone images to be detected. The preprocessing can include data padding, adaptive anchor frame calculation, etc., to ensure the uniform scaling of RGB images, so as to meet the input size requirements of the backbone network layer. Among them, data augmentation is a commonly used technique in machine learning, especially in the field of deep learning, which is used to improve the generalization ability of the model by increasing the diversity of training data. In the model application stage, data augmentation (such as brightness fine-tuning, small-angle rotation) can simulate some real perturbations and improve the model's adaptability to domain shift. Since in the dynamic scenario of drones, sudden changes in light, motion blur, and changes in shooting angles are normal. For example: light change: During the alternation of day and night, the imaging of the same target under strong light and weak light is significantly different. Motion blur: When flying at high speed, the propeller or small components may be blurred due to jitter. Therefore, introducing data augmentation in the application stage also helps to improve the accuracy of subsequent target recognition.
[0089] Adaptive Anchor Calculation, that is, adaptive anchor box calculation, refers to automatically calculating and adjusting the shape, size, and number of anchor boxes according to the characteristics of the dataset and the requirements of the algorithm in the object detection task, so as to improve the accuracy and efficiency of object detection.
[0090] When preprocessing the drone images to be detected in the embodiment of this application, through the preprocessing method that combines data augmentation and adaptive anchor frame calculation, the enhancement at the data level and the collaborative optimization of the model architecture are realized, so that the improved YOLOv7-tiny shows stronger adaptability to multi-scale targets, complex lighting, and changes in shooting angles in scenarios such as drone inspection, and can improve the detection accuracy of small targets.
[0091] Furthermore, as Figure 1 a specific implementation of the method, the embodiment of this application provides a drone target detection device, as Figure 8 shown, the device includes:
[0092] An acquisition module, used to acquire a dataset;
[0093] A design module, used to design an improved YOLOv7-tiny model;
[0094] A training module for training the improved YOLOv7-tiny model using the dataset to obtain a trained improved YOLOv7-tiny model, where the improved YOLOv7-tiny model includes an input layer, a backbone network layer, a neck network layer, and a head network layer. The backbone network layer includes a C-ELAN module, and the C-ELAN module is generated based on a convolutional block attention unit and an ELAN sub-module. The neck network layer includes an SCFM module and the C-ELAN module, and the SCFM module includes a spatial filtering sub-module and a channel filtering sub-module;
[0095] A detection module for inputting an unmanned aerial vehicle image to be detected into the input layer of the trained improved YOLOv7-tiny model for image preprocessing and outputting a preprocessed image; inputting the preprocessed image into the backbone network layer, performing feature extraction based on the C-ELAN module, and outputting a feature extraction image; inputting the feature extraction image into the neck network layer, performing feature fusion based on the SCFM module and the C-ELAN module, and outputting a feature fusion image; inputting the feature fusion image into the head network layer for detection and outputting a target detection image.
[0096] Optionally, the detection module is specifically configured to:
[0097] Input the preprocessed image into a first CBL module, a second CBL module, a first C-ELAN module, a first MP module, and a second C-ELAN module in sequence, and output a first feature extraction image;
[0098] Input the first feature extraction image into a second MP module and a third C-ELAN module in sequence, and output a second feature extraction image;
[0099] Input the second feature extraction image into a third MP module, a fourth C-ELAN module, a third CBL module, an SPP module, and a fourth CBL module in sequence, and output a third feature extraction image.
[0100] Optionally, the detection module is specifically configured to:
[0101] Input the first feature extraction image into a first SCFM module and a first Conv module in sequence to generate a first image;
[0102] Stitch the first image with the upsampled image corresponding to the first feature extraction image to obtain a first stitched image;
[0103] Input the first stitched image into a fifth C-ELAN module, and output a first feature fusion image;
[0104] Input the second feature extraction image into the second SCFM module and the second Conv module in sequence to generate a second image;
[0105] Stitch the second image and the upsampled image corresponding to the third feature extraction image to obtain a second stitched image;
[0106] Input the second stitched image into the sixth C-ELAN module to generate a third image;
[0107] Input the first feature fusion image into the fifth CBL module to generate a fourth image;
[0108] Stitch the third image and the fourth image to obtain a third stitched image;
[0109] Input the third stitched image into the seventh C-ELAN module to output a second feature fusion image;
[0110] Input the second feature fusion image into the sixth CBL module to generate a fifth image;
[0111] Stitch the third feature extraction image and the fifth image to obtain a fourth stitched image;
[0112] Input the fourth stitched image into the eighth C-ELAN module to output a third feature fusion image.
[0113] Optionally, the head network layer includes a first head detection module, a second head detection module, and a third head detection module, and the first head detection module, the second head detection module, and the third head detection module are respectively used to detect targets of different scales; the detection module is specifically used for:
[0114] Input the first feature fusion image into the first head detection module to output a first target detection image;
[0115] Input the second feature fusion image into the second head detection module to output a second target detection image;
[0116] Input the third feature fusion image into the third head detection module to output a third target detection image.
[0117] Optionally, the ELAN sub-module includes a stitching unit and multiple CBL units, and the C-ELAN module is obtained by combining the ELAN sub-module with the convolutional block attention unit; the detection module is further used for:
[0118] After receiving the first input image, the C-ELAN module inputs the first input image into the first CBL unit and the second CBL unit in the C-ELAN module respectively, and outputs a first convolution result corresponding to the first CBL unit and a second convolution result corresponding to the second CBL unit;
[0119] The first convolution result is sequentially input into the convolutional block attention unit and the third CBL unit in the C-ELAN module, and a third convolution result is output;
[0120] The third convolution result is input into the fourth CBL unit in the C-ELAN module, and a fourth convolution result is output;
[0121] The first convolution result, the second convolution result, the third convolution result and the fourth convolution result are concatenated to generate a target concatenation result;
[0122] The target concatenation result is input into the fifth CBL unit in the C-ELAN module to obtain the output result corresponding to the C-ELAN module.
[0123] Optionally, the SCFM module further includes a Conv sub-module, a first product sub-module, a second product sub-module and a summation sub-module; the detection module is further configured to:
[0124] After receiving the second input image, the SCFM module inputs the second input image into the Conv sub-module and outputs a fifth convolution result;
[0125] The fifth convolution result is respectively input into the spatial filtering sub-module and the channel filtering sub-module, and a first filtering result corresponding to the spatial filtering sub-module and a second filtering result corresponding to the channel filtering sub-module are output;
[0126] The first filtering result and the fifth convolution result are jointly input into the first product sub-module to obtain a first product result, and the second filtering result and the fifth convolution result are jointly input into the second product sub-module to obtain a second product result;
[0127] The first product result and the second product result are jointly input into the summation sub-module to obtain the output result corresponding to the SCFM module.
[0128] Optionally, the spatial filtering sub-module includes a Log_Softmax activation function, and the channel filtering sub-module includes an average pooling unit, a max pooling unit, a plurality of convolutional units, a Hardswish activation function, a plurality of upsampling units and a summation unit; the detection module is further configured to:
[0129] After receiving the third input image, the channel filtering sub-module sequentially inputs the third input image into an average pooling unit, a first convolutional unit, a Hardswish activation function, a second convolutional unit, and a first upsampling unit to output a first upsampling result. Additionally, the third input image is sequentially input into a max pooling unit, a third convolutional unit, a Hardswish activation function, a fourth convolutional unit, and a second upsampling unit to output a second upsampling result;
[0130] The first upsampling result and the second upsampling result are jointly input into the summation unit to obtain the output result corresponding to the channel filtering sub-module.
[0131] Optionally, the positioning loss function adopted by the improved YOLOv7-tiny model is the Wise-IoU loss function. The Wise-IoU loss function is specifically:
[0132] ;
[0133] ;
[0134] ;
[0135] In the formula, represents the Wise-IoU loss function, represents the first hyperparameter, represents the second hyperparameter, represents the third hyperparameter, represents the penalty term, represents the IoU loss function, exp represents the exponential function, represents the coordinates of the center point of the predicted bounding box, represents the coordinates of the center point of the ground truth bounding box, represents the width of the smallest enclosing box of the predicted bounding box and the ground truth bounding box, represents the height of the smallest enclosing box of the predicted bounding box and the ground truth bounding box, represents the intersection over union.
[0136] Optionally, the detection module is specifically configured to:
[0137] Input the drone image to be detected into the input layer of the trained improved YOLOv7-tiny model. The input layer performs data augmentation and adaptive anchor frame calculation on the drone image to be detected and outputs the preprocessed image.
[0138] It should be noted that for other corresponding descriptions of each functional unit involved in the drone target detection device provided in the embodiments of the present application, reference can be made to Figures 1 to 7The corresponding descriptions in the method will not be elaborated here.
[0139] The embodiments of the present application further provide a computer device, which can specifically be a personal computer, a server, a network device, etc. For example, Figure 9 As shown, the computer device includes a bus, a processor, a memory, and a communication interface, and may further include an input / output interface and a display device. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store location information. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements the steps in the method embodiments.
[0140] Those skilled in the art can understand that Figure 9 the structure shown in [[ ]] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component layout.
[0141] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium may be non-volatile or volatile, and has a computer program stored thereon. When the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0142] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0143] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties.
[0144] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., without limitation.
[0145] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0146] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for detecting a drone target, characterized in that: include: Get the dataset; Design and improve the YOLOv7-tiny model; The improved YOLOv7-tiny model is trained using the data set to obtain a trained improved YOLOv7-tiny model, wherein the improved YOLOv7-tiny model includes an input layer, a backbone network layer, a neck network layer and a head network layer, the backbone network layer includes a C-ELAN module, the C-ELAN module is generated based on a convolutional block attention unit and an ELAN submodule, the neck network layer includes an SCFM module and the C-ELAN module, and the SCFM module includes a spatial filtering submodule and a channel filtering submodule; Input the drone image to be detected into the input layer of the trained improved YOLOv7-tiny model for image preprocessing, and output the preprocessed image; Input the preprocessed image into the backbone network layer, perform feature extraction based on the C-ELAN module, and output a feature-extracted image; Input the feature extraction image into the neck network layer, perform feature fusion based on the SCFM module and the C-ELAN module, and output a feature fusion image; Input the feature fusion image into the head network layer for detection, and output a target detection image; The C-ELAN module includes a plurality of CBL units, a convolutional block attention unit and a splicing unit, wherein the convolutional block attention unit is used to assign weights to the spatial dimension and channel dimension of the input feature; The SCFM module further includes a Conv submodule, a first product submodule, a second product submodule and a summation submodule; after the SCFM module receives the second input image, the method further includes: Input the second input image into the Conv submodule, and output a fifth convolution result; Input the fifth convolution result into the spatial filtering submodule and the channel filtering submodule respectively, and output a first filtering result corresponding to the spatial filtering submodule and a second filtering result corresponding to the channel filtering submodule; Inputting the first filtering result and the fifth convolution result into the first product submodule to obtain a first product result, and inputting the second filtering result and the fifth convolution result into the second product submodule to obtain a second product result; The first multiplication result and the second multiplication result are input into the summing submodule to obtain an output result corresponding to the SCFM module.
2. The drone target detection method according to claim 1, characterized in that: The step of inputting the preprocessed image into the backbone network layer, performing feature extraction based on the C-ELAN module, and outputting a feature-extracted image specifically includes: The preprocessed image is sequentially input into a first CBL module, a second CBL module, a first C-ELAN module, a first MP module, and a second C-ELAN module, and a first feature extraction image is output; Inputting the first feature extraction image into a second MP module and a third C-ELAN module in sequence, and outputting a second feature extraction image; The second feature extraction image is sequentially input into the third MP module, the fourth C-ELAN module, the third CBL module, the SPP module and the fourth CBL module, and a third feature extraction image is output.
3. The method for detecting a drone target according to claim 2, characterized in that: The step of inputting the feature extraction image into the neck network layer, performing feature fusion based on the SCFM module and the C-ELAN module, and outputting a feature fusion image specifically includes: Inputting the first feature extraction image into a first SCFM module and a first Conv module in sequence to generate a first image; splicing the first image with an upsampled image corresponding to the first feature extraction image to obtain a first spliced image; Input the first stitched image into a fifth C-ELAN module, and output a first feature fused image; Inputting the second feature extraction image into a second SCFM module and a second Conv module in sequence to generate a second image; splicing the second image with an upsampled image corresponding to the third feature extraction image to obtain a second spliced image; Inputting the second stitched image into a sixth C-ELAN module to generate a third image; Inputting the first feature fusion image into a fifth CBL module to generate a fourth image; splicing the third image and the fourth image to obtain a third spliced image; Input the third stitched image into a seventh C-ELAN module, and output a second feature fused image; Inputting the second feature fusion image into a sixth CBL module to generate a fifth image; splicing the third feature extraction image and the fifth image to obtain a fourth spliced image; The fourth stitched image is input into the eighth C-ELAN module, and a third feature fused image is output.
4. The method for detecting a drone target according to claim 3, characterized in that: The head network layer includes a first head detection module, a second head detection module and a third head detection module, wherein the first head detection module, the second head detection module and the third head detection module are respectively used to detect targets of different scales; the inputting the feature fusion image into the head network layer for detection and outputting the target detection image specifically includes: Inputting the first feature fusion image into a first head detection module, and outputting a first target detection image; Inputting the second feature fusion image into a second head detection module, and outputting a second target detection image; The third feature fusion image is input into a third head detection module, and a third target detection image is output.
5. The drone target detection method according to claim 1, characterized in that: The ELAN submodule includes a splicing unit and a plurality of CBL units, and the C-ELAN module is obtained by combining the ELAN submodule with the convolutional block attention unit; After the C-ELAN module receives the first input image, the method further includes: Inputting the first input image into a first CBL unit and a second CBL unit in the C-ELAN module respectively, and outputting a first convolution result corresponding to the first CBL unit and a second convolution result corresponding to the second CBL unit; Inputting the first convolution result into the convolution block attention unit and the third CBL unit in the C-ELAN module in sequence, and outputting the third convolution result; Input the third convolution result into the fourth CBL unit in the C-ELAN module, and output the fourth convolution result; Splicing the first convolution result, the second convolution result, the third convolution result, and the fourth convolution result to generate a target splicing result; The target splicing result is input into the fifth CBL unit in the C-ELAN module to obtain the output result corresponding to the C-ELAN module.
6. The drone target detection method according to claim 1, characterized in that: The spatial filtering submodule includes a Log_Softmax activation function, and the channel filtering submodule includes an average pooling unit, a maximum pooling unit, a plurality of convolution units, a Hardswish activation function, a plurality of upsampling units, and a summing unit; After the channel filtering submodule receives the third input image, the method further includes: Inputting the third input image sequentially into an average pooling unit, a first convolution unit, a Hardswish activation function, a second convolution unit, and a first upsampling unit, and outputting a first upsampling result; and inputting the third input image sequentially into a maximum pooling unit, a third convolution unit, a Hardswish activation function, a fourth convolution unit, and a second upsampling unit, and outputting a second upsampling result; The first up-sampling result and the second up-sampling result are input into the summing unit to obtain an output result corresponding to the channel filtering submodule.
7. The method for detecting a drone target according to claim 1, characterized in that: The positioning loss function used by the improved YOLOv7-tiny model is the Wise-IoU loss function, and the Wise-IoU loss function is specifically: ; ; ; In the formula, represents the Wise-IoU loss function, represents the first hyperparameter, represents the second hyperparameter, represents the third hyperparameter, represents the penalty term, represents the IoU loss function, exp represents the exponential function, Represents the coordinates of the center point of the prediction box, Represents the coordinates of the center point of the real frame, Indicates the width of the minimum bounding box between the predicted box and the real box, Indicates the height of the minimum bounding box between the predicted box and the real box, Represents the intersection and union ratio.
8. The method for detecting a drone target according to claim 1, characterized in that: The method of inputting the image of the drone to be detected into the input layer of the trained improved YOLOv7-tiny model for image preprocessing and outputting the preprocessed image specifically includes: The drone image to be detected is input into the input layer of the trained improved YOLOv7-tiny model, and data expansion and adaptive anchor frame calculation are performed on the drone image to be detected through the input layer, and the preprocessed image is output.
9. A drone target detection device, characterized in that: The device comprises: The acquisition module is used to obtain the data set; Design module for designing and improving the YOLOv7-tiny model; A training module, used to train the improved YOLOv7-tiny model using the data set to obtain a trained improved YOLOv7-tiny model, wherein the improved YOLOv7-tiny model includes an input layer, a backbone network layer, a neck network layer and a head network layer, the backbone network layer includes a C-ELAN module, the C-ELAN module is generated based on a convolutional block attention unit and an ELAN submodule, the neck network layer includes an SCFM module and the C-ELAN module, and the SCFM module includes a spatial filtering submodule and a channel filtering submodule; The detection module is used to input the drone image to be detected into the input layer of the trained improved YOLOv7-tiny model for image preprocessing, and output the preprocessed image; input the preprocessed image into the backbone network layer, perform feature extraction based on the C-ELAN module, and output a feature extraction image; input the feature extraction image into the neck network layer, perform feature fusion based on the SCFM module and the C-ELAN module, and output a feature fusion image; input the feature fusion image into the head network layer for detection, and output a target detection image; The C-ELAN module includes a plurality of CBL units, a convolutional block attention unit and a splicing unit, wherein the convolutional block attention unit is used to assign weights to the spatial dimension and channel dimension of the input feature; The SCFM module further includes a Conv submodule, a first product submodule, a second product submodule and a summation submodule; the detection module is further used for: After receiving the second input image, the SCFM module inputs the second input image into the Conv submodule and outputs a fifth convolution result; Input the fifth convolution result into the spatial filtering submodule and the channel filtering submodule respectively, and output a first filtering result corresponding to the spatial filtering submodule and a second filtering result corresponding to the channel filtering submodule; Inputting the first filtering result and the fifth convolution result into the first product submodule to obtain a first product result, and inputting the second filtering result and the fifth convolution result into the second product submodule to obtain a second product result; The first multiplication result and the second multiplication result are input into the summing submodule to obtain an output result corresponding to the SCFM module.
Citation Information
Patent Citations
Remote sensing target detection method and device, electronic equipment and storage medium
CN117671509A
Industrial product surface defect detection method
CN118411339A