Unmanned aerial vehicle target detection method and device

By improving the convolutional block attention unit of the YOLOv7-tiny model and introducing the SCFM module, the problem of insufficient detection accuracy of small targets and performance degradation in complex contexts in UAV target detection is solved, and stronger recognition ability and robustness of small targets are achieved.

CN120047677AActive Publication Date: 2025-05-27XIAN ORDNANCE IND TECH IND DEV CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510520251.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The YOLOv7-tiny algorithm has problems with insufficient small target detection accuracy and performance degradation in complex contexts in UAV target detection.

Method used

By improving the convolutional block attention unit in the YOLOv7-tiny model, the weight allocation of spatial dimensions and channel dimensions is performed to enhance the network's attention to drone feature information. At the same time, an SCFM module is introduced to filter out interference information of the feature map from both the channel and the space, and enhance the small-target feature information.

Benefits of technology

It improves the recognition ability of small targets and its robustness in dense contexts, and enhances the accuracy and efficiency of drone target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047677A_ABST
    Figure CN120047677A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of target detection, and particularly discloses an unmanned aerial vehicle target detection method and device, and the method comprises the steps: obtaining a data set; an improved YOLOv7-tiny model is designed; a data set is utilized to train an improved YOLOv7-tiny model to obtain a trained improved YOLOv7-tiny model, a backbone network layer of the improved YOLOv7-tiny model comprises a C-ELAN module, the C-ELAN module is generated based on a convolution block attention unit and an ELAN sub-module, a neck network layer comprises an SCFM module and a C-ELAN module, and the SCFM module comprises a spatial filtering sub-module and a channel filtering sub-module; and inputting a to-be-detected unmanned aerial vehicle image into the trained improved YOLOv7-tiny model, and outputting a target detection image. According to the invention, the C-ELAN module and the SCFM module are introduced, so that the detection precision of a small target can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of target detection, and in particular to a method and device for detecting a target of an unmanned aerial vehicle. Background Art

[0002] Unmanned aerial vehicles (UAVs) are increasingly used in target detection, especially in surveillance, search and rescue, and military applications. Target detection algorithms play a central role in real-time data processing and decision support for UAVs.

[0003] Traditional drone target detection methods have problems such as window redundancy and poor feature robustness. Since the emergence of deep learning, target detection has made great breakthroughs, including two-stage algorithms represented by Faster R-CNN and single-stage target detection algorithms represented by YOLO and SSD series. Among them, the YOLO (You Only Look Once) series of algorithms have become the mainstream choice for target detection due to their high efficiency and accuracy. YOLOv7-tiny is a simplified version of the YOLOv7 model. While maintaining detection performance, it improves processing speed and computing efficiency by reducing network complexity and parameter volume, and adapts to the needs of embedded systems and edge devices. However, in drone target detection applications, the YOLOv7-tiny algorithm faces specific challenges. Drones usually fly in complex environments, and the targets to be detected are often small or highly occluded. Although YOLOv7-tiny has efficient real-time processing capabilities, it has problems such as insufficient detection accuracy of small targets and decreased performance in complex backgrounds. Summary of the invention

[0004] In view of this, the present application provides a method and device for detecting drone targets. By improving the convolutional block attention unit in the YOLOv7-tiny model, the spatial dimension and channel dimension of the input feature can be weighted, which increases the network's attention to the drone feature information, thereby reducing the impact of complex background on drone recognition to a certain extent; through the SCFM module, the interference information of the feature map can be effectively filtered out from both the channel and space aspects, and the feature information of small targets can be enhanced. Therefore, the embodiments of the present application can enhance the recognition ability of small targets and the robustness in dense backgrounds.

[0005] According to one aspect of the present application, a method for detecting a drone target is provided, comprising: Get the dataset; Design and improve the YOLOv7-tiny model; The improved YOLOv7-tiny model is trained using the data set to obtain a trained improved YOLOv7-tiny model, wherein the improved YOLOv7-tiny model includes an input layer, a backbone network layer, a neck network layer and a head network layer, the backbone network layer includes a C-ELAN module, the C-ELAN module is generated based on a convolutional block attention unit and an ELAN submodule, the neck network layer includes an SCFM module and the C-ELAN module, and the SCFM module includes a spatial filtering submodule and a channel filtering submodule; Input the drone image to be detected into the input layer of the trained improved YOLOv7-tiny model for image preprocessing, and output the preprocessed image; Input the preprocessed image into the backbone network layer, perform feature extraction based on the C-ELAN module, and output a feature-extracted image; Input the feature extraction image into the neck network layer, perform feature fusion based on the SCFM module and the C-ELAN module, and output a feature fusion image; The feature fusion image is input into the head network layer for detection, and a target detection image is output.

[0006] According to another aspect of the present application, a drone target detection device is provided, the device comprising: The acquisition module is used to obtain the data set; Design module for designing and improving the YOLOv7-tiny model; A training module, used to train the improved YOLOv7-tiny model using the data set to obtain a trained improved YOLOv7-tiny model, wherein the improved YOLOv7-tiny model includes an input layer, a backbone network layer, a neck network layer and a head network layer, the backbone network layer includes a C-ELAN module, the C-ELAN module is generated based on a convolutional block attention unit and an ELAN submodule, the neck network layer includes an SCFM module and the C-ELAN module, and the SCFM module includes a spatial filtering submodule and a channel filtering submodule; The detection module is used to input the drone image to be detected into the input layer of the trained improved YOLOv7-tiny model for image preprocessing, and output the preprocessed image; input the preprocessed image into the backbone network layer, perform feature extraction based on the C-ELAN module, and output a feature extracted image; input the feature extracted image into the neck network layer, perform feature fusion based on the SCFM module and the C-ELAN module, and output a feature fused image; input the feature fused image into the head network layer for detection, and output a target detection image.

[0007] According to another aspect of the present application, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned drone target detection method is implemented.

[0008] According to another aspect of the present application, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor implements the above-mentioned drone target detection method when executing the program.

[0009] By means of the above technical scheme, the present application provides a method and device for detecting unmanned aerial vehicle targets, a storage medium, and a computer device. First, a data set can be obtained. Then, an improved YOLOv7-tiny model can be designed. Among them, the improved YOLOv7-tiny model may include an input layer, a backbone network layer, a neck network layer, and a head network layer. Specifically, the backbone network layer includes a C-ELAN module formed by combining a convolutional block attention unit and an ELAN submodule, and the neck network layer includes an SCFM module and the above-mentioned C-ELAN module. Afterwards, the obtained data set can be used to train the improved YOLOv7-tiny model to obtain a trained improved YOLOv7-tiny model. In the model application stage, the image of the unmanned aerial vehicle to be detected can be obtained, and the image of the unmanned aerial vehicle to be detected can be input into the trained improved YOLOv7-tiny model. After inputting into the trained improved YOLOv7-tiny model, first, the input layer of the model can perform data preprocessing on the drone image to be detected, and then the backbone network layer of the model can use multiple C-ELAN modules therein to extract image features, thereby obtaining a feature extraction image. Further, the neck network layer of the model can use SCFM modules, C-ELAN modules, etc. to fuse image features, thereby obtaining a feature fusion image. Finally, the head network layer of the model performs detection based on the received feature fusion image and outputs a target detection image. The embodiment of the present application can allocate weights to the spatial dimension and channel dimension of the input feature by improving the convolutional block attention unit in the YOLOv7-tiny model, thereby improving the network's attention to the drone feature information, thereby reducing the impact of complex background on drone recognition to a certain extent; through the SCFM module, the interference information of the feature map can be effectively filtered out from both the channel and space aspects, and the feature information of small targets can be enhanced. Therefore, the embodiment of the present application can enhance the recognition ability of small targets and the robustness in dense backgrounds.

[0010] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings: Figure 1 A schematic diagram of a process flow of a drone target detection method provided in an embodiment of the present application is shown; Figure 2 A schematic diagram of the structure of an improved YOLOv7-tiny model provided in an embodiment of the present application is shown; Figure 3 A schematic diagram of the structure of a CBL provided in an embodiment of the present application is shown; Figure 4 A schematic diagram of the structure of a C-ELAN module provided in an embodiment of the present application is shown; Figure 5 A schematic diagram of the structure of an SPP module provided in an embodiment of the present application is shown; Figure 6 A schematic diagram of the structure of a SCFM module provided in an embodiment of the present application is shown; Figure 7 A schematic diagram of the structure of a channel filtering submodule provided in an embodiment of the present application is shown; Figure 8 A schematic diagram of the structure of a drone target detection device provided in an embodiment of the present application is shown; Fig. 9 A schematic diagram of the device structure of a computer device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0012] The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other without conflict.

[0013] In drone target detection applications, the YOLOv7-tiny algorithm faces specific challenges, such as insufficient detection accuracy for small targets and degraded performance in complex backgrounds. Drones often fly in complex environments, and the targets to be detected are often small or highly occluded. Although YOLOv7-tiny has efficient real-time processing capabilities, its standard configuration may not perform well in these scenarios. Therefore, it is particularly important to improve the YOLOv7-tiny algorithm to enhance its recognition of small targets and robustness in dense backgrounds.

[0014] In this embodiment, a method for detecting a drone target is provided. Figure 1 As shown, the method includes: Step 101, obtaining a data set.

[0015] Step 102, design an improved YOLOv7-tiny model.

[0016] Step 103, using the data set to train the improved YOLOv7-tiny model to obtain a trained improved YOLOv7-tiny model, wherein the improved YOLOv7-tiny model includes an input layer, a backbone network layer, a neck network layer and a head network layer, the backbone network layer includes a C-ELAN module, the C-ELAN module is generated based on a convolutional block attention unit and an ELAN submodule, the neck network layer includes an SCFM module and the C-ELAN module, and the SCFM module includes a spatial filtering submodule and a channel filtering submodule.

[0017] The embodiment of the present application provides a method for detecting drone targets, which is implemented based on an improved YOLOv7-tiny model. First, a data set can be obtained. Here, the training samples and test samples in the data set can be acquired based on an image acquisition device, and can specifically include drone images of different scenes (cities, mountains, seas) and different lighting conditions (daytime, dusk, cloudy days) that have been annotated. For example, drone images that have been annotated with bounding boxes and have annotation information such as target categories and coordinates.

[0018] In one embodiment, the data set can also be expanded in the following ways to increase the diversity of training samples and test samples. The data set expansion method may include: (1) Geometric transformation. Geometric transformation may include rotation transformation, translation and scaling, mirror flipping, etc. Rotation transformation: Rotate at any angle based on the center of the image to increase the diversity of the target angle. Translation and scaling: Simulate the distance change of the drone, for example, by random cropping (such as randomly cutting out an 800x800 area from a 1024x1024 image) to increase the diversity of the target position. Mirror flipping: Use horizontal mirroring to increase the diversity of shooting directions. (2) Color enhancement. Color enhancement may include brightness adjustment, noise injection, dynamic combination strategy, etc. Brightness adjustment: Increase the diversity of light intensity changes in different time periods (such as the difference in light between dawn and dusk) through Gamma transformation or histogram equalization. Noise injection: Add Gaussian noise to simulate sensor interference to improve the robustness of the model in haze or electromagnetic interference environments. Dynamic combination strategy: Using online enhancement, 3-5 transformation combinations are randomly applied to each input image, such as "rotate 15 degrees + increase brightness by 20% + add salt and pepper noise", to ensure that the model continues to learn the characteristics of different scenes.

[0019] Next, an improved YOLOv7-tiny model can be designed. The improved YOLOv7-tiny model can include an input layer, a backbone network layer, a neck network layer, and a head network layer. Specifically, the input layer can receive a 3-channel RGB image and preprocess the RGB image; the backbone network layer is a feature extraction layer. In order to solve the problem that the background of the constructed drone dataset is complex and interferes with drone identification, the convolutional block attention unit (CBAM, Convolutional BlockAttention Module) and the ELAN submodule in the backbone network layer of the original YOLOv7-Tiny model are combined to form a C-ELAN module. The neck network layer is a feature fusion layer, and the neck network layer includes an SCFM module and the above-mentioned C-ELAN module. The SCFM module includes a spatial filter submodule and a channel filter submodule. This structure can effectively filter out the interference information of the feature map from both the channel and space aspects, and enhance the feature information of small targets.

[0020] Afterwards, the acquired data set can be used to train the improved YOLOv7-tiny model to obtain the trained improved YOLOv7-tiny model.

[0021] Step 104, input the drone image to be detected into the input layer of the trained improved YOLOv7-tiny model for image preprocessing, and output the preprocessed image.

[0022] Step 105: input the preprocessed image into the backbone network layer, perform feature extraction based on the C-ELAN module, and output a feature-extracted image.

[0023] Step 106: input the feature extraction image into the neck network layer, perform feature fusion based on the SCFM module and the C-ELAN module, and output a feature fusion image.

[0024] Step 107: input the feature fusion image into the head network layer for detection, and output a target detection image.

[0025] In this embodiment, in the model application stage, the drone image to be detected can be obtained, and the drone image to be detected can be input into the trained improved YOLOv7-tiny model. After inputting into the trained improved YOLOv7-tiny model, first, the input layer of the model can perform data preprocessing on the drone image to be detected, for example, standardizing the input size of the drone image to be detected, data enhancement, etc., to obtain the preprocessed image. Subsequently, the backbone network layer of the model can extract image features using multiple C-ELAN modules, etc., to obtain a feature extraction image. Further, the neck network layer of the model can fuse image features using SCFM modules, C-ELAN modules, etc., to obtain a feature fusion image. Finally, the head network layer of the model performs detection based on the received feature fusion image and outputs a target detection image. Among them, the target detection image can include information such as bounding box coordinates, target confidence, and category probability.

[0026] By applying the technical solution of this embodiment, first, a data set can be obtained. Then, an improved YOLOv7-tiny model can be designed. Among them, the improved YOLOv7-tiny model may include an input layer, a backbone network layer, a neck network layer and a head network layer. Specifically, the backbone network layer includes a C-ELAN module formed by combining a convolutional block attention unit and an ELAN submodule, and the neck network layer includes an SCFM module and the above-mentioned C-ELAN module. Afterwards, the improved YOLOv7-tiny model can be trained using the acquired data set to obtain a trained improved YOLOv7-tiny model. In the model application stage, the drone image to be detected can be obtained, and the drone image to be detected can be input into the trained improved YOLOv7-tiny model. After inputting into the trained improved YOLOv7-tiny model, first, the input layer of the model can perform data preprocessing on the drone image to be detected, and then the backbone network layer of the model can extract image features using multiple C-ELAN modules, etc., to obtain a feature extraction image. Furthermore, the neck network layer of the model can use the SCFM module, C-ELAN module, etc. to fuse image features, and then obtain a feature fusion image. Finally, the head network layer of the model performs detection based on the received feature fusion image and outputs a target detection image. The embodiment of the present application improves the convolutional block attention unit in the YOLOv7-tiny model to assign weights to the spatial dimension and channel dimension of the input feature, thereby increasing the network's attention to the feature information of the drone, thereby reducing the impact of complex backgrounds on drone recognition to a certain extent; through the SCFM module, the interference information of the feature map can be effectively filtered out from both the channel and space aspects, and the feature information of small targets can be enhanced. Therefore, the embodiment of the present application can enhance the recognition ability of small targets and the robustness in dense backgrounds.

[0027] In the embodiment of the present application, optionally, step 105 specifically includes: inputting the preprocessed image into the first CBL module, the second CBL module, the first C-ELAN module, the first MP module, and the second C-ELAN module in sequence, and outputting a first feature extraction image; inputting the first feature extraction image into the second MP module and the third C-ELAN module in sequence, and outputting a second feature extraction image; inputting the second feature extraction image into the third MP module, the fourth C-ELAN module, the third CBL module, the SPP module, and the fourth CBL module in sequence, and outputting a third feature extraction image.

[0028] In this embodiment, if Figure 2As shown, the backbone network layer may include a first CBL module, a second CBL module, a first C-ELAN module, a first MP module, a second C-ELAN module, a second MP module, a third C-ELAN module, a third MP module, a fourth C-ELAN module, a third CBL module, an SPP module, and a fourth CBL module. Among them, the structures of the first CBL module, the second CBL module, the third CBL module, and the fourth CBL module may be the same; the structures of the first C-ELAN module, the second C-ELAN module, the third C-ELAN module, and the fourth C-ELAN module may be the same; and the structures of the first MP module, the second MP module, and the third MP module may be the same.

[0029] In a specific embodiment, Figure 3 As shown in the figure, each CBL module can include Conv (Convolutional Layer), BN (Batch Normalization) and LeakyReLU activation function. Conv is a convolutional layer used to extract local features; BN is a batch normalization layer used to normalize the output of the convolutional layer; LeakyReLU activation function is used to introduce nonlinear factors to enhance the expression ability of the model. The convolution kernel of Conv can be determined according to actual needs.

[0030] In an embodiment of the present application, optionally, the ELAN submodule includes a splicing unit and a plurality of CBL units, and the C-ELAN module is obtained by combining the ELAN submodule with the convolutional block attention unit; after the C-ELAN module receives the first input image, the method further includes: inputting the first input image into the first CBL unit and the second CBL unit in the C-ELAN module respectively, outputting a first convolution result corresponding to the first CBL unit, and a second convolution result corresponding to the second CBL unit; inputting the first convolution result into the convolutional block attention unit and the third CBL unit in the C-ELAN module in sequence, and outputting a third convolution result; inputting the third convolution result into the fourth CBL unit in the C-ELAN module, and outputting a fourth convolution result; splicing the first convolution result, the second convolution result, the third convolution result and the fourth convolution result to generate a target splicing result; inputting the target splicing result into the fifth CBL unit in the C-ELAN module to obtain the output result corresponding to the C-ELAN module.

[0031] like Figure 4As shown, the C-ELAN module may include multiple CBL units, convolutional block attention units and splicing units. Among them, the convolutional block attention unit increases the network's attention to the drone feature information by allocating weights to the spatial dimension and channel dimension of the input features, thereby reducing the impact of complex background on drone recognition to a certain extent. The structure of the CBL unit may be the same as that of the CBL module. "Unit" and "module" are only used to distinguish the level to which the CBL belongs. The splicing unit may specifically be a CONCAT unit. It should be noted that no matter which C-ELAN module's input is, it can be uniformly referred to as the first input image. In the C-ELAN module, the first CBL unit and the second CBL unit are two parallel CBL units. In the C-ELAN module, the first convolution result is processed by the convolution block attention unit. On the one hand, the channel importance weight can be learned through the channel attention in the convolution block attention unit, and on the other hand, the spatial importance weight can be learned through the spatial attention. This can not only enhance the contribution of important feature channels, but also highlight the spatial position of the target area. While improving the model's sensitivity to key features, it can suppress background noise interference, thereby maintaining a high inference speed while effectively improving the model's ability to detect multi-scale targets.

[0032] In a specific embodiment, the MP module may be a maximum pooling layer.

[0033] In a specific embodiment, Figure 5 As shown, the SPP (Spatial Pyramid Pooling) module may include a first CBL submodule, three maximum pooling submodules, a first splicing submodule, a second CBL submodule, and a second splicing submodule. Among them, the structures of the first CBL submodule and the second CBL submodule may be the same, and the size of the convolution kernel may be determined according to the requirements; the structures of the first splicing submodule and the second splicing submodule may be the same. It should be noted that the structure of the CBL submodule here may be the same as that of the aforementioned CBL module, and "submodule" and "module" are only used to distinguish the level to which the CBL belongs. The maximum pooling submodule here may be the Maxpool layer, and the splicing submodule may be CONCAT. SPP completes feature fusion by fusing feature map information of different scales through convolution kernel pooling operations of different sizes.

[0034] In the embodiment of the present application, optionally, step 106 specifically includes: inputting the first feature extraction image into the first SCFM module and the first Conv module in sequence to generate a first image; splicing the first image with the upsampled image corresponding to the first feature extraction image to obtain a first spliced ​​image; inputting the first spliced ​​image into the fifth C-ELAN module to output a first feature fusion image; inputting the second feature extraction image into the second SCFM module and the second Conv module in sequence to generate a second image; splicing the second image with the upsampled image corresponding to the third feature extraction image to obtain a second spliced ​​image; inputting the second spliced ​​image into the sixth C-ELAN module to generate a third image; inputting the first feature fusion image into the fifth CBL module to generate a fourth image; splicing the third image and the fourth image to obtain a third spliced ​​image; inputting the third spliced ​​image into the seventh C-ELAN module to output a second feature fusion image; inputting the second feature fusion image into the sixth CBL module to generate a fifth image; splicing the third feature extraction image and the fifth image to obtain a fourth spliced ​​image; inputting the fourth spliced ​​image into the eighth C-ELAN module to output a third feature fusion image.

[0035] In this embodiment, if Figure 2 As shown, the neck network layer is the feature fusion stage, and the neck network layer includes multiple SCFM modules, multiple Conv modules, multiple CONCAT modules, multiple C-ELAN modules, multiple CBL modules and multiple upsampling modules. The neck network layer adopts Path Aggregation Network (PANet) as the feature fusion part of the network. By adding bottom-up path enhancement on the basis of Feature Pyramid Network (FPN), the entire feature hierarchy is enhanced using accurate low-level positioning signals, thereby shortening the information path between low-level and top-level features. The embodiment of the present application introduces a spatial filtering submodule and a channel filtering submodule in the SCFM module. This structure can effectively filter out interference information of the feature map from the channel and space, and enhance the feature information of small targets.

[0036] It should be noted that before upsampling the first feature extraction image and the third feature extraction image, they can also be processed by the CBL module before upsampling. The CBL module here has the same structure as the aforementioned CBL module, wherein the size of the convolution kernel can be determined according to the requirements; the splicing processing operation here is implemented by each CONCAT module.

[0037] In an embodiment of the present application, optionally, the SCFM module also includes a Conv submodule, a first product submodule, a second product submodule and a summation submodule; after the SCFM module receives the second input image, the method also includes: inputting the second input image into the Conv submodule and outputting a fifth convolution result; inputting the fifth convolution result into the spatial filtering submodule and the channel filtering submodule respectively, outputting a first filtering result corresponding to the spatial filtering submodule and a second filtering result corresponding to the channel filtering submodule; inputting the first filtering result and the fifth convolution result into the first product submodule together to obtain a first product result, and inputting the second filtering result and the fifth convolution result into the second product submodule together to obtain a second product result; inputting the first product result and the second product result into the summation submodule together to obtain the output result corresponding to the SCFM module.

[0038] In this embodiment, if Figure 6 As shown in the figure, the SCFM module includes a Conv submodule, a first product submodule, a second product submodule, a summation submodule, a spatial filter submodule and a channel filter submodule. First, the second input image is compressed in the channel direction by the Conv submodule, and then input into the spatial filter submodule and the channel filter submodule respectively. Among them, the spatial filter submodule is used to capture local spatial patterns (such as edges and textures) and enhance the feature expression in the spatial dimension; the channel filter submodule is used to learn the relationship between channels and enhance the feature expression in the channel dimension. Through these two submodules, the small target features can be effectively enhanced. Here, the spatial filter submodule processes the compressed feature map using the Log_Softmax activation function to filter the feature map in space to enhance the feature information. The first product submodule is used to fuse the spatial filter feature (that is, the first filter result) and the compression feature (that is, the fifth convolution result) to obtain the spatial enhancement feature (that is, the first product result); the second product submodule is used to fuse the channel filter feature (that is, the second filter result) and the compression feature to obtain the channel enhancement feature (that is, the second product result). The summation submodule is used to integrate the spatial enhancement feature and the channel enhancement feature. The embodiment of the present application enhances the features in the spatial and channel dimensions so that the improved YOLOv7-tiny model can more effectively capture the multi-scale features and complex contextual information of the target, and is particularly suitable for target detection tasks with large target scale variations and complex backgrounds.

[0039] It should be noted that no matter which SCFM module's input is, it can be uniformly referred to as the second input image. The structure of the Conv submodule can be the same as the structure of the aforementioned Conv module, and "submodule" and "module" are only used to distinguish the level to which Conv belongs. The first product module and the second product module can be Multipy layers.

[0040] In a specific embodiment, the spatial filtering submodule can capture local context information through a large convolution kernel (5x5), and the channel filtering submodule can learn the nonlinear relationship between channels through 1x1 convolution.

[0041] In an embodiment of the present application, optionally, the spatial filtering submodule includes a Log_Softmax activation function, and the channel filtering submodule includes an average pooling unit, a maximum pooling unit, multiple convolution units, a Hardswish activation function, multiple upsampling units and a summation unit; after the channel filtering submodule receives the third input image, the method also includes: inputting the third input image into the average pooling unit, the first convolution unit, the Hardswish activation function, the second convolution unit and the first upsampling unit in sequence, and outputting a first upsampling result, and inputting the third input image into the maximum pooling unit, the third convolution unit, the Hardswish activation function, the fourth convolution unit and the second upsampling unit in sequence, and outputting a second upsampling result; the first upsampling result and the second upsampling result are input into the summation unit together to obtain the output result corresponding to the channel filtering submodule.

[0042] In this embodiment, the spatial filtering submodule processes the compressed feature map (ie, the fifth convolution result) using the Log_Softmax activation function to filter the feature map spatially to enhance feature information.

[0043] The spatial filter submodule includes a Log_Softmax activation function, which performs a Log operation on the result based on the Softmax function to generate relative weights of all positions relative to the channel.

[0044] The Softmax function formula is: ; The Log_Softmax function formula is: ; Where: represents the Softmax function, Z represents the input vector, represents the jth element, represents the i-th element, e represents the exponential function, represents the summation function, represents the Log_Softmax function, ln Represents a logarithmic function.

[0045] From the algorithm formulas of the two, we can see that although Log_Softmax and Softmax are both monotonic, they have different effects on the relative values ​​of the loss function. Using the Log_Softmax function can make the algorithm converge faster.

[0046] like Figure 7 As shown, a channel filtering submodule provided by an embodiment of the present application is shown. The channel filtering submodule includes an average pooling unit, a maximum pooling unit, multiple convolution units, a Hardswish activation function, and multiple upsampling units. After the third input image is input into the channel filtering submodule, it passes through the branch corresponding to the average pooling unit and the branch corresponding to the maximum pooling unit respectively, and the combination of the results of the subsequent two branches will obtain more detailed global features; two convolution units and Hardswish activation functions are introduced in each branch to enhance the features of small targets, and finally the feature map size is restored through the nearest upsampling operation, and then the results of the two branches are added to obtain the final output of the channel filtering submodule.

[0047] Hardswish is introduced as the activation function here because it not only has the advantages of the Swish function, that is, it helps prevent the gradient from gradually approaching 0 and causing saturation during slow training, smoothness plays an important role in optimization and generalization, and the derivative is always greater than 0, but also uses a combination of common operators on the basis of the Swish function, which greatly reduces the computational complexity of the algorithm model while achieving similar effects.

[0048] It should be noted that the average pooling unit here can be an AvgPool layer, the maximum pooling unit can be a MaxPool layer, the first convolution unit, the second convolution unit, the third convolution unit and the fourth convolution unit can be Conv2d with the same structure, and the parameter settings can be determined according to actual needs. The first upsampling unit and the second upsampling unit can be Upsample layers.

[0049] In an embodiment of the present application, optionally, the head network layer includes a first head detection module, a second head detection module and a third head detection module, and the first head detection module, the second head detection module and the third head detection module are respectively used to detect targets of different scales; step 107 specifically includes: inputting the first feature fusion image into the first head detection module, and outputting a first target detection image; inputting the second feature fusion image into the second head detection module, and outputting a second target detection image; inputting the third feature fusion image into the third head detection module, and outputting a third target detection image.

[0050] In this embodiment, if Figure 2 As shown, the head network layer is the detection stage, and the head network layer may include a first head detection module, a second head detection module, and a third head detection module, each of which may be a YOLO HEAD module. Different head detection modules are used to detect targets of different scales. Among them, the first head detection module is used to process shallow high-resolution features, the second head detection module is used to process middle-resolution features, and the third head detection module is used to process deep low-resolution features. Using three head detection modules to output target detection images of three scales respectively, large, medium, and small objects can be detected.

[0051] In the embodiment of the present application, optionally, the positioning loss function used by the improved YOLOv7-tiny model is a Wise-IoU loss function, and the Wise-IoU loss function is specifically: ; ; ; In the formula, represents the Wise-IoU loss function, represents the first hyperparameter, represents the second hyperparameter, represents the third hyperparameter, represents the penalty term, represents the IoU loss function, exp represents the exponential function, Represents the coordinates of the center point of the prediction box, Represents the coordinates of the center point of the real box, Indicates the width of the minimum bounding box between the predicted box and the real box, Indicates the height of the minimum bounding box between the predicted box and the real box, Represents the intersection and union ratio.

[0052] In this embodiment, the loss function of the model is further optimized. The loss function in the traditional YOLOv7-tiny model consists of confidence loss, classification loss and positioning loss. The positioning loss uses the CIoU loss function, which has the disadvantage of not considering the mismatch between the predicted box and the real box, resulting in slow convergence. Therefore, the embodiment of the present application introduces Wise-IoU as a new positioning loss function. In the Wise-IoU loss function, It represents the intersection-over-union ratio, which is a common measure of the overlap between the predicted box and the true box. and Usually set to 1.9 and 3.0. When .

[0053] By introducing outliers To describe the quality of the anchor box, Negatively correlated with the quality of the anchor box. The outlier degree calculation formula is: ; In the formula, is the gradient gain of the monotonic focusing coefficient, and The definition is the same, here This means that changes are continuously calculated during training based on the detection of each target; is the sliding average of momentum m, and introduce This means that the maximum gradient gain can be dynamically adjusted according to the training process. The calculation formula for momentum m is: ; Where: t is the value of epoch, n is the value of batch size. The significance of introducing momentum m is that after t rounds of training, WIoU assigns small gradient gains to low-quality anchor boxes to reduce harmful gradients.

[0054] In an embodiment of the present application, optionally, step 104 specifically includes: inputting the drone image to be detected into the input layer of the trained improved YOLOv7-tiny model, performing data expansion and adaptive anchor frame calculation on the drone image to be detected through the input layer, and outputting the preprocessed image.

[0055] In this embodiment, the input layer of the improved YOLOv7-tiny model can preprocess the drone image to be detected, and the preprocessing may include data padding, adaptive anchor frame calculation, etc. to ensure uniform scaling of the RGB image, thereby meeting the input size requirements of the backbone network layer. Among them, data augmentation is a commonly used technique in machine learning, especially in the field of deep learning, which is used to improve the generalization ability of the model by increasing the diversity of training data. In the model application stage, data augmentation (such as brightness fine-tuning, small angle rotation) can simulate some real disturbances and improve the model's adaptability to domain shift. Because in the dynamic scene of the drone, sudden changes in illumination, motion blur, and changes in shooting angles are normal. For example: illumination changes: when the morning and evening alternate, the imaging of the same target in strong light and weak light is significantly different. Motion blur: When flying at high speed, propellers or small parts may be blurred due to jitter. Therefore, introducing data augmentation in the application stage can also help improve the accuracy of subsequent target recognition.

[0056] Adaptive Anchor Calculation refers to the automatic calculation and adjustment of the shape, size and number of anchor frames according to the characteristics of the dataset and the requirements of the algorithm in the target detection task, so as to improve the accuracy and efficiency of target detection.

[0057] When preprocessing the drone image to be detected, the embodiment of the present application realizes the coordinated optimization of data-level enhancement and model architecture through a preprocessing method that combines data expansion and adaptive anchor frame calculation, so that the improved YOLOv7-tiny can show stronger adaptability to multi-scale targets, complex lighting and shooting angle changes in scenarios such as drone inspections, and can improve the detection accuracy of small targets.

[0058] Further, as Figure 1 The specific implementation of the method, the embodiment of the present application provides a drone target detection device, such as Figure 8 As shown, the device comprises: The acquisition module is used to obtain the data set; Design module for designing and improving the YOLOv7-tiny model; A training module, used to train the improved YOLOv7-tiny model using the data set to obtain a trained improved YOLOv7-tiny model, wherein the improved YOLOv7-tiny model includes an input layer, a backbone network layer, a neck network layer and a head network layer, the backbone network layer includes a C-ELAN module, the C-ELAN module is generated based on a convolutional block attention unit and an ELAN submodule, the neck network layer includes an SCFM module and the C-ELAN module, and the SCFM module includes a spatial filtering submodule and a channel filtering submodule; The detection module is used to input the drone image to be detected into the input layer of the trained improved YOLOv7-tiny model for image preprocessing, and output the preprocessed image; input the preprocessed image into the backbone network layer, perform feature extraction based on the C-ELAN module, and output a feature extracted image; input the feature extracted image into the neck network layer, perform feature fusion based on the SCFM module and the C-ELAN module, and output a feature fused image; input the feature fused image into the head network layer for detection, and output a target detection image.

[0059] Optionally, the detection module is specifically used to: The preprocessed image is sequentially input into a first CBL module, a second CBL module, a first C-ELAN module, a first MP module, and a second C-ELAN module, and a first feature extraction image is output; Inputting the first feature extraction image into a second MP module and a third C-ELAN module in sequence, and outputting a second feature extraction image; The second feature extraction image is sequentially input into the third MP module, the fourth C-ELAN module, the third CBL module, the SPP module and the fourth CBL module, and a third feature extraction image is output.

[0060] Optionally, the detection module is specifically used to: Inputting the first feature extraction image into a first SCFM module and a first Conv module in sequence to generate a first image; splicing the first image with an upsampled image corresponding to the first feature extraction image to obtain a first spliced ​​image; Input the first stitched image into a fifth C-ELAN module, and output a first feature fused image; Inputting the second feature extraction image into a second SCFM module and a second Conv module in sequence to generate a second image; splicing the second image with an upsampled image corresponding to the third feature extraction image to obtain a second spliced ​​image; Inputting the second stitched image into a sixth C-ELAN module to generate a third image; Inputting the first feature fusion image into a fifth CBL module to generate a fourth image; splicing the third image and the fourth image to obtain a third spliced ​​image; Input the third stitched image into a seventh C-ELAN module, and output a second feature fused image; Inputting the second feature fusion image into a sixth CBL module to generate a fifth image; splicing the third feature extraction image and the fifth image to obtain a fourth spliced ​​image; The fourth stitched image is input into the eighth C-ELAN module, and a third feature fused image is output.

[0061] Optionally, the head network layer includes a first head detection module, a second head detection module and a third head detection module, wherein the first head detection module, the second head detection module and the third head detection module are respectively used to detect objects of different scales; the detection module is specifically used to: Inputting the first feature fusion image into a first head detection module, and outputting a first target detection image; Inputting the second feature fusion image into a second head detection module, and outputting a second target detection image; The third feature fusion image is input into a third head detection module, and a third target detection image is output.

[0062] Optionally, the ELAN submodule includes a splicing unit and a plurality of CBL units, and the C-ELAN module is obtained by combining the ELAN submodule with the convolutional block attention unit; and the detection module is further used for: After receiving the first input image, the C-ELAN module inputs the first input image into a first CBL unit and a second CBL unit in the C-ELAN module respectively, and outputs a first convolution result corresponding to the first CBL unit and a second convolution result corresponding to the second CBL unit; Inputting the first convolution result into the convolution block attention unit and the third CBL unit in the C-ELAN module in sequence, and outputting the third convolution result; Input the third convolution result into the fourth CBL unit in the C-ELAN module, and output the fourth convolution result; Splicing the first convolution result, the second convolution result, the third convolution result, and the fourth convolution result to generate a target splicing result; The target splicing result is input into the fifth CBL unit in the C-ELAN module to obtain the output result corresponding to the C-ELAN module.

[0063] Optionally, the SCFM module further includes a Conv submodule, a first product submodule, a second product submodule and a summation submodule; the detection module is further used for: After receiving the second input image, the SCFM module inputs the second input image into the Conv submodule and outputs a fifth convolution result; Input the fifth convolution result into the spatial filtering submodule and the channel filtering submodule respectively, and output a first filtering result corresponding to the spatial filtering submodule and a second filtering result corresponding to the channel filtering submodule; Inputting the first filtering result and the fifth convolution result into the first product submodule to obtain a first product result, and inputting the second filtering result and the fifth convolution result into the second product submodule to obtain a second product result; The first multiplication result and the second multiplication result are input into the summing submodule to obtain an output result corresponding to the SCFM module.

[0064] Optionally, the spatial filtering submodule includes a Log_Softmax activation function, the channel filtering submodule includes an average pooling unit, a maximum pooling unit, multiple convolution units, a Hardswish activation function, multiple upsampling units and a summation unit; the detection module is further used to: After receiving the third input image, the channel filtering submodule sequentially inputs the third input image into the average pooling unit, the first convolution unit, the Hardswish activation function, the second convolution unit and the first upsampling unit, and outputs the first upsampling result; and sequentially inputs the third input image into the maximum pooling unit, the third convolution unit, the Hardswish activation function, the fourth convolution unit and the second upsampling unit, and outputs the second upsampling result; The first up-sampling result and the second up-sampling result are input into the summing unit to obtain an output result corresponding to the channel filtering submodule.

[0065] Optionally, the positioning loss function used by the improved YOLOv7-tiny model is a Wise-IoU loss function, and the Wise-IoU loss function is specifically: ; ; ; In the formula, represents the Wise-IoU loss function, represents the first hyperparameter, represents the second hyperparameter, represents the third hyperparameter, represents the penalty term, represents the IoU loss function, exp represents the exponential function, Represents the coordinates of the center point of the prediction box, Represents the coordinates of the center point of the real box, Indicates the width of the minimum bounding box between the predicted box and the real box, Indicates the height of the minimum bounding box between the predicted box and the real box, Represents the intersection and union ratio.

[0066] Optionally, the detection module is specifically used to: The drone image to be detected is input into the input layer of the trained improved YOLOv7-tiny model, and data expansion and adaptive anchor frame calculation are performed on the drone image to be detected through the input layer, and the preprocessed image is output.

[0067] It should be noted that for other corresponding descriptions of the functional units involved in the drone target detection device provided in the embodiment of the present application, reference can be made to Figures 1 to 7 The corresponding description in the method will not be repeated here.

[0068] The present application also provides a computer device, which may be a personal computer, a server, a network device, etc. Fig. 9 As shown, the computer device includes a bus, a processor, a memory and a communication interface, and may also include an input and output interface and a display device. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store location information. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the steps in each method embodiment are implemented.

[0069] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0070] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium may be non-volatile or volatile, and stores a computer program thereon. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0071] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0072] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0073] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0074] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0075] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for detecting a drone target, characterized in that: include: Get the dataset; Design and improve the YOLOv7-tiny model; The improved YOLOv7-tiny model is trained using the data set to obtain a trained improved YOLOv7-tiny model, wherein the improved YOLOv7-tiny model includes an input layer, a backbone network layer, a neck network layer and a head network layer, the backbone network layer includes a C-ELAN module, the C-ELAN module is generated based on a convolutional block attention unit and an ELAN submodule, the neck network layer includes an SCFM module and the C-ELAN module, and the SCFM module includes a spatial filtering submodule and a channel filtering submodule; Input the drone image to be detected into the input layer of the trained improved YOLOv7-tiny model for image preprocessing, and output the preprocessed image; Input the preprocessed image into the backbone network layer, perform feature extraction based on the C-ELAN module, and output a feature-extracted image; Input the feature extraction image into the neck network layer, perform feature fusion based on the SCFM module and the C-ELAN module, and output a feature fusion image; The feature fusion image is input into the head network layer for detection, and a target detection image is output.

2. The drone target detection method according to claim 1, characterized in that: The step of inputting the preprocessed image into the backbone network layer, performing feature extraction based on the C-ELAN module, and outputting a feature-extracted image specifically includes: The preprocessed image is sequentially input into a first CBL module, a second CBL module, a first C-ELAN module, a first MP module, and a second C-ELAN module, and a first feature extraction image is output; Inputting the first feature extraction image into a second MP module and a third C-ELAN module in sequence, and outputting a second feature extraction image; The second feature extraction image is sequentially input into the third MP module, the fourth C-ELAN module, the third CBL module, the SPP module and the fourth CBL module, and a third feature extraction image is output.

3. The method for detecting a drone target according to claim 2, characterized in that: The step of inputting the feature extraction image into the neck network layer, performing feature fusion based on the SCFM module and the C-ELAN module, and outputting a feature fusion image specifically includes: Inputting the first feature extraction image into a first SCFM module and a first Conv module in sequence to generate a first image; splicing the first image with an upsampled image corresponding to the first feature extraction image to obtain a first spliced ​​image; Input the first stitched image into a fifth C-ELAN module, and output a first feature fused image; Inputting the second feature extraction image into a second SCFM module and a second Conv module in sequence to generate a second image; splicing the second image with an upsampled image corresponding to the third feature extraction image to obtain a second spliced ​​image; Inputting the second stitched image into a sixth C-ELAN module to generate a third image; Inputting the first feature fusion image into a fifth CBL module to generate a fourth image; splicing the third image and the fourth image to obtain a third spliced ​​image; Input the third stitched image into a seventh C-ELAN module, and output a second feature fused image; Inputting the second feature fusion image into a sixth CBL module to generate a fifth image; splicing the third feature extraction image and the fifth image to obtain a fourth spliced ​​image; The fourth stitched image is input into the eighth C-ELAN module, and a third feature fused image is output.

4. The method for detecting a drone target according to claim 3, characterized in that: The head network layer includes a first head detection module, a second head detection module and a third head detection module, wherein the first head detection module, the second head detection module and the third head detection module are respectively used to detect targets of different scales; the inputting the feature fusion image into the head network layer for detection and outputting the target detection image specifically includes: Inputting the first feature fusion image into a first head detection module, and outputting a first target detection image; Inputting the second feature fusion image into a second head detection module, and outputting a second target detection image; The third feature fusion image is input into a third head detection module, and a third target detection image is output.

5. The drone target detection method according to claim 1, characterized in that: The ELAN submodule includes a splicing unit and a plurality of CBL units, and the C-ELAN module is obtained by combining the ELAN submodule with the convolutional block attention unit; After the C-ELAN module receives the first input image, the method further includes: Inputting the first input image into a first CBL unit and a second CBL unit in the C-ELAN module respectively, and outputting a first convolution result corresponding to the first CBL unit and a second convolution result corresponding to the second CBL unit; Inputting the first convolution result into the convolution block attention unit and the third CBL unit in the C-ELAN module in sequence, and outputting the third convolution result; Input the third convolution result into the fourth CBL unit in the C-ELAN module, and output the fourth convolution result; Splicing the first convolution result, the second convolution result, the third convolution result, and the fourth convolution result to generate a target splicing result; The target splicing result is input into the fifth CBL unit in the C-ELAN module to obtain the output result corresponding to the C-ELAN module.

6. The drone target detection method according to claim 1, characterized in that: The SCFM module also includes a Conv submodule, a first product submodule, a second product submodule and a summation submodule; After the SCFM module receives the second input image, the method further includes: Input the second input image into the Conv submodule, and output a fifth convolution result; Input the fifth convolution result into the spatial filtering submodule and the channel filtering submodule respectively, and output a first filtering result corresponding to the spatial filtering submodule and a second filtering result corresponding to the channel filtering submodule; Inputting the first filtering result and the fifth convolution result into the first product submodule to obtain a first product result, and inputting the second filtering result and the fifth convolution result into the second product submodule to obtain a second product result; The first multiplication result and the second multiplication result are input into the summing submodule to obtain an output result corresponding to the SCFM module.

7. The method for detecting a drone target according to claim 6, characterized in that: The spatial filtering submodule includes a Log_Softmax activation function, and the channel filtering submodule includes an average pooling unit, a maximum pooling unit, a plurality of convolution units, a Hardswish activation function, a plurality of upsampling units, and a summing unit; After the channel filtering submodule receives the third input image, the method further includes: Inputting the third input image sequentially into an average pooling unit, a first convolution unit, a Hardswish activation function, a second convolution unit, and a first upsampling unit, and outputting a first upsampling result; and inputting the third input image sequentially into a maximum pooling unit, a third convolution unit, a Hardswish activation function, a fourth convolution unit, and a second upsampling unit, and outputting a second upsampling result; The first up-sampling result and the second up-sampling result are input into the summing unit to obtain an output result corresponding to the channel filtering submodule.

8. The method for detecting a drone target according to claim 1, characterized in that: The positioning loss function used by the improved YOLOv7-tiny model is the Wise-IoU loss function, and the Wise-IoU loss function is specifically: ; ; ; In the formula, represents the Wise-IoU loss function, represents the first hyperparameter, represents the second hyperparameter, represents the third hyperparameter, represents the penalty term, represents the IoU loss function, exp represents the exponential function, Represents the coordinates of the center point of the prediction box, Represents the coordinates of the center point of the real frame, Indicates the width of the minimum bounding box between the predicted box and the real box, Indicates the height of the minimum bounding box between the predicted box and the real box, Represents intersection and union ratio.

9. The method for detecting a drone target according to claim 1, characterized in that: The method of inputting the image of the drone to be detected into the input layer of the trained improved YOLOv7-tiny model for image preprocessing and outputting the preprocessed image specifically includes: The drone image to be detected is input into the input layer of the trained improved YOLOv7-tiny model, and data expansion and adaptive anchor frame calculation are performed on the drone image to be detected through the input layer, and the preprocessed image is output.

10. A drone target detection device, characterized in that: The device comprises: Acquisition module, used to acquire data sets; Design module for designing and improving the YOLOv7-tiny model; A training module, used to train the improved YOLOv7-tiny model using the data set to obtain a trained improved YOLOv7-tiny model, wherein the improved YOLOv7-tiny model includes an input layer, a backbone network layer, a neck network layer and a head network layer, the backbone network layer includes a C-ELAN module, the C-ELAN module is generated based on a convolutional block attention unit and an ELAN submodule, the neck network layer includes an SCFM module and the C-ELAN module, and the SCFM module includes a spatial filtering submodule and a channel filtering submodule; The detection module is used to input the drone image to be detected into the input layer of the trained improved YOLOv7-tiny model for image preprocessing, and output the preprocessed image; input the preprocessed image into the backbone network layer, perform feature extraction based on the C-ELAN module, and output a feature extracted image; input the feature extracted image into the neck network layer, perform feature fusion based on the SCFM module and the C-ELAN module, and output a feature fused image; input the feature fused image into the head network layer for detection, and output a target detection image.

Citation Information

Patent Citations

  • Remote sensing satellite image target detection method based on deep learning

    CN117456376A

  • Remote sensing target detection method and device, electronic equipment and storage medium

    CN117671509A

  • Industrial product surface defect detection method

    CN118411339A

  • Multi-camera vision system facilitating authentication and secure data transfer

    US20240111897A1