Unmanned aerial vehicle image target detection method and device, electronic equipment and readable storage medium
By replacing the module in the backbone network of the drone image object detection model and combining the attention mechanism and small-objective dense area calculation module, the problem of low detection accuracy in drone aerial photography scenarios is solved, and higher detection accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510214237.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-10
AI Technical Summary
The existing drone image target detection methods have low detection accuracy in drone aerial photography scenarios, especially for small targets, dense targets and multi-scale targets, which affects the accuracy and robustness of the detection.
By replacing the C2PSA module as the main branch and side branch module in the backbone network of the drone image object detection model, combining the spatial attention mechanism and channel attention mechanism, basic features are extracted, and the small target dense area calculation module YOLOv11 model is connected at both ends of the C3K2 module to enhance the feature information of dense areas.
It reduces the computational complexity of the model, reduces information loss, significantly improves the generalization ability of the model and the ability to detect drone image targets, and improves the detection accuracy.
Smart Images

Figure CN120126035A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of drone image processing, and in particular to a drone image target detection method, device, electronic device and readable storage medium. Background Art
[0002] With the rapid development of computer technology and drone aerial photography, the number and application scope of drone aerial images are constantly expanding. Object detection in drone aerial images is widely used in the fields of intelligent traffic management, infrastructure, inspection and maintenance, disaster prevention, search and rescue, crop management and analysis, ecological protection and monitoring. Drone aerial photography plays a key role in object detection, which aims to detect, accurately locate and classify various objects of interest, such as pedestrians, bicycles, trucks and tricycles, from complex aerial images.
[0003] In recent years, the rapid development of deep learning technology has gradually shifted target detection technology from traditional manual feature extraction methods to automated technology based on deep learning. The target detection algorithm based on deep learning has significantly improved the detection accuracy, processing speed and stability compared with traditional methods. According to the complexity of algorithm processing, deep learning target detection algorithms can be divided into two categories: two-stage detection algorithms and one-stage detection algorithms. Two-stage detection algorithms such as R-CNN, Fast R-CNN and Faster R-CNN have complex network structures and high computing resource requirements. These characteristics make them face many challenges in actual deployment and real-time detection. In contrast, single-stage detection algorithms such as FCOS, RetinaNet, SSD and YOLO series have significant advantages in detection accuracy, simplicity of network structure, detection speed and model deployment.
[0004] However, these methods face many problems in drone aerial photography scenarios. For example, as the network depth increases, the model may lose detailed information, resulting in low detection accuracy for small targets, dense targets, multi-scale targets, etc., affecting the accuracy and robustness of detection. Therefore, it is necessary to develop a detection algorithm that meets the detection accuracy requirements to meet the needs of practical applications.
[0005] In order to solve the above problems, it is necessary to develop a detection algorithm that meets the detection accuracy requirements to meet the needs of drone image target detection applications. Summary of the invention
[0006] In this embodiment, a method, device, electronic device and storage medium for detecting a target in an unmanned aerial vehicle image are provided to solve the problem that the detection accuracy in the related art needs to be improved.
[0007] In the first aspect, the present application provides a method for detecting targets in drone images, comprising:
[0008] Obtain the UAV image to be detected and the UAV image target detection model. The UAV image target detection model is a YOLOv11 model in which the C2PSA module in the backbone network is replaced by a main branch for extracting basic features and a side branch for fusing the spatial attention mechanism and the channel attention mechanism according to the basic features, and small target dense area calculation modules are connected to both ends of any C3K2 module in the backbone network;
[0009] Input the UAV image to be detected into the UAV image target detection model to obtain the UAV image target detection result.
[0010] In some embodiments, the main branch first processes the input features through 1×1 convolution to increase the channel dimension and fuse the information of each channel to generate a first feature map, then extracts multi-scale context information through dilated convolution, and then generates the basic features after passing through an activation function. Then, fuse the third feature map obtained by processing through the spatial attention mechanism in the side branch with the basic features to generate a fourth feature map. Then, process the fourth feature map through depthwise separable convolution to obtain local features. Then, fuse the fourth feature map with the local features through a residual structure to generate a fifth feature map. Based on the fifth feature map and the sixth feature map obtained by processing through the channel attention mechanism in the side branch, fuse them to generate a seventh feature map. Finally, fuse the seventh feature map with the input features to output an eighth feature map.
[0011] In some embodiments, perform global average pooling operation on the input features through the channel attention mechanism, then perform 1×1 convolution and ReLU operations in sequence, then perform 1×1 convolution operation again, and finally perform Sigmoid operation to obtain the sixth feature map.
[0012] In some embodiments, perform pooling operation on the input features through the spatial attention mechanism, and then perform 7×7 convolution and Sigmoid operations in sequence to obtain the third feature map.
[0013] In some embodiments, the small target dense area calculation module uses the clustering algorithm DBSCAN to calculate the target dense area with a large number of targets in the grid, and then uses bilinear interpolation to upsample the target dense area.
[0014] In some embodiments, the UAV image target detection model is trained according to the target loss function, and the target loss function is calculated by the following formula:
[0015] L AShape-IoU =1 - α * IoU - β * distanceshape -0.5*Ω shape
[0016]
[0017]
[0018] Among them, L AShape-IoU represents the target loss function; IoU is used to calculate the overlapping degree between the predicted bounding box and the ground truth bounding box, B is the area of the predicted bounding box, and B gt is the area of the ground truth bounding box; α is an adaptive parameter to dynamically adjust the weight of IoU, β is used to dynamically adjust the distance weight, and distance shape is the distance between the two weighted bounding boxes, and β w and β h are the weight coefficients in the horizontal and vertical directions respectively, c is the hypotenuse length of the detection bounding box, is the coordinate of the center point of the ground truth bounding box, (x c , y c ) is the coordinate of the center point of the predicted bounding box; Ω shape is the shape loss, which is used to measure the difference in shape between the predicted bounding box and the ground truth bounding box, B 1 , B 2 represent two anchor boxes respectively, min(B 1 , B 2 ) represents taking the smaller value of the areas of the two anchor boxes, max(B 1 , B 2 ) represents taking the larger value of the areas of the two anchor boxes, and the center coordinates of the anchor boxes are (x 1 , y 1 ) and (x 2 , y 2 ) respectively. The center distance is defined as the Euclidean distance d, (w 1 , h 1 ), (w 2 , h 2 ) are the widths and heights of the two anchor boxes respectively.
[0019] Second, in this application, a drone image target detection device is provided, including:
[0020] An acquisition unit, configured to acquire a drone image to be detected and a drone image target detection model, where the drone image target detection model is a YOLOv11 model in which the C2PSA module in the backbone network is replaced by a main branch for extracting basic features and a side branch for fusing a spatial attention mechanism and a channel attention mechanism, and small target dense area calculation modules are connected to both ends of any C3K2 module in the backbone network;
[0021] The detection unit is configured to input the drone image to be detected into the drone image target detection model to obtain the drone image target detection result.
[0022] In a third aspect, the present application provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to implement the drone image target detection method described in the first aspect above.
[0023] In a fourth aspect, the present application provides a readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the drone image target detection method described in the first aspect above.
[0024] Compared with the related art, the present application provides a drone image target detection method, including: obtaining a drone image to be detected and a drone image target detection model. The drone image target detection model is a YOLOv11 model in which the C2PSA module in the backbone network is replaced by a main branch for extracting basic features and a side branch for fusing the spatial attention mechanism and the channel attention mechanism according to the basic features, and small target dense area calculation modules are connected to both ends of any C3K2 module in the backbone network; inputting the drone image to be detected into the drone image target detection model to obtain the drone image target detection result. The drone image target detection method proposed by this method can reduce the computational complexity of the model, reduce information loss, and significantly improve the generalization ability of the model by replacing the C2PSA module in the backbone network with a main branch for extracting basic features and a side branch for fusing the spatial attention mechanism and the channel attention mechanism according to the basic features; small target dense area calculation modules are connected to both ends of any C3K2 module in the backbone network to enhance the feature information of the dense area and enhance the detection ability of the model for drone image targets.
[0025] Details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions thereof of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0027] Figure 1 is a flowchart of the drone image target detection method of the present invention;
[0028] Figure 2 is a network architecture diagram of the improved Yolov11 model of the present invention;
[0029] Figure 3 is the processing flowchart of the DAFA module of the present invention;
[0030] Figure 4 is the network architecture diagram of C2PSA_MSCA of the present invention;
[0031] Figure 5 is the structural block diagram of the drone image target detection device of the present invention;
[0032] Figure 6 is the architecture diagram of the electronic device of the present invention. Specific embodiments
[0033] To more clearly understand the purpose, technical solution and advantages of the present application, the present application will be described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0034] Unless otherwise defined, the technical terms or scientific terms involved in the present application shall have the general meaning understood by those with ordinary skills in the technical field to which the present application belongs. In the present application, words such as "a", "one", "a kind of", "the", "these" and the like do not indicate a limitation in quantity, and they can be singular or plural. The terms "include", "comprise", "have" and any variants thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in the present application do not limit to physical or mechanical connections, but may include electrical connections, whether directly or indirectly. The term "plurality" involved in the present application means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third" and the like involved in the present application are only used to distinguish similar objects and do not represent a specific sorting of the objects.
[0035] The present application can be used in many general or special computing device environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor devices, distributed computing environments including any of the above devices or equipment, etc.
[0036] Next, in combination with Figure 1A detailed introduction to the UAV image target detection method of this application is as follows:
[0037] Step S11: Obtain the UAV image to be detected and the UAV image target detection model;
[0038] The UAV image target detection model is a YOLOv11 model in which the C2PSA module in the backbone network is replaced by a main branch for extracting basic features and a side branch for fusing the spatial attention mechanism and the channel attention mechanism according to the basic features, and small target dense area calculation modules are connected to both ends of any C3K2 module in the backbone network.
[0039] Specifically, the improved Yolov11 model network architecture is as Figure 2 shown. The C2PSA module in the backbone network of the Yolov11 model is replaced by the C2PSA_MSCA module, and small target dense area calculation modules DAFA are connected to both ends of any C3K2 module in the backbone network. By replacing the C2PSA module in the backbone network of the model with the C2PSA_MSCA module, the computational complexity of the model is reduced, information loss is reduced, and at the same time, the generalization ability of the model is significantly improved; small target dense area calculation modules DAFA are connected to both ends of any C3K2 module in the backbone network to enhance the feature information of the dense area and enhance the detection ability of the model for UAV image targets.
[0040] The processing process of the main branch for extracting basic features includes: the main branch first processes the input features through 1×1 convolution to increase the channel dimension and fuse the information of each channel to generate the first feature map, then extracts multi-scale context information through dilated convolution, and then generates the basic features through activation function processing; the processing process of fusing the spatial attention mechanism and the channel attention mechanism according to the basic features includes: fusing the third feature map obtained by processing through the spatial attention mechanism in the side branch with the basic features to generate the fourth feature map, then processing the fourth feature map through depthwise separable convolution to obtain local features, then fusing the fourth feature map with the local features through a residual structure to generate the fifth feature map, based on the fifth feature map and the sixth feature map obtained by processing through the channel attention mechanism in the side branch to generate the seventh feature map, and finally fusing the seventh feature map with the input features to output the eighth feature map.
[0041] Specifically, the network architecture of the C2PSA_MSCA module is as Figure 4As shown, the main branch first processes the input features through a 1×1 convolution to increase the channel dimension and fuse the information of each channel to generate a first feature map. Then, it extracts multi-scale context information through a dilated convolution DConv3×3×3 (representing a convolution kernel size of 3 and a dilation rate of 3), and then generates the base feature through the ReLU activation function. The third feature map obtained by processing through the channel attention mechanism (channel Aggregation) in the side branch is fused with the base feature to generate a fourth feature map. Then, the fourth feature map is processed through a depthwise separable convolution DWConv3×3 to obtain local features. Next, the fourth feature map is fused with the local features through a residual structure to generate a fifth feature map. Based on the fifth feature map, it is fused with the sixth feature map obtained by processing through the spatial attention mechanism (spatial Aggregation) in the side branch to generate a seventh feature map. Finally, the seventh feature map is fused with the input feature to output an eighth feature map.
[0042] The C2PSA_MSCA module integrates enhancement strategies in both channel and spatial dimensions, and at the same time maintains the original input features through cross-layer connections, avoiding information loss and promoting gradient propagation. The finally generated features have both global context information and fine-grained spatial details. Applying the C2PSA_MSCA module to Yolov11 enables the model to effectively improve the ability for drone image target detection and enhance the generalization ability of the model with lower computational complexity.
[0043] Further, as Figure 4 shown, the input features are subjected to global average pooling (GAP) operation through the channel attention mechanism, followed by 1×1 convolution and ReLU operations in sequence, then another 1×1 convolution operation, and finally a Sigmoid operation to obtain the sixth feature map.
[0044] In some embodiments, as Figure 4 shown, the input features are pooled (Pool) through the spatial attention mechanism, followed by 7×7 convolution and Sigmoid operations in sequence to obtain the third feature map.
[0045] As Figure 3 shown, the small target dense area calculation module DAFA uses the clustering algorithm DBSCAN to calculate the target dense areas with a large number of targets in the grid, and then uses bilinear interpolation to upsample (UpSample) the target dense areas.
[0046] Specifically, the processing flow of the DAFA module includes: extracting the response value (i.e., activation value) of each pixel point through the feature map. Then, taking the position and response value of each pixel point of the feature map as three-dimensional data (spatial coordinates and response value), constructing the input data set of DBSCAN. Next, applying the DBSCAN algorithm, by setting appropriate parameters, identifying the dense regions in the feature map, which usually correspond to the positions where the target objects are located. DBSCAN will classify the points in the feature map into different clusters according to the density. By screening the number of points included in each cluster, if the number of points in a certain cluster is greater than the preset value, then it is considered that the cluster is the target dense region. After obtaining the target dense region, bilinear interpolation is used to upsample the target dense region to improve the resolution of this region, and the obtained feature map is detected for the second time.
[0047] The DAFA module calculates the target dense region in the feature map, magnifies the region, strengthens the target feature information of the feature map, enables the model to focus on the complete context information of the target, and improves the model's detection ability for dense small targets.
[0048] Step S12, input the drone image to be detected into the drone image target detection model to obtain the drone image target detection result.
[0049] As can be seen from the above technical solutions, the drone image target detection method provided by this application can use the drone image target detection model to perform target detection on the drone image. Among them, the drone image target detection model is a YOLOv11 model in which the C2PSA module in the backbone network is replaced by a main branch for extracting basic features and a side branch for fusing the spatial attention mechanism and the channel attention mechanism according to the basic features, and small target dense region calculation modules are connected to both ends of any C3K2 module in the backbone network. The drone image target detection method proposed in this application reduces the computational complexity of the model, reduces information loss, and significantly improves the generalization ability of the model by replacing the C2PSA module in the backbone network of the model with a main branch for extracting basic features and a side branch for fusing the spatial attention mechanism and the channel attention mechanism according to the basic features; small target dense region calculation modules are connected to both ends of any C3K2 module in the backbone network to enhance the feature information of the dense region and enhance the model's detection ability for drone image targets.
[0050] The drone image target detection model is trained according to the target loss function, and the target loss function is calculated by the following formula:
[0051] L AShape-IoU = 1 - α * IoU - β * distance shape - 0.5 * Ω shape
[0052]
[0053] Among them, L AShape-IoU represents the target loss function; IoU is used to calculate the overlap degree between the predicted bounding box and the ground truth bounding box, B is the area of the predicted bounding box, and B gt is the area of the ground truth bounding box; α is used as an adaptive parameter to dynamically adjust the weight of IoU, β is used to dynamically adjust the distance weight, and distance shape is the distance between the two weighted bounding boxes, and β w and β h are the weight coefficients in the horizontal and vertical directions respectively, c is the hypotenuse length of the detection bounding box, is the coordinate of the center point of the ground truth bounding box, (x c , y c ) is the coordinate of the center point of the predicted bounding box; Ω shape is the shape loss, which is used to measure the difference in shape between the predicted bounding box and the ground truth bounding box, B 1 , B 2 represent two anchor boxes respectively, min(B 1 , B 2 ) represents taking the smaller value of the areas of the two anchor boxes, and max(B 1 , B 2 ) represents taking the larger value of the areas of the two anchor boxes. The center coordinates of the anchor boxes are (x 1 , y 1 ) and (x 2 , y 2 ) respectively. The center distance is defined as the Euclidean distance d, (w 1 , h 1 ), (w 2 , h 2 ) are the widths and heights of the two anchor boxes respectively.
[0054] The existing YOLO11 model bounding box regression loss uses the CIoU loss function, which is calculated based on the aspect ratio of the predicted box and the true box. When the aspect ratio of the two is the same, the CIoU result will be the same. However, when the drone is flying at a certain altitude, the scale of the object may change greatly, and the sample may be affected by environmental factors such as lighting conditions, occlusion, and perspective changes. Some low-quality samples are inevitably included. If the original CIoU is continued to be used, the geometric factors will increase the penalty for these low-quality samples, resulting in a decrease in the generalization ability of the model. By introducing AShape-IoU to replace CIoU, the influence of the shape and size of the bounding box itself on the bounding box regression is considered to improve the accuracy of the regression loss calculation; at the same time, by providing a variety of shape intersection-over-union calculation methods to meet the different requirements for similarity in different application scenarios. For example, in some scenarios, you may pay more attention to specific aspects such as the center distance of the box and the similarity ratio of the shape, and a specific shape intersection-over-union calculation method cannot meet the needs.
[0055] In AShape-IoU, the weight parameters are automatically adjusted according to the area ratio and center distance ratio of the boxes. If the area ratio of the two boxes is small, it means that their sizes are quite different. In this case, the weight of IoU is reduced to reduce the penalty for distance and shape costs. If the center distance ratio of the two boxes is large, it means that they are far apart. In this case, the weight is also adjusted appropriately to adapt to this situation. Therefore, α is used as an adaptive parameter to dynamically adjust the IoU weight, and β is used to dynamically adjust the distance weight. Limiting β to [0,1] can clearly express the adjustment range of the weight and avoid nonlinear amplification effects on the adjustment of the distance weight, which can lead to model instability.
[0056] During the model training process, the AShape-IoU loss function replaces the CIoU loss function and is used for the positioning loss calculation of the target box regression, which effectively improves the detection accuracy and the generalization ability of the model.
[0057] In a second aspect, the present invention provides a drone image target detection device 500, Figure 5 is a structural block diagram of the drone image target detection device of the present invention, such as Figure 5 As shown, the device comprises:
[0058] An acquisition unit 501 is used to acquire a drone image to be detected and a drone image target detection model, wherein the drone image target detection model is a YOLOv11 model in which a C2PSA module in a backbone network is replaced with a main branch for extracting basic features and a module for fusing a side branch of a spatial attention mechanism and a channel attention mechanism according to the basic features, and a small target dense area calculation module is connected to both ends of any C3K2 module in the backbone network;
[0059] The detection unit 502 is configured to input the drone image to be detected into the drone image target detection model to obtain a drone image target detection result.
[0060] In this device, the C2PSA module in the backbone network of the model is replaced by a main branch for extracting basic features and a side branch for fusing the spatial attention mechanism and the channel attention mechanism according to the basic features, reducing the computational complexity of the model, reducing information loss, and significantly improving the generalization ability of the model; both ends of any C3K2 module in the backbone network are connected with a small target dense area calculation module to enhance the feature information of the dense area and enhance the detection ability of the model for drone image targets.
[0061] It should be noted that the above-mentioned each module can be a functional module or a program module, and can be implemented either by software or by hardware. For the modules implemented by hardware, the above-mentioned each module can be located in the same processor; or the above-mentioned each module can also be located in different processors in any combined form.
[0062] The following embodiments take the application of the above method to a computer device as an example for illustration. It can be understood that the computer device can be any device with computing and processing functions, and can be, but is not limited to, a server or a personal laptop computer, etc. In one embodiment, the computer device can be an application server, and the application server can be a server for running an application program to be tested.
[0063] In a third aspect, refer to Figure 6 , which shows a hardware structure block diagram of an electronic device. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present application described and / or claimed herein.
[0064] As Figure 6 shown, the electronic device includes: at least one processor 1, at least one communication interface 2, at least one memory 3, and at least one communication bus 4;
[0065] In the embodiment of the present application, the number of the processor 1, the communication interface 2, the memory 3, and the communication bus 4 is at least one, and the processor 1, the communication interface 2, and the memory 3 complete mutual communication through the communication bus 4;
[0066] The processor 1 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, etc.;
[0067] The memory 3 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory;
[0068] Wherein, the memory stores a program, and the processor can call the program stored in the memory, and the program is used for: implementing each processing flow of the aforementioned UAV image target detection method.
[0069] In a fourth aspect, an embodiment of the present invention further provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the aforementioned UAV image target detection method is implemented.
[0070] In a fifth aspect, an embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the aforementioned UAV image target detection method is implemented.
[0071] The above embodiments have described the present invention in particular detail with respect to possible scenarios. Those skilled in the art will recognize that the present invention can be practiced through other embodiments. The specific naming of components, the case of terms, attributes, data structures, or any other programming or structural aspects are not mandatory or important. The mechanisms or features for implementing the present invention can have different names, forms, or procedures. The system can be implemented through a combination of hardware and software (as described), entirely through hardware elements, or entirely through software elements. The specific division of functions between the various system components described in the text is exemplary, not mandatory; on the contrary, the functions performed by a single system component can be performed by multiple components, or the functions performed by multiple components can be performed by a single component.
[0072] Those skilled in the art should understand that each step of the above disclosed method can be implemented by a general-purpose computing device. They can be concentrated on a single computing device, or distributed over a network composed of multiple computing devices. Optionally, they can be implemented with program code executable by the computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the disclosure of the embodiments of the present invention is not limited to any specific combination of hardware and software.
[0073] Programs (also referred to as programs, software, software applications, or code) executable by these computing devices include machine instructions for a programmable processor and can implement these computing programs using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0074] Certain aspects of the present invention include the process steps and instructions described herein in the form of algorithms. It should be noted that the process steps and instructions of the present invention can be implemented in software, firmware, and / or hardware. When implemented by software, it can be downloaded and thus stored on different platforms used by various operating systems and operated from these platforms.
[0075] Those skilled in the art can understand that the structures shown in the drawings are only block diagrams of some of the structures related to the solution of the present application and do not constitute a limitation on the terminal devices to which the solution of the present application is applied. The specific terminal devices may include more or fewer components than those shown in the drawings, or combine some components, or have different component arrangements.
[0076] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "possible design" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0077] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0078] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting targets in drone images, characterized in that: include: Obtain a drone image to be detected and a drone image target detection model, wherein the drone image target detection model is a YOLOv11 model in which the C2PSA module in the backbone network is replaced with a main branch for extracting basic features and a module for fusing side branches of a spatial attention mechanism and a channel attention mechanism according to the basic features, and small target dense area calculation modules are connected to both ends of any C3K2 module in the backbone network; The drone image to be detected is input into the drone image target detection model to obtain a drone image target detection result.
2. The detection method according to claim 1, characterized in that: The main branch first processes the input features through 1×1 convolution to increase the channel dimension and fuses the information of each channel to generate a first feature map, then extracts multi-scale context information through dilated convolution, and then generates the basic features through activation function processing, and then fuses the third feature map obtained by processing the spatial attention mechanism in the side branch with the basic features to generate a fourth feature map, and then processes the fourth feature map through depth-wise separable convolution to obtain local features, and then fuses the fourth feature map with the local features through a residual structure to generate a fifth feature map, and fuses the fifth feature map with the sixth feature map obtained by processing the channel attention mechanism in the side branch to generate a seventh feature map, and finally fuses the seventh feature map with the input features to output an eighth feature map.
3. The detection method according to claim 2, characterized in that: A global average pooling operation is performed on the input features through the channel attention mechanism, followed by 1×1 convolution and ReLU operations in sequence, and then a 1×1 convolution operation is performed again, and finally a Sigmoid operation is performed to obtain the sixth feature map.
4. The detection method according to claim 3, characterized in that: The input features are pooled using the spatial attention mechanism, and then 7×7 convolution and Sigmoid operations are performed in sequence to obtain the third feature map.
5. The detection method according to claim 1, characterized in that: The small target dense area calculation module uses the clustering algorithm DBSCAN to calculate the target dense area with a large number of targets in the grid, and then uses bilinear interpolation to upsample the target dense area.
6. The detection method according to claim 1, characterized in that: The drone image target detection model is trained according to the target loss function, and the target loss function is calculated by the following formula: L AShape-IoU =1-α*IoU-β*distance shape -0.5*Ω shape Among them, L AShape-IoU represents the target loss function; IoU is used to calculate the overlap between the predicted box and the real box, B is the area of the predicted box, and B gt is the area of the real box; α is used as an adaptive parameter to dynamically adjust the IoU weight, and β is used to dynamically adjust the distance weight. shape is the weighted distance between the two bounding boxes, β w and β h are the weight coefficients in the horizontal and vertical directions, c is the length of the hypotenuse of the detection frame, is the coordinate of the center point of the real frame, (x c ,y c ) is the coordinate of the center point of the prediction box; Ω shape is the shape loss, which is used to measure the difference between the predicted box and the true box shape. B1 and B2 represent two anchor boxes respectively. min(B1, B2) means taking the smaller area value of the two anchor boxes, and max(B1, B2) means taking the larger area value of the two anchor boxes. The center coordinates of the anchor boxes are (x1, y1) and (x2, y2) respectively. The center distance is defined as the Euclidean distance d. (w1, h1) and (w2, h2) are the width and height of the two anchor boxes respectively.
7. A drone image target detection device, characterized in that: include: An acquisition unit is used to acquire a drone image to be detected and a drone image target detection model, wherein the drone image target detection model is a YOLOv11 model in which a C2PSA module in a backbone network is replaced with a main branch for extracting basic features and a module for fusing a side branch of a spatial attention mechanism and a channel attention mechanism according to the basic features, and a small target dense area calculation module is connected to both ends of any C3K2 module in the backbone network; The detection unit is used to input the drone image to be detected into the drone image target detection model to obtain the drone image target detection result.
8. An electronic device, comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to implement the drone image target detection method according to any one of claims 1 to 6.
9. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for detecting targets in drone images according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Water surface garbage identification method and system based on unmanned aerial vehicle
CN120708103A
An unmanned aerial vehicle-based water surface garbage identification method and system
CN120708103B