Multi-pedestrian target detection method based on unmanned aerial vehicle

By building a YOLOv8 optimization network, combining multi-scale feature fusion, channel reweighting and spatial context perception, the instability and occlusion problems of drone pedestrian detection are solved, high-precision real-time detection is achieved, and intelligent traffic management and public safety monitoring are supported.

CN120472342APending Publication Date: 2025-08-12HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510551975.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The image quality of the drone is unstable when shooting at high altitudes, and is disturbed by the external environment, resulting in inaccurate pedestrian detection. Pedestrian detection is easily affected by occlusion and behavioral diversity, making it difficult to provide reliable data support in intelligent transportation systems.

Method used

A pedestrian detection optimization network based on YOLOv8 is built, including feature enhancement module, channel reweighting module and spatial context perception module. Combining position loss, category loss and confidence loss, the model is trained through the gradient descent method to achieve high-precision real-time detection.

Benefits of technology

In complex scenarios, the robustness and accuracy of pedestrian detection are significantly improved, reliable data support is provided, and traffic management efficiency and safety are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472342A_ABST
    Figure CN120472342A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-pedestrian target detection method based on an unmanned aerial vehicle, and the method comprises the steps: (1) collecting pedestrian video data of the unmanned aerial vehicle, carrying out the frame extraction preprocessing, and marking a pedestrian bounding box, a type, and a confidence coefficient; (2) constructing a YOLOv8 optimization network, wherein a neck network of the YOLOv8 optimization network is integrated with a feature enhancement module, a channel reweighting module and a spatial context sensing module; (3) designing a ternary loss function fusing intersection-to-union ratio, center distance and length-width ratio constraints, wherein the ternary loss function comprises position, category and confidence loss; (4) training the network to converge by adopting a gradient descent method; and (5) deploying the model to the unmanned aerial vehicle to realize real-time detection. According to the method, by improving the network structure and the loss function, the pedestrian detection precision in a complex scene is remarkably improved, efficient data support is provided for an intelligent traffic system, and pedestrian safety guarantee and traffic management efficiency are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of visual detection technology, and in particular to a multi-pedestrian target detection method based on an unmanned aerial vehicle (UAV). Background Art

[0002] In intelligent transportation systems, pedestrian detection is a crucial component for improving traffic efficiency. With the advancement of drone technology, using drones equipped with advanced sensors and computer vision techniques for pedestrian detection has become a research hotspot in recent years. Due to their flexibility and high-altitude viewing angles, drones can capture dynamic pedestrian information that traditional ground-based monitoring equipment cannot, providing more comprehensive data support for traffic management departments. However, in practical applications, using drone platforms for pedestrian detection still faces numerous challenges and technical difficulties.

[0003] First, due to their small size and poor stability when shooting at high altitudes, drones often capture images that are subject to interference from the external environment, resulting in unstable image quality. This instability includes, but is not limited to, image blur, jitter, or partial loss of focus due to factors such as fluctuating wind speeds, varying flight altitudes, and camera angle adjustments, which can affect the computer vision algorithm's ability to accurately identify pedestrians. Intelligent transportation systems require high real-time performance, and significant image processing delays can impact traffic management's response to real-time traffic conditions and even increase the risk of traffic accidents.

[0004] Secondly, due to the high angle and altitude of drone photography, pedestrians are often located in different visual positions. Under these conditions, pedestrian detection is susceptible to partial occlusion by ground obstacles. The presence of objects such as buildings, trees, and road signs may partially or completely obscure pedestrians, making traditional pedestrian detection algorithms perform poorly in such complex scenarios. Furthermore, the diversity of pedestrian behavior is also a factor that cannot be ignored. For example, when crossing the road or waiting at a traffic light, pedestrians may remain in a fixed position or exhibit complex dynamic changes. These dynamic changes often make traditional pedestrian detection algorithms difficult to cope with, resulting in detection failures or inaccurate results.

[0005] Therefore, how to efficiently and accurately detect pedestrian targets based on the image information obtained by drones, and then provide reliable data support for the intelligent transportation system, ensure the safety of pedestrians, and improve the level and efficiency of urban traffic management has become a technical problem that needs to be solved urgently. Summary of the Invention

[0006] The main purpose of this invention is to provide a multi-pedestrian target detection method based on drones, which aims to detect pedestrian targets efficiently and accurately based on the image information obtained by drones, thereby providing reliable data support for intelligent transportation systems, ensuring the safety of pedestrians, and improving the level and efficiency of urban traffic management.

[0007] In order to achieve the above objectives, the present invention proposes a multi-pedestrian target detection method based on a drone, comprising the following steps: (1) Collect a video dataset containing pedestrians from a drone platform, extract frames and preprocess them to obtain a sequence of annotated pedestrian image frames. The annotations include the coordinates of the pedestrian bounding box, category labels, and confidence information. (2) Constructing a pedestrian detection optimization network based on YOLOv8, the network includes a backbone network and an improved neck network; wherein the improved neck network includes a feature enhancement module, a channel reweighting module and a spatial context perception module connected in sequence; the feature enhancement module fuses multi-scale features through a multi-branch convolutional structure, the channel reweighting module dynamically weights feature channels through a spatial attention mechanism, and the spatial context perception module combines global statistical features and pixel-level attention to enhance the saliency of the target area; (3) Constructing a loss function, including position loss, category loss, and confidence loss, where the confidence loss integrates the intersection-over-union ratio, center point distance, and aspect ratio constraints; (4) training the optimization network using the gradient descent method, and updating the network parameters by minimizing the loss function until the model converges; (5) Deploy the trained model to the UAV platform to perform pedestrian target detection in real-time video streams.

[0008] In one embodiment of the present application, the feature enhancement module includes: The first branch uses 1×1 convolution followed by 3×3 convolution; The second branch uses 1×1 convolution followed by 1×3 convolution and dilated convolution; The third branch uses 1×1 convolution followed by 3×1 convolution and dilated convolution; Concatenate the output of each branch with the input features along the channel dimension.

[0009] In one embodiment of the present application, the channel reweighting module is implemented by the following steps: The input features are compressed into a channel weight vector of n×c×1×1 through the convolution block; The weight vector is broadcast multiplied with the original input feature to dynamically adjust the contribution of each channel, where n represents the batch size, c represents the number of channels, and 1×1 represents the spatial size.

[0010] In one embodiment of the present application, the spatial context perception module calculates the enhanced features of the pixel points in the feature map using the following formula:

[0011] Represents the enhanced eigenvalue of the jth pixel in the i-th feature map; Represents the original eigenvalue of the jth pixel in the i-th feature map; Represents the attention weight of the j-th pixel; represents the linear transformation weight matrix, q represents the query vector, and k represents the key vector; represents the linear transformation weight matrix, and v represents the value vector; Represents the total number of pixels in the i-th feature map.

[0012] In one embodiment of the present application, the position loss is calculated by the mean square error between the coordinates of the upper left corner and the lower right corner of the predicted box and the real box:

[0013] in, Represents the position loss value of the mth pedestrian in the nth image frame; , Represents the coordinate of the upper left corner of the m-th pedestrian bounding box in the n-th frame predicted by the model, and the superscript 1 represents the upper left corner of the bounding box; , represents the coordinate of the upper left corner of the mth pedestrian bounding box in the nth frame of the annotation; ori represents the actual annotation value; and Represents the coordinates of the lower right corner of the m-th pedestrian bounding box in the n-th frame predicted by the model, and the superscript 2 represents the lower right corner of the bounding box; , Represents the lower right corner coordinate of the mth pedestrian bounding box in the ground-truth nth frame.

[0014] In one embodiment of the present application, the calculation formula of the confidence loss is:

[0015] in, Represents the confidence loss value of the mth pedestrian in the nth image frame; It represents the ratio of the intersection area of the predicted box and the true box to the union area, that is, the intersection-union ratio; The Euclidean distance between the center point of the predicted box and the true box; Represents the center point coordinates of the mth pedestrian prediction box in the nth frame; Represents the coordinates of the center point of the mth pedestrian real box in the nth frame; Represents the diagonal length of the minimum enclosing rectangle that surrounds the predicted box and the true box; Indicates the balanced aspect ratio difference term The weight coefficient of .

[0016] In one embodiment of the present application, the aspect ratio difference term Calculated by the following formula:

[0017] in, Represents the aspect ratio difference of the mth pedestrian in the nth frame image; Indicates the width of the real box; Indicates the height of the real frame; Indicates the width of the prediction box; Indicates the height of the prediction box.

[0018] In one embodiment of the present application, the expansion rate of the dilation convolution of the second branch and the third branch is set to 1.

[0019] In one embodiment of the present application, the backbone network includes a combination of ConvBlock, Bottleneck, C2f, and SPPF modules, wherein the Bottleneck module controls whether to retain input features through the hyperparameter add during residual connection.

[0020] In one embodiment of the present application, the intersection-over-union ratio The calculation of adopts segmented threshold processing, specifically:

[0021] Where d represents the lower threshold and u represents the upper threshold.

[0022] By adopting the above technical solution, the present invention integrates multi-scale features through the multi-branch convolution structure of the feature enhancement module, effectively solving the problems of large scale changes and loss of details of pedestrian targets from the perspective of the drone; the channel reweighting module dynamically adjusts the channel weights through the spatial attention mechanism to suppress background interference and enhance the target feature expression; the spatial context perception module combines global statistical features and pixel-level attention mechanism to significantly improve the detection robustness of occluded targets in complex scenes; the confidence loss function integrates the intersection-over-union ratio, center point distance and aspect ratio constraints to optimize positioning accuracy and reduce false detections and missed detections; finally, high-precision, real-time multi-pedestrian target detection is achieved on the drone platform, providing reliable technical support for intelligent traffic management, public safety monitoring and disaster relief. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The present invention will be described in detail below with reference to specific embodiments and accompanying drawings, wherein: Figure 1This is a schematic structural diagram of the first embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the following specific embodiments are only used to explain the present invention and do not constitute a limitation of the present invention.

[0025] like Figure 1 As shown, in order to achieve the above purpose, the present invention proposes a multi-pedestrian target detection method based on a drone, comprising the following steps: (1) Collect a video dataset containing pedestrians from a drone platform, extract frames and preprocess them to obtain a sequence of annotated pedestrian image frames. The annotations include the coordinates of the pedestrian bounding box, category labels, and confidence information. (2) Constructing a pedestrian detection optimization network based on YOLOv8, the network includes a backbone network and an improved neck network; wherein the improved neck network includes a feature enhancement module, a channel reweighting module and a spatial context perception module connected in sequence; the feature enhancement module fuses multi-scale features through a multi-branch convolutional structure, the channel reweighting module dynamically weights feature channels through a spatial attention mechanism, and the spatial context perception module combines global statistical features and pixel-level attention to enhance the saliency of the target area; (3) Constructing a loss function, including position loss, category loss, and confidence loss, where the confidence loss integrates the intersection-over-union ratio, center point distance, and aspect ratio constraints; (4) training the optimization network using the gradient descent method, and updating the network parameters by minimizing the loss function until the model converges; (5) Deploy the trained model to the UAV platform to perform pedestrian target detection in real-time video streams.

[0026] Specifically, we first collect a video dataset containing pedestrians from a drone platform, extract frames from the video data to obtain a raw image sequence, and annotate the pedestrian targets in each frame. The annotation information includes the coordinates of the upper left and lower right corners of the pedestrian's bounding box, the category label, and the confidence level, forming a labeled pedestrian image frame sequence. Then, a pedestrian detection optimization network based on YOLOv8 is constructed. The backbone network of the network consists of Conv Block module, Bottleneck module, C2f module, SPPF module and Detect module. The Conv Block module processes input features through two-dimensional convolution, batch normalization and SiLU activation function. The Bottleneck module selectively fuses features according to residual connection. The C2f module enhances feature diversity by splitting channels and connecting multiple Bottleneck branches in series. The SPPF module fuses contextual information of different receptive fields through multi-scale maximum pooling operation. The Detect module outputs the coordinates, confidence and category of the predicted box; the improved neck network adds feature enhancement module, channel reweighting module and spatial context perception module in sequence on the basis of the original structure of YOLOv8. The feature enhancement module processes input features through three parallel branches. The first branch uses 1×1 convolution and 3×3 convolution to extract local features. The second branch uses 1×1 convolution, 1×3 convolution and void convolution to expand the receptive field. The branch uses 1×1 convolution, 3×1 convolution and hole convolution to capture long-distance dependencies, and finally splices the features of each branch with the original input features along the channel dimension; the channel reweighting module generates a channel weight matrix through the spatial attention mechanism, and broadcasts and multiplies the input features with the weight matrix to highlight important channels; the spatial context perception module extracts global statistical features through global average pooling and global maximum pooling, and calculates the weight of each pixel in combination with the pixel-level attention mechanism to dynamically enhance the saliency of the target area; the constructed loss function includes position loss, category loss and confidence loss. The position loss is obtained by calculating the mean square error of the coordinates of the predicted box and the true box, and the category loss is calculated by the mean square error of the predicted category and the true label. The confidence loss integrates the intersection-over-union ratio, the Euclidean distance between the center point of the predicted box and the true box, and the aspect ratio consistency constraint. The specific formula is:

[0027] in, Represents the confidence loss value of the mth pedestrian in the nth image frame; It represents the ratio of the intersection area of the predicted box and the true box to the union area, that is, the intersection-union ratio; The Euclidean distance between the center point of the predicted box and the true box; Represents the center point coordinates of the mth pedestrian prediction box in the nth frame; Represents the coordinates of the center point of the mth pedestrian real box in the nth frame; Represents the diagonal length of the minimum enclosing rectangle that surrounds the predicted box and the true box; Indicates the balanced aspect ratio difference term The balance parameter is dynamically adjusted based on the difference between the intersection-over-union ratio and the aspect ratio. The optimization network is trained using the gradient descent method, and the network parameters are updated by minimizing the total loss function through backpropagation until the position loss, category loss, and confidence loss all converge. Finally, the trained model is deployed on a drone platform to extract frames, preprocess, and infer the real-time video stream, and output the pedestrian detection results.

[0028] By adopting the above technical solution, the present invention integrates multi-scale features through the multi-branch convolution structure of the feature enhancement module, effectively solving the problems of large scale changes and loss of details of pedestrian targets from the perspective of the drone; the channel reweighting module dynamically adjusts the channel weights through the spatial attention mechanism to suppress background interference and enhance the target feature expression; the spatial context perception module combines global statistical features and pixel-level attention mechanism to significantly improve the detection robustness of occluded targets in complex scenes; the confidence loss function integrates the intersection-over-union ratio, center point distance and aspect ratio constraints to optimize positioning accuracy and reduce false detections and missed detections; finally, high-precision, real-time multi-pedestrian target detection is achieved on the drone platform, providing reliable technical support for intelligent traffic management, public safety monitoring and disaster relief.

[0029] In one embodiment of the present application, the feature enhancement module includes: The first branch uses 1×1 convolution followed by 3×3 convolution; The second branch uses 1×1 convolution followed by 1×3 convolution and dilated convolution; The third branch uses 1×1 convolution followed by 3×1 convolution and dilated convolution; Concatenate the output of each branch with the input features along the channel dimension.

[0030] Specifically, define the input features for each step , Indicates the number of input channels, Indicates the batch size, represents the width of the input feature map, Represents the height of the input feature map. Through three parallel branches, the first branch performs 1×1 convolution and 3×3 convolution operations in sequence, specifically: , where k is the convolution kernel size, s is the stride, p is the padding, and c is the number of output channels; the second branch performs 1×1 convolution, 1×3 convolution, and dilated convolution operations in sequence, specifically: , and then through the dilated convolution The dilation rate of the dilated convolution is 1. The third branch performs 1×1 convolution, 3×1 convolution, and dilated convolution operations in sequence, specifically: ; Then through the dilated convolution ; Arous() represents the dilated convolution, and finally the three branches are fused along the channel dimension and added to the residual module Complete multi-scale feature enhancement.

[0031] By adopting the above technical solution, the present invention adopts a three-branch design of the feature enhancement module. The 1×1 convolution and 3×3 convolution of the first branch extract local detail features, the 1×3 convolution of the second branch is combined with the dilated convolution to capture long-distance context information in the horizontal direction, and the 3×1 convolution of the third branch is combined with the dilated convolution to enhance the vertical feature perception. The dilated convolution expands the receptive field without increasing the number of parameters. The original input features and the multi-branch outputs are spliced along the channel, retaining the original information while fusing multi-scale features, significantly improving the detail expression ability of pedestrian targets from the perspective of the drone and the robustness to complex backgrounds, effectively solving the problem of feature loss and occlusion in small target detection.

[0032] In one embodiment of the present application, the channel reweighting module is implemented by the following steps: The input features are compressed into a channel weight vector of n×c×1×1 through the convolution block; The weight vector is broadcast multiplied with the original input feature to dynamically adjust the contribution of each channel, where n represents the batch size, c represents the number of channels, and 1×1 represents the spatial size.

[0033] This technical solution generates channel weight vectors through 1×1 convolution compression, automatically learning the importance of different channels. This strengthens feature channels that are effective for detecting objects (such as pedestrians) and suppresses irrelevant or noisy channels, improving the model's feature discrimination capabilities. Requiring only simple 1×1 convolution and channel-by-channel multiplication operations, it eliminates the need for complex structures (such as fully connected layers), resulting in low computational overhead and suitable for deployment on resource-constrained drone platforms.

[0034] In one embodiment of the present application, the spatial context perception module calculates the enhanced features of the pixel points in the feature map using the following formula:

[0035] Represents the enhanced eigenvalue of the jth pixel in the i-th feature map; Represents the original eigenvalue of the jth pixel in the i-th feature map; Represents the attention weight of the j-th pixel; represents the linear transformation weight matrix, q represents the query vector, and k represents the key vector; represents the linear transformation weight matrix, and v represents the value vector; Represents the total number of pixels in the i-th feature map.

[0036] Using the above technical solution, by calculating the attention weights between all pixels The module can capture the global dependencies of feature maps and enhance the model's ability to understand long-range feature associations. It is especially suitable for scenes where pedestrians are sparsely distributed or blocked from the perspective of drones. Adaptively weight local features to highlight the contribution of important areas (such as pedestrian targets), suppress irrelevant background interference, and improve the sensitivity of small target detection.

[0037] In one embodiment of the present application, the position loss is calculated by the mean square error between the coordinates of the upper left corner and the lower right corner of the predicted box and the real box:

[0038] in, Represents the position loss value of the mth pedestrian in the nth image frame; , Represents the coordinate of the upper left corner of the m-th pedestrian bounding box in the n-th frame predicted by the model, and the superscript 1 represents the upper left corner of the bounding box; , represents the coordinate of the upper left corner of the mth pedestrian bounding box in the nth frame of the annotation; ori represents the actual annotation value; and Represents the coordinates of the lower right corner of the m-th pedestrian bounding box in the n-th frame predicted by the model, and the superscript 2 represents the lower right corner of the bounding box; , Represents the lower right corner coordinate of the mth pedestrian bounding box in the ground-truth nth frame.

[0039] This technical solution directly regresses the coordinates of the top-left and bottom-right corners of the bounding box. This approach has clear physical meaning and is completely consistent with the definition of a bounding box in object detection tasks, making it easy to understand and implement. It requires only a simple squared error calculation, without complex mathematical operations (such as trigonometric functions or exponential operations). This method is fast and suitable for drone detection systems with high real-time requirements. The mean squared error (MSE) loss function has a smooth gradient, providing a stable gradient signal during training, which facilitates model convergence and avoids exploding or vanishing gradients.

[0040] In one embodiment of the present application, the confidence loss is calculated as follows:

[0041] in, Represents the confidence loss value of the mth pedestrian in the nth image frame; It represents the ratio of the intersection area of the predicted box and the true box to the union area, that is, the intersection-union ratio; The Euclidean distance between the center point of the predicted box and the true box; Represents the center point coordinates of the mth pedestrian prediction box in the nth frame; Represents the coordinates of the center point of the mth pedestrian real box in the nth frame; Represents the diagonal length of the minimum enclosing rectangle that surrounds the predicted box and the true box; Indicates the balanced aspect ratio difference term The weight coefficient of .

[0042] The above technical solution is adopted and the loss function is made insensitive to the target scale through normalization design, which is suitable for multi-scale pedestrian detection taken by drones at different heights.

[0043] In one embodiment of the present application, the aspect ratio difference term Calculated by the following formula:

[0044] in, Represents the aspect ratio difference of the mth pedestrian in the nth frame image; Indicates the width of the real box; Indicates the height of the real frame; Indicates the width of the prediction box; Indicates the height of the prediction box.

[0045] In one embodiment of the present application, the expansion rate of the dilation convolution of the second branch and the third branch is set to 1.

[0046] Using the above technical solution, when the dilation rate = 1, the sampling points of the convolution kernel are continuous and without gaps, avoiding the loss of local information caused by interval sampling in the void convolution.

[0047] In one embodiment of the present application, the backbone network includes a combination of ConvBlock, Bottleneck, C2f, and SPPF modules, wherein the Bottleneck module controls whether to retain input features through the hyperparameter add during residual connection.

[0048] In one embodiment of the present application, the intersection-over-union ratio The calculation of adopts segmented threshold processing, specifically:

[0049] Where d represents the lower threshold and u represents the upper threshold.

[0050] The detailed steps are: Step 1: Collect the video dataset of pedestrians on the UAV platform and extract and preprocess the frames to obtain the pedestrian image frame sequence. ,in, Indicates the Pedestrian image frames; , Indicates the total number of image frames after frame extraction.

[0051] right Mark the position of each pedestrian in the pedestrian image frames, denoted as , thus obtaining the labeled pedestrian image frame sequence ,Will Middle The bounding box coordinates and confidence level of each pedestrian are recorded as ; Step 2: Build a pedestrian detection optimization network based on the yolov8 backbone network. The basic network modules of yolov8 include Conv Block module, Bottleneck module, C2f module, Detect module and SPPF module in sequence.

[0052] Step 2.1: First define the Conv Block module, Bottleneck module, C2f module, Detect module and SPPF module.

[0053] For each step of input features , Indicates the number of input channels, Indicates the batch size, represents the width of the input feature map, Indicates the height of the input feature map.

[0054] Step 2.1.1: Conv Block module first transforms the features Perform a two-dimensional convolution operation:

[0055] Represents the output feature map after the two-dimensional convolution operation, represents a two-dimensional convolution operation, represents the convolution kernel size, represents stride length, Represents filling, Represents the number of output channels. Then batch normalization is performed , represents the output feature map of the batch normalization operation, Represents a 2D batch normalization operation.

[0056] Finally, the features after processing Input to the activation function:

[0057] Among them, Y represents the output feature map of the SiLU activation function; represents the activation function; Represents the Sigmoid function.

[0058] Step 2.1.2: Bottleneck module first transforms the features Input to the first Conv Block module, and then the features Input to the second Conv Block module and introduce hyperparameters on this basis ,when When , the residual structure is added, otherwise it is not added. The purpose of introducing the residual structure is to prevent the model from losing some features during the training process.

[0059] To sum up, the logic of the entire module is as follows:

[0060] Step 2.1.3: The C2f module first inputs the feature X into the Conv Block module output Then split it into two along the channel dimension ,in, Represents an equal split operation along the channel dimension; Represents the two sub-feature maps after segmentation.

[0061] Then input it into N Bottleneck modules, define , then connect them along the channel dimension , and finally input it into the Conv Block module ; Step 2.1.4: The SPPF module first inputs it into the Conv Block module and outputs , the purpose is to adjust the number of channels. Then perform the maximum pooling operation on them respectively , , Then concatenate them along the channel dimension , and finally enter ; Step 2.1.5: The Detect module first inputs X into the Conv Block module, and its respective convolution kernel is 3*3, and the output result is , , then , Each performs a convolution operation. , ,in is the number of candidate boxes, num is the type. Finally, it is passed to the loss function for feedback and update; Step 2.2: Combine the defined modules to form the backbone network structure of yolov8; Step 3: Mainly improve the neck of the yolov8 backbone network, and add three improved modules on this basis, namely FEM module, CRC module and SCAM module; Step 3.1: Define the characteristics of each input , Indicates the number of input channels, Indicates the batch size, represents the width of the input feature map, Indicates the height of the input feature map.

[0062] Step 3.1.1: The FEM module first inputs X into different branches. The first branch: ; Second branch: ; ; The third branch: ; ; Arous() represents the dilated convolution, and finally the three branches are fused along the channel dimension and added to the residual module. .

[0063] Step 3.1.2: Based on the neck, the CRC module is introduced. This module mainly introduces the spatial attention mechanism, adds learning weights to each channel, and outputs the input feature X through the Conv Block module as an n*c*1*1 feature, which is then broadcast multiplied by X.

[0064] Step 3.1.3: Create SCAM module and define For the The feature map pixels, is the global maximum pooling, is the global average pooling, 、 It is a linear formula, and its logical formula for feature transformation is as follows:

[0065] in, Represents the output value of the jth pixel of the i-th feature map (enhanced feature); represents the spatial attention weight; represents the linear transformation weight matrix, q represents the query vector, and k represents the key vector; represents the linear transformation weight matrix, and v represents the value vector; Represents the total number of pixels in the i-th feature map.

[0066] The calculation formula of spatial attention weight is: ; in, represents the spatial attention weight of the jth pixel of the i-th feature map; represents the input features of the i-th feature map; Represents the nth pixel value of the i-th feature map; represents the total number of pixels of the i-th feature map (H×W); Step 3.1.4: Fuse the defined module with the neck part of yolov8.

[0067] Step 4: Construct a loss function, including position loss, confidence loss, and category loss; Step 4.1: Use formula (2) to construct Middle Pedestrian position loss :

[0068] in, Represents the confidence loss value of the mth pedestrian in the nth image frame; It represents the ratio of the intersection area of the predicted box and the true box to the union area, that is, the intersection-union ratio; The Euclidean distance between the center point of the predicted box and the true box; Represents the center point coordinates of the mth pedestrian prediction box in the nth frame; Represents the coordinates of the center point of the mth pedestrian real box in the nth frame; Represents the diagonal length of the minimum enclosing rectangle that surrounds the predicted box and the true box; Indicates the balanced aspect ratio difference term The weight coefficient of . 、 For the Ground truth boxes of pedestrians The coordinates of the upper left corner and lower right corner; Step 4.2: Use formula (3) to construct Middle Pedestrian category loss : ; in, Represents the category probability predicted by the model for the mth pedestrian; express Middle The true category labels of pedestrians.

[0069] Step 4.3: Use the formula in 4.2 to construct Middle Confidence loss of each pedestrian :

[0070] in, Represents the confidence loss value of the mth pedestrian in the nth image frame; It represents the ratio of the intersection area of the predicted box and the true box to the union area, that is, the intersection-union ratio; The Euclidean distance between the center point of the predicted box and the true box; Represents the center point coordinates of the mth pedestrian prediction box in the nth frame; Represents the coordinates of the center point of the mth pedestrian real box in the nth frame; Represents the diagonal length of the minimum enclosing rectangle that surrounds the predicted box and the true box; Indicates the balanced aspect ratio difference term The weight coefficient of .

[0071] express Middle The intersection of the predicted box and the real box of a pedestrian:

[0072] Where d represents the lower threshold and u represents the upper threshold.

[0073] Balance aspect ratio differences The calculation formula is: in, and express Middle The width and height of the pedestrian's ground truth box, and express Middle The width and height of the prediction box of each person;

[0074] Step 4.4: Use the gradient descent method to train the optimized YOLOV8 prediction network, and calculate the position loss, confidence loss and category loss to update the network parameters until the position loss, confidence loss and category loss converge, thereby obtaining the trained optimized YOLOV8 inference model; Step 5: Put the optimized YOLOV8 inference model into the integrated software pycharm for training and retain the best training results.

[0075] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. All equivalent structural transformations made by using the contents of the present invention description and drawings under the inventive concept of the present invention, or direct / indirect application in other related technical fields are included in the patent protection scope of the present invention.

Claims

1. A multi-pedestrian target detection method based on drone, characterized in that: The following steps are involved: (1) Collect a video dataset containing pedestrians from a drone platform, extract frames and preprocess them to obtain a sequence of annotated pedestrian image frames. The annotations include the coordinates of the pedestrian bounding box, category labels, and confidence information. (2) Constructing a pedestrian detection optimization network based on YOLOv8, the network includes a backbone network and an improved neck network; wherein the improved neck network includes a feature enhancement module, a channel reweighting module and a spatial context perception module connected in sequence; the feature enhancement module fuses multi-scale features through a multi-branch convolutional structure, the channel reweighting module dynamically weights feature channels through a spatial attention mechanism, and the spatial context perception module combines global statistical features and pixel-level attention to enhance the saliency of the target area; (3) Constructing a loss function, including position loss, category loss, and confidence loss, where the confidence loss integrates the intersection-over-union ratio, center point distance, and aspect ratio constraints; (4) training the optimization network using the gradient descent method, and updating the network parameters by minimizing the loss function until the model converges; (5) Deploy the trained model to the UAV platform to perform pedestrian target detection in real-time video streams.

2. The multi-pedestrian target detection method based on a drone as claimed in claim 1, characterized in that: The feature enhancement module includes: The first branch uses 1×1 convolution followed by 3×3 convolution; The second branch uses 1×1 convolution followed by 1×3 convolution and dilated convolution; The third branch uses 1×1 convolution followed by 3×1 convolution and dilated convolution; Concatenate the output of each branch with the input features along the channel dimension.

3. The multi-pedestrian target detection method based on a drone as claimed in claim 1, characterized in that: The channel reweighting module is implemented by the following steps: The input features are compressed into a channel weight vector of n×c×1×1 through the convolution block; The weight vector is broadcast multiplied with the original input feature to dynamically adjust the contribution of each channel, where n represents the batch size, c represents the number of channels, and 1×1 represents the spatial size.

4. The method for detecting multiple pedestrians based on a drone as claimed in claim 1, wherein: The spatial context perception module calculates the enhanced features of the pixels in the feature map using the following formula: Represents the enhanced eigenvalue of the jth pixel in the i-th feature map; Represents the original eigenvalue of the jth pixel in the i-th feature map; Represents the attention weight of the j-th pixel; represents the linear transformation weight matrix, q represents the query vector, and k represents the key vector; represents the linear transformation weight matrix, and v represents the value vector; Represents the total number of pixels in the i-th feature map.

5. The method for detecting multiple pedestrians based on a drone as claimed in claim 1, wherein: The position loss is calculated by the mean square error between the upper left corner and lower right corner coordinates of the predicted box and the real box: in, Represents the position loss value of the mth pedestrian in the nth image frame; , Represents the coordinate of the upper left corner of the m-th pedestrian bounding box in the n-th frame predicted by the model, and the superscript 1 represents the upper left corner of the bounding box; , represents the coordinate of the upper left corner of the mth pedestrian bounding box in the nth frame of the annotation; ori represents the actual annotation value; and Represents the coordinates of the lower right corner of the m-th pedestrian bounding box in the n-th frame predicted by the model, and the superscript 2 represents the lower right corner of the bounding box; , Represents the lower right corner coordinate of the mth pedestrian bounding box in the ground-truth nth frame.

6. The method for detecting multiple pedestrians based on a drone as claimed in claim 1, wherein: The confidence loss is calculated as follows: in, Represents the confidence loss value of the mth pedestrian in the nth image frame; It represents the ratio of the intersection area of the predicted box and the true box to the union area, that is, the intersection-union ratio; The Euclidean distance between the center point of the predicted box and the true box; Represents the center point coordinates of the mth pedestrian prediction box in the nth frame; Represents the coordinates of the center point of the mth pedestrian real box in the nth frame; Represents the diagonal length of the minimum enclosing rectangle that surrounds the predicted box and the true box; Indicates the balanced aspect ratio difference term The weight coefficient of .

7. The method for detecting multiple pedestrians based on a drone as claimed in claim 6, wherein: The aspect ratio difference term Calculated by the following formula: in, Represents the aspect ratio difference of the mth pedestrian in the nth frame image; Indicates the width of the real box; Indicates the height of the real frame; Indicates the width of the prediction box; Indicates the height of the prediction box.

8. The method for detecting multiple pedestrians based on a drone as claimed in claim 2, wherein: The dilation rate of the dilated convolution of the second branch and the third branch is set to 1.

9. The method for detecting multiple pedestrians based on a drone as claimed in claim 1, wherein: The backbone network includes a combination of ConvBlock, Bottleneck, C2f, and SPPF modules, wherein the Bottleneck module controls whether to retain input features through the hyperparameter add during residual connection.

10. The method for detecting multiple pedestrians based on a drone as claimed in claim 1, wherein: The intersection-over-intersection ratio The calculation of adopts segmented threshold processing, specifically: Where d represents the lower threshold and u represents the upper threshold.

Citation Information

Cited By

  • Light-weight aerial photography target detection method and system based on recursive attention fusion

    CN121767890A

  • Lightweight aerial target detection method and system based on recursive attention fusion

    CN121767890B