Multi-Person Detection and Pose Estimation Method from the Perspective of Ultra-Low Altitude UAVs

By constructing a data set suitable for ultra-low altitude drones and using YOLOv5 and HRNet algorithms, the multi-peer detection and attitude estimation problems in the perspective of ultra-low altitude drones are solved, and high-precision detection and attitude estimation are achieved.

CN116152682BActive Publication Date: 2025-07-04SHENYANG AEROSPACE UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310136199.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-20
Publication Date
2025-07-04
Estimated Expiration
2043-02-20

AI Technical Summary

Technical Problem

The prior art multi-peer detection and attitude estimation methods from the perspective of ultra-low altitude drone have problems such as redundant data sets and inaccurate labeling, resulting in uncertainty in detection results, and small pedestrian targets are difficult to accurately detect.

Method used

By building a data set suitable for ultra-low altitude drones, data preprocessing and enhancement are performed, the target detection network is trained using the YOLOv5 algorithm, and the single-person pose estimation algorithm HRNet is used to perform multi-peer detection and pose estimation, and data enhancement technology and deep learning algorithms are used to improve detection accuracy.

Benefits of technology

Multi-peer detection and attitude estimation from the perspective of ultra-low altitude drone is realized, solving the uncertainty of data set quality on detection results, and improving the detection accuracy and accuracy of attitude estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116152682B_ABST
    Figure CN116152682B_ABST
Patent Text Reader

Abstract

This invention patent belongs to the field of computer vision and specifically provides a method for multi-person detection and pose estimation from the perspective of ultra-low-altitude drones, including collecting image data using drones, preprocessing the dataset, enhancing the dataset, training an object detection network using the enhanced data with the YOLOv5 algorithm, performing multi-person detection using the trained object detection network, and performing pose estimation using a single-person pose estimation algorithm. This invention solves the problems of dataset redundancy and inaccurate annotation, avoiding the uncertain influence of dataset quality on detection results; solves the problem of difficult detection caused by small pedestrian targets, and improves detection accuracy using deep learning algorithms and data augmentation techniques; and realizes the pose estimation of pedestrians from the perspective of ultra-low-altitude drones.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention patent belongs to the field of computer vision, and specifically provides a multi-person detection and pose estimation method from the perspective of ultra-low altitude drones. Background Art

[0002] Currently, sudden public safety accidents usually only rely on surveillance videos, and most of the investigations and early warnings rely on manual labor, which not only consumes a lot of manpower and material resources, but often leads to untimely discovery of problems, thus causing serious consequences. Since drones are equipped with image or video acquisition devices such as pan-tilt heads, they can obtain real-time high-resolution images and videos through wireless transmission technology, and take pictures of target objects such as transportation hubs, key monitoring areas, and geomorphic objects to obtain image information of the target objects. Therefore, in the field of public safety, drones can be relied on to conduct regular inspections of public areas to fill the inspection loopholes caused by factors such as insufficient police force. The drone inspection system can conduct inspections on forestry areas, power lines, etc. to improve the inspection efficiency, but the multi-person detection and pose estimation method from the perspective of ultra-low altitude drones is still in development. When the drone executes a flight mission, it has a certain flight speed, resulting in deviations between the collected images or videos and the real scene, and there are problems such as motion blur, video defocus, and pose occlusion in the collected pedestrian pictures. At the same time, problems such as small human targets, high similarity, and frequent interactions in the pictures have all posed obstacles to the multi-person detection and pose estimation tasks. Summary of the Invention

[0003] To solve the above problems, this invention provides a multi-person detection and pose estimation method from the perspective of ultra-low altitude drones, including,

[0004] Using the drone to collect image data to form a dataset, and performing data preprocessing on the dataset;

[0005] Performing data augmentation on the dataset;

[0006] Using the YOLOv5 algorithm to train the target detection network with the augmented data;

[0007] Using the trained target detection network to perform multi-person detection, and using a single-person pose estimation algorithm to perform pose estimation to obtain the human pose.

[0008] Furthermore, the data set is screened and labeled according to the following tags: including target ID, the x coordinate of the upper left corner of the annotation box, the y coordinate of the upper left corner of the annotation box, the x coordinate of the lower right corner of the annotation box, the y coordinate of the lower right corner of the annotation box, the frame number corresponding to this annotation, whether the target is lost, whether the target is occluded, whether it is a generated annotation, and the category corresponding to the annotation. Among them, if the target is lost, it is labeled as 1; if the target is occluded, it is labeled as 1; if this annotation is generated by automatic interpolation, it is labeled as 1; within 180 frames, those with the same ID are the same person; in the next 180 frames, there will be a new ID, and a new ID will be generated if a person is absent for 90 frames; the upper left of the annotation box is the minimum xy coordinate; the coordinates of the lower right of the annotation box are the maximum xy coordinates.

[0009] Furthermore, the following rules are used to preprocess and eliminate invalid information from the screened tag information:

[0010] 1) If there is duplicate information for the same pedestrian in multiples of 180 frames, the target ID information corresponding to the larger frame number is selected as the tag information for this frame;

[0011] 2) Tags with the loss situation labeled as 1 are discarded, and tags with the loss situation labeled as 0 are retained;

[0012] 3) Tags with the occlusion situation labeled as 1 are retained, and tags with the occlusion situation labeled as 0 are retained;

[0013] 4) Tags with the interpolation generation situation labeled as 1 are discarded, and tags with the interpolation generation situation labeled as 0 are retained. Furthermore, the screened valid tag information is expressed as:

[0014] (track ID, x min , y min , x max , y max , label),

[0015] Among them, track ID represents the target ID, x min is the x coordinate of the upper left corner of the annotation box, y min is the y coordinate of the upper left corner of the annotation box, x max is the x coordinate of the lower right corner of the annotation box, y max is the y coordinate of the lower right corner of the annotation box, and label is the category; (track ID, x min , y min , x max , y max, (label) is converted to (cls_id, x, y, w, h) and normalized, which means (target category, x value of the center coordinate of the target, y value of the center coordinate of the target, width of the target annotation box, height of the target annotation box). Since there is only one category 'person' in the label, cls_id is all 0; W is the image width, H is the image height, and the normalization formula is as follows,

[0016]

[0017] And use labelimg to correct the filtered dataset, establish a database using the corrected dataset, and divide it into a training set, a validation set, and a test set with a ratio of 7:2:1.

[0018] Furthermore, perform data augmentation on the dataset, including mosaic data augmentation, affine transformation, image flipping, HSV random augmentation, and cutout data augmentation in sequence.

[0019] Furthermore, use the YOLOv5 algorithm to train the dataset. Specifically, the anchor boxes of YOLOv5 automatically re-learn the sizes of the anchor boxes based on the training data. When the best possible recall (BPR) is less than 0.98, update the anchor in the model file. Use rectangular training, fill it as a square and input it into the network for training to reduce redundant information, reduce the number of meaningless boxes generated by the network, and speed up the network training speed. Use the binary cross-entropy loss function and the Focal loss function to calculate the localization loss and classification loss. The expression form of the binary cross-entropy loss function:

[0020]

[0021] Among them, represents the probability that the model predicts the i-th sample to be a certain class, and y( i ) represents the label;

[0022] The expression form of the Focal loss function:

[0023]

[0024] Among them, p represents the probability that the predicted sample belongs to 1, y represents the label, and γ represents the focusing parameter;

[0025] Set the weight file path, number of iterations, batch size for each gradient update, number of input images, and number of working cores; load the pre-trained weights until the network model converges, use the validation set for validation during the training process, do not perform gradient backpropagation, and save the obtained weight file.

[0026] Further, the weight file is used to participate in multi-person detection. Specifically, anchor boxes are applied to the feature map to generate an output vector of class probabilities, object scores, and preselected box information. First, it is determined whether the prediction confidence of each preselected box is within the threshold to obtain the approximate position of the target. Then, the non-maximum suppression algorithm is used to filter the preselected boxes to eliminate redundant preselected boxes. The class to which the selected preselected box belongs is the target class, and multi-person detection is completed.

[0027] The multi-person detection result is input into a single-person pose estimation algorithm, which is the HRNet algorithm, to obtain human skeleton key points. Based on the heatmap method, a heatmap is predicted for each key point of the human body. The positioning of the key point is the position where the point with the largest predicted value in the heatmap offsets 1 / 4 towards the point with the second largest predicted value, and 17 human skeleton key points are identified and the key point connection relationship is defined to represent the skeleton, thereby obtaining the human pose.

[0028] The present invention proposes a method for multi-person detection and pose estimation from the perspective of an ultra-low altitude unmanned aerial vehicle. The present invention mainly has the following advantages: solving the problems of redundancy and inaccurate annotation of the data set, and avoiding the uncertain influence of the data set quality on the detection result; solving the problem that it is difficult to detect due to the small size of the pedestrian target, and using deep learning algorithms and data augmentation techniques to improve the detection accuracy; realizing the pose estimation of pedestrians from the perspective of an ultra-low altitude unmanned aerial vehicle. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a flow schematic diagram of the present invention;

[0030] Figure 2 It is a schematic diagram of image annotation;

[0031] Figure 3 It is a schematic diagram of image enhancement;

[0032] Figure 4 It is a schematic diagram of multi-person pose estimation Figure 1 ;

[0033] Figure 5 It is a schematic diagram of multi-person pose estimation Figure 2 ;

[0034] Figure 6 It is a schematic diagram of multi-person pose estimation Figure 3 ;

[0035] Figure 7 It is a schematic diagram of multi-person pose estimation in a certain campus Figure 1 ;

[0036] Figure 8 It is a schematic diagram of multi-person pose estimation in a certain campus Figure 2 ;

[0037] Figure 9 Schematic diagram of multi-person pose estimation in a certain campus Figure 3 . Specific implementation manner

[0038] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0039] In order to achieve multi-person detection and pose estimation from the perspective of ultra-low altitude drones, a dataset that meets the flight requirements of ultra-low altitude drones and a deep learning algorithm suitable for this task need to be designed to achieve. Currently, open-source drone datasets at home and abroad include: HighD dataset, small aircraft target detection dataset, Stanford Drone dataset, University-1652 dataset, UAV123 dataset, etc. Their functions are mostly vehicle detection, drone detection, vehicle detection, building positioning; while the COCO dataset and MPII dataset that can achieve pedestrian detection and pose estimation are difficult to meet the shooting requirements of ultra-low altitude drones. There are still redundant and inaccurate annotation situations in the Okutama-Action dataset. In addition, object detection algorithms also face more challenges in this task, such as fewer available features, high positioning accuracy requirements, small target occupancy, sample imbalance, etc. At the same time, object detection algorithms often pay more attention to the detection performance of large and medium targets, and the detection accuracy of small targets is low.

[0040] Embodiment 1:

[0041] Refer to Figures 1 - 9 , the objective of the present invention is precisely to solve the problems of redundant and inaccurate annotation of the above-mentioned drone dataset and inaccurate object detection, and propose a multi-person detection and pose estimation method from the perspective of ultra-low altitude drones, overcome the uncertainty influence of the drone dataset on the detection results, achieve multi-person detection from the perspective of ultra-low altitude drones and input the results into a single-person pose estimation network to obtain human poses.

[0042] 1. Obtain an ultra-low altitude drone dataset

[0043] Search for relevant materials and research status at home and abroad, and select open-source datasets at home and abroad that meet the requirements of ultra-low altitude drone flight, including but not limited to the COCO dataset, Okutama-Action dataset (the specific content of the Okutama-Action dataset can be found in http: / / okutama - action.org / )。The mission altitude of the ultra-low altitude drone is between 0 and 100 meters. The shooting angle and the camera view are both more flexible, enabling a more comprehensive understanding of the shooting scene. The road scene of a certain campus is photographed using the on-board camera of the drone, obtaining image data containing the positions and postures of pedestrians, and the Okutama-Action dataset is acquired. The data includes the Okutama-Action dataset and the image data obtained by photographing the road scene of a certain campus with the drone.

[0044] The Okutama-Action dataset provides 43 fully annotated sequences, which can be used for training and evaluating models. The shooting location is at a baseball field in Japan, using a DJI Phantom 4 drone. To ensure the clarity of pedestrian movements, the flight altitude is from 10 meters to 45 meters, and the camera angle is from 45° to 90°. Each frame image of the drone in the dataset corresponds to a label containing eleven categories, and the required label information is selected. The required label information is: target ID, x coordinate of the upper left corner of the annotation box, y coordinate of the upper left corner of the annotation box, x coordinate of the lower right corner of the annotation box, y coordinate of the lower right corner of the annotation box, frame number corresponding to this annotation, whether the target is lost (if lost, marked as 1), whether the target is occluded (if occluded, marked as 1), whether it is a generated annotation (if the annotation is generated by automatic interpolation, marked as 1), and the category corresponding to the annotation. It should be noted that within 180 frames, those with the same ID are the same person; in the next 180 frames, there will be new IDs, and a new ID will be generated if a person is absent for 90 frames. The upper left corner of the annotation box is the minimum value of the xy coordinates; the coordinates of the lower right corner of the annotation box are the maximum values of the xy coordinates.

[0045] The above-selected label information is preprocessed according to the following rules to eliminate invalid information.

[0046] 1) There is a problem of duplicate information for the same pedestrian in multiples of 180 frames. Select the target ID information corresponding to the larger frame number as the label information for this frame;

[0047] 2) Discard the labels with the loss situation marked as 1, and retain the labels with the loss situation marked as 0;

[0048] 3) Retain the labels with the occlusion situation marked as 1, and retain the labels with the occlusion situation marked as 0;

[0049] 4) Discard the labels with the interpolation generation situation marked as 1, and retain the labels with the interpolation generation situation marked as 0.

[0050] Furthermore, the valid information after the above operations is retained as (track ID, x min , y min , x max , y max, (label), meaning (target ID, x coordinate of the upper left corner of the bounding box, y coordinate of the upper left corner of the bounding box, x coordinate of the lower right corner of the bounding box, y coordinate of the lower right corner of the bounding box, category)

[0051] Convert (track ID, x min , y min , x max , y max , label) to (cls_id, x, y, w, h) and normalize (the formula is as follows), meaning (target category, x value of the center coordinate of the target, y value of the center coordinate of the target, width of the target bounding box, height of the target bounding box). Since only the category "person" is included in label, cls_id is all 0; W is the image width and H is the image height.

[0052]

[0053] Use labelimg to correct the filtered dataset. Create a folder named VOC2007, and create folders Annotations and JPEGImages in it to store annotations and images to be annotated respectively. Create predefined_classes.txt to store the category names to be annotated. In the present invention, only the category "person" needs to be annotated.

[0054] Use the corrected Okutama - Action dataset and the dataset obtained by shooting a certain campus road scene with a drone to establish a database, and divide it into a training set, a validation set and a test set, with a ratio of 7:2:1.

[0055] 2. Data augmentation of the ultra - low - altitude drone dataset

[0056] Data augmentation for the training set in the dataset includes mosaic data augmentation, affine transformation, image flipping, HSV random augmentation, and cutout data augmentation. For mosaic data augmentation, multiple images are combined into one image according to a certain ratio, enabling the model to recognize targets within a smaller range, improving the detection performance of small targets, enriching the background of the detected objects, and calculating the data of four images during BN calculation. The specific method is as follows: randomly read four images each time and randomly select the image splicing reference point; perform operations such as flipping and scaling on the four images and place them in the upper left, upper right, lower left, and lower right positions of the reference point; map the operations of each image to the corresponding label information of the image; and perform image splicing. Affine transformation includes random rotation, translation, scaling, and shearing operations. Image flipping includes flipping up and down and flipping left and right. HSV random augmentation generates corresponding Look-Up Table (LUT) lookup tables for hue, saturation, and value respectively, and uses the cv2.LUT method to perform transformations using the Look-Up Tables just generated for hue, saturation, and value. For cutout data augmentation, a part of the image is randomly removed during training, which can improve the robustness of the model. Generating some objects similar to being occluded through cutout can not only make the model perform better when encountering occlusion problems but also make the model consider the environment more when making decisions.

[0057] 3. Use the YOLOv5 algorithm for multi-person detection

[0058] Train the object detection network using the YOLOv5 algorithm. The specific operations are as follows:

[0059] Configure the network environment, select the Windows system and the Pytorch framework for training, choose Pycharm as the integrated development environment, and put in the training set and the validation set.

[0060] Customize the data configuration file: create a person.yaml configuration file in the data folder, and write the path information of the training set and the test set during the training process, the number of categories to be detected, and the names of the categories to be recognized.

[0061] Customize the model configuration file: create yolov5_person.yaml in the model folder, write the number of categories to be detected, the parameters controlling the depth and width of the model, set three sizes of feature maps, and each feature map has three sizes of anchors. The anchor boxes of YOLOv5 can automatically re-learn the sizes of the anchor boxes based on the training data. When the Best Possible Recall (BPR) is less than 0.98, update the anchor in the model file.

[0062] Use rectangular training, fill it to a square and input it into the network for training to reduce redundant information, reduce the number of meaningless boxes generated by the network, and speed up the network training speed.

[0063] Use the binary cross-entropy loss function and the Focal loss function to calculate the localization loss and the classification loss. The expression form of the binary cross-entropy loss function is:

[0064]

[0065] where, represents the probability that the model predicts the i-th sample to be a certain class, and y( i ) represents the label. The expression form of the Focal loss function is:

[0066]

[0067] where, p represents the probability when the predicted sample belongs to 1, y represents the label, and γ represents the focusing parameter.

[0068] Set the weight file path, the number of iterations, the batch size for each gradient update, the number of input images, and the number of working cores. Load the pre-trained weights until the network model converges, use the validation set for validation during the training process, do not perform gradient backpropagation, and save the obtained weight file.

[0069] At the end of data training, two weight files are generated, the weight file of the last round and the best weight file. Use the best weight file to participate in multi-person detection. Apply anchor boxes on the feature map to generate an output vector of class probabilities, object scores, and preselected box information. First, judge whether the prediction confidence of each preselected box is within the threshold to obtain the approximate position of the target; then use the non-maximum suppression algorithm to filter the preselected boxes and eliminate redundant preselected boxes; the class to which the selected preselected box belongs is the target class.

[0070] 4. Use a single-person pose estimation algorithm for pose estimation

[0071] Input the multi-person detection results into the single-person pose estimation algorithm, and the single-person pose estimation algorithm is the HRNet algorithm to obtain the human body skeleton key points. The human body skeleton key point information is defined in the form of the COCO dataset, with a total of 17 key points, corresponding numbers:

[0072] 0 - 16: ["nose", "left_eye", "right_eye", "left_ear", "right_ear", "left_shoulder", "right_shoulder", "left_elbow", "right_elbow", "left_wrist", "right_wrist", "left_hip", "right_hip", "left_knee", "right_knee", "left_ankle", "right_ankle"], that is, [nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, right ankle].

[0073] The HRNet algorithm uses a heatmap - based method to predict a heatmap for each key point of the human body. The localization of the key point is the position where the point with the largest predicted value offsets 1 / 4 towards the point with the second - largest predicted value. The loss of the single - person pose estimation algorithm uses the mean - square error. By constructing a 2D Gaussian - distributed GT heatmap, different weights are used to calculate the loss of each key point. HRNet maintains a high resolution throughout the process. Therefore, the predicted key - point heatmap is more accurate.

[0074] Input the multiple pedestrian pre - selection boxes of each frame of the image into the single - person pose estimation algorithm, perform pose estimation on each pedestrian respectively, identify 17 human - body skeletal key points and define the key - point connection relationships to represent the skeleton. The connection relationships between the skeleton and the key points are as follows:

[0075]

[0076] Draw the key points and the skeleton on the original image. The line shapes are all anti - aliased, and the visualization effect is better.

[0077] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A multi - pedestrian detection and pose estimation method from the perspective of ultra - low - altitude unmanned aerial vehicles, characterized in that: include, Use drones to collect image data and perform data preprocessing on the data set; Perform data augmentation on the dataset; Use the YOLOv5 algorithm to train the target detection network using the enhanced data; Use the trained target detection network to detect multiple pedestrians, and use the single-person pose estimation algorithm to estimate poses; The YOLOv5 algorithm is used to train the data set. Specifically, the anchor box of YOLOv5 automatically re-learns the size of the anchor box based on the training data, uses rectangular training, fills it with a square and passes it into the network for training, and uses the binary cross entropy loss function and focal loss loss function to calculate the positioning loss and classification loss. The binary cross entropy loss function is expressed as: Among them, represents the probability that the model predicts the i-th sample to be a certain class, and y( i ) represents the label; Focal loss loss function expression: Among them, p represents the probability of predicting that the sample belongs to 1, y represents the label, and γ represents the focusing parameter; Set the weight file path, number of iterations, batch size for each gradient update, number of input images, and number of working cores; load the pre-trained weights until the network model converges, use the validation set for validation during training, do not perform gradient backpropagation, and save the weight file; Use the weight file to participate in multi-pedestrian detection. Specifically, apply anchor boxes on the feature map to generate output vectors of category probability, object score and pre-selected box information. First, determine whether the prediction confidence of each pre-selected box is within the threshold to obtain the approximate location of the target. Then use the non-maximum suppression algorithm to filter the pre-selected boxes and remove redundant pre-selected boxes. The category of the filtered pre-selected boxes is the target category, and multi-pedestrian detection is completed. The multi-pedestrian detection results are input into the single-person posture estimation algorithm, which is the HRNet algorithm. The key points of the human skeleton are obtained. Based on the heatmap method, a heat map is predicted for each key point of the human body. The positioning of the key point is the position where the point with the largest predicted value in the heat map is offset by 1 / 4 from the point with the second largest predicted value. 17 key points of the human skeleton are identified and the connection relationship of the key points is defined to represent the skeleton, so as to obtain the human posture.

2. The multi-pedestrian detection and pose estimation method from the perspective of an ultra-low altitude unmanned aerial vehicle according to claim 1, characterized in that: Filter the dataset according to the following labels: Including target ID, x-coordinate of the upper left corner of the annotation box, y-coordinate of the upper left corner of the annotation box, x-coordinate of the lower right corner of the annotation box, y-coordinate of the lower right corner of the annotation box, frame number corresponding to the annotation, whether the target is lost, whether the target is occluded, whether it is a generated annotation and the category corresponding to the annotation; if the target is lost, it is marked as 1; if the target is occluded, it is marked as 1; if the annotation is generated by automatic interpolation, it is marked as 1; within 180 frames, the person with the same ID is one person; there will be a new ID in the next 180 frames, and a person will have a new ID if he is absent for 90 frames; the upper left of the annotation box is the minimum xy coordinate; the lower right of the annotation box is the maximum xy coordinate.

3. The multi-person detection and pose estimation method from the perspective of an ultra-low altitude unmanned aerial vehicle according to claim 2, wherein: The filtered label information is preprocessed using the following rules to remove invalid information: 1) If the same pedestrian has repeated information in multiple frames of 180 frames, the target ID information corresponding to the larger frame number is selected as the label information of this frame; 2) The labels marked as 1 for loss are discarded, and the labels marked as 0 for loss are retained; 3) Labels with occlusion status marked as 1 are retained, and labels with occlusion status marked as 0 are retained; 4) Labels with interpolation generation status marked as 1 are discarded, and labels with interpolation generation status marked as 0 are retained.

4. The multi-pedestrian detection and pose estimation method from the perspective of an ultra-low altitude unmanned aerial vehicle according to claim 3, wherein: The filtered valid label information is represented as: (track ID, x min , y min , x max , y max , label), Among them, track ID represents the target ID, x min is the x-coordinate of the upper left corner of the annotation box, y min is the y-coordinate of the upper left corner of the annotation box, x max is the x-coordinate of the lower right corner of the annotation box, y max is the y-coordinate of the lower right corner of the annotation box, label is the category; convert (track ID, x min , y min , x max , y max , label) to (cls_id, x, y, w, h) and normalize it, which means (target category, x value of the center coordinate of the target, y value of the center coordinate of the target, width of the target annotation box, height of the target annotation box). Since there is only one category "person" in label, so cls_id is all 0; W is the image width, H is the image height, and the normalization formula is as follows, And use labelimg to correct the filtered dataset, establish a database using the corrected dataset, and divide it into training set, validation set and test set with a ratio of 7:2:

1.

5. The multi - pedestrian detection and pose estimation method under the perspective of ultra - low - altitude unmanned aerial vehicle according to claim 4, wherein: Perform data augmentation on the dataset, including mosaic data augmentation, affine transformation, image flipping, HSV random augmentation, and cutout data augmentation in sequence.

Citation Information

Patent Citations

  • Multi-person behavior recognition method and system based on attitude estimation and double classification

    CN114399838A

  • Unmanned aerial vehicle small target detection method based on improved multi-head self-attention

    CN114863302A