Method, System, Medium and Device for Improving Accuracy of Pedestrian Quantity Detection
The detection box is obtained through the human body and head branch calculation channels of the pedestrian detection model, and the matching box is used to supplement the undetected pedestrians, which solves the problem of inaccurate detection of pedestrians in the prior art, achieving higher detection accuracy and lower missed detection rate.
Patent Information
- Application Number
- CN202110758659.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-05
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-07-05
AI Technical Summary
The existing pedestrian number detection technology has the problems of high missed detection rate and insufficient accuracy, especially when pedestrians are dense and human bodies are blocked.
Through the calculation channels of the human branch and head branch of the pedestrian detection model, the real human detection box and the real human head detection box were obtained respectively, and the undetected pedestrians were supplemented by matching the human detection box. By crossing and comparing the correlation, the total number of pedestrians was finally counted.
It significantly improves the accuracy of pedestrian count detection, reduces the missed detection rate, enhances the detection ability of pedestrians being blocked, and increases the calculation amount by a very small increase.
Smart Images

Figure CN113361479B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and specifically provides a method, a system, a medium and a device for improving the accuracy of pedestrian quantity detection. Background Art
[0002] In the task of fully structured pedestrian detection, the detection model system is usually deployed at relatively high camera positions such as traffic arteries, building entrances and exits, stadiums, etc., so it will inevitably encounter occasions where pedestrians are dense. For pedestrians who are close to each other, most of their bodies are often blocked, which brings considerable difficulty to human body detection.
[0003] In the existing detection methods, the human body and the human head are detected separately, and then they are corresponded one by one. Finally, the number of pairs obtained is the detected quantity. In this case, the missed detection rate is relatively high, and the accuracy of the obtained pedestrian quantity is insufficient.
[0004] Correspondingly, there is a need in the art for a new method to solve the problem of inaccurate existing pedestrian quantity detection. Summary of the Invention
[0005] The present invention aims to solve the above technical problems, that is, to solve the problem of inaccurate existing pedestrian quantity detection.
[0006] In a first aspect, the present invention provides a method for improving the accuracy of pedestrian quantity detection, characterized in that the method includes:
[0007] S01. Input the to-be-recognized picture with pedestrians into a pedestrian detection model;
[0008] S02. Obtain the "true human body detection frame" of the pedestrian through the human body branch calculation channel of the pedestrian detection model;
[0009] S03. Obtain the "true human head detection frame + matching human body detection frame" of the pedestrian through the human head branch calculation channel of the pedestrian detection model, where the matching human body detection frame is an analog attached frame corresponding one by one to the true human head detection frame;
[0010] S04. Perform intersection over union (IoU) association on the true human body detection frame and the matching human body detection frame, and remove the duplicate matching human body detection frames;
[0011] S05. Count the sum of the true human body detection frame and the de-duplicated matching human body detection frames as the total number of pedestrians in the to-be-recognized picture.
[0012] In a preferred technical solution of the above method, the training method of the pedestrian detection model includes:
[0013] Annotate the ground truth boxes for the heads and bodies of the images in the training set respectively, and associate the ground truth boxes of the heads and the ground truth boxes of the bodies belonging to the same pedestrian;
[0014] Set the output channels of the body branch calculation channels of the original pedestrian detection model to 4*A, and increase the output channels of the head branch calculation channels from 4*A to 8*A;
[0015] Input the images in the training set and the body ground truth box annotations into the body branch of the original pedestrian detection model for training, output the real body detection boxes, and compare them with the body ground truth boxes to improve the accuracy of the body branch of the original pedestrian detection model;
[0016] Input the images in the training set, the head ground truth boxes and the associated body ground truth boxes into the head branch of the original pedestrian detection model for training, output "real head detection boxes + matching body detection boxes", and compare them with the head ground truth boxes and the associated body ground truth boxes to improve the accuracy of the head branch of the original pedestrian detection model;
[0017] Obtain the trained pedestrian detection model;
[0018] Among them, 4 or 8 represents the number of coordinate information, and A represents the number of anchor boxes.
[0019] In the preferred technical solution of the above method, step S02 further includes:
[0020] S021. Obtain multiple groups of body detection boxes through the body branch calculation channels of the pedestrian detection model;
[0021] S022. Remove duplicates from the multiple groups of body detection boxes to obtain the "real body detection boxes" of the pedestrians;
[0022] And / or
[0023] Step S03 further includes:
[0024] S031. Obtain multiple groups of head detection boxes through the head branch calculation channels of the pedestrian detection model;
[0025] S032. Remove duplicates from the multiple groups of head detection boxes;
[0026] S033. Add the attached matching body detection boxes to each group of head detection boxes after duplicate removal to obtain the "real head detection boxes + matching body detection boxes" of the pedestrians.
[0027] In a second aspect, the present invention provides a method for improving the accuracy of pedestrian number detection, characterized in that the method includes:
[0028] S06. Input the image to be recognized with pedestrians into the pedestrian detection model;
[0029] S07. Obtain the "true human detection box + matching head detection box" of the pedestrian through the human body branch calculation channel of the pedestrian detection model, where the matching head detection box is an analog accessory box corresponding one-to-one to the true human detection box;
[0030] S08. Obtain the "true head detection box" of the pedestrian through the head branch calculation channel of the pedestrian detection model;
[0031] S09. Perform intersection over union (IoU) association on the true head detection box and the matching head detection box, and remove the duplicate matching head detection boxes;
[0032] S10. Count the sum of the true head detection box and the de-duplicated matching head detection boxes as the total number of pedestrians in the image to be recognized.
[0033] In a third aspect, the present invention provides a system for improving the accuracy of pedestrian number detection, characterized in that the system includes:
[0034] Image input module: Input the image to be recognized with pedestrians into the pedestrian detection model;
[0035] Human body branch output module: Obtain the "true human detection box" of the pedestrian through the human body branch calculation channel of the pedestrian detection model;
[0036] Head branch output module: Obtain the "true head detection box + matching human detection box" of the pedestrian through the head branch calculation channel of the pedestrian detection model, where the matching human detection box is an analog accessory box corresponding one-to-one to the true head detection box;
[0037] Matching human detection box de-duplication module: Perform intersection over union (IoU) association on the true human detection box and the matching human detection box, and remove the duplicate matching human detection boxes;
[0038] Total pedestrian number statistical output module: Count the sum of the true human detection box and the de-duplicated matching human detection boxes as the total number of pedestrians in the image to be recognized.
[0039] In a preferred technical solution of the above system, the structure of training the pedestrian detection model specifically includes:
[0040] Annotation module: Perform true value box annotation on the heads and human bodies in the training set images respectively, and perform association processing on the head true value boxes and human body true value boxes belonging to the same pedestrian;
[0041] Output channel adjustment module: Set the output channels of the human body branch calculation channels of the original pedestrian detection model to 4*A, and increase the output channels of the head branch calculation channels from 4*A to 8*A;
[0042] Model training module: Input the pictures in the training set and the human body ground truth box annotations into the human body branch of the original pedestrian detection model for training, output the real human body detection box, and compare it with the human body ground truth box to improve the accuracy of the human body branch of the original pedestrian detection model; Input the pictures in the training set, the head ground truth box and the associated human body ground truth box into the head branch of the original pedestrian detection model for training, output the "real head detection box + matching human body detection box", and compare it with the head ground truth box and the associated human body ground truth box to improve the accuracy of the head branch of the original pedestrian detection model; Obtain the trained pedestrian detection model;
[0043] Among them, 4 or 8 represents the number of coordinate information, and A represents the number of anchor boxes.
[0044] In the preferred technical solution of the above system, the human body branch output module further includes:
[0045] Human body detection box acquisition module: Obtain multiple groups of human body detection boxes through the human body branch calculation channels of the pedestrian detection model;
[0046] Real human body detection box output module: Remove duplicates from multiple groups of human body detection boxes to obtain the "real human body detection box" of the pedestrian.
[0047] In the preferred technical solution of the above system, the head branch output module further includes:
[0048] Head detection box acquisition module: Obtain multiple groups of head detection boxes through the head branch calculation channels of the pedestrian detection model;
[0049] Head detection box duplicate removal module: Remove duplicates from multiple groups of head detection boxes;
[0050] Real head detection box and matching human body detection box output module: Add the associated matching human body detection box to each group of head detection boxes after duplicate removal to obtain the "real head detection box + matching human body detection box" of the pedestrian.
[0051] In the fourth aspect, the present invention provides a system for improving the accuracy of pedestrian quantity detection, characterized in that the system includes:
[0052] Picture input module: Input the picture to be recognized with pedestrians into the pedestrian detection model;
[0053] Human body branch output module: Obtain the "true human body detection box + matching head detection box" of a pedestrian through the human body branch calculation channel of the pedestrian detection model, where the matching head detection box is a simulated accessory box corresponding one-to-one to the true human body detection box;
[0054] Head branch output module: Obtain the "true head detection box" of a pedestrian through the head branch calculation channel of the pedestrian detection model;
[0055] Matching head detection box duplicate removal module: Perform intersection-over-union association on the true head detection box and the matching head detection box, and remove duplicate matching head detection boxes;
[0056] Total pedestrian count output module: Count the sum of the true head detection box and the de-duplicated matching head detection boxes as the total number of pedestrians in the image to be recognized.
[0057] In a fifth aspect, the present invention provides a computer-readable storage medium, in which multiple program codes are stored. It is characterized in that the program codes are adapted to be loaded and run by a processor to execute the method for improving the accuracy of pedestrian count detection in any one of the above technical solutions.
[0058] In a sixth aspect, the present invention provides a control device, which includes a processor and a memory. The memory is adapted to store multiple program codes. It is characterized in that the program codes are adapted to be loaded and run by the processor to execute the method for improving the accuracy of pedestrian count detection in any one of the above technical solutions.
[0059] Those skilled in the art can understand that in the technical solution of the present invention, the method for improving the accuracy of pedestrian count detection includes the following steps:
[0060] S01. Input the image to be recognized with pedestrians into the pedestrian detection model;
[0061] S02. Obtain the "true human body detection box" of a pedestrian through the human body branch calculation channel of the pedestrian detection model;
[0062] S03. Obtain the "true head detection box + matching human body detection box" of a pedestrian through the head branch calculation channel of the pedestrian detection model, where the matching human body detection box is a simulated accessory box corresponding one-to-one to the true head detection box;
[0063] S04. Perform intersection-over-union association on the true human body detection box and the matching human body detection box, and remove duplicate matching human body detection boxes;
[0064] S05. Count the sum of the true human body detection box and the de-duplicated matching human body detection boxes as the total number of pedestrians in the image to be recognized.
[0065] In the case of adopting the above technical solution, the present invention can complete the pedestrians with only the real human head detection frame being blocked. By adding a matching human detection frame as an attached frame to each real human head detection frame one by one, each recognized human head in the captured image has a matching human detection frame. At this time, for a single image, we have two sets of human detection frames (real human detection frames and matching human detection frames). Obviously, not every matching human detection frame is necessary because some of them overlap with the real human detection frames obtained from the human branch calculation channels.
[0066] Therefore, in step S04, the real human detection frames (from the human branch) and the matching human detection frames (from the human head branch) are associated by intersection over union. When they are considered to be the same human body, the matching human detection frames will be suppressed, and the de-duplicated matching human detection frames will be retained.
[0067] Finally, the real human detection frames (including those with "human head + human body" and those with "only human body, human head blocked" in the picture) and the de-duplicated matching human detection frames (including those with "only human head, human body blocked" in the picture, and those with "human head + human body" in the picture have been de-duplicated), the sum of the two can basically cover most of the blocked pedestrians. Compared with the existing detection method where both the human body and the human head must be detected and corresponding to be considered a pedestrian, the method of the present invention only slightly increases the computational amount, only adding a 1*1 convolution operation with an input channel of 256 and an output channel of 4A (A is the anchor, anchor box) quantity. However, the improvement in accuracy is significant. Unless the human head is completely blocked and most of the body is blocked, or simply both the human body and the human head are completely blocked, others can be recognized and then counted into the total number of pedestrians, improving the accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings, in which:
[0069] Figure 1 is the first embodiment of the method for improving the accuracy of pedestrian number detection of the present invention;
[0070] Figure 2 is Figure 1 the detailed flowchart of step S02 in
[0071] Figure 3 is Figure 1 the detailed flowchart of step S03 in
[0072] Figure 4 is the second embodiment of the method for improving the accuracy of pedestrian number detection of the present invention;
[0073] Figure 5 This is the first implementation mode of the system for improving the accuracy of pedestrian quantity detection of the present invention;
[0074] Figure 6 This is the specific structure used in the training process of the pedestrian detection model used for improving the accuracy of pedestrian quantity detection of the present invention;
[0075] Figure 7 This is the second implementation mode of the system for improving the accuracy of pedestrian quantity detection of the present invention.
[0076] List of reference numerals:
[0077] 1. System for improving the accuracy of pedestrian quantity detection; 11. Image input module; 12. Human body branch output module; 121. Human body detection frame acquisition module; 122. True human body detection frame output module; 13. Human head branch output module; 14. Matching human body detection frame duplicate removal module; 15. Total pedestrian quantity statistical output module;
[0078] 16. Image input module; 17. Human body branch output module; 18. Human head branch output module; 19. Matching human head detection frame duplicate removal module; 20. Total pedestrian quantity statistical output module;
[0079] 21. Annotation module; 22. Output channel adjustment module; 23. Model training module. Specific implementation mode
[0080] To facilitate the understanding of the present invention, the present invention will be described more comprehensively and meticulously below in conjunction with the accompanying drawings of the specification and embodiments. However, those skilled in the art should understand that these implementation modes are only used to explain the technical principle of the present invention and are not intended to limit the protection scope of the present invention.
[0081] In the description of the present invention, a "module" and a "processor" may include hardware, software, or a combination of both. A module may include a hardware circuit, various suitable sensors, communication ports, a memory, and may also include a software part, such as program code, or may be a combination of software and hardware. A processor may be a central processing unit, a microprocessor, an image processor, a digital signal processor, or any other suitable processor. The processor has data and / or signal processing functions. The processor may be implemented in software, in hardware, or in a combination of both. A non-transitory computer-readable storage medium includes any suitable medium that can store program code, such as a magnetic disk, a hard disk, an optical disk, a flash memory, a read-only memory, a random access memory, and so on. The term "A and / or B" represents all possible combinations of A and B, such as only A, only B, or A and B. The term "at least one of A or B" or "at least one of A and B" has a meaning similar to "A and / or B" and may include only A, only B, or A and B. The singular terms "a" and "the" may also include the plural form.
[0082] Embodiment 1:
[0083] To solve the problem of inaccurate detection of the number of pedestrians in the prior art, as Figures 1-3 shown, in Embodiment 1, the present invention provides a method for improving the accuracy of pedestrian number detection, characterized in that the method includes:
[0084] S01. Input the to-be-recognized picture of pedestrians into the pedestrian detection model;
[0085] Specifically, the to-be-recognized picture can be taken by, for example, a camera at an intersection, such as for detecting the pedestrian flow at the intersection. The pedestrian detection model can be obtained by training common models such as the YOLO model and the RetinaNet model. Here, the YOLO model is taken as an example for illustration. There are multiple branches in the YOLO model for detecting pedestrians. For example, there are usually three branches: the human body branch, the human head branch, and the human face branch. By inputting the picture into different branches, different recognition detection frames can be obtained. Since this application mainly focuses on improving the recognition accuracy of the number of pedestrians, the data such as the human face branch that will only be used in subsequent processing has a lower importance degree for this application, and whether the human face is blocked has a smaller impact on the total number statistics. Therefore, the improvement is mainly aimed at the human face branch and the human body branch.
[0086] Among them, the training method for the pedestrian detection model includes:
[0087] Label the true value frames of the human head and the human body respectively for the pictures in the training set, and perform association processing on the true value frames of the human head and the human body belonging to the same pedestrian;
[0088] Set the output channels of the human body branch calculation channel of the original pedestrian detection model to 4*A, and increase the output channels of the human head branch calculation channel from 4*A to 8*A;
[0089] Input the images in the training set and the human body ground truth box annotations into the human body branch of the original pedestrian detection model for training, output the real human body detection box, and compare it with the human body ground truth box to improve the accuracy of the human body branch of the original pedestrian detection model;
[0090] Input the images in the training set, the human head ground truth box, and the associated human body ground truth box into the human head branch of the original pedestrian detection model for training, output the "real human head detection box + matching human body detection box", and compare it with the human head ground truth box and the associated human body ground truth box to improve the accuracy of the human head branch of the original pedestrian detection model;
[0091] Obtain the trained pedestrian detection model;
[0092] Among them, 4 or 8 represents the number of coordinate information, and A represents the number of anchor boxes.
[0093] In the above training method, the specific annotation method can be manual annotation, or it can be judged by using the logic of the simple intersection over union function and the size relationship, etc., as long as the correct annotation can be achieved. Because finally, different numbers of detection boxes need to be obtained through the human body branch and the human head branch. The human body branch only needs to obtain one real human body detection box, and 4*A is enough. The human head branch needs to obtain one real human head detection box and also virtualize an associated matching human body detection box. Therefore, it needs to be set to 8*A. Finally, after repeated training, a pedestrian detection model with a controllable accuracy is obtained.
[0094] S02. Obtain the "real human body detection box" of the pedestrian through the human body branch calculation channel of the pedestrian detection model; among them, the output channel of the human body branch calculation channel of the pedestrian detection model is 4*A, 4 represents the number of coordinate information, and A represents the number of anchor boxes.
[0095] Furthermore, step S02 further includes:
[0096] S021. Obtain multiple groups of human body detection boxes through the human body branch calculation channel of the pedestrian detection model;
[0097] S022. Perform duplicate removal processing on the multiple groups of human body detection boxes to obtain the "real human body detection box" of the pedestrian.
[0098] For single-stage detection algorithms such as YOLO and RetinaNet that are anchor-based (a type of object detection algorithm), the number of channels in the human body branch used to predict the detection box is 4*A, where A represents the number of anchor boxes, and 4 represents the 4 coordinate information predicted (the final detection box is obtained through conversion with the anchor box coordinates). Obviously, in step S02, only the "true human detection box" needs to be obtained. Therefore, after determining which one the human detection box belongs to, only 4 coordinate information needs to be obtained to obtain the human detection box. In this way, multiple groups of human detection boxes can be obtained for each picture, and then these human detection boxes are de-duplicated, which can be done through NMS (Non-Maximum Suppression). The remaining ones are the "true human detection boxes" of pedestrians, which are the determined human detection boxes directly recognized by the YOLO model in the picture.
[0099] S03. Calculate the channels of the human head branch of the pedestrian detection model to obtain the "true human head detection box + matching human detection box" of the pedestrian. Among them, the matching human detection box is an analog attached box corresponding one-to-one to the true human head detection box; among them, the output channels of the human head branch calculation channels of the pedestrian detection model are 8*A.
[0100] Furthermore, step S03 further includes:
[0101] S031. Calculate the channels of the human head branch of the pedestrian detection model to obtain multiple groups of human head detection boxes;
[0102] S032. De-duplicate multiple groups of human head detection boxes;
[0103] S033. Add the attached matching human detection box to each group of human head detection boxes after de-duplication to obtain the "true human head detection box + matching human detection box" of the pedestrian.
[0104] On the human head branch used to predict the detection box, since the matching human detection box is added, that is, after detecting the human head, an additional virtual matching human detection box is supplemented through the YOLO model. At this time, if both the human head and the human body appear in the image, the human body of this person will have both the true human detection box obtained by the human body branch and the matching human detection box obtained by the human head branch. If only an isolated human head appears in the image, the human body of this person will at least have the virtual matching human detection box obtained by the human head branch.
[0105] Since the human head branch outputs both the real human head detection boxes and the matching human body detection boxes simultaneously, and the two are in one-to-one correspondence, the number of channels of the human head branch for predicting the detection boxes is 8*A. Obviously, in addition to the 4 coordinate information of itself for predicting the real human head detection boxes, the remaining 4 coordinate information is used to match the human body detection boxes. The method of the present invention has a minimal increase in computational complexity, only adding a 1*1 convolution operation with an input channel of 256 and an output channel of 4A (A is the anchor, the anchor box).
[0106] Since the matching human body detection boxes are virtual and in one-to-one correspondence with the real human head detection boxes, there is no need to remove duplicates for the matching human body detection boxes. Only the real human head detection boxes need to be de-duplicated, and the NMS method can also be used. What remains is the "real human head detection boxes + matching human body detection boxes".
[0107] S04. Calculate the intersection over union (IoU) between the real human body detection boxes and the matching human body detection boxes, and remove the duplicate matching human body detection boxes;
[0108] S05. Count the sum of the real human body detection boxes and the de-duplicated matching human body detection boxes as the total number of pedestrians in the image to be recognized.
[0109] As already mentioned in the explanation of step S03, if both the human head and the human body in the image are complete, then for the people who appear completely, there will be three boxes in the image, namely: the real human body detection box output by the human body branch, and the real human head detection box + matching human body detection box output by the human head branch. At this time, this person has two detection boxes, which is obviously not correct. Therefore, in step S04, since the real human body detection box actually exists, if the IoU between the real human body detection box and the matching human body detection box exceeds the preset value, it can be determined that the matching human body detection box at this time is redundant, and this person is not only the head detected, but the body is blocked. At this time, the duplicate matching human body detection boxes can be removed. If there is only a human body in the image, the real human body detection box can still be circled. If there is only a human head in the image, the human body output branch will no longer be able to output the real human body detection box, while the human head output branch can output both the real human head detection box and the matching human body detection box to make up for the situation where pedestrians are not counted in this occlusion case.
[0110] After removing the redundant matching human body detection boxes in step S04, the remaining matching human body detection boxes can actually represent all the pedestrians with isolated human heads detected in the image. Adding them to the real human body detection boxes (which can detect whether there is a human head or not) can obtain a more accurate number of pedestrians.
[0111] Embodiment 2:
[0112] In the first embodiment, a specific implementation method has been introduced in detail. That is, the entire human body and head are detected, and then virtual human bodies are supplemented based on the detected heads. Next, redundant matching human body detection frames are removed. Finally, the number of remaining human body detection frames is regarded as the total number of pedestrians within the visible range. Compared with the traditional method, a significant improvement has been achieved.
[0113] The difference between the second embodiment and the first embodiment is that in the second embodiment, the human body branch is set to 8*A channels, and the head branch is set to 4*A channels. Additional matching head detection frames are virtualized on the human body branch, and then they are de-duplicated with the real head detection frames obtained from the head branch. Finally, the sum is obtained as the total number of pedestrians within the visible range. That is, in the first embodiment, human bodies are virtualized based on visible heads, while in the second embodiment, heads can be virtualized based on visible human bodies.
[0114] As Figure 4 shown, the specific method of the second embodiment is as follows:
[0115] S06. Input the image to be recognized with pedestrians into the pedestrian detection model;
[0116] S07. Calculate the channels of the human body branch of the pedestrian detection model to obtain the "real human body detection frame + matching head detection frame" of the pedestrian, where the matching head detection frame is an analog attached frame corresponding one-to-one to the real human body detection frame;
[0117] S08. Calculate the channels of the head branch of the pedestrian detection model to obtain the "real head detection frame" of the pedestrian;
[0118] S09. Perform intersection over union (IoU) association on the real head detection frame and the matching head detection frame, and remove the duplicate matching head detection frames;
[0119] S10. Statistically count the sum of the real head detection frame and the de-duplicated matching head detection frames as the total number of pedestrians in the image to be recognized.
[0120] In different scenarios, the accuracies of Example 1 and Example 2 are different. For example, at a crossroads, pedestrians are in motion. At this time, the proportion of human body occlusion is low. Through actual experiments, the total number of pedestrians in the test set is 13,310. Using the conventional method of separately detecting human heads and human bodies and then pairing them, the actual number of detected pedestrians is 7,638 pairs. By using the method of supplementing the human body with the human head in Example 1, the obtained value is 8,936 pairs, which is a 10% increase in accuracy compared to the conventional solution. Further, due to the low proportion of human body occlusion in this working condition, by using the method of supplementing the human head with the human body in Example 2, the obtained value is 9,488 pairs, which is about a 5% increase in accuracy on the basis of Example 1. This is only the statistics in the scenario of the crossroads. Experiments have found that in a stadium, due to the low proportion of human head occlusion, which is exactly the opposite of the crossroads, the data of both Example 1 and Example 2 are still better than the prior art. At this time, the better solution is Example 1
[0121] It should be noted that the above embodiments are only used to illustrate the principle of the present invention and are not intended to limit the protection scope of the present invention. Without departing from the principle of the present invention, those skilled in the art can adjust the above structure so that the present invention can be applied to more specific application scenarios.
[0122] The above has described two methods, namely Example 1 and Example 2, for improving the accuracy of pedestrian number detection of the present invention. The present invention also separately proposes a system for each of the above two methods.
[0123] Refer to Figure 5 , the system 1 for improving the accuracy of pedestrian number detection corresponding to Example 1 includes:
[0124] Image input module 11: Input the to-be-recognized image of pedestrians into the pedestrian detection model;
[0125] Human body branch output module 12: Obtain the "true human body detection frame" of pedestrians through the calculation channel of the human body branch of the pedestrian detection model;
[0126] Human head branch output module 13: Obtain the "true human head detection frame + matching human body detection frame" of pedestrians through the calculation channel of the human head branch of the pedestrian detection model, where the matching human body detection frame is an analog accessory frame corresponding one-to-one to the true human head detection frame;
[0127] Matching human body detection frame deduplication module 14: Perform intersection over union (IoU) association on the true human body detection frame and the matching human body detection frame, and remove the duplicate matching human body detection frames;
[0128] Total pedestrian number statistical output module 15: Statistically count the sum of the true human body detection frame and the deduplicated matching human body detection frames as the total number of pedestrians in the to-be-recognized image.
[0129] Reference Figure 6 , the structure for training the pedestrian detection model specifically includes:
[0130] Annotation module 21: Perform true value box annotation on the heads and human bodies in the training set images respectively, and perform association processing on the true value boxes of the heads and the true value boxes of the human bodies belonging to the same pedestrian;
[0131] Output channel adjustment module 22: Set the output channel of the human body branch calculation channel of the original pedestrian detection model to 4*A, and increase the output channel of the head branch calculation channel from 4*A to 8*A;
[0132] Model training module 23: Input the images in the training set and the human body true value box annotations into the human body branch of the original pedestrian detection model for training, output the true human body detection box, and compare it with the human body true value box to improve the accuracy of the human body branch of the original pedestrian detection model; Input the images in the training set, the head true value box and the associated human body true value box into the head branch of the original pedestrian detection model for training, output the "true head detection box + matching human body detection box", and compare it with the head true value box and the associated human body true value box to improve the accuracy of the head branch of the original pedestrian detection model; Obtain the trained pedestrian detection model;
[0133] Among them, 4 or 8 represents the number of coordinate information, and A represents the number of anchor boxes.
[0134] Return to the reference again Figure 5 , in the preferred technical solution of the above system, the human body branch output module 12 further includes:
[0135] Human body detection box acquisition module 121: Obtain multiple groups of human body detection boxes through the human body branch calculation channel of the pedestrian detection model;
[0136] True human body detection box output module 122: Perform duplicate removal processing on multiple groups of human body detection boxes to obtain the "true human body detection box" of the pedestrian.
[0137] In the preferred technical solution of the above system, the head branch output module 13 further includes:
[0138] Head detection box acquisition module 131: Obtain multiple groups of head detection boxes through the head branch calculation channel of the pedestrian detection model;
[0139] Head detection box duplicate removal module 132: Perform duplicate removal processing on multiple groups of head detection boxes;
[0140] True head detection box and matching human body detection box output module 133: Add an attached matching human body detection box to each group of head detection boxes after duplicate removal processing to obtain the "true head detection box + matching human body detection box" of pedestrians.
[0141] The specific working process of the system 1 for improving the accuracy of pedestrian quantity detection is basically the same as the method proposed in the first embodiment, and will not be elaborated here.
[0142] Refer to Figure 7 , the system 1 for improving the accuracy of pedestrian quantity detection corresponding to the second embodiment includes:
[0143] Image input module 16: Input the to-be-recognized image of pedestrians into the pedestrian detection model;
[0144] Human body branch output module 17: Obtain the "true human body detection box + matching head detection box" of pedestrians through the human body branch calculation channel of the pedestrian detection model, where the matching head detection box is an analog attached box corresponding one-to-one to the true human body detection box;
[0145] Head branch output module 18: Obtain the "true head detection box" of pedestrians through the head branch calculation channel of the pedestrian detection model;
[0146] Duplicate matching head detection box removal module 19: Perform intersection over union association on the true head detection box and the matching head detection box, and remove the duplicate matching head detection boxes;
[0147] Total pedestrian quantity statistical output module 20: Statistically count the sum of the true head detection box and the non-duplicate matching head detection boxes as the total number of pedestrians in the to-be-recognized image.
[0148] In addition, the present invention also provides a computer-readable storage medium, in which multiple program codes are stored, and it is characterized in that the program codes are suitable for being loaded and run by a processor to execute the method for improving the accuracy of pedestrian quantity detection in any one of the above technical solutions.
[0149] In addition, the present invention also provides a control device, which includes a processor and a memory, and the memory is suitable for storing multiple program codes, and it is characterized in that the program codes are suitable for being loaded and run by the processor to execute the method for improving the accuracy of pedestrian quantity detection in any one of the above technical solutions.
[0150] Those skilled in the art can understand that the above-mentioned system and device for improving the accuracy of pedestrian number detection also include some other well-known structures, such as processors, controllers, memories, etc. Among them, the memory includes but is not limited to random access memory, flash memory, read-only memory, programmable read-only memory, volatile memory, non-volatile memory, serial memory, parallel memory or registers, etc. The processor includes but is not limited to CPLD / FPGA, DSP, ARM processor, MIPS processor, etc. In order not to unnecessarily obscure the embodiments of the present disclosure, these well-known structures are not shown in the drawings.
[0151] Although the steps are described in the above order in the above embodiments, those skilled in the art can understand that in order to achieve the effects of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reversed order. For example, steps S02 and S03 are carried out synchronously without a particular order, and the same applies to steps S07 and S08.
[0152] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.
Claims
1. A method for improving the accuracy of pedestrian quantity detection, characterized in that The method includes: S01. Input the to-be-recognized image with a pedestrian into a pedestrian detection model; S02. Obtain the "true human detection box" of the pedestrian through the human body branch calculation channel of the pedestrian detection model; S03. Obtain the "true head detection box + matching human detection box" of the pedestrian through the head branch calculation channel of the pedestrian detection model, where the matching human detection box is an analog accessory box corresponding one-to-one to the true head detection box; S04. Perform intersection over union (IoU) association on the true human detection box and the matching human detection box, and remove the duplicate matching human detection boxes; S05. Count the sum of the true human detection box and the deduplicated matching human detection boxes as the total number of pedestrians in the to-be-recognized image; Among them, the training method of the pedestrian detection model includes: Label the ground truth boxes for the heads and human bodies respectively in the training set images, and perform association processing on the head ground truth box and the human body ground truth box belonging to the same pedestrian; Set the output channel of the human body branch calculation channel of the original pedestrian detection model to 4*A, and increase the output channel of the head branch calculation channel of the original pedestrian detection model from 4*A to 8*A; Input the training set images and the human body ground truth box labels into the human body branch of the original pedestrian detection model for training, output the true human detection box, and compare it with the human body ground truth box to improve the accuracy of the human body branch of the original pedestrian detection model; Input the training set images, the head ground truth box, and the associated human body ground truth box into the head branch of the original pedestrian detection model for training, output the "true head detection box + matching human detection box", and compare it with the head ground truth box and the associated human body ground truth box to improve the accuracy of the head branch of the original pedestrian detection model; Obtain the trained pedestrian detection model; Among them, 4 or 8 represents the number of coordinate information, and A represents the number of anchor boxes.
2. The method according to claim 1, wherein Step S02 further includes: S021. Obtain multiple groups of human detection boxes through the human body branch calculation channel of the pedestrian detection model; S022. Perform deduplication processing on the multiple groups of human detection boxes to obtain the "true human detection box" of the pedestrian; And / or Step S03 further includes: S031. Obtain multiple groups of head detection boxes through the head branch calculation channel of the pedestrian detection model; S032. Perform deduplication processing on the multiple groups of head detection boxes; S033. Add the associated matching human detection boxes to each group of head detection boxes after deduplication processing to obtain the "true head detection box + matching human detection box" of the pedestrian.
3. A method for improving the accuracy of pedestrian number detection, characterized in that, The method includes: S06. Input the to-be-recognized image with a pedestrian into a pedestrian detection model; S07. Obtain the "true human detection box + matching head detection box" of the pedestrian through the human body branch calculation channel of the pedestrian detection model, where the matching head detection box is an analog accessory box corresponding one-to-one to the true human detection box; S08. Obtain the "true head detection box" of the pedestrian through the head branch calculation channel of the pedestrian detection model; S09. Perform intersection over union (IoU) association on the true head detection box and the matching head detection box, and remove the duplicate matching head detection boxes; S10. Count the sum of the real head detection boxes and the deduplicated matching head detection boxes as the total number of pedestrians in the image to be recognized; Among them, the training method of the pedestrian detection model includes: Label the true value boxes for the heads and bodies in the training set images respectively, and associate the head true value boxes and body true value boxes belonging to the same pedestrian; Set the output channels of the body branch calculation channels of the original pedestrian detection model to 4*A, and increase the output channels of the head branch calculation channels from 4*A to 8*A; Input the training set images and the body true value box labels into the body branch of the original pedestrian detection model for training, output the real body detection boxes, and compare them with the body true value boxes to improve the accuracy of the body branch of the original pedestrian detection model; Input the training set images, the head true value boxes, and the associated body true value boxes into the head branch of the original pedestrian detection model for training, output "real head detection box + matching body detection box", and compare them with the head true value boxes and the associated body true value boxes to improve the accuracy of the head branch of the original pedestrian detection model; Obtain the trained pedestrian detection model; Among them, 4 or 8 represents the number of coordinate information, and A represents the number of anchor boxes.
4. A system for improving the accuracy of pedestrian quantity detection, characterized in that, The system includes: Image input module: Input the image to be recognized with pedestrians into the pedestrian detection model; Body branch output module: Obtain the "real body detection box" of the pedestrian through the body branch calculation channels of the pedestrian detection model; Head branch output module: Obtain the "real head detection box + matching body detection box" of the pedestrian through the head branch calculation channels of the pedestrian detection model, where the matching body detection box is an analog attached box corresponding one-to-one to the real head detection box; Matching body detection box deduplication module: Perform intersection over union (IoU) association on the real body detection boxes and the matching body detection boxes, and remove the duplicate matching body detection boxes; Total pedestrian count output module: Count the sum of the real body detection boxes and the deduplicated matching body detection boxes as the total number of pedestrians in the image to be recognized; Among them, the structure of training the pedestrian detection model specifically includes: Labeling module: Label the true value boxes for the heads and bodies in the training set images respectively, and associate the head true value boxes and body true value boxes belonging to the same pedestrian; Output channel adjustment module: Set the output channels of the body branch calculation channels of the original pedestrian detection model to 4*A, and increase the output channels of the head branch calculation channels from 4*A to 8*A; Model training module: Input the pictures in the training set and the human body ground truth box annotations into the human body branch of the original pedestrian detection model for training, output the real human body detection box, and compare it with the human body ground truth box to improve the accuracy of the human body branch of the original pedestrian detection model; Input the pictures in the training set, the head ground truth box and the associated human body ground truth box into the head branch of the original pedestrian detection model for training, output the "real head detection box + matching human body detection box", and compare it with the head ground truth box and the associated human body ground truth box to improve the accuracy of the head branch of the original pedestrian detection model; Obtain the trained pedestrian detection model. Among them, 4 or 8 represents the number of coordinate information, and A represents the number of anchor boxes.
5. The system according to claim 4, wherein The human body branch output module further includes: Human body detection box acquisition module: Obtain multiple groups of human body detection boxes through the calculation channel of the human body branch of the pedestrian detection model; Real human body detection box output module: Remove duplicates from multiple groups of human body detection boxes to obtain the "real human body detection box" of the pedestrian; And / or The head branch output module further includes: Head detection box acquisition module: Obtain multiple groups of head detection boxes through the calculation channel of the head branch of the pedestrian detection model; Head detection box duplicate removal module: Remove duplicates from multiple groups of head detection boxes; Real head detection box and matching human body detection box output module: Add the associated matching human body detection box to each group of head detection boxes after duplicate removal to obtain the "real head detection box + matching human body detection box" of the pedestrian.
6. A system for improving the accuracy of pedestrian number detection, characterized in that, The system includes: Picture input module: Input the picture to be recognized with pedestrians into the pedestrian detection model; Human body branch output module: Obtain the "real human body detection box + matching head detection box" of the pedestrian through the calculation channel of the human body branch of the pedestrian detection model, where the matching head detection box is an analog associated box corresponding one by one to the real human body detection box; Head branch output module: Obtain the "real head detection box" of the pedestrian through the calculation channel of the head branch of the pedestrian detection model; Matching head detection box duplicate removal module: Perform intersection over union (IoU) association on the real head detection box and the matching head detection box, and remove the duplicate matching head detection boxes; Pedestrian total number statistics output module: Count the sum of the real head detection box and the de-duplicated matching head detection boxes as the total number of pedestrians in the picture to be recognized; Among them, the structure of training the pedestrian detection model specifically includes: Annotation module: Annotate the head and human body of the pictures in the training set with ground truth boxes respectively, and perform association processing on the head ground truth box and the human body ground truth box belonging to the same pedestrian; Output channel adjustment module: Set the output channel of the calculation channel of the human body branch of the original pedestrian detection model to 4*A, and increase the output channel of the calculation channel of the head branch from 4*A to 8*A; Model training module: Input the images in the training set and the human ground truth box annotations into the human body branch of the original pedestrian detection model for training, output the real human body detection box, and compare it with the human body ground truth box to improve the accuracy of the human body branch of the original pedestrian detection model; Input the images in the training set, the human head ground truth box, and the associated human body ground truth box into the human head branch of the original pedestrian detection model for training, output "real human head detection box + matching human body detection box", and compare it with the human head ground truth box and the associated human body ground truth box to improve the accuracy of the human head branch of the original pedestrian detection model; Obtain the trained pedestrian detection model; Among them, 4 or 8 represents the number of coordinate information, and A represents the number of anchor boxes.
7. A computer-readable storage medium, in which multiple program codes are stored, characterized in that, The program code is suitable for being loaded and run by a processor to execute the method for improving the accuracy of pedestrian number detection according to any one of claims 1-3.
8. A control device, the control device comprising a processor and a memory, the memory being adapted to store a plurality of program codes, characterized in that, The program code is suitable for being loaded and run by the processor to execute the method for improving the accuracy of pedestrian number detection according to any one of claims 1-3.
Citation Information
Patent Citations
Crowd counting method, device and equipment and storage medium
CN112380960A
Target detection method and device and electronic system
CN112613540A