Multi-scale Pedestrian Detection Method and Device Based on the Combination of Head and Overall Information

By constructing the Faster R-CNN network model, integrating the improved feature extraction and occlusion overlap rate discrimination strategy, generating head and overall candidate boxes, and optimizing the loss function, the problem of low pedestrian detection accuracy in complex and dense scenarios is solved, and higher detection accuracy and lower missed detection rate are achieved.

CN119904894BActive Publication Date: 2025-07-25CHANGCHUN INST OF OPTICS FINE MECHANICS & PHYSICS CHINESE ACAD OF SCI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510409873.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-25
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The existing pedestrian detection algorithm has low detection accuracy for small-scale pedestrian targets and occluded overlapping pedestrian targets in complex and dense scenarios, with high missed detection rates, and lacks effective combination of head and overall information, resulting in insufficient detection performance.

Method used

The Faster R-CNN network model is built, and the improved feature extraction network is integrated. Through the non-uniform and difficult sample mining strategy of dense connection and occlusion overlap rate discrimination, the head and overall candidate boxes are generated, and the joint loss function optimization post-processing link is constructed to reduce redundant boxes and false detection and missed detection.

Benefits of technology

The detection accuracy of multi-scale pedestrian targets and obstructed pedestrian targets has been improved, the missed detection rate and false detection rate have been reduced, and the detection capability in complex crowd-intensive scenarios have been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119904894B_ABST
    Figure CN119904894B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer vision technology, and specifically provides a multi-scale pedestrian detection method and device based on the joint of head and overall information. By densely connecting features at different levels, the sensitivity of the network to multi-scale pedestrian targets is improved; secondly, the sampling method of the region proposal network is optimized, and by calculating the occlusion overlap rate of each sample in the sample set, the adaptability of the model to occluded pedestrian targets is improved; then a joint detection framework for pedestrian head and overall information is constructed to reduce the adverse impact of pedestrian body occlusion on detection. The post-processing link and the loss function module are optimized to weaken the interference caused by adjacent pedestrian targets to detection, and at the same time improve the intelligence and rationality of screening redundant boxes, further reducing the false negative rate and false positive rate of pedestrian detection. It can have stronger detection ability for multi-scale pedestrian targets and occluded pedestrian targets in complex crowded scenes, and can reduce the false negative rate of pedestrians.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and particularly provides a multi-scale pedestrian detection method and device based on the combination of head and overall information. Background Art

[0002] In the field of computer vision, pedestrian detection technology is one of the popular research directions with wide applications. The pedestrian detection task mainly refers to accurately detecting and positioning pedestrian target instances in image, video, and video stream data through computer vision-related technologies. This task is essentially a classification and regression process. In real life, pedestrian detection plays an important role and has high application value in fields such as autonomous driving, intelligent monitoring, intelligent robots, and human-computer interaction. Among them, in the fields of autonomous driving and assisted safe driving, pedestrian detection technology can be used to detect in real time whether there are pedestrians suddenly breaking into the front of the vehicle, and then make timely adjustments according to the actual situation to ensure pedestrian safety and driving safety. Especially in congested sections and crowded scenarios, the auxiliary role of pedestrian detection technology in autonomous driving safety is more obvious. In the field of intelligent monitoring, most public places nowadays use cameras to monitor the entire scene and statistically analyze the crowd flow data in real time. Especially during the epidemic, most large public places will strictly control the pedestrian flow density. Using pedestrian detection technology can accurately count the pedestrian flow in the venue in real time, and by analyzing and predicting these data, it helps the management to take corresponding adjustment measures. In the field of intelligent robots, sensors such as cameras transmit corresponding environmental scene signals to intelligent robots. The pedestrian detection algorithm, as an important thinking and perception network in the brain of intelligent robots, can help intelligent robots quickly and accurately perceive pedestrian targets and make corresponding decisions in a timely manner for adjustment. In the field of human-computer interaction, devices such as intelligent food delivery vehicles and intelligent express delivery vehicles in campuses and restaurants integrate multiple functions such as pedestrian detection algorithms, and achieve the purpose of better serving people's daily lives by interacting with pedestrians.

[0003] Based on the detection idea, pedestrian detection algorithms based on deep learning can be divided into single-stage pedestrian detection algorithms and two-stage pedestrian detection algorithms. Among them, the single-stage pedestrian detection algorithms are mainly represented by the YOLO series network models, and the two-stage pedestrian detection algorithms are represented by the R-CNN series network models. However, both of these two types of algorithms lack research and design for specific detection scenarios, and their detection capabilities for small-scale pedestrian targets and severely occluded and overlapping pedestrian targets at a relatively far viewing angle in dense scenarios are weak, mainly manifested in low detection accuracy and high miss rate. In addition, both of these two types of algorithms ignore an important issue: although the body parts of pedestrians are easily occluded and overlapped in complex and dense scenarios, their head regions are often less occluded or even completely unoccluded. Even if the body parts of pedestrians are severely occluded, the head regions can still provide important feature information, which is very important for detecting pedestrian targets in dense scenarios. However, the overall scale of the head region is small, and it is easy to generate ambiguities on hands, elbows, and surrounding small objects, thus causing false detections. It can be seen that it is difficult to achieve accurate pedestrian detection only relying on overall detection and head detection. It is necessary to effectively combine head detection and overall detection, and more rich and detailed multi-scale feature information needs to be obtained in the feature extraction link, so as to better play the respective advantages of head detection and overall detection and improve the accuracy of pedestrian detection. Therefore, it is very necessary to design a multi-scale pedestrian detection algorithm based on the combination of head and overall information, which is of great significance for improving the detection accuracy of occluded pedestrian targets and multi-scale pedestrian targets in dense scenarios.

[0004] In response to the above requirements, there are also many relevant solutions at home and abroad currently.

[0005] The invention patent application with the Chinese patent publication number CN111767882A, publication date of October 13, 2020, and patent name "A Multi-modal Pedestrian Detection Method Based on an Improved YOLO Model" improves the pedestrian detection effect by integrating the CBAM attention mechanism and optimizing the loss function. However, there are also some problems. For example, (1) in the feature extraction stage, the sensitivity to multi-scale pedestrian targets is not high enough, especially the ability to learn the feature information of small-scale pedestrians at a farther viewing angle is weak. (2) There is still room for improvement in the missed detection phenomenon of pedestrian targets with severe occlusion and overlap. The invention patent application with the Chinese patent publication number CN113989939A, publication date of January 28, 2022, and patent name "A Small Target Pedestrian Detection System Based on an Improved YOLO Algorithm". However, in actual detection scenarios, there will inevitably be crowded pedestrians, so the detection ability of this method for occluded pedestrian targets is weak, and there are certain limitations in actual applications. The invention patent application with the Chinese patent publication number CN114882527A, publication date of August 9, 2022, and patent name "A Pedestrian Detection Method and System Based on Dynamic Group Convolution" mainly uses group convolution to achieve pedestrian detection, but lacks consideration of specific pedestrian detection scenarios, and there are certain limitations in the execution efficiency in complex crowded scenarios.

[0006] For existing pedestrian detection technologies, more solutions rely on classic object detection algorithms to detect pedestrian targets in the scene. According to the detection ideas, they can be divided into single-stage detection algorithms represented by the YOLO series network models and two-stage detection algorithms represented by the Faster R-CNN network model. However, most classic object detection algorithms lack specific consideration for the detection scene. Especially in complex crowded scenarios, the robustness of these mainstream detection algorithms is affected, mainly including the following two points:

[0007] (1) The inconsistent scales of pedestrian targets affect the comprehensive performance of the detector. Since most current pedestrian detection datasets are obtained based on camera shooting and calibration processing, and there is a rule of "objects appear larger when closer and smaller when farther" when the camera shoots. Pedestrian targets at a closer viewing angle have a larger overall scale, while pedestrian targets at a farther viewing angle have a smaller overall scale. Since the resolution of small-scale pedestrian targets is relatively low, during the feature extraction process of the algorithm, problems such as limited learned feature information or weak feature expression ability are likely to occur, making it difficult to have high sensitivity to pedestrian targets of different scales, and thus prone to missed detection or false detection phenomena.

[0008] (2) The comprehensive performance of the detector is affected due to serious occlusion of pedestrian targets. In the pedestrian detection task in complex and crowded scenarios, pedestrian targets are often subject to a certain degree of occlusion. By analyzing the images in the pedestrian detection dataset, it can be seen that the pedestrian occlusion problem mainly includes two situations: intra-class occlusion and inter-class occlusion. Among them, intra-class occlusion refers to the occlusion between pedestrian targets. Inter-class occlusion refers to the interference of pedestrian targets by background information, which mainly includes buildings, trees, vehicles, items carried by pedestrians themselves, and items carried by other nearby pedestrians, etc. Intra-class occlusion and inter-class occlusion lead to a decrease in the proportion of the visible area of the whole body of pedestrians, which causes difficulties in the feature extraction link of the algorithm, and the information required for the inference of the algorithm detection module also decreases accordingly. In addition, it will also affect the accuracy of pedestrian target positioning, thereby affecting the comprehensive performance of the pedestrian detection algorithm. Summary of the Invention

[0009] To solve the above problems, the present invention provides a multi-scale pedestrian detection method and device based on the combination of head and overall information, which improves the detection accuracy of the detector for multi-scale pedestrian targets and occluded pedestrian targets in complex and crowded scenarios and reduces its missed detection rate.

[0010] In the first aspect, the present invention provides a multi-scale pedestrian detection method based on the combination of head and overall information, including:

[0011] Construct a Faster R-CNN network model;

[0012] Fuse the backbone network of the Faster R-CNN network model with the improved feature extraction network, and input the image to be detected into the Faster R-CNN model of the fused and improved feature extraction network for feature extraction to obtain the extracted feature map;

[0013] Improve the sampling method of the region proposal network, construct a non-uniform hard sample mining strategy based on occlusion overlap rate discrimination, calculate the occlusion overlap rate of each sample in the sample set, and assign a higher weight to the sample with a higher occlusion overlap rate, and use the region proposal network to simultaneously generate a set of head candidate boxes and overall candidate boxes for all pedestrian instances in the scene;

[0014] Construct a pedestrian head detection branch module and a pedestrian overall detection branch module, and obtain preliminary target detection results through the pedestrian head detection branch module and the pedestrian overall detection branch module. The preliminary target detection results include pedestrian head detection boxes and pedestrian overall detection boxes;

[0015] Post-process the obtained pedestrian head detection boxes and pedestrian overall detection boxes, and screen out redundant detection boxes generated during the joint detection process to obtain the final pedestrian detection results;

[0016] To further suppress the false detection and missed detection situations that occur in the post-processing link, a joint loss function is constructed in the loss function part. The joint loss function is used to punish the false detection situations that are not correctly non-maximum suppressed and the missed detection situations that are wrongly non-maximum suppressed, so that the final detection result is more accurate.

[0017] As a preferred solution, the backbone network of the Faster R-CNN network model is fused with the improved feature extraction network, and the image to be detected is input into the Faster R-CNN model of the fused improved feature extraction network for feature extraction to obtain the extracted feature map, including:

[0018] Taking the ResNet50 network as the backbone network of the Faster R-CNN network model;

[0019] Using the backbone network to learn the feature information of the image to be detected;

[0020] Performing feature fusion on the feature information obtained by the backbone network, combining the idea of dense connection, and improving the feature splicing strategy of the Feature Pyramid Network (FPN), that is, making the features of all scales participate in the calculation of the output feature information, absorbing the advantages of the large-scale feature map expressing high-level semantic information and the small-scale feature map expressing details such as underlying textures, and enhancing the network's fusion of multi-scale features. Then the predicted values corresponding to the features of each layer scale ~ Specifically, as shown in Formula 1 and Formula 2;

[0021] (1)

[0022] (2)

[0023] Among them, are all weight parameters, is the feature mapping function, ~ are the features of each layer scale obtained by the feature extraction network. In the specific training process, the parameter participates in the backpropagation of the gradient and is learned and updated to the most appropriate value through the training of the model.

[0024] As a preferred solution, the sampling method of the Region Proposal Network is improved, and a non-uniform hard sample mining strategy based on the occlusion overlap rate discrimination is constructed. By calculating the occlusion overlap rate of each sample in the sample set and assigning a higher weight to the sample with a higher occlusion overlap rate, the Region Proposal Network is used to simultaneously generate the head candidate box and the overall candidate box set for all pedestrian instances in the scene, including:

[0025] The non-uniform hard sample mining strategy based on occlusion overlap rate discrimination divides the sample set into two categories, i.e., the hard set ( ) and the normal set ( ) according to the average occlusion overlap rate ) by introducing a decision threshold . The probability of each sample being selected is defined as , and the specific calculation method is shown in Formula 3;

[0026] (3)

[0027] where represents the occlusion coefficient of the -th sample, which is used to reflect the degree of occlusion of the -th sample;

[0028] For the occlusion coefficient of the -th sample, it is further expressed as:

[0029] (4)

[0030] where represents the number of samples required for detection, represents the total number of candidate samples, and respectively represent the occlusion overlap rate of the -th sample in the sample set and the average occlusion overlap rate of the entire sample set. The decision threshold can be set to different values according to the overall occlusion degree of the data set;

[0031] For the occlusion overlap rate of the -th sample in the sample set and the average occlusion overlap rate of the entire sample set, they are further expressed as:

[0032] (5)

[0033] (6)

[0034] As a preferred solution, the post-processing of the obtained pedestrian head detection frame and the pedestrian overall detection frame, and the screening of redundant detection frames generated during the joint detection to obtain the final pedestrian detection result include:

[0035] For the head detection frame and the pedestrian overall detection frame obtained by the joint detection framework, penalty factors with decreasing confidence are introduced respectively. By gradually reducing the confidence scores of the overlapping detection frames, while not overly suppressing the overlapping frames, their competition is reduced. The Score of the overall pedestrian detection box and the score of the head detection box are expressed as:

[0036] (7)

[0037] (8)

[0038] (9)

[0039] (10)

[0040] Among them, and are respectively the overall box and the head box with the highest scores in the previous iteration, and are respectively the NMS thresholds corresponding to overall detection and head detection, is a small fixed value set initially;

[0041] The confidence scores of the obtained head boxes and overall boxes are weighted and summed to obtain the joint confidence score , which is specifically expressed as:

[0042] (11)

[0043] Among them, represents the weight of the head detection score. If the head overlap or overall overlap is greater than the preset threshold, the detection box with a lower confidence is suppressed.

[0044] As a preferred solution, to further suppress false detections and missed detections in the post-processing link, a joint loss function is constructed in the loss function part. The joint loss function is used to punish false detections that are not correctly non-maximum suppressed and missed detections that are wrongly non-maximum suppressed, making the final detection result more accurate, including:

[0045] The joint loss function is specifically expressed as:

[0046] (12)

[0047] Among them, the coefficients , , and are all weights for balancing the loss, and are used to pull the misdetected head boxes and overall boxes closer for elimination, and Used to push the undetected head bounding boxes and overall bounding boxes that are wrongly non-maximum suppressed far away, so that the head bounding boxes and overall bounding boxes correctly correspond to the target bounding boxes;

[0048] For the And Loss function, which is further expressed as:

[0049] (13)

[0050] (14)

[0051] For the And Loss function, which is further expressed as:

[0052] (15)

[0053] (16)

[0054] Wherein, And Are respectively the NMS thresholds corresponding to the head detection branch and the overall detection branch. After removing the redundant bounding boxes generated by the head detection branch and the overall detection branch, the final pedestrian detection result is obtained.

[0055] In a second aspect, the present invention provides a multi-scale pedestrian detection device based on the joint information of the head and the whole, including:

[0056] A construction unit for constructing a Faster R-CNN network model;

[0057] An extraction unit for fusing and improving the feature extraction network of the backbone network of the Faster R-CNN network model, and inputting the image to be detected into the Faster R-CNN model of the fused and improved feature extraction network for feature extraction to obtain an extracted feature map;

[0058] A generation unit for improving the sampling method of the region proposal network, constructing a non-uniform hard sample mining strategy based on the occlusion overlap rate discrimination, calculating the occlusion overlap rate of each sample in the sample set, and assigning a higher weight to the sample with a higher occlusion overlap rate, and using the region proposal network to simultaneously generate a set of head candidate bounding boxes and a set of overall candidate bounding boxes for all pedestrian instances in the scene;

[0059] The preliminary detection unit is used to construct a pedestrian head detection branch module and a pedestrian overall detection branch module, and obtain preliminary target detection results through the pedestrian head detection branch module and the pedestrian overall detection branch module. The preliminary target detection results include a pedestrian head detection box and a pedestrian overall detection box;

[0060] The post-processing unit is used to post-process the obtained pedestrian head detection box and the pedestrian overall detection box, and screen out redundant detection boxes generated during the joint detection process to obtain the final pedestrian detection result;

[0061] The output unit is used to further suppress false detections and missed detections that occur in the post-processing link. A joint loss function is constructed in the loss function part. The joint loss function is used to penalize false detections that are not correctly non-maximum suppressed and missed detections that are wrongly non-maximum suppressed, so that the final detection result is more accurate.

[0062] As a preferred solution, the extraction unit is specifically used for:

[0063] Taking the ResNet50 network as the backbone network of the Faster R-CNN network model;

[0064] Using the backbone network to learn the feature information of the image to be detected;

[0065] Performing feature fusion on the feature information obtained by the backbone network, combining the idea of dense connection, and improving the feature splicing strategy of the Feature Pyramid Network (FPN), that is, making all scales of features participate in calculating the output feature information, absorbing the advantages of large-scale feature maps expressing high-level semantic information and small-scale feature maps expressing low-level texture and other detailed information, and enhancing the network's fusion of multi-scale features. Then the predicted values corresponding to the feature scales of each layer ~ Are specifically shown in Formula 1 and Formula 2;

[0066] (1)

[0067] (2)

[0068] Among them, Are all weight parameters, Is a feature mapping function, ~ Are the feature scales of each layer obtained by the feature extraction network. During the specific training process, the parameter Participates in the backpropagation of the gradient and is learned and updated to the most appropriate value through the training of the model.

[0069] As a preferred solution, the generation unit is specifically used for:

[0070] The non-uniform hard sample mining strategy based on occlusion overlap rate discrimination divides the sample set into two categories, the hard set ( ) and the normal set ( ) according to the average occlusion overlap rate by introducing a decision threshold, and defines the probability of each sample being selected as , and the specific calculation method is shown in formula (3); ) by introducing a decision threshold, defines the probability of each sample being selected as , and the specific calculation method is as shown in formula 3;

[0071] (3)

[0072] Wherein, represents the occlusion coefficient of the th sample, which is used to reflect the occlusion degree of the th sample;

[0073] For the occlusion coefficient of the th sample, it is further expressed as:

[0074] (4)

[0075] Wherein, represents the number of samples required for detection, represents the total number of candidate samples, and respectively represent the occlusion overlap rate of the th sample in the sample set and the average occlusion overlap rate of the entire sample set, and the decision threshold can be set to different values according to the overall occlusion degree of the data set;

[0076] For the occlusion overlap rate of the th sample in the sample set and the average occlusion overlap rate of the entire sample set, they are further expressed as:

[0077] (5)

[0078] (6)

[0079] As a preferred solution, the post-processing unit is specifically configured to:

[0080] Introduce penalty factors with decreasing confidence for the head detection box and the pedestrian overall detection box obtained by the joint detection framework respectively. By gradually reducing the confidence scores of the overlapping detection boxes, while not overly suppressing the overlapping boxes, reduce their competition. The score of the th pedestrian overall detection box and the Score of a head detection box It is expressed as:

[0081] (7)

[0082] (8)

[0083] (9)

[0084] (10)

[0085] Among them, and are respectively the overall box and the head box with the highest scores in the previous iteration, and are respectively the NMS thresholds corresponding to the overall detection and the head detection, is a small fixed value set initially;

[0086] The confidence scores of the obtained head boxes and overall boxes are weighted and summed to obtain the combined confidence score , which is specifically expressed as:

[0087] (11)

[0088] Among them, represents the weight of the head detection score. If the head overlap or the overall overlap is greater than the preset threshold, the detection boxes with lower confidence are suppressed.

[0089] As a preferred solution, the output unit is specifically used for:

[0090] The combined loss function is specifically expressed as:

[0091] (12)

[0092] Among them, the coefficients 、 、 and are all weights for balancing the loss, and are used to pull the misdetected head boxes and overall boxes closer for elimination, and are used to push the undetected head boxes and overall boxes that are wrongly suppressed by non-maximum suppression farther away, so that the head boxes and overall boxes correctly correspond to the target boxes;

[0093] For the and used to push the undetected head boxes and overall boxes that are wrongly suppressed by non-maximum suppression farther awayThe loss function is further expressed as:

[0094] (13)

[0095] (14)

[0096] For the purpose of bringing the falsely detected head frame and the overall frame closer together for elimination and The loss function is further expressed as:

[0097] (15)

[0098] (16)

[0099] in, and are the NMS thresholds corresponding to the head detection branch and the overall detection branch respectively. After removing the redundant frames generated by the head detection branch and the overall detection branch, the final pedestrian detection result is obtained.

[0100] Compared with the prior art, the present invention can achieve the following beneficial effects:

[0101] In an embodiment of the present invention, a multi-scale pedestrian detection method and device based on the combination of head and overall information are proposed. First, by densely connecting features at different levels, that is, making all scale features participate in the calculation of the final output features, each level of features finally obtained can cover semantic information and detail texture information, thereby improving the network's sensitivity to multi-scale pedestrian targets; secondly, the present invention optimizes the sampling method of the region proposal network, by calculating the occlusion overlap rate of each sample in the sample set and assigning a higher weight to samples with a higher occlusion overlap rate, thereby strengthening the learning of severely occluded samples in the difficult sample set during training, thereby improving the model's detection ability for occluded pedestrian targets; then, the present invention constructs a joint detection framework for pedestrian heads and overall information, aiming to use head detection to assist pedestrian detection, thereby reducing the adverse effects of pedestrian body occlusion on detection. In addition, the post-processing link and loss function module are optimized to reduce the interference of adjacent pedestrian targets on detection, and at the same time improve the intelligence and rationality of screening out redundant frames, so that the detection frames in dense areas are not overly suppressed while eliminating the false detection frames generated by the two detection branches, thereby further reducing the missed detection rate and false detection rate of pedestrian detection. Therefore, compared with previous algorithms, the pedestrian detection algorithm provided by the present invention has a stronger detection capability for multi-scale pedestrian targets and occluded pedestrian targets in complex crowd-dense scenes, and can reduce the missed detection rate of pedestrian targets. BRIEF DESCRIPTION OF THE DRAWINGS

[0102] Figure 1It is a schematic flowchart of a multi-scale pedestrian detection method based on the combination of head and overall information provided by an embodiment of the present invention.

[0103] Figure 2 It is a schematic overall logic diagram of a multi-scale pedestrian detection method based on the combination of head and overall information provided by an embodiment of the present invention.

[0104] Figure 3 It is a schematic structural diagram of the ResNet50 network in a multi-scale pedestrian detection method based on the combination of head and overall information provided by an embodiment of the present invention.

[0105] Figure 4 It is a structural diagram of a feature extraction network in a multi-scale pedestrian detection method based on the combination of head and overall information provided by an embodiment of the present invention.

[0106] Figure 5 It is a pseudocode flowchart of a joint post-processing algorithm in a multi-scale pedestrian detection method based on the combination of head and overall information provided by an embodiment of the present invention.

[0107] Figure 6 It is a comprehensive performance comparison chart of a multi-scale pedestrian detection method based on the combination of head and overall information provided by an embodiment of the present invention and some mainstream detection algorithms on the CrowdHuman dataset.

[0108] Figure 7 It is a comprehensive performance comparison chart of a multi-scale pedestrian detection method based on the combination of head and overall information provided by an embodiment of the present invention and some mainstream detection algorithms on the CityPersons dataset.

[0109] Figure 8a It is the algorithm performance comparison result of a multi-scale pedestrian detection method based on the combination of head and overall information provided by an embodiment of the present invention and some mainstream detection algorithms in the TJU-Ped-campus subset.

[0110] Figure 8b It is the algorithm performance comparison result of a multi-scale pedestrian detection method based on the combination of head and overall information provided by an embodiment of the present invention and some mainstream detection algorithms in the TJU-Ped-traffic subset.

[0111] Figure 9 It is a structural block diagram of a multi-scale pedestrian detection device based on the combination of head and overall information provided by an embodiment of the present invention. Detailed implementation manners

[0112] In the following, embodiments of the present invention will be described with reference to the accompanying drawings. In the following description, the same modules are denoted by the same reference numerals. In the case of the same reference numerals, their names and functions are also the same. Therefore, their detailed descriptions will not be repeated.

[0113] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and do not constitute a limitation to the present invention.

[0114] Combined with Figure 1 As shown, an embodiment of the present invention provides a multi-scale pedestrian detection method based on the combination of head and overall information, including:

[0115] S101. Construct a Faster R-CNN network model;

[0116] S102. Integrate the backbone network of the Faster R-CNN network model with the improved feature extraction network, and input the image to be detected into the Faster R-CNN model of the integrated and improved feature extraction network for feature extraction to obtain an extracted feature map;

[0117] S103. Improve the sampling method of the region proposal network, construct a non-uniform hard sample mining strategy based on the occlusion overlap rate discrimination, calculate the occlusion overlap rate of each sample in the sample set, assign a higher weight to the sample with a higher occlusion overlap rate, and use the region proposal network to simultaneously generate a set of head candidate boxes and overall candidate boxes for all pedestrian instances in the scene;

[0118] S104. Construct a pedestrian head detection branch module and a pedestrian overall detection branch module, and obtain preliminary object detection results through the pedestrian head detection branch module and the pedestrian overall detection branch module. The preliminary object detection results include a pedestrian head detection box and a pedestrian overall detection box;

[0119] S105. Post-process the obtained pedestrian head detection box and the pedestrian overall detection box, and filter out redundant detection boxes generated during the joint detection process to obtain the final pedestrian detection result;

[0120] S106. To further suppress false detections and missed detections that occur in the post-processing link, a joint loss function is constructed in the loss function part. The joint loss function is used to punish false detections that are not correctly non-maximum suppressed and missed detections that are wrongly non-maximum suppressed, so that the final detection result is more accurate.

[0121] In S102, fusing the backbone network of the Faster R-CNN network model with the improved feature extraction network, and inputting the image to be detected into the Faster R-CNN model with the fused and improved feature extraction network for feature extraction to obtain the extracted feature map, including:

[0122] Taking the ResNet50 network as the backbone network of the Faster R-CNN network model;

[0123] Using the backbone network to learn the feature information of the image to be detected;

[0124] Performing feature fusion on the feature information obtained by the backbone network, combining the idea of dense connection, and improving the feature splicing strategy of the Feature Pyramid Network (FPN), that is, making all-scale features participate in calculating the output feature information, absorbing the advantages of large-scale feature maps expressing high-level semantic information and small-scale feature maps expressing low-level texture and other detailed information, and enhancing the network's fusion of multi-scale features. Then, the predicted values corresponding to the features of each layer scale ~ Specifically, as shown in Formula 1 and Formula 2;

[0125] (1)

[0126] (2)

[0127] Among them, are all weight parameters, is the feature mapping function, ~ are the features of each layer scale obtained by the feature extraction network. In the specific training process, the parameter participates in the backpropagation of the gradient and is learned and updated to the most appropriate value through the training of the model.

[0128] Furthermore, in S103, improving the sampling method of the Region Proposal Network, constructing a non-uniform hard sample mining strategy based on the occlusion overlap rate discrimination, by calculating the occlusion overlap rate of each sample in the sample set, and assigning a higher weight to the sample with a higher occlusion overlap rate, and using the Region Proposal Network to simultaneously generate the head candidate box and the overall candidate box set for all pedestrian instances in the scene, including:

[0129] The non-uniform hard sample mining strategy based on the occlusion overlap rate discrimination divides the sample set into two categories, a hard set ( ) and a normal set ( ) according to the average occlusion overlap rate ) by introducing a discrimination threshold . Defining the probability of each sample being extracted as , the specific calculation method is shown in Formula 3;

[0130] (3)

[0131] Among them, represents the occlusion coefficient of the th sample, which is used to reflect the degree of occlusion of the th sample;

[0132] For the occlusion coefficient of the th sample, it is further expressed as:

[0133] (4)

[0134] Among them, represents the number of samples required for detection, represents the total number of candidate samples, and respectively represent the occlusion overlap rate of the th sample in the sample set and the average occlusion overlap rate of the entire sample set. The determination threshold can be set to different values according to the overall occlusion degree of the data set;

[0135] For the occlusion overlap rate of the th sample in the sample set and the average occlusion overlap rate of the entire sample set, it is further expressed as:

[0136] (5)

[0137] (6)

[0138] Furthermore, in S105, the post-processing of the obtained pedestrian head detection frame and the pedestrian overall detection frame is performed, and the redundant detection frames generated during the joint detection process are screened out to obtain the final pedestrian detection result, including:

[0139] Penalty factors with decreasing confidence are introduced for the head detection frame and the pedestrian overall detection frame obtained by the joint detection framework. By gradually reducing the confidence scores of the overlapping detection frames, while not overly suppressing the overlapping frames, their competition is reduced. The score of the th pedestrian overall detection frame and the score of the

[0140] (7)

[0141] (8)

[0142] (9)

[0143] (10)

[0144] Among them, and are respectively the overall box and the head box with the highest scores in the previous iteration, and are respectively the NMS thresholds corresponding to the overall detection and the head detection, is a small fixed value set initially;

[0145] The confidence scores of the obtained head boxes and overall boxes are weighted and summed to obtain the joint confidence score , specifically expressed as:

[0146] (11)

[0147] Among them, represents the weight of the head detection score. If the head overlap or the overall overlap is greater than the preset threshold, the detection boxes with lower confidence are suppressed.

[0148] Furthermore, in S106, to further suppress the false detections and missed detections that occur in the post-processing link, a joint loss function is constructed in the loss function part. The joint loss function is used to punish the false detection cases that are not correctly non-maximum suppressed and the missed detection cases that are wrongly non-maximum suppressed, so that the final detection result is more accurate, including:

[0149] The joint loss function is specifically expressed as:

[0150] (12)

[0151] Among them, the coefficients , , and are all weights for balancing the loss. and are used to pull the misdetected head boxes and overall boxes closer for elimination. and are used to push the missed detection head boxes and overall boxes that are wrongly non-maximum suppressed farther away, so that the head boxes and overall boxes correctly correspond to the target boxes;

[0152] For and the loss function used to push the missed detection head boxes and overall boxes that are wrongly non-maximum suppressed farther away, it is further expressed as:

[0153] (13)

[0154] (14)

[0155] For the and loss function used to pull the misdetected head bounding boxes and overall bounding boxes closer for elimination, it is further expressed as:

[0156] (15)

[0157] (16)

[0158] Wherein, and are the NMS thresholds corresponding to the head detection branch and the overall detection branch respectively. After eliminating the redundant bounding boxes generated by the head detection branch and the overall detection branch, the final pedestrian detection result is obtained.

[0159] In the embodiment of the present invention, a multi-scale pedestrian detection method based on the joint of head and overall information is proposed. First, by densely connecting different-level features, that is, by making all scale features participate in the calculation of the final output feature, so that each level of feature obtained finally can cover semantic information and detailed texture information, thereby improving the sensitivity of the network to multi-scale pedestrian targets; Secondly, the present invention optimizes the sampling method of the region proposal network. By calculating the occlusion overlap rate of each sample in the sample set and assigning a higher weight to the sample with a higher occlusion overlap rate, the learning of the seriously occluded samples in the difficult sample set during training is strengthened, thereby improving the detection ability of the model for occluded pedestrian targets; Then, the present invention constructs a joint detection framework for pedestrian head and overall information, aiming to use head detection to assist pedestrian detection, thereby reducing the adverse impact of the occlusion of the pedestrian body on detection. And the post-processing link and the loss function module are optimized, weakening the interference caused by adjacent pedestrian targets to detection, and at the same time improving the intelligence and rationality of screening redundant bounding boxes, so that the detection bounding boxes in the dense area are not overly suppressed while eliminating the misdetected bounding boxes generated by the two detection branches, so as to further reduce the miss rate and false alarm rate of pedestrian detection. Therefore, compared with the previous algorithms, the pedestrian detection algorithm provided by the present invention has stronger detection ability for multi-scale pedestrian targets and occluded pedestrian targets in complex crowded scenes, and can reduce the miss rate of pedestrian targets.

[0160] Combined with Figure 2 、 3, as shown in FIGS. 4 and 5, the present invention addresses the problem that the existing object detection algorithms have a decrease in the detection accuracy of small-scale pedestrian targets and occluded pedestrian targets and a high false omission rate in complex and dense scenarios. To further facilitate the understanding of the solution of the present invention, the following describes a multi-scale pedestrian detection method based on the joint head and overall information under an embodiment of the present invention. The overall method flowchart is as Figure 2 shown, and the specific steps include:

[0161] Step 1: Construct a Faster R-CNN network model;

[0162] Step 2: Integrate the backbone network of the Faster R-CNN network model with the improved feature extraction network, and input the image to be detected into the Faster R-CNN network model that integrates the improved feature extraction network for feature extraction to obtain the extracted feature map;

[0163] Step 2.1: Use the ResNet50 network as the backbone network of the Faster R-CNN network model. The structure diagram of the ResNet50 network is as Figure 3 shown;

[0164] Step 2.2: Integrate the improved feature extraction network with the Faster R-CNN network model. The overall structure of the improved feature extraction network is as Figure 4 shown. The specific implementation steps are as follows:

[0165] Step 2.2.1: Use the backbone network ResNet50 to learn the feature information of the image to be detected;

[0166] Step 2.2.2: Perform feature fusion on the feature information obtained by the backbone network. Specifically, in combination with the idea of dense connection, improve the feature splicing strategy of the Feature Pyramid Network (FPN), that is, make all-scale features participate in the calculation of the output feature information, so as to simultaneously absorb the advantages of large-scale feature maps expressing high-level semantic information and small-scale feature maps expressing details such as underlying textures, and enhance the network's fusion of multi-scale features. Then the predicted values corresponding to the feature scales of each layer ~ are specifically shown in Formulas (1) and (2).

[0167] (1)

[0168] (2)

[0169] Among them, are all weight parameters, is the feature mapping function, ~ The scale features of each layer obtained by the feature extraction network. During the specific training process, the parameter participates in the backpropagation of the gradient and is updated to the most appropriate value through the training of the model. This method enables the model to learn which connections are helpful for the detection module to make predictions and which connections are ineffective computations, thereby improving the performance of the feature extraction network.

[0170] Step 3: Improve the sampling method of the Region Proposal Network (RPN), and construct a non-uniform hard sample mining strategy based on the occlusion overlap rate discrimination. By calculating the occlusion overlap rate of each sample in the sample set and assigning a higher weight to the sample with a higher occlusion overlap rate, the learning of the severely occluded samples in the hard sample set during training is strengthened, thereby improving the model's detection ability for occluded pedestrian targets. Then, the region proposal network is used to simultaneously generate the head candidate box and the overall candidate box set for all pedestrian instances in the scene.

[0171] Step 3.1: For the non-uniform hard sample mining strategy based on the occlusion overlap rate discrimination constructed in the present invention, by introducing a determination threshold , the sample set is divided into two categories: a hard set ( ) and a normal set ( ) according to its average occlusion overlap rate . Then, the probability of each sample being selected is defined as , and the specific calculation method is shown in formula (3).

[0172] (3)

[0173] where represents the occlusion coefficient of the th sample, which can reflect the degree of occlusion of the th sample.

[0174] Step 3.2: For the occlusion coefficient of the th sample in Step 3.1, it can be further expressed as:

[0175] (4)

[0176] where represents the number of samples required for detection, represents the total number of candidate samples. and respectively represent the occlusion overlap rate of the th sample in the sample set and the average occlusion overlap rate of the entire sample set, and the threshold Different values can be set according to the overall occlusion degree of the dataset.

[0177] Step 3.3: For the occlusion overlap rate of the th sample in the sample set in Step 3.2 and the average occlusion overlap rate of the entire sample set , it can be further expressed as:

[0178] (5)

[0179] (6)

[0180] Step 4: Construct a pedestrian head detection branch and a pedestrian overall detection branch, and obtain preliminary object detection results through the pedestrian head detection branch module and the pedestrian overall detection branch module.

[0181] Step 5: Post-process the obtained pedestrian head detection boxes and pedestrian overall detection boxes, and filter out redundant detection results generated during the joint detection process. The pseudo-code of the overall process is as Figure 5 shown. The specific execution steps are as follows:

[0182] Step 5.1: Introduce penalty factors with decreasing confidence levels for the head detection boxes and pedestrian overall detection boxes obtained from the joint detection framework respectively. By gradually reducing the confidence scores of overlapping detection boxes, while not overly suppressing the overlapping boxes, their competitiveness is reduced. The score of the th pedestrian overall detection box and the score of the

[0183] (7)

[0184] (8)

[0185] (9)

[0186] (10)

[0187] Among them, and are the overall box and the head box with the highest scores in the previous iteration respectively, and are the NMS thresholds corresponding to overall detection and head detection respectively, is a small fixed value set initially.

[0188] Step 5.2: Perform weighted summation on the confidence scores of the head boxes and overall boxes obtained in Step 5.1 to obtain the joint confidence score , which can be specifically expressed as:

[0189] (11)

[0190] Among them, represents the weight of the head detection score. If the head overlap or overall overlap is greater than the preset threshold, the detection box with a lower confidence will be suppressed.

[0191] Step 6: For the false detection and missed detection situations that occur in the non-maximum suppression process, the present invention constructs a combined loss function in the loss function part , aiming to punish the false detection situations that are not correctly non-maximum suppressed and the missed detection situations that are wrongly non-maximum suppressed. Specifically expressed as:

[0192] (12)

[0193] Among them, the coefficients , , and are all weights for balancing the loss. and are used to "pull closer" the false detected head boxes and overall boxes for easy elimination, and are used to "push away" the missed detected head boxes and overall boxes that are wrongly non-maximum suppressed, so that they can correctly correspond to the target boxes.

[0194] Step 6.1: For the and loss function used to "push away" the missed detected head boxes and overall boxes that are wrongly non-maximum suppressed, it can be further expressed as:

[0195] (13)

[0196] (14)

[0197] Step 6.2: For the and loss function used to "pull closer" the false detected head boxes and overall boxes for easy elimination, it can be further expressed as:

[0198] (15)

[0199] (16)

[0200] Among them, and They are the NMS thresholds corresponding to the head detection branch and the overall detection branch respectively. After removing the redundant bounding boxes generated by the head detection branch and the overall detection branch, the final pedestrian detection results are obtained.

[0201] The technical solutions involved in the present invention can be preferably applied to the fields such as intelligent monitoring, assisted safe driving, video monitoring in large places, and intelligent robots.

[0202] The effectiveness of the present invention can be verified and illustrated by the following experiments:

[0203] (I) Experimental Conditions and Experimental Contents

[0204] 1. The ImageNet dataset is used as the pre-training dataset for the pedestrian detection model, and then the pedestrian detection model is trained respectively based on the CrowdHuman, CityPersons, and TJU-DHD-pedestrian datasets. Some enhancement strategies are used for the datasets during the training process, including random cropping, random horizontal flipping, and random interference of image brightness. During the actual training process, the stochastic gradient descent optimizer is used for training for 35 epochs, the momentum factor momentum is set to 0.9, and the learning rate adopts the warm up strategy. The initial learning rate is set to a relatively low value, which is set to in the present invention. From the 1st iteration to the 800th iteration of each epoch, the learning rate linearly increases to and then remains unchanged. By the 24th epoch, the learning rate is reduced to 10% of the original, and by the 28th epoch, the learning rate becomes 1% of the original.

[0205] 2. The average precision (AP) and the logarithmic mean miss rate (MR -2 ) are used as evaluation metrics to evaluate the comprehensive performance of the method involved in the present invention. Among them, the specific calculation methods of the average precision and the logarithmic mean miss rate can be expressed as:

[0206]

[0207]

[0208] Among them, represents the accuracy rate, and its value is the ratio of the number of samples with the predicted result being a pedestrian target to the number of samples actually marked as pedestrian targets in the image; represents the recall rate, and its value is the proportion of the number of correctly predicted pedestrian targets to all predicted results during the pedestrian target detection process; represents the miss rate, represents the average number of false detections in each image. The higher the AP value and the lower the MR-2 The lower the value, the better the detection performance of the algorithm.

[0209] (2) Experimental results

[0210] As Figure 6 and 7 shown in Figure 8, on the CrowdHuman, CityPersons, and TJU-DHD-pedestrian pedestrian detection datasets, compared with some relatively popular pedestrian detection algorithms and the baseline Faster R-CNN algorithm before improvement, the average detection accuracy value of the present invention is higher and the logarithmic average miss rate value is lower. Therefore, the experimental results can fully demonstrate that the present invention is more excellent than some popular pedestrian detection algorithms in performance and can better detect multi-scale pedestrian targets and heavily occluded pedestrian targets.

[0211] The multi-scale pedestrian detection method provided by the present invention has the following characteristics compared with the prior art:

[0212] 1. Compared with the existing YOLO series network models and R-CNN series network models, the method involved in the present invention, in the feature extraction link, through the strategy of densely connecting different-level features and enabling all-scale features to participate in the calculation of the final output features, enables each level of the finally obtained features to cover semantic information and detailed texture information, thereby enhancing the network's detection ability for multi-scale pedestrian targets. Analyzed from the perspective of information flow, the feature extraction network constructed by the present invention can not only learn rich and detailed multi-scale features, but also fuse deep semantic information in shallow features and shallow detailed texture information in deep features, which is more helpful for improving the network's sensitivity to multi-scale pedestrian targets. Analyzed from the perspective of parameter optimization, when using the strategy of dense connection to fuse multi-scale features, during the actual training stage, the gradient can be more efficiently backpropagated, thereby accelerating the convergence speed of the network model without overly increasing the network complexity and being more conducive to obtaining a high-quality global optimal solution.

[0213] 2. Compared with the existing R-CNN series network models, the present invention optimizes the sampling method of the region proposal network. By calculating the occlusion overlap rate of each sample in the sample set and assigning a higher weight to the sample with a higher occlusion overlap rate, the learning of severely occluded samples in the difficult sample set during training is strengthened. And as the training progresses, the model will gradually improve its ability to detect occluded pedestrian targets, thus showing high performance in complex crowded scenes.

[0214] 3. Compared with the existing YOLO series network models and R-CNN series network models, the method involved in the present invention designs a dual detection branch of the head and the whole in the core part of the model for joint detection, thereby making full use of the head detection to assist pedestrian detection and reducing the impact of the occlusion of the pedestrian body on the performance of the detector.

[0215] 4. Compared with the existing YOLO series network models and R-CNN series network models, the method involved in the present invention optimizes the post-processing link and the loss function module, weakens the interference caused by adjacent pedestrian targets to detection, and at the same time improves the intelligence and rationality of screening redundant boxes, so that the detection boxes in dense areas are not overly suppressed while eliminating the misdetection boxes generated by the two detection branches, thereby further reducing the missed detection rate and misdetection rate of pedestrian detection.

[0216] Correspondingly, as shown in Figure 9 the present invention provides a multi-scale pedestrian detection device based on the joint information of the head and the whole in an embodiment, including:

[0217] A construction unit 901, configured to construct a Faster R-CNN network model;

[0218] An extraction unit 902, configured to fuse the backbone network of the Faster R-CNN network model with an improved feature extraction network, and input the image to be detected into the Faster R-CNN model of the fused and improved feature extraction network for feature extraction to obtain an extracted feature map;

[0219] A generation unit 903, configured to improve the sampling method of the region proposal network, construct a non-uniform hard sample mining strategy based on the occlusion overlap rate discrimination, calculate the occlusion overlap rate of each sample in the sample set, and assign a higher weight to the sample with a higher occlusion overlap rate, and use the region proposal network to simultaneously generate a set of head candidate boxes and a set of whole candidate boxes for all pedestrian instances in the scene;

[0220] A preliminary detection unit 904, configured to construct a pedestrian head detection branch module and a pedestrian whole detection branch module, and obtain a preliminary target detection result through the pedestrian head detection branch module and the pedestrian whole detection branch module, where the preliminary target detection result includes a pedestrian head detection box and a pedestrian whole detection box;

[0221] A post-processing unit 905, configured to perform post-processing on the obtained pedestrian head detection box and the pedestrian whole detection box, and screen out redundant detection boxes generated during the joint detection to obtain a final pedestrian detection result;

[0222] An output unit 906 is used to construct a joint loss function in the loss function part to further suppress false detections and missed detections in the post - processing link. The joint loss function is used to penalize false detections that are not correctly non - maximum suppressed and missed detections that are wrongly non - maximum suppressed, making the final detection result more accurate.

[0223] As a preferred solution, the extraction unit 902 is specifically configured to:

[0224] Take the ResNet50 network as the backbone network of the Faster R - CNN network model;

[0225] Use the backbone network to learn the feature information of the image to be detected;

[0226] Perform feature fusion on the feature information obtained by the backbone network. Combining the idea of dense connections, improve the feature splicing strategy of the Feature Pyramid Network (FPN), that is, make all scale features participate in calculating the output feature information, absorb the advantages of large - scale feature maps expressing high - level semantic information and small - scale feature maps expressing low - level texture and other detailed information, and enhance the network's fusion of multi - scale features. Then the predicted values corresponding to each layer scale feature ~ Are specifically shown in Formula 1 and Formula 2;

[0227] (1)

[0228] (2)

[0229] Where, Are all weight parameters, Is a feature mapping function, ~ Are the features of each layer scale obtained by the feature extraction network. In the specific training process, the parameter Participates in the backpropagation of the gradient and is updated to the most appropriate value through the training of the model.

[0230] As a preferred solution, the generation unit 903 is specifically configured to:

[0231] The non - uniform hard sample mining strategy based on the occlusion overlap rate discriminant divides the sample set into two categories, a hard set ( ) and a normal set ( ) according to the average occlusion overlap rate ) by introducing a discrimination threshold ). Define the probability of each sample being drawn as , and the specific calculation method is shown in Formula 3;

[0232] (3)

[0233] Among them, represents the occlusion coefficient of the -th sample, which is used to reflect the degree of occlusion of the -th sample;

[0234] For the occlusion coefficient of the -th sample, it is further expressed as:

[0235] (4)

[0236] Among them, represents the number of samples required for detection, represents the total number of candidate samples, and respectively represent the occlusion overlap rate of the -th sample in the sample set and the average occlusion overlap rate of the entire sample set. The determination threshold can be set to different values according to the overall occlusion degree of the data set;

[0237] For the occlusion overlap rate of the -th sample in the sample set and the average occlusion overlap rate of the entire sample set, it is further expressed as:

[0238] (5)

[0239] (6)

[0240] As a preferred solution, the post-processing unit 905 is specifically configured to:

[0241] Introduce penalty factors with decreasing confidence for the head detection box and the pedestrian overall detection box obtained by the joint detection framework. By gradually reducing the confidence scores of the overlapping detection boxes, while not overly suppressing the overlapping boxes, reduce their competition. The score of the -th pedestrian overall detection box and the score of the -th head detection box are expressed as:

[0242] (7)

[0243] (8)

[0244] (9)

[0245] (10)

[0246] Among them, and are respectively the overall box and the head box with the highest scores in the previous iteration, and are respectively the NMS thresholds corresponding to overall detection and head detection, is a small fixed value set initially;

[0247] The confidence scores of the obtained head boxes and overall boxes are weighted and summed to obtain the joint confidence score , specifically expressed as:

[0248] (11)

[0249] Among them, represents the weight of the head detection score. If the head overlap or overall overlap is greater than the preset threshold, the detection box with a lower confidence is suppressed.

[0250] As a preferred solution, the output unit 906 is specifically used for:

[0251] The joint loss function is specifically expressed as:

[0252] (12)

[0253] Among them, the coefficients , , and are all weights for balancing the loss. and are used to pull the misdetected head boxes and overall boxes closer for elimination. and are used to push the undetected head boxes and overall boxes that are wrongly suppressed by non-maximum suppression farther away, so that the head boxes and overall boxes correctly correspond to the target boxes;

[0254] For and the loss function used to push the undetected head boxes and overall boxes that are wrongly suppressed by non-maximum suppression farther away, it is further expressed as:

[0255] (13)

[0256] (14)

[0257] For and the loss function used to pull the misdetected head boxes and overall boxes closer for elimination, it is further expressed as:

[0258] (15)

[0259] (16)

[0260] Among them, and are the NMS thresholds corresponding to the head detection branch and the overall detection branch respectively. After removing the redundant bounding boxes generated by the head detection branch and the overall detection branch, the final pedestrian detection result is obtained.

[0261] In an embodiment of the present invention, a multi-scale pedestrian detection device based on the joint information of the head and the whole is proposed. First, by densely connecting different-level features, that is, by making all scale features participate in the calculation of the final output feature, semantic information and detailed texture information can be covered in each level of the finally obtained feature, thereby improving the sensitivity of the network to multi-scale pedestrian targets. Secondly, the present invention optimizes the sampling method of the region proposal network. By calculating the occlusion overlap rate of each sample in the sample set and assigning a higher weight to the sample with a higher occlusion overlap rate, the learning of severely occluded samples in the difficult sample set during training is strengthened, thereby improving the detection ability of the model for occluded pedestrian targets. Then, the present invention constructs a joint detection framework for pedestrian head and whole information, aiming to use head detection to assist pedestrian detection, so as to reduce the adverse impact of the occlusion of the pedestrian body on detection. And the post-processing link and the loss function module are optimized, weakening the interference caused by adjacent pedestrian targets to detection, and at the same time improving the intelligence and rationality of removing redundant bounding boxes, so that the detection bounding boxes in the dense area are not overly suppressed while removing the misdetected bounding boxes generated by the two detection branches, so as to further reduce the miss rate and false alarm rate of pedestrian detection. Therefore, compared with previous algorithms, the pedestrian detection algorithm provided by the present invention has stronger detection ability for multi-scale pedestrian targets and occluded pedestrian targets in complex crowded scenes, and can reduce the miss rate of pedestrian targets.

[0262] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

[0263] The above specific implementation manners of the present invention do not constitute a limitation to the protection scope of the present invention. Any other corresponding changes and deformations made according to the technical concept of the present invention should be included in the protection scope of the claims of the present invention.

Claims

1. A multi-scale pedestrian detection method based on the combination of head and overall information, characterized in that, Including: Constructing a Faster R-CNN network model; Fusing the backbone network of the Faster R-CNN network model with the improved feature extraction network, and inputting the image to be detected into the Faster R-CNN model with the fused and improved feature extraction network for feature extraction to obtain an extracted feature map, including: Taking the ResNet50 network as the backbone network of the Faster R-CNN network model; Using the backbone network to learn the feature information of the image to be detected; Perform feature fusion on the feature information obtained by the backbone network. Combining the idea of dense connection, improve the feature splicing strategy of the Feature Pyramid Network (FPN), that is, let the features of all scales participate in the calculation of the output feature information, absorb the advantages of large-scale feature maps expressing high-level semantic information and small-scale feature maps expressing underlying texture detail information, and enhance the network's fusion of multi-scale features. Then, the predicted values corresponding to the features of each layer scale ~ are specifically shown in Formula 1 and Formula 2; (1) (2) Among them, are all weight parameters, is a feature mapping function, ~ are the scale features of each layer obtained by the feature extraction network. During the specific training process, the parameter participates in the backpropagation of the gradient and is learned and updated to the most appropriate value through the training of the model; Improving the sampling method of the region proposal network, constructing a non-uniform hard sample mining strategy based on occlusion overlap rate discrimination, calculating the occlusion overlap rate of each sample in the sample set, and assigning a higher weight to the sample with a higher occlusion overlap rate, and using the region proposal network to simultaneously generate head candidate boxes and an overall candidate box set for all pedestrian instances in the scene, including: The non-uniform hard sample mining strategy based on occlusion overlap rate discriminant divides the sample set into a hard set and a normal set by introducing a decision threshold and according to the average occlusion overlap rate . The probability of each sample being extracted is defined as , and the specific calculation method is shown in Formula 3; (3) Among them, represents the occlusion coefficient of the th sample, which is used to reflect the degree of occlusion of the th sample; For the occlusion coefficient of the th sample, it is further expressed as: (4) Among them, represents the number of samples required for detection, represents the total number of candidate samples, and respectively represent the occlusion overlap rate of the th sample in the sample set and the average occlusion overlap rate of the entire sample set. The determination threshold can be set to different values according to the overall occlusion degree of the data set; For the occlusion overlap rate of the th sample in the sample set and the average occlusion overlap rate of the entire sample set , it is further expressed as: (5) (6); Constructing a pedestrian head detection branch module and a pedestrian overall detection branch module, and obtaining preliminary object detection results through the pedestrian head detection branch module and the pedestrian overall detection branch module, where the preliminary object detection results include pedestrian head detection boxes and pedestrian overall detection boxes; Performing post-processing on the obtained pedestrian head detection boxes and pedestrian overall detection boxes, and screening out redundant detection boxes generated during the joint detection process to obtain the final pedestrian detection results; To further suppress false detections and missed detections that occur in the post-processing link, a joint loss function is constructed in the loss function part. The joint loss function is used to penalize false detection cases that are not correctly suppressed by non-maximum suppression and missed detection cases that are wrongly suppressed by non-maximum suppression, making the final detection results more accurate.

2. The multi-scale pedestrian detection method based on the combination of head and overall information according to claim 1, wherein The performing post-processing on the obtained pedestrian head detection boxes and pedestrian overall detection boxes, and screening out redundant detection boxes generated during the joint detection process to obtain the final pedestrian detection results, including: For the head detection boxes and pedestrian overall detection boxes obtained by the joint detection framework, penalty factors with decreasing confidence are introduced respectively. By gradually reducing the confidence scores of overlapping detection boxes, while not overly suppressing the overlapping boxes, their competition is reduced. The score of the th pedestrian overall detection box and the score of the th head detection box are expressed as: (7) (8) (9) (10) Among them, and are respectively the overall box and the head box with the highest scores in the previous iteration, and are respectively the NMS thresholds corresponding to the overall detection and the head detection, is a fixed value set initially; Perform a weighted sum of the confidence scores of the obtained head box and the overall box to obtain a combined confidence score , specifically expressed as: (11) Among them, represents the weight of the head detection score. If the head overlap or overall overlap is greater than a preset threshold, the detection boxes with lower confidence are suppressed.

3. The multi-scale pedestrian detection method based on the combination of head and overall information according to claim 2, wherein The to further suppress false detections and missed detections that occur in the post-processing link, a joint loss function is constructed in the loss function part. The joint loss function is used to penalize false detection cases that are not correctly suppressed by non-maximum suppression and missed detection cases that are wrongly suppressed by non-maximum suppression, making the final detection results more accurate, including: The combined loss function Specifically expressed as: (12) Among them, the coefficients , , and are all weights for balancing losses, and are used to pull the misdetected head boxes and overall boxes closer for elimination, and are used to push the undetected head boxes and overall boxes that are wrongly suppressed by non-maximum suppression farther away, so that the head boxes and overall boxes correctly correspond to the target boxes; For the and loss function used to push away the undetected head boxes and overall boxes that are wrongly non-maximum suppressed, it is further expressed as: (13) (14) For the purpose of bringing the falsely detected head frame and the overall frame closer together for elimination and The loss function is further expressed as: (15) (16) Among them, and are the NMS thresholds corresponding to the head detection branch and the overall detection branch respectively. After removing the redundant bounding boxes generated by the head detection branch and the overall detection branch, the final pedestrian detection result is obtained.

4. A multi-scale pedestrian detection device based on the joint of head and overall information, characterized in that, Including: A construction unit for constructing a Faster R-CNN network model; An extraction unit for fusing the backbone network of the Faster R-CNN network model with the improved feature extraction network, and inputting the image to be detected into the Faster R-CNN model with the fused and improved feature extraction network for feature extraction to obtain an extracted feature map; the extraction unit is specifically used for: Taking the ResNet50 network as the backbone network of the Faster R-CNN network model; Using the backbone network to learn the feature information of the image to be detected; Perform feature fusion on the feature information obtained by the backbone network. Combining the idea of dense connection, improve the feature splicing strategy of the Feature Pyramid Network (FPN), that is, let the features of all scales participate in the calculation of the output feature information, absorb the advantages of large-scale feature maps expressing high-level semantic information and small-scale feature maps expressing low-level texture detail information, and enhance the network's fusion of multi-scale features. Then, the predicted values corresponding to the features of each layer scale ~ are specifically shown in Formula 1 and Formula 2; (1) (2) Among them, are all weight parameters, is a feature mapping function, ~ are the scale features of each layer obtained by the feature extraction network. During the specific training process, the parameter participates in the backpropagation of the gradient and is learned and updated to the most appropriate value through the training of the model; A generating unit, which is used to improve the sampling method of the region proposal network, construct a non-uniform hard sample mining strategy based on occlusion overlap rate discrimination, calculate the occlusion overlap rate of each sample in the sample set, assign higher weights to samples with higher occlusion overlap rates, and use the region proposal network to simultaneously generate head candidate boxes and a set of overall candidate boxes for all pedestrian instances in the scene; specifically, the generating unit is used for: The non-uniform hard sample mining strategy based on occlusion overlap rate discrimination divides the sample set into a hard set and a normal set by introducing a decision threshold and based on the average occlusion overlap rate . The probability of each sample being selected is defined as , and the specific calculation method is shown in Formula 3; (3) Among them, represents the occlusion coefficient of the th sample, which is used to reflect the degree of occlusion of the th sample; For the occlusion coefficient of the th sample, it is further expressed as: (4) Among them, represents the number of samples required for detection, represents the total number of candidate samples, and respectively represent the occlusion overlap rate of the th sample in the sample set and the average occlusion overlap rate of the entire sample set. The determination threshold can be set to different values according to the overall occlusion degree of the data set; For the occlusion overlap rate of the th sample in the sample set and the average occlusion overlap rate of the entire sample set , it is further expressed as: (5) (6) A preliminary detection unit, which is used to construct a pedestrian head detection branch module and a pedestrian overall detection branch module, and obtain preliminary object detection results through the pedestrian head detection branch module and the pedestrian overall detection branch module, where the preliminary object detection results include pedestrian head detection boxes and pedestrian overall detection boxes; A post-processing unit, which is used to post-process the obtained pedestrian head detection boxes and the pedestrian overall detection boxes, and filter out redundant detection boxes generated during the joint detection process to obtain the final pedestrian detection results; An output unit, which is used to further suppress false detections and missed detections that occur in the post-processing link. A joint loss function is constructed in the loss function part. The joint loss function is used to punish false detections that are not correctly non-maximum suppressed and missed detections that are wrongly non-maximum suppressed, so that the final detection results are more accurate.

5. The multi-scale pedestrian detection device based on the joint head and overall information according to claim 4, characterized in that Specifically, the post-processing unit is used for: For the head detection box and the pedestrian overall detection box obtained by the joint detection framework, penalty factors with decreasing confidence are introduced respectively. By gradually reducing the confidence scores of overlapping detection boxes, while not overly suppressing the overlapping boxes, their competition is reduced. The score of the th pedestrian overall detection box and the score of the th head detection box are expressed as: (7) (8) (9) (10) Among them, and are respectively the overall box and the head box with the highest scores in the previous iteration, and are respectively the NMS thresholds corresponding to the overall detection and the head detection, is a fixed value set initially; Perform a weighted sum of the confidence scores of the obtained head box and the overall box to obtain the joint confidence score , which is specifically expressed as: (11) Among them, represents the weight of the head detection score. If the head overlap or overall overlap is greater than a preset threshold, the detection boxes with lower confidence are suppressed.

6. The multi-scale pedestrian detection device based on the combination of head and overall information according to claim 5, wherein Specifically, the output unit is used for: The combined loss function Specifically expressed as: (12) Among them, the coefficients , , and are all weights for balancing losses, and are used to pull the misdetected head boxes and overall boxes closer for elimination, and are used to push the undetected head boxes and overall boxes that are wrongly suppressed by non-maximum suppression further away, so that the head boxes and overall boxes correctly correspond to the target boxes; For the and loss function used to push away the missed detection head boxes and overall boxes that are wrongly non-maximum suppressed, it is further expressed as: (13) (14) For the purpose of bringing the falsely detected head frame and the overall frame closer together for elimination and The loss function is further expressed as: (15) (16) Among them, and are the NMS thresholds corresponding to the head detection branch and the overall detection branch respectively. After removing the redundant bounding boxes generated by the head detection branch and the overall detection branch, the final pedestrian detection result is obtained.

Citation Information

Patent Citations

  • Multi-modal pedestrian detection method based on improved YOLO model

    CN111767882A

  • Small target pedestrian detection system based on improved YOLO algorithm

    CN113989939A

  • Pedestrian detection method and system based on dynamic grouping convolution

    CN114882527A

  • Traffic image automatic labeling method and device based on difficult sample mining

    CN117765534A

  • Pedestrian detection method and device capable of resisting shielding overlapping and scale change

    CN119027986A