A pedestrian detection method and system applicable to crowded scenarios
Through the improved pedestrian detection method, the improved model and mask technology are used to detect pedestrians in crowded scenarios, solving the problem of difficulty in feature extraction and NMS threshold setting, significantly reducing the missed detection rate and improving detection accuracy.
Patent Information
- Application Number
- CN202111515400.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-12-13
AI Technical Summary
In crowded scenarios, pedestrian detection has problems such as difficulty in feature extraction and NMS threshold setting, resulting in a high missed detection rate.
The improved pedestrian detection method is used to detect through a pre-trained improved model, and the pedestrian prediction box, instance segmentation diagram and number of human key points are obtained, pedestrian visibility is calculated and masked, and then the obstructed pedestrians are detected to bypass the NMS limitation.
The missed detection rate of pedestrian detection in crowded scenarios is significantly reduced, the feature extraction of crowded pedestrians is enhanced, the accuracy of detection results is improved, and the detection time is reduced compared to constructing all image masks.
Smart Images

Figure CN114170570B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a pedestrian detection method and system applicable to crowded scenarios, belonging to the technical field of image detection. Background Art
[0002] Pedestrian detection is a classic problem in the field of computer vision, which is characterized by wide application ranges such as unmanned driving, robots, intelligent monitoring, human behavior analysis, amblyopia assistance technology, etc. Traditional pedestrian detection methods mainly use HOG (Histogram of Oriented Gradient) to extract pedestrian features and then use SVM (Support Vector Machine) for classification. However, HOG can only describe pedestrian features from gradients or textures, with poor discriminative power. At the same time, SVM is also not suitable for the increasingly large pedestrian detection data sets. In recent years, with the development of deep convolutional neural networks, the accuracy of pedestrian detection has been greatly improved, but there are still difficulties in pedestrian detection in crowded scenarios.
[0003] There are mainly two difficulties in pedestrian detection in crowded scenarios. One is the high similarity between pedestrians. At present, the object detection models based on deep learning focus on extracting overall features, which will make it difficult for the model to distinguish highly overlapping pedestrians. The other is the limitation in the post-processing method of prediction boxes. For example, object detection frameworks such as Faster R-CNN, YOLOv3, and SSD sample on the feature map to generate dense prediction boxes. For a large number of prediction boxes, NMS (Non-Maximum Suppression) is used for screening. However, when applied to crowded pedestrian scenarios, it is very difficult to set the NMS threshold. If the NMS threshold is too low, a large number of missed detections will occur. If the NMS threshold is set too high, a large number of false detections will occur.
[0004] In practical applications, it is very common for group pedestrians to form crowded scenarios. Therefore, how to strengthen the feature extraction of crowded pedestrians and improve the limitations of NMS is of great significance for pedestrian detection in crowded scenarios, and it can also provide a basis for application fields such as intelligent monitoring and unmanned driving.
[0005] "Ps-rcnn: Detecting secondary human instances in a crowd via primary object suppression" published by Zheng Ge et al. in the 2020 "IEEE International Conference on Multimedia and Expo" first uses P-RCNN to detect less crowded pedestrians, and artificially constructs a mask to cover these pedestrians, and then uses S-RCNN to detect the remaining crowded targets (both P-RCNN and S-RCNN are based on the Faster-RCNN architecture). By constructing a mask, the model is forced to pay attention to crowded targets, but constructing masks for all detected images will significantly increase the detection time.
[0006] In "Adaptive nms: Refining pedestrian detection in a crowd" published by Songtao Liu et al. in the 2019 "Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition", a branch is added to the detection network to predict the density of each box, and the predicted density is used to replace the NMS threshold to achieve dynamic adjustment of the NMS threshold. However, there are still difficulties in density prediction itself, and it is still doubtful whether the density can represent the optimal NMS threshold setting. Moreover, the predicted boxes often do not exactly match the ground truth boxes, which will lead to the inconsistency between the IOU (Intersection-over-Union) between the predicted boxes and the predicted density, thus affecting the prediction results. Summary of the Invention
[0007] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a pedestrian detection method and system applicable to crowded scenarios, which can solve the problems of difficult pedestrian feature extraction and difficult NMS threshold setting in crowded scenarios, and effectively reduce the missed detection rate of pedestrian detection in crowded scenarios. To achieve the above purpose, the present invention is implemented by the following technical solutions:
[0008] In the first aspect, the present invention provides a pedestrian detection method applicable to crowded scenarios, and the method includes:
[0009] Obtain a to-be-detected image in a crowded scenario;
[0010] Input the obtained to-be-detected image into a pre-trained improved model for detection to obtain pedestrian prediction boxes, instance segmentation maps, and the number of human key points of each pedestrian;
[0011] Calculate the pedestrian visibility of the image to be detected based on the number of human key points of each pedestrian. For images with visibility less than the preset threshold, there is a phenomenon of mutual occlusion between pedestrians. Construct a mask on the image to be detected according to the instance segmentation map;
[0012] Input the image to be detected with the constructed mask into a pre-trained improved model for detection to obtain the prediction bounding boxes of occluded pedestrians;
[0013] Merge the pedestrian prediction bounding boxes and the prediction bounding boxes of occluded pedestrians, and output the pedestrian detection results.
[0014] Combined with the first aspect, further, the improved model is trained through the following steps:
[0015] Obtain a pedestrian dataset in a crowded scene with annotations, and construct pseudo instance segmentation annotations according to the head annotation bounding box information and the visible part annotation bounding box information of pedestrians in the pedestrian dataset;
[0016] Input the images in the crowded scene with annotations into the pre-constructed improved model to obtain the prediction training results;
[0017] Calculate the loss function between the prediction training results and the pseudo instance segmentation annotations, calculate the gradient using the backpropagation algorithm, and update the parameters of the pre-constructed improved model;
[0018] When the value of the loss function no longer continues to decrease, the training is completed, and the pre-trained improved model is obtained.
[0019] Combined with the first aspect, further, it also includes: before training the improved model, pre-train the improved model using the COCO human key point dataset so that the improved model has the ability to detect human key points.
[0020] Combined with the first aspect, further, the pre-constructed improved model includes: adding an SFPN module and an MKFRCNN module to the Mask R-CNN model;
[0021] The SFPN module is used to obtain the feature map and semantic segmentation map of the image to be detected;
[0022] The MKFRCNN module is used to obtain the pedestrian prediction bounding boxes, the corresponding instance segmentation map, and the human key points of each pedestrian according to the proposed bounding boxes.
[0023] Combined with the first aspect, further, the MKFRCNN module does not output the human key points of each pedestrian during the training of the improved model.
[0024] Combined with the first aspect, further, the loss function is a multi-task loss function, which is expressed by the following formula:
[0025] (1)
[0026] (2)
[0027] (3)
[0028] In formulas (1) to (3), Loss The multi-task loss function is composed of four parts; L cls is the classification loss of the prediction box, L box is the localization loss of the prediction box, L mask is the instance segmentation loss of each prediction box, L seg is the semantic segmentation loss; i is the index of the proposal box; p i is the prediction probability that the prediction box corresponding to the proposal box is a pedestrian. If the proposal box is marked as positive, it is 1, otherwise it is 0; is the offset of the proposal box relative to the ground truth box, t i is the offset of the prediction box corresponding to the proposal box relative to the ground truth box. The ground truth box refers to the position annotation box of pedestrians in the dataset.
[0029] Combined with the first aspect, further, the visibility of each pedestrian is calculated by the following formula:
[0030] (4)
[0031] In formula (4), N represents the number of detected pedestrians; k j represents the j th number of human key points detected for a pedestrian; K represents the number of annotations of human key points in the dataset used to train human key points; represents the visibility of each pedestrian. The detection result is the score of each key point. If the score of a certain key point is greater than 0, the key point is detected successfully. The number of annotations of human key points in different datasets used to train human key points is different.
[0032] In the second aspect, the present invention provides a pedestrian detection system applicable to crowded scenarios, including:
[0033] An acquisition module: used to acquire the image to be detected in a crowded scenario;
[0034] The first prediction module: It is used to input the acquired image to be detected into a pre-trained improved model for detection, and obtain pedestrian prediction boxes, instance segmentation maps, and the number of human key points for each pedestrian.
[0035] The processing module: It is used to calculate the visibility of pedestrians in the image to be detected according to the number of human key points for each pedestrian. For an image with a visibility less than a preset threshold, there is a phenomenon of mutual occlusion among pedestrians. A mask is constructed on the image to be detected according to the instance segmentation map.
[0036] The second prediction module: It inputs the image to be detected with the constructed mask into a pre-trained improved model for detection, and obtains the prediction boxes of occluded pedestrians.
[0037] The output module: It is used to merge the pedestrian prediction boxes and the prediction boxes of occluded pedestrians, and output the pedestrian detection results.
[0038] In a third aspect, the present invention provides a computer device, including a processor and a storage medium;
[0039] The storage medium is used to store instructions;
[0040] The processor is used to operate according to the instructions to execute the steps of the method described in the first aspect.
[0041] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method described in the first aspect.
[0042] Compared with the prior art, the beneficial effects achieved by the pedestrian detection method and system applicable to crowded scenes provided by the embodiments of the present invention include:
[0043] The present invention acquires an image to be detected in a crowded scene; inputs the acquired image to be detected into a pre-trained improved model for detection, and obtains pedestrian prediction boxes, instance segmentation maps, and the number of human key points for each pedestrian; calculates the visibility of pedestrians in the image to be detected according to the number of human key points for each pedestrian. For an image with a visibility less than a preset threshold, there is a phenomenon of mutual occlusion among pedestrians. A mask is constructed on the image to be detected according to the instance segmentation map; inputs the image to be detected with the constructed mask into a pre-trained improved model for detection, and obtains the prediction boxes of occluded pedestrians. After constructing the mask, the present invention inputs it into the improved model again, bypassing the limitations of NMS, and can detect occluded pedestrians and pedestrians filtered out due to not meeting the NMS threshold, significantly reducing the missed detection rate in crowded crowds; the present invention constructs a mask for some images according to the instance segmentation map, which can greatly reduce the detection time compared with constructing masks for all images.
[0044] Merge the predicted bounding boxes of pedestrians and the predicted bounding boxes of occluded pedestrians, and output the pedestrian detection results; the present invention strengthens the feature extraction of crowded pedestrians, and the detection results are accurate. Description of the Drawings
[0045] Figure 1 is a flowchart of a pedestrian detection method applicable to crowded scenarios provided in Embodiment 1 of the present invention;
[0046] Figure 2 is a schematic diagram of the overall improved model in a pedestrian detection method applicable to crowded scenarios provided in Embodiment 1 of the present invention;
[0047] Figure 3 is a schematic diagram of pseudo instance segmentation annotation in a pedestrian detection method applicable to crowded scenarios provided in Embodiment 1 of the present invention;
[0048] Figure 4 is a schematic diagram of the SFPN module of a pedestrian detection method applicable to crowded scenarios provided in Embodiment 1 of the present invention;
[0049] Figure 5 is a schematic diagram of the MKFRCNN module of a pedestrian detection method applicable to crowded scenarios provided in Embodiment 1 of the present invention. Detailed Embodiment
[0050] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present invention.
[0051] Embodiment 1:
[0052] As Figure 1 , the embodiment of the present invention provides a pedestrian detection method applicable to crowded scenarios, including: training of the improved model and application of the improved model.
[0053] The training of the improved model occurs before the application of the improved model, and its function is to iteratively train the improved model through the backpropagation algorithm to strengthen the ability of the improved model to extract pedestrian features in crowded scenarios.
[0054] The training of the improved model includes:
[0055] Obtain a labeled pedestrian dataset in a crowded scenario, and construct pseudo instance segmentation annotation according to the head annotation bounding box information and pedestrian visible part annotation bounding box information in the pedestrian dataset;
[0056] Input the labeled images in the crowded scenario into the pre-constructed improved model to obtain the predicted training results;
[0057] Calculate the loss function between the predicted training result and the pseudo-instance segmentation annotation, calculate the gradient using the backpropagation algorithm, and update the parameters of the pre-constructed improved model;
[0058] When the value of the loss function no longer continues to decrease, the training is completed, and the pre-trained improved model is obtained.
[0059] The specific steps include:
[0060] Step 1: Obtain the labeled pedestrian dataset in crowded scenes, and construct pseudo-instance segmentation annotations according to the head annotation box information and the visible body part annotation box information in the pedestrian dataset.
[0061] Since there is no human keypoint annotation in the crowded pedestrian dataset, in order to enable the model to have the ability to detect human keypoints, first pre-train the improved model using the COCO human keypoint dataset, so that the improved model has the ability to detect human keypoints.
[0062] The labeled pedestrian dataset in crowded scenes can be the CrowdHuman dataset.
[0063] As Figure 3 shown in the schematic diagram of pseudo-instance segmentation annotation, since there is no pixel-level annotation in the pedestrian dataset used for training and the pixel-level annotation cost is high, by combining the head annotation box information of pedestrians and the visible body part annotation box information of pedestrians to construct pseudo-instance segmentation annotations, the annotation cost can be significantly reduced, and the ability of the model to extract pedestrian edge features can also be improved.
[0064] Constructing pseudo-instance segmentation annotations includes: assuming that the upper left corner coordinates, length, and width of the head annotation box of a certain pedestrian are , and the upper left corner coordinates, length, and width of the visible body part annotation box of the pedestrian are . The polygon constructed by using these four coordinates to construct eight coordinates is the pseudo-instance segmentation annotation, and the horizontal and vertical coordinates are represented by P and Q respectively. The calculation process is as follows:
[0065] (1)
[0066] It should be noted that after the pseudo-instance segmentation map is labeled, a pseudo-semantic segmentation map can also be obtained. The difference lies in the pixel values of the segmented parts. Instance segmentation assigns different pixel values to each instance, and semantic segmentation assigns the same pixel values to the targets belonging to a certain category.
[0067] Step 2: Input the images in the labeled crowded scenes into the pre-constructed improved model to obtain the predicted training result.
[0068] The pre - constructed improved model includes adding an SFPN module and an MKFRCNN module to the Mask R - CNN model.
[0069] As Figure 4 shown, the SFPN module is used to extract pedestrian features to obtain the feature map of the image to be tested and generate a semantic segmentation map.
[0070] The specific meaning of SFPN is the Feature Pyramid Network with a semantic segmentation branch added, which is an extension of the FPN proposed in 2017. Since the FPN structure is similar to the encoding - decoding structure of the classic semantic segmentation network U - Net, it can be conveniently extended with a semantic segmentation branch.
[0071] As Figure 4 shown, Figure 4 the numbers above each bar chart in are the number of channels. First, select ResNet50 pre - trained on the ImageNet dataset as the basic network structure, extract the feature maps obtained after conv1 7 convolution and the feature maps output by the last group of residual blocks in each layer of conv2, conv3, conv4, conv5, and name them C1, C2, C3, C4, C5 respectively. Then, first perform convolution on C5 to get M5, upsample M5 (using bilinear interpolation) to the same resolution as C4 and then add C4 after convolution to get M4, and so on to get M3, M2. Then, pass M5, M4, M3, M2 through convolution to obtain P5, P4, P3, P2 feature maps. The feature maps are used to generate proposal boxes in the RPN (Region Proposal Network) stage. RPN is the region proposal network proposed in 2015, which can generate proposal boxes in an end - to - end form. The establishment of the semantic segmentation branch starts from P2. First, upsample P2 to get S1, then perform convolution on S1 and then pass through the Relu activation function to get S2 with the same number of channels as C1. Introducing the Relu activation function increases the non - linear fitting ability of the model and speeds up model convergence. Then, add S2 and C1 and perform convolution for feature aggregation to get S3, and finally obtain the probability distribution map through the Sigmoid function. Here, we do not first perform convolution on C1 to expand the number of channels to 256 and then add S1 because this method occupies more video memory during the backpropagation calculation of gradients and does not significantly improve the detection performance. The structure of the present invention can reduce the video memory usage and save computing resources.
[0072] As Figure 5As shown, the MKFRCNN module is used to obtain pedestrian prediction boxes, corresponding instance segmentation maps, and human key points of each pedestrian according to the proposed boxes.
[0073] Figure 5 It is a schematic structural diagram of the MKFRCNN module, which has a total of three branches: Box, Mask, and Keypoint, predicting the position of pedestrians, instance segmentation maps, and human key points respectively. The numbers within the square patterns represent the resolution and number of channels, such as indicating that the resolution of the feature map is and the number of channels is 256. The numbers within the rectangular patterns represent the number of nodes in the fully connected layer. The numbers on the arrows represent the size of the convolutional kernel and the number of convolutional operations. For example, indicating 4 times of convolution. K represents the number of human key points to be detected, which is determined by the annotation of the pre-trained dataset. During training, only the Box and Mask branches are enabled. During testing, all three branches are enabled, but after constructing the binary mask, the Mask and Keypoint branches need to be closed to improve the detection speed.
[0074] The present invention improves the instance segmentation branch of MKFRCNN. The upsampling method in the instance segmentation branch is changed from the original transposed convolution to first performing bilinear interpolation and then aggregating features through convolution. This is because the pattern of the pseudo-instance segmentation annotation used for training is relatively fixed, and using transposed convolution may cause overfitting, thus affecting the detection performance. It is easier to retain the spatial structure of the target through the bilinear interpolation method.
[0075] Step 3: Calculate the loss function between the predicted training result and the pseudo-instance segmentation annotation, calculate the gradient using the backpropagation algorithm, and update the parameters of the pre-constructed improved model.
[0076] The loss function is composed of classification loss, bounding box regression loss, instance segmentation loss, and semantic segmentation loss. Among them, the classification loss, instance segmentation loss, and semantic segmentation loss all use the cross-entropy loss function. The difference lies in that the objects for calculating the loss are the image category scores and pixel category scores respectively. The loss function is a multi-task loss function, which is expressed by the following formula:
[0077] (2)
[0078] (3)
[0079] (4)
[0080] In formulas (2) to (4), Loss The multi-task loss function is composed of four parts; L clsis the classification loss of the prediction box, L box is the localization loss of the prediction box, L mask is the instance segmentation loss for each prediction box, L seg is the semantic segmentation loss; i is the index of the proposal box; p i is the prediction probability that the prediction box corresponding to the proposal box is a pedestrian. If the proposal box is marked as positive, it is 1, otherwise it is 0; is the offset of the proposal box relative to the ground truth box, t i is the offset of the prediction box corresponding to the proposal box relative to the ground truth box. The ground truth box refers to the position annotation box of pedestrians in the dataset.
[0081] Step 5: When the value of the loss function no longer decreases, the training is completed, and a pre-trained improved model is obtained.
[0082] As Figure 1 shown, the application of the improved model includes:
[0083] Obtain the image to be detected in a crowded scene;
[0084] Input the obtained image to be detected into the pre-trained improved model for detection to obtain pedestrian prediction boxes, instance segmentation maps, and the number of human key points for each pedestrian;
[0085] Calculate the visibility of pedestrians in the image to be detected according to the number of human key points for each pedestrian. If the visibility is less than a preset threshold, there is a phenomenon of mutual occlusion among pedestrians in the image. Construct a mask on the image to be detected according to the instance segmentation map;
[0086] Input the image to be detected with the constructed mask into the pre-trained improved model for detection to obtain the prediction boxes of occluded pedestrians;
[0087] Merge the pedestrian prediction boxes and the prediction boxes of occluded pedestrians, and output the pedestrian detection results.
[0088] Among them, the visibility of each pedestrian is calculated by the following formula:
[0089] (5)
[0090] In formula (5), N represents the number of detected pedestrians; k i represents j the number of human key points detected for the KIndicates the number of annotations for human key points in the dataset used for training human key points; Indicates the visibility of each pedestrian. The detection result is the score of each key point. If the score of a certain key point is greater than 0, the detection of that key point is successful. The number of annotations for human key points in different datasets used for training human key points is different.
[0091] The SFPN and MKFRCNN modules constructed in the present invention effectively enhance the feature extraction of crowded pedestrians, and screen images with a high pedestrian density according to the rule of estimating the body visibility of pedestrians in the image based on human key points. Thus, after adding a binary mask, the image can be input into the detection network again to detect pedestrians who are occluded or filtered out due to not meeting the NMS threshold, significantly reducing the missed detection rate in crowded crowds.
[0092] Embodiment 2:
[0093] The embodiment of the present invention provides a pedestrian detection system applicable to crowded scenes, including:
[0094] An acquisition module: used to acquire an image to be detected in a crowded scene;
[0095] A first prediction module: used to input the acquired image to be detected into a pre-trained improved model for detection, and obtain pedestrian prediction frames, an instance segmentation map, and the number of human key points of each pedestrian;
[0096] A processing module: used to calculate the visibility of pedestrians in the image to be detected according to the number of human key points of each pedestrian. For images with a visibility less than a preset threshold, there is a phenomenon of mutual occlusion between pedestrians, and a mask is constructed on the image to be detected according to the instance segmentation map;
[0097] A second prediction module: input the image to be detected with the constructed mask into a pre-trained improved model for detection, and obtain the prediction frames of occluded pedestrians;
[0098] An output module: used to merge the pedestrian prediction frames and the prediction frames of occluded pedestrians, and output the pedestrian detection result.
[0099] Embodiment 3:
[0100] The embodiment of the present invention provides a computer device, including a processor and a storage medium;
[0101] The storage medium is used to store instructions;
[0102] The processor is used to operate according to the instructions to execute the steps of the method described in Embodiment 1.
[0103] Embodiment 4:
[0104] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in Embodiment 1 are implemented.
[0105] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0106] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0107] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0109] The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A pedestrian detection method applicable to crowded scenarios, characterized in that, Including: Obtain a to-be-detected image in a crowded scene; Input the obtained to-be-detected image into a pre-trained improved model for detection to obtain pedestrian prediction boxes, an instance segmentation map, and the number of human key points of each pedestrian; wherein, the improved model includes: adding an SFPN module and an MKFRCNN module to the Mask R-CNN model; The SFPN module is used to obtain a feature map and a semantic segmentation map of the to-be-detected image; The MKFRCNN module is used to obtain pedestrian prediction boxes, corresponding instance segmentation maps, and human key points of each pedestrian according to proposal boxes; The Mask R-CNN model is a mask region convolutional neural network, the SFPN module is a feature pyramid network module with a semantic segmentation branch added, and the MKFRCNN module is a mask and key point fast region convolutional neural network module; Calculate the pedestrian visibility of the to-be-detected image according to the number of human key points of each pedestrian. If the visibility is less than a preset threshold, there is a phenomenon of mutual occlusion between pedestrians in the image. Construct a mask on the to-be-detected image according to the instance segmentation map; wherein, calculate the visibility of each pedestrian through the following formula: (4) In formula (4), N represents the number of detected pedestrians; k j represents the j number of human key points detected for the K n-th pedestrian; represents the number of annotations of human key points in the dataset for training human key points; represents the visibility of each pedestrian. The detection result is the score of each key point. If the score of a certain key point is greater than 0, the key point is successfully detected. The number of annotations of human key points in different datasets for training human key points is different. Input the to-be-detected image with the constructed mask into the pre-trained improved model for detection to obtain prediction boxes of occluded pedestrians; Merge the pedestrian prediction boxes and the prediction boxes of occluded pedestrians, and output the pedestrian detection result.
2. The pedestrian detection method applicable to crowded scenarios according to claim 1, wherein, The improved model is trained through the following steps: Obtain a pedestrian dataset in a crowded scene with annotations, and construct pseudo instance segmentation annotations according to the head annotation box information and pedestrian visible part annotation box information in the pedestrian dataset; Input the images in the crowded scene with annotations into the pre-constructed improved model to obtain prediction training results; Calculate the loss function between the prediction training results and the pseudo instance segmentation annotations, calculate the gradient using the backpropagation algorithm, and update the parameters of the pre-constructed improved model; When the value of the loss function no longer continues to decrease, the training is completed, and a pre-trained improved model is obtained.
3. The pedestrian detection method applicable to crowded scenarios according to claim 2, wherein Also including: Before training the improved model, pre-train the improved model using the COCO human key point dataset so that the improved model has the ability to detect human key points.
4. The pedestrian detection method applicable to crowded scenarios according to claim 3, wherein The MKFRCNN module does not output the human key points of each pedestrian when training the improved model.
5. The pedestrian detection method applicable to crowded scenarios according to claim 2, wherein The loss function is a multi-task loss function, which is represented by the following formula: (1) (2) (3) In formulas (1) to (3), Loss The multi-task loss function consists of four parts; L cls is the classification loss of the prediction box, L box is the localization loss of the prediction box, L mask is the instance segmentation loss of each prediction box, L seg is the semantic segmentation loss; i is the index of the proposal box; p i is the prediction probability that the prediction box corresponding to the proposal box is a pedestrian. If the proposal box is labeled positive, it is 1, otherwise it is 0; is the offset of the proposal box relative to the ground truth box, t i is the offset of the prediction box corresponding to the proposal box relative to the ground truth box. The ground truth box refers to the position annotation box of pedestrians in the dataset, is the prediction probability that the proposal box is a pedestrian.
6. A pedestrian detection system applicable to crowded scenarios, characterized in that, Including: An acquisition module: used to obtain a to-be-detected image in a crowded scene; A first prediction module: used to input the obtained to-be-detected image into a pre-trained improved model for detection to obtain pedestrian prediction boxes, an instance segmentation map, and the number of human key points of each pedestrian; wherein, the improved model includes: adding an SFPN module and an MKFRCNN module to the Mask R-CNN model; The SFPN module is used to obtain a feature map and a semantic segmentation map of the to-be-detected image; The MKFRCNN module is used to obtain pedestrian prediction boxes, corresponding instance segmentation maps, and human key points of each pedestrian according to proposal boxes; The Mask R-CNN model is a Mask Region Convolutional Neural Network, the SFPN module is a Feature Pyramid Network module with a semantic segmentation branch added, and the MKFRCNN module is a Mask and Keypoint Fast Region Convolutional Neural Network module; Processing module: used to calculate the visibility of pedestrians in the image to be detected according to the number of human keypoints of each pedestrian. If the visibility of an image is less than a preset threshold, there is a phenomenon of mutual occlusion between pedestrians. A mask is constructed on the image to be detected according to the instance segmentation map; among them, the visibility of each pedestrian is calculated through the following formula: (4) In formula (4), N represents the number of detected pedestrians; k j represents the j number of human key points detected for the K nth pedestrian; represents the number of annotations of human key points in the dataset used to train human key points; represents the visibility of each pedestrian. The detection result is the score of each key point. If the score of a certain key point is greater than 0, the key point is successfully detected. The number of annotations of human key points in different datasets used to train human key points is different; Second prediction module: Input the image to be detected with the constructed mask into a pre-trained improved model for detection to obtain the prediction box of the occluded pedestrian; Output module: used to merge the pedestrian prediction box and the prediction box of the occluded pedestrian and output the pedestrian detection result.
7. A computer device, characterized in that, It includes a processor and a storage medium; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method described in any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Instance segmentation method based on key points
CN111507334A