A Pedestrian Search Method Based on Dynamic RoI Feature Extraction

By adopting dynamic RoI feature extraction method in pedestrian search, RoI features are adaptively extracted and geometric transformation is processed, which solves the problems of intricate positioning and large calculations in the prior art, and improves the accuracy and performance indicators of pedestrian search.

CN116052214BActive Publication Date: 2025-06-27TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310059406.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-18
Publication Date
2025-06-27
Estimated Expiration
2043-01-18

AI Technical Summary

Technical Problem

When the existing pedestrian search methods deal with problems such as pedestrian posture changes, perspective changes, and occlusion, there are problems such as intricate positioning and large calculations.

Method used

Using the dynamic RoI feature extraction method, the RoI features are adaptively extracted by generating offsets on the RoI region, adding an internal mechanism for processing geometric transformation.

Benefits of technology

The accuracy of pedestrian search and Top-1 performance indicators are improved, which can better deal with pedestrian posture changes, perspective changes, occlusion and other problems, and the calculation amount increases less.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116052214B_ABST
    Figure CN116052214B_ABST
Patent Text Reader

Abstract

The present invention relates to a pedestrian search method based on dynamic RoI feature extraction. The pedestrian search network adopted uses a one-step network framework based on candidate box generation, including a backbone network, a neck network, and a head network. The head network is composed of an object detection head network and a pedestrian re-identification head network connected in series; the head network adopts a multi-stage cascade architecture, that is, the output result of the previous-level network is used as the input of the next-level network; the input of each level of the head network includes the feature map corresponding to the neck network, candidate box position information, and candidate box feature vectors, and the following steps are included: preparing an image set containing different pedestrians, annotating the annotation information of pedestrians in each image of the image set, including the identity information and annotation box information of pedestrians; dividing the image set into a training set, a validation set, and a test set; setting relevant hyperparameters in the training stage; and training the pedestrian search network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a robust and effective pedestrian search method in the field of computer vision such as pedestrian tracking, intelligent video surveillance, and intelligent transportation, and specifically relates to a pedestrian search method based on a deep convolutional neural network. Background Art

[0002] The task of pedestrian search is to determine whether there are pedestrians in an image library or a video sequence and locate and identify the target person. Pedestrian search has a very wide range of applications in the field of computer applications, such as intelligent video surveillance, aerial images, human-computer interaction systems, motion analysis, etc. Figure 1 An example of the application of pedestrian search in an intelligent monitoring system is given. As Figure 1 shown are the pictures taken by a monitoring camera at different times and perspectives. The boxes in the first column of the figures represent the target pedestrians to be searched, and the second and third columns are the images to be searched. The intelligent monitoring system needs to accurately find and locate the target pedestrians from the set of images to be searched (the boxes in the second column represent the found target pedestrians). Since pedestrians are easily affected by clothing, scale, occlusion, posture, and perspective, etc., pedestrian search has become a research topic with both research value and great challenges in the field of computer vision.

[0003] As a joint task of object detection and re-identification (re-id), pedestrian search not only needs to handle the challenges existing in these two separate subtasks, but also needs to jointly optimize the different objectives of the two subtasks. Existing person search methods can be mainly divided into two-step methods and one-step methods. The two-step methods [1, 2, 3] respectively use two independent deep convolutional neural networks for detection and re-identification. In the first step, a detection deep convolutional neural network is used to detect people from the image. Then, another deep convolutional neural network is used for re-identification based on the cropped people. The one-step methods [4, 5] aim to perform object detection and re-identification in a single unified deep convolutional neural network. The one-step method based on candidate box generation [6, 7] is a representative method with advanced performance, usually based on modern object detection frameworks such as Faster R-CNN [8]. They first predict several candidate boxes from dense detection boxes, and then perform detection and re-identification based on the features extracted from the candidate boxes.

[0004] RoI (Region of Interest) feature extraction is an important step in the pedestrian search method based on candidate box generation. It converts an RoI region of any size into a unified fixed-size feature. Here, the region of interest corresponds to the region on the feature map corresponding to the candidate box. RoI pooling is the RoI feature extraction method used by Faster R-CNN [8] and is also a conventional RoI feature extraction method. As Figure 2As shown, the entire background is the feature map, and the bounding box is the RoI region. First, the RoI region in the feature map is divided into k×k sub-regions (shown as 3×3 sub-regions in the figure). No pooling is performed within each sub-region, and the RoI features with an output size of k×k are obtained. To extract more accurate features, RoI Align Pooling improves RoI Pooling. It uniformly samples within each sub-region in a uniform sampling manner and uses average pooling to fuse the feature values of different sampling points to generate the features of this sub-region. Here, the feature values of the sampling points are calculated through bilinear interpolation operations.

[0005] RoI Pooling divides the RoI into fixed spatial regions and lacks an internal mechanism for handling geometric transformations. Since different positions may correspond to objects with different scales or deformations, it is necessary to adaptively determine the scale or receptive field size to achieve visual recognition with fine localization. The deformable RoI Pooling method was proposed in Deformable Convolutional Networks[9], enhancing the ability of deep convolutional neural networks to model geometric transformations. As Figure 3 shown, it learns the offset vectors from the feature map region that has undergone one RoI Align Pooling through an additional fully connected layer, adds an offset to each sub-region of RoI Pooling, thereby achieving adaptive localization of objects with different shapes. However, the method of deformable RoI Align Pooling requires introducing an additional RoI Align Pooling calculation, significantly increasing the computational amount.

[0006] References

[0007] [1] Chen D, Zhang S, Ouyang W, et al. Person search via a mask-guided two-stream cnn model. Proceedings of the European Conference on Computer Vision. 2018:734 - 750.

[0008] [2] Lan X, Zhu X, Gong S. Person search by multi-scale matching. Proceedings of the European Conference on Computer Vision. 2018:536 - 552.

[0009] [3] Zheng L, Zhang H, Sun S, et al. Person re-identification in the wild. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017:1367-1376.

[0010] [4] Xiao T, Li S, Wang B, et al. Joint detection and identification feature learning for person search. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2017:3415-3424.

[0011] [5] Yan Y, Li J, Qin J, et al. Anchor-free person search. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021:7690-7699.

[0012] [6] Chen D, Zhang S, Yang J, et al. Norm-aware embedding for efficient person search. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2020:12615-12624.

[0013] [7] Li Z, Miao D. Sequential end-to-end network for efficient person search. Proceedings of the AAAI Conference on Artificial Intelligence. 2021, 35(3):2011-2019.

[0014] [8]Ren S, He K, Girshick R, et al. Faster r-cnn: Towards real-time object detection with region proposal networks[J]. Advances in Neural Information Processing Systems, 2015, 28.

[0015] [9]Dai J, Qi H, Xiong Y, et al. Deformable convolutional networks. Proceedings of the IEEE International Conference on Computer Vision. 2017:764-773.

[0016]

[10] Sun P, Zhang R, Jiang Y, et al. Sparse R-CNN: End-to-end object detection with learnable proposals. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2021:14454-14463. Summary of the Invention

[0017] The present invention provides a pedestrian search method based on dynamic RoI feature extraction to achieve pedestrian search with fine positioning. By using the dynamic RoI feature extraction method proposed by the present invention, RoI features can be adaptively extracted for each sub-region divided in the RoI region, which can better handle problems such as pedestrian pose change, view change, occlusion, etc., and improve two common performance indicators of pedestrian search, namely accuracy and Top-1. The technical solutions are as follows:

[0018] A pedestrian search method based on dynamic RoI feature extraction, the pedestrian search network used adopts a one-step network framework based on candidate box generation, including a backbone network, a neck network and a head network. The head network is composed of an object detection head network and a pedestrian re-identification head network connected in series; the head network adopts a multi-stage cascade architecture, that is, the output result of the previous-level network is used as the input of the next-level network; the input of each level of the head network includes the corresponding feature map of the neck network, candidate box position information, and candidate box feature vectors, and the following steps are included:

[0019] Step 1: Prepare an image set containing different pedestrians, and annotate the annotation information of pedestrians in each image of the image set, including the identity information and annotation box information of pedestrians;

[0020] Step 2: Divide the image set into a training set, a validation set, and a test set;

[0021] Step 3: Set relevant hyperparameters in the training stage;

[0022] Step 4: Train a pedestrian search network, which is divided into the following sub-steps:

[0023] Sub-step 1: Initialize relevant convolutional weights using an ImageNet pre-trained model;

[0024] Sub-step 2: The input image passes through the backbone network and the neck network to generate a feature map, and the feature map is input into the head network;

[0025] Sub-step 3: In the head network, first input the candidate box feature vector of the previous level into the fully connected layer to calculate the offset vectors of different sub-regions in the region of interest, where the candidate box feature vector of the first level is randomly initialized; assume that the RoI region is divided into k×k sub-regions in the RoI alignment pooling, then the size of the offset vector is 2×k×k;

[0026] Sub-step 4: Design a dynamic RoI alignment pooling layer: The dynamic RoI pooling layer uses the candidate box feature vector as the input, predicts the offsets of each sub-region in the RoI region using the fully connected layer, calculates the positions of the actual sampling points of different sub-regions based on the offsets, and then generates the RoI feature vector using the RoI alignment pooling operation;

[0027] Sub-step 5: Input the obtained RoI feature vector and the candidate box feature vector of the previous level into the object detection head network to obtain the updated candidate box category and position information, as well as the updated candidate box feature vector. Input the RoI feature vector into the person re-identification head network to obtain the re-identification result; the updated candidate box and candidate box feature vector are used as the candidate box and candidate box feature vector of the next level input;

[0028] Sub-step 6: Set the loss function in the training stage, which includes the loss function of object detection and the loss function of re-identification; update the weight parameters of the network through the backpropagation algorithm commonly used in deep convolutional neural networks;

[0029] Sub-step 7: When the number of iterations ends, the learned weight parameters are the final network parameters, and the model training is completed.

[0030] The present invention generates offsets on the RoI region to extract dynamic RoI features, improving the previous method that could only extract features of static and fixed RoI regions, adding an internal mechanism capable of handling geometric transformations, and being able to adaptively extract features for each sub-region in the RoI region, better handling issues such as pedestrian pose changes, perspective changes, and occlusions. At the same time, obtaining the offset vector only introduces an additional fully connected layer, increasing only a small amount of computational complexity. On this premise, pedestrian search with fine positioning is achieved, improving the performance metrics of pedestrian search. Description of the Drawings

[0031] Figure 1 Schematic diagram of pedestrian search application in intelligent monitoring scenarios

[0032] Figure 2 RoI Pooling

[0033] Figure 3 Deformable RoI Pooling

[0034] Figure 4 Dynamic RoI Feature Extraction

[0035] Figure 5 Specific implementation method of the method proposed by the present invention Detailed implementation

[0036] To solve the above problems such as RoI pooling dividing the RoI into fixed spatial regions and lacking an internal mechanism for handling geometric transformations. Without significantly increasing the computational complexity, the present invention proposes a pedestrian search method based on dynamic RoI feature extraction by introducing offsets for each sub-region to achieve pedestrian search with fine positioning. Using the dynamic RoI feature extraction method proposed by the present invention, it is possible to adaptively extract RoI features for each sub-region divided in the RoI region, better handle issues such as pedestrian pose changes, perspective changes, and occlusions, and improve two common performance metrics of pedestrian search, namely accuracy and Top-1. The following introduces the pedestrian search method proposed by the present invention. It mainly includes the following steps:

[0037] Step 1: Prepare an image set containing different pedestrians, and label the annotation information of pedestrians in each image of the image set, including pedestrian identity information and annotation box information.

[0038] Step 2: Divide the image set into a training set, a validation set, and a test set. The training set is used to train the deep convolutional neural network, the validation set is used to select the best training model, and the test set is used for subsequent testing of the model effect.

[0039] Step 3: Set the relevant hyperparameters in the training stage, including the number of iterations, the initial learning rate, the change of the learning rate, the number of images in each training batch, etc.

[0040] Step 4: Design a pedestrian search network. The pedestrian search network adopts the one-step network framework based on candidate box feature vectors described in the background art. The backbone network uses ResNet50, the neck network uses FPN, and the head network is composed of an object detection head network and a pedestrian re-identification head network connected in series. The object detection head network uses Sparse R-CNN

[10] , and the pedestrian re-identification head network can be composed of convolutional layers, activation layers, normalization layers, etc. The head network adopts a multi-level cascaded architecture, that is, the output result of the previous-level network is used as the input of the next-level network.

[0041] Step 5: Train the pedestrian search network, which can be divided into the following sub-steps:

[0042] Sub-step 1: Initialize the relevant convolutional weights using the ImageNet pre-trained model.

[0043] Sub-step 2: The input image passes through the backbone network and the neck network to generate a feature map, and the feature map is input into the head network.

[0044] Sub-step 3: In the head network, first, the candidate box feature vector of the previous level is input into the fully connected layer to calculate the offset vectors of different sub-regions of the region of interest, where the candidate box feature vector of the first level is randomly initialized. Assume that the RoI region is divided into k×k sub-regions in the RoI alignment pooling, then the size of the offset vector is 2×k×k. Since each offset includes offsets in the x and y directions, the number of parameters is twice the number of sub-regions. The meaning of this offset vector is the adaptive offset generated for each divided sub-region.

[0045] Sub-step 4: Design a dynamic RoI alignment pooling layer. Figure 3 The basic structure of the dynamic RoI alignment pooling layer is given. Inputting the feature map and the offset vector into the dynamic RoI alignment pooling layer can obtain the RoI feature vector. Specifically, the RoI pooling layer uses average pooling to fuse the feature values of different sampling points in each divided sub-region. The dynamic RoI pooling layer uses the candidate box feature vector as the input, and uses the fully connected layer to predict the offsets of each sub-region in the RoI region (including horizontal and vertical offsets). Based on this offset, the positions of the actual sampling points in different sub-regions can be calculated, and then the RoI feature vector is generated using the RoI alignment pooling operation;. Generally, the position coordinates of the sampling points take non-integer values, so the features of the sampling points are obtained by bilinear interpolation of the values of the adjacent four integer points.

[0046] Sub-step 5: Input the obtained RoI feature vectors and the candidate box feature vectors of the previous level into the object detection head network, and the updated candidate box category and location information, as well as the updated candidate box feature vectors, can be obtained. Input the RoI feature vectors into the person re-identification head network to obtain the re-identification results; use the updated candidate boxes and the updated candidate box feature vectors as the input candidate boxes and candidate box feature vectors of the next level.

[0047] Sub-step 6: Set the loss function in the training phase, which includes the loss function for object detection and the loss function for re-identification. Update the weight parameters of the network through the backpropagation algorithm commonly used in deep convolutional neural networks.

[0048] Sub-step 7: When the number of iterations ends, the learned weight parameters are the final network parameters, and the model training is completed.

[0049] Step 6: Test the trained model. Input the images in the test set into the model, and the results of pedestrian search can be obtained through the above sub-steps 2-5, and then the performance indicators of the model can be obtained. Among them, the trained dynamic RoI feature extraction module can extract fine and adaptively located dynamic RoI feature vectors.

[0050] Experimental results on the PRW dataset show that by introducing dynamic RoI feature extraction, the accuracy rate has increased from 43.8% to 46.9%, and Top-1 has increased from 75.9% to 77.5%.

Claims

1. A pedestrian search method based on dynamic RoI feature extraction. The pedestrian search network adopted uses a one-step network framework based on candidate box generation, including a backbone network, a neck network, and a head network. The head network is composed of an object detection head network and a pedestrian re-identification head network connected in series; the head network adopts a multi-level cascaded architecture, that is, the output result of the previous-level network is used as the input of the next-level network; the input of each level of the head network includes the feature map corresponding to the neck network, candidate box position information, and candidate box feature vectors, and the following steps are included: Step 1: Prepare an image set containing different pedestrians, and label the annotation information of pedestrians in each image of the image set, including the identity information and annotation box information of pedestrians; Step 2: Divide the image set into a training set, a validation set, and a test set; Step 3: Set relevant hyperparameters in the training stage; Step 4: Train the pedestrian search network, which is divided into the following sub-steps: Sub-step 1: Initialize the relevant convolutional weights using the ImageNet pre-trained model; Sub-step 2: The input image passes through the backbone network and the neck network to generate a feature map, and the feature map is input into the head network; Sub-step 3: In the head network, first input the candidate box feature vector of the previous level into the fully connected layer to calculate the offset vectors of different sub-regions in the region of interest, where the candidate box feature vector of the first level is randomly initialized; assume that the RoI region is divided into k×k sub-regions in the RoI alignment pooling, then the size of the offset vector is 2×k×k; Sub-step 4: Design a dynamic RoI alignment pooling layer: The dynamic RoI pooling layer uses the candidate box feature vector as the input, predicts the offset of each sub-region in the RoI region using the fully connected layer, calculates the positions of the actual sampling points of different sub-regions based on this offset, and then generates the RoI feature vector using the RoI alignment pooling operation; Sub-step 5: Input the obtained RoI feature vector and the candidate box feature vector of the previous level into the object detection head network to obtain the updated candidate box category and position information, as well as the updated candidate box feature vector. Input the RoI feature vector into the pedestrian re-identification head network to obtain the re-identification result; the updated candidate box and candidate box feature vector are used as the candidate box and candidate box feature vector of the next-level input; Sub-step 6: Set the loss function in the training stage, and this loss function includes the loss function of object detection and the loss function of re-identification; update the weight parameters of the network through the backpropagation algorithm commonly used in deep convolutional neural networks; Sub-step 7: When the number of iterations ends, the learned weight parameters are the final network parameters, and the model training is completed.

Citation Information

Patent Citations

  • Fast pedestrian detection method and device

    CN108399362A

  • A pedestrian search method and device based on a priori candidate box selection strategy

    CN109165540A