A High-Performance Pedestrian Retrieval and Re-Identification Method and Device

By inserting Ghost lightweight module into the pedestrian detection model and adopting channel-level sparse pruning method, the deployment problem of pedestrian re-identification on devices with limited hardware resources in the prior art is solved, and high-performance and low-power pedestrian detection and re-identification effects are achieved.

CN115063831BActive Publication Date: 2025-05-30ZHEJIANG GONGSHANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210409679.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-19
Publication Date
2025-05-30
Estimated Expiration
2042-04-19

AI Technical Summary

Technical Problem

The existing pedestrian recognition technology is difficult to achieve high-performance and low-power deployment on devices with limited hardware resources and tight power budgets, and the separate pedestrian recognition model cannot meet the application requirements of large-scale video surveillance systems.

Method used

By inserting the Ghost lightweight module into the pedestrian detection model and using the channel-level sparse pruning method, the model size is reduced and the calculation efficiency is reduced. Then quantitatively deploy it on Guanfeng SC5 and cloud AI computing acceleration card to build a pedestrian search system to meet the needs of high performance and low power consumption.

Benefits of technology

It realizes efficient deployment of pedestrian detection and re-identification models on devices with limited hardware resources, ensuring detection accuracy and real-time performance, while reducing the memory usage and time-consuming of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115063831B_ABST
    Figure CN115063831B_ABST
Patent Text Reader

Abstract

The present invention discloses a high-performance pedestrian retrieval and re-identification method and device. The method includes: respectively obtaining pedestrian data in single-view and multi-view real monitoring scenarios, and performing data annotation on the pedestrian data. Among them, the single-view pedestrian data and the pedestrian part in the COCO dataset jointly construct a pedestrian detection dataset, and the multi-view pedestrian data constructs a pedestrian re-identification dataset; using the pedestrian detection dataset to train a network model with the YOLOv5 pedestrian detection algorithm improved based on the Ghost lightweight model; using the pedestrian re-identification dataset to train a pedestrian re-identification model; building a pedestrian search system. The present invention realizes a low-cost and high-performance pedestrian re-identification system by means of the collaborative optimization of deep model compression and algorithm computing power, and a top-down method from algorithm to hardware to optimize the efficiency of deep learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of pedestrian detection and pedestrian re-identification, and particularly relates to a high-performance pedestrian retrieval and re-identification method and device. Background Art

[0002] Object detection technology is mainly used to find objects of specific categories in a given image, and at the same time detect the class labels and coordinates of the objects. The algorithm based on deep convolutional neural network has become the mainstream algorithm in the field of object detection. According to different classification criteria, it can be divided into Two-stage algorithms and One-stage algorithms at present. The R-CNN series are such Two-stage algorithms. First, candidate boxes are generated, and then the candidate boxes are classified and their positions are fine-tuned. One-stage algorithms such as Yolo and SSD do not need to pre-generate candidate boxes, and directly perform regression and classification on each position of the image. Two-stage algorithms have higher accuracy but slower speed. The improvement of the algorithm is accompanied by the improvement of speed. One-stage algorithms are fast, and the improvement of the algorithm is accompanied by the improvement of accuracy.

[0003] Pedestrian re-identification technology is the core technology of long-term and cross-domain multi-object tracking. The main goal is to re-identify the same pedestrian across cameras. Existing video analysis systems are a collection of multiple task algorithms, which have extremely high requirements for computing chips. Pedestrian search includes the processes of searching (object detection) and matching (pedestrian re-identification). Pedestrian re-identification models are all based on pedestrian images obtained by preprocessing surveillance videos, which requires a lot of preparatory work. For real-world scenarios, a single pedestrian re-identification model cannot meet the application requirements.

[0004] Different from the single-scene of academic datasets, since the pedestrian search algorithm inputs the result images of the pedestrian detection algorithm into the pedestrian re-identification module, the detection effect of the pedestrian detection model is the key step to correctly identify pedestrians. For large-scale video surveillance systems, intelligent device platforms with high performance and low power consumption are often required. In the pedestrian re-identification task, the method based on deep learning is currently the best algorithm. However, complex deep learning models usually have good detection effects and re-identification capabilities, and it is difficult to be deployed on devices with limited hardware resources and tight power budgets while ensuring accuracy and real-time performance. Summary of the Invention

[0005] In view of the deficiencies of the above-mentioned existing technologies, the present invention provides a high-performance pedestrian retrieval and re-identification method. During the deployment process of the model, it mainly faces difficulties such as model size, memory occupation during operation, and computing efficiency. Therefore, the core of the present invention is to insert a Ghost lightweight module into the detection model and adopt a channel-level sparse pruning method, so as to reduce the size of the original search network model and keep the accuracy benchmark equivalent. And the trained model is quantized on the Suanfeng SC5 and computing cards, and then deployed to the hardware lightly and quickly, meeting the needs of large-scale video surveillance systems for intelligent device platforms with high performance and low power consumption hardware, ensuring accuracy and real-time performance while being deployed on devices with limited hardware resources and tight power budgets.

[0006] The object of the present invention is achieved through the following technical solutions:

[0007] According to the first aspect of this specification, a high-performance pedestrian retrieval and re-identification method is provided, including the following steps:

[0008] S1, obtain single-view pedestrian data and multi-view pedestrian data in the actual monitoring scenario respectively. Among them, the pedestrian detection dataset is jointly constructed by using the single-view pedestrian data and the pedestrian part data in the COCO dataset, and the multi-view pedestrian data is used to construct the pedestrian re-identification dataset;

[0009] S2, use the pedestrian detection dataset in S1 to train a pedestrian detection model based on the YOLOv5 pedestrian detection algorithm improved by the Ghost lightweight model;

[0010] S3, use the pedestrian re-identification dataset in S1 to train the pedestrian re-identification network after channel-level sparse pruning to obtain a pedestrian re-identification model;

[0011] S4, use the trained pedestrian detection model in S2 and the pedestrian re-identification model in S3 to perform quantization deployment based on the Suanfeng SC5 and the cloud AI computing acceleration card, and build a pedestrian search system.

[0012] Further, the pedestrian detection model in step S2 includes four modules, namely an input end module, a backbone network module, a neck network module, and an output end module, and the input is an image in the pedestrian detection dataset;

[0013] The image is first input into the backbone network module through the input end module to extract the pedestrian feature image, and the pedestrian feature image is sent into the neck network module. The neck network module will enhance the detection of the pedestrian feature image at different scaling scales by the pedestrian detection model, and the enhanced pedestrian feature image is sent to the output end. The output end predicts the pedestrian feature image to generate the bounding box and the category in the predicted pedestrian feature image.

[0014] Furthermore, the backbone network module includes three types of modules, namely the Focus sub-module, the CBL sub-module, and the GhostCSP sub-module;

[0015] The Focus sub-module performs slicing operations on the input image and downsampling operations at every other pixel. The CBL sub-module performs convolutional operations on the input image. The GhostCSP sub-module is generated by replacing with the Ghost network, where the Ghost network with a stride of 1 replaces the residual components in the CSP structure, and the Ghost network with a stride of 2 replaces the convolutional layer in the CSP structure, serving the role of downsampling.

[0016] Furthermore, the neck network module performs multiple feature extractions on the pedestrian feature image extracted by the backbone network module, generates pedestrian feature images at scales of 8, 16, and 32, calculates the loss based on the pedestrian feature images at scales of 8, 16, and 32 to obtain a loss value, and the pedestrian detection model is trained and updated according to the loss value to obtain a trained enhanced pedestrian detection model.

[0017] Furthermore, the person re-identification network in step S3 includes a ResNet50 network and a BNNeck module, and the input is an image from a person re-identification dataset;

[0018] For the input image, it is randomly cropped to different sizes and aspect ratios, scaled to the same size, and randomly erased. A rectangular box filled with random values is used to occlude the image to obtain an enhanced image;

[0019] The enhanced image is input into the ResNet50 network. The ResNet50 network is pre-trained with the ImageNet dataset to extract pedestrian image features, and global pooling is performed on the extracted features to obtain the pedestrian global feature F global ;

[0020] The BNNeck module separates the person re-identification loss into two different feature spaces for optimization to complete one learning.

[0021] Furthermore, the loss function Loss of the person re-identification network in step S3 is:

[0022]

[0023] where: n is the number of samples, x i is the input image, y i is its class label, o(y i |x i ) represents the predicted probability that x i is recognized as y i ; dp is the distance between the homogeneous image and the input image, d n is the distance between the heterogeneous image and the input image, α and β are hyperparameters for balancing the loss, and max(*) is to take the maximum distance; represents the features before the fully connected layer, represents the feature center of the yi-th category, is the L2 norm.

[0024] Furthermore, the ResNet50 network is processed by a channel-level sparsification pruning method. A scaling factor α is introduced for each channel. First, the connectivity is learned through normal network training. During the training process, these scale factors are sparsified and regularized to automatically identify the importance of the channels. Finally, the channels with lower scaling factors obtained from the training are pruned.

[0025] Furthermore, the objective function for pruning the person re-identification model is as follows:

[0026]

[0027] where (x, y) are the training input and target. The first term as a whole represents the original loss function of the unpruned network. The second term is the penalty term on the scaling factor. A represents the trainable parameters in the network, α is the scaling factor, β is the balance factor for the two terms, and |·| is the L1 norm.

[0028] Furthermore, the specific steps of step S4 are as follows:

[0029] S41, Select 3000 pictures from the pedestrian detection training set constructed in step S1 and 1500 pictures from the person re-identification training set constructed in step S1, and convert them into lmdb datasets for subsequent calibration and quantization;

[0030] S42, Use the BMNNSDK2 SDK tool to convert the pedestrian detection model and the person re-identification model trained in S2 and S3 into fp32umodel files and corresponding prototxt files, where fp32umodel is a format private to the bit platform;

[0031] S43, Use the calibration_use_pb quantization tool to convert the fp32umodel converted in step S42 into an intermediate temporary model int8umodel private to the bit, and use the lmdb dataset in S41 as the quantization calibration, where int8umodel is a network coefficient file in the int8 format generated by quantization;

[0032] S44. Use the calibration visualization analysis tool to check the network error of the int8u model after conversion in S43. Use the mean absolute percentage error and cosine function as the error evaluation criteria, which are defined as follows:

[0033]

[0034]

[0035] Among them, Actual i represents the true value, Forecast i represents the predicted value, and n is the number of samples;

[0036] S45. After confirming that the quantization accuracy is normal through the error evaluation criteria, use the BMNETU tool provided by the BMNNSDK2 SDK. Use the int8u del model in step S43 as the input, compile it into the files required by BMRuntime, and during the compilation, calculate and compare the results of each layer of the NPU model with the CPU calculation results to obtain the int8b model for pedestrian detection and re-identification;

[0037] S46. Based on the int8b model completed in quantization in step S45, input multiple monitoring video streams, create each frame of the video on the specified chip, and use the int8b model for pedestrian detection completed in quantization in step S45 to perform pedestrian detection on each frame of the video in each video stream to obtain the pedestrian bounding box and confidence;

[0038] S47. Use the generated pedestrian bounding box and confidence to filter through the DIoU-NMS method. If the final pedestrian confidence is less than the confidence threshold, suppress it to obtain the filtered video surveillance pedestrian image. The formula is as follows:

[0039]

[0040] Among them, ε is the NMS threshold, N i is the classification confidence, M is the detection box with the highest confidence, B i is the bounding box, and R DIoU is the central distance between two bounding boxes;

[0041] S48. Crop each frame of the pedestrian selected in step S47 in the video frame as the pedestrian library picture. When the set batch processing quantity is reached, send each batch of pedestrians to be recognized in the library into the int8b model for pedestrian re-identification completed in quantization in step S45 to extract the pedestrian picture features to obtain the candidate set features, and input a pedestrian image to be queried to obtain the query set features;

[0042] S49 calculates the Euclidean distance feature between the candidate set features and the query set features in S48, thereby obtaining the pedestrian similarity value, and determines whether the pedestrian similarity value is greater than a preset pedestrian threshold to obtain the re-identification result of the given target pedestrian.

[0043] According to the second aspect of this specification, a high-performance pedestrian retrieval and re-identification device is provided, including a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, it is used to implement the high-performance pedestrian retrieval and re-identification method as described in the first aspect.

[0044] The beneficial effects of the present invention are as follows: The present invention constructs an improved Yolov5 pedestrian detection network, replaces the original CSP structure with a Ghost module, reduces the network calculation amount, and improves the detection efficiency. At the same time, DIoU-NMS is introduced in the inference stage to reduce missed detection cases and improve the detection accuracy. A residual network ResNet50 is constructed to extract global features from pedestrian image features, and combined with triplet loss, center loss, and Identity loss for training, effectively reducing the overfitting degree of the model and improving the generalization ability of the model. The channel-level sparsification pruning method is adopted to prune the convolutional network of the pedestrian re-identification network, reduce the size and number of parameters of the model, and reduce the memory occupancy and time consumption during model operation. A pedestrian search system is built, and quantization of the model is realized by combining Sophon SC5 and cloud AI computing acceleration cards, further reducing the model size and accelerating the inference speed. Use SC5 + specific interfaces to deploy the model, quickly and accurately retrieve specific pedestrians in surveillance videos, and optimize the efficiency of deep learning through a top-down method from algorithms to hardware. Description of the Drawings

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0046] Figure 1 It is a schematic diagram of the yolov5-ghost network framework in an embodiment of the present invention;

[0047] Figure 2 It is a schematic diagram of the ghost bottleneck framework in an embodiment of the present invention;

[0048] Figure 3 It is a schematic diagram of the pedestrian re-identification network framework in an embodiment of the present invention;

[0049] Figure 4 It is the implementation flowchart of model quantization in an embodiment of the present invention;

[0050] Figure 5 It is the implementation flowchart of the pedestrian search system in an embodiment of the present invention;

[0051] Figure 6 It is the structural diagram of a high-performance pedestrian retrieval and re-identification device in an embodiment of the present invention. Specific embodiments

[0052] To better understand the technical solution of the present application, the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0053] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.

[0054] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms of "a", "an", and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0055] The present invention provides a high-performance pedestrian retrieval and re-identification method, including the following steps:

[0056] S1. Obtain single-view pedestrian data and multi-view pedestrian data in the actual monitoring scenario respectively. Among them, a pedestrian detection dataset is jointly constructed by using the single-view pedestrian data and the pedestrian part data in the COCO dataset, and a pedestrian re-identification dataset is constructed with the multi-view pedestrian data;

[0057] In one embodiment, 64,115 pedestrian images with the label of the person category in the COCO dataset are extracted;

[0058] Video surveillance images are collected by several fixed cameras at different angles, and 7,537 pedestrian labels and coordinates are marked for the images; the collected pedestrian images and the MS COCO dataset are used to construct a pedestrian detection dataset according to a ratio of 4:1.

[0059] Collect at least 50,000 images with a pixel size of not less than 64*128. The images must be captured by at least two cameras and must include pedestrian images. A pedestrian re-identification dataset is constructed according to a ratio of 4:1.

[0060] S2. Utilize the pedestrian detection dataset in S1 to train a pedestrian detection model based on the YOLOv5 pedestrian detection algorithm improved by the Ghost lightweight model;

[0061] As Figure 1 shown, in one embodiment, the pedestrian detection model in step S2 includes four modules, namely an input end module, a backbone network module, a neck network module, and an output end module, and the input is an image in a pedestrian detection dataset;

[0062] The image is first input into the backbone network module through the input end module to extract the pedestrian feature image, and the pedestrian feature image is sent into the neck network module. The neck network module will enhance the detection of the pedestrian feature image at different scaling scales, and the enhanced pedestrian feature image is sent to the output end. The output end predicts the pedestrian feature image to generate the bounding box and the category in the predicted pedestrian feature image.

[0063] In one embodiment, the backbone network module includes three modules, namely a Focus sub-module, a CBL sub-module, and a GhostCSP sub-module;

[0064] The image input in the backbone network module is self-duplicated and sliced through the Focus sub-module, and downsampling is performed at every other pixel to reduce the network calculation amount and speed up the speed of feature extraction for candidate regions. The sliced image undergoes convolution normalization through the CBL sub-module for feature extraction. During the extraction process, the GhostCSP sub-module optimizes the network gradient information, speeds up the network calculation speed, and reduces the calculation amount. Finally, the SPP module transforms the inputs of different sizes into the same-sized output to solve the problem of inconsistent input image sizes. The GhostCSP sub-module is generated by replacing with a Ghost network, where the Ghost network with a stride of 1 replaces the residual component in the CSP structure, and the Ghost network with a stride of 2 replaces the convolutional layer in the CSP structure to play the role of downsampling

[0065] The GhostCSP module includes CBL processing, Ghost networks with strides of 1 and 2, batch normalization processing, and rectified linear unit processing.

[0066] Figure 2The Ghost network used in this example has two bottleneck structures. For the case where the stride is 1, it consists of two Ghost modules. The first Ghost module performs an expansion process to increase the number of channels of the input feature map. The second Ghost module is used to reduce the number of channels of the feature map to match the diameter structure in the network and connect the information transfer of the two modules according to the diameter structure. The ReLU activation function is used after the first Ghost module, and only normalization processing is used for the second Ghost module. Such a structure enables the model to effectively reduce the number of parameters and computational complexity and optimize the feature map. For the case where the stride is 2, the two Ghost modules are connected by a depth convolution with a stride of 2, and a downsampling layer is used in the diameter connection path to halve the size of the feature map, further reducing the computational complexity.

[0067] The Ghost network convolves the input feature map of size c*h*w with n groups of k*k convolutional kernels to generate an intrinsic feature map with n channels and a size of h′*w′, and then applies a linear transformation (Φ a ) to further generate new similar feature maps ghost. Finally, the two sets of feature information are combined to obtain all the feature information, as follows:

[0068] Y′ = X*f′ + b

[0069]

[0070] where Y′ ∈ R h′*w′*n represents the output intrinsic feature map, X′ ∈ R h*w*c represents the input feature map, and f′ ∈ R c*k*k*n represents the convolutional kernel of this convolution. y′ i (1 ≤ i ≤ m) are the n feature maps of Y′, and each feature map has to go through a lightweight linear operation Φ i,j (1 ≤ j ≤ s) to obtain s similar feature maps ghost.

[0071] The neck network module adopts the FPN and PAN network structures. By upsampling and fusing the features of different layers, it utilizes the high resolution of the underlying features and the semantic information of the high-level features to output three different detection scale feature maps.

[0072] Based on the above three feature maps, the loss value is calculated, and the pedestrian detection model is trained and updated according to the loss value to obtain a trained model.

[0073] Among them, the loss function includes the bounding box regression loss, the class prediction loss, and the confidence loss. These losses are added according to specific weights to obtain the total object detection loss.

[0074] The bounding box regression is expressed as follows:

[0075]

[0076] Among them, IOU is the intersection over union of the predicted bounding box and the ground truth bounding box, Distance_C is the diagonal distance, Distance_2 is the Euclidean distance between the centers of the two bounding boxes, and v is a parameter to measure the aspect ratio consistency.

[0077] The class prediction loss and the confidence loss are represented by binary cross-entropy loss as follows:

[0078]

[0079] N represents the total number of classes, and y i is the probability of the current class obtained after passing through the activation function is the ground truth value (0 or 1) of the current class.

[0080] In one embodiment, the neck network module performs multiple feature extractions on the pedestrian feature image extracted by the backbone network module, generates pedestrian feature images at scales of 8, 16, and 32, calculates the loss based on the pedestrian feature images at scales of 8, 16, and 32 to obtain a loss value, and the pedestrian detection model is trained and updated according to the loss value to obtain a trained enhanced pedestrian detection model.

[0081] S3. Use the pedestrian re-identification dataset in S1 to train the pedestrian re-identification network after channel-level sparsification pruning to obtain a pedestrian re-identification model;

[0082] In one embodiment, the pedestrian re-identification network in step S3 includes a ResNet50 network and a BNNeck module, and the input is an image in the pedestrian re-identification dataset;

[0083] For the input image, randomly crop it to different sizes and aspect ratios, scale it to the same size, and perform random erasing. Use a rectangular box filled with random values to occlude the image to obtain an enhanced image;

[0084] Input the enhanced image into the ResNet50 network. The ResNet50 network is pre-trained with the ImageNet dataset to extract pedestrian image features, and global pooling is performed on the extracted features to obtain the pedestrian global feature F global ;

[0085] The BNNeck module separates the pedestrian re-identification loss into two different feature spaces for optimization to complete one learning.

[0086] Such as Figure 3As shown, the 50-layer Residual Network ResNet50 built in this step can be divided into seven parts. The first part does not contain residual modules and mainly performs convolution, regularization, and max pooling calculations on the input. The second, third, fourth, and fifth parts all contain residual modules. Each residual module has three convolutions. After the convolution calculations in these five parts, the pooling layer converts them into feature vectors, and finally the classifier performs calculations and outputs class probabilities.

[0087] Among them, the sixth part normalizes the obtained feature vectors through the BN layer, balances the dimensions of each feature, and reduces the constraint of the ID loss on the features.

[0088] In one embodiment, the ResNet50 network is processed using the channel-level sparsification pruning method. A scaling factor α is introduced for each channel. First, the connectivity is learned through normal network training. During the training process, these scale factors are sparsified and regularized to automatically identify the importance of channels. Finally, the channels with lower scaling factors obtained from training are pruned.

[0089] For example, using the channel-level sparsification pruning method, some batch normalization layers in the network of the person re-identification model are pruned. The BN performs the following transformation:

[0090]

[0091] n out = αN + β

[0092] where u b represents the mean of a certain feature map in the minimum batch, n in and n out are the input and output of the batch normalization layer respectively, α is the scaling factor, δ and β are adjustment hyperparameters. The connectivity is learned through normal network training. During the training process, these scale factors are sparsified and regularized to automatically identify the importance of channels. A global threshold for the network layer is set to prune the channels with lower scaling factors.

[0093] In one embodiment, the objective function for pruning the person re-identification model is as follows:

[0094]

[0095] where (x, y) are the training input and target. The first term as a whole represents the original loss function of the unpruned network. The second term is the penalty term on the scaling factor. A represents the trainable parameters in the network, α is the scaling factor, β is the balance factor for the two terms, and |·| is the L1 norm.

[0096] In one embodiment, the loss function for training the person re-identification network adopts triplet loss, center loss, and identity loss. These three losses are added according to specific weights to obtain the total re-identification loss. The loss function Loss of the person re-identification network in step S3 is as follows:

[0097]

[0098] where: n is the number of samples, x i is the input image, y i is its class label, p(y i |x i ) represents the predicted probability that x i is recognized as y i after softmax classification; d p is the distance between the same-class image and the input image, d n is the distance between the different-class image and the input image, α and β are hyperparameters for balancing the loss, and max(·) is to take the maximum distance; represents the feature before the fully connected layer, represents the feature center of the yi-th class, is the L2 norm.

[0099] S4. Use the trained person detection model in S2 and the person re-identification model in S3 for quantization deployment based on the computing power SC5 and the cloud AI computing acceleration card to build a person search system.

[0100] Figure 4 This is the implementation flowchart of the model quantization of the present invention. Step S4 is specifically as follows:

[0101] S41. Select 3000 pictures from the person detection training set constructed in step S1 and 1500 pictures from the person re-identification training set constructed in step S1. Use the convert_imageset interface in the BMNNSDK tool to convert them into lmdb datasets respectively. Set the size of the person detection pictures to 640*640 and the size of the person re-identification pictures to 256*128. Set the shuffle parameter to randomly disrupt the order of the pictures and labels for subsequent calibration quantization;

[0102] S42. Use the BMNNSDK2 SDK tool to convert the trained person detection model and person re-identification model in S2 and S3 into fp32umodel files and corresponding prototxt files. Among them, fp32umodel is a format private to the bit platform; for the prototxt generated in S42, according to the corresponding preprocessing operations on the images in S2 and S3, add the corresponding preprocessing parameters in the data layer to ensure that the data sent to the network is consistent with the original framework;

[0103] S43. Use the calibration_use_pb quantization tool to convert the fp32umodel converted in step S42 into a bit-private intermediate temporary model int8umodel, using the lmdb dataset in S41 as the quantization calibration and the prototxt file in S43 as the network layer, where int8umodel is the network coefficient file in int8 format generated by quantization;

[0104] S44. Use the calibration visualization analysis tool to check the network error of the int8umodel converted in S43, and use the mean absolute percentage error and cosine function as the error evaluation criteria, which are defined as follows:

[0105]

[0106]

[0107] where, Actual i represents the true value, Forecast i represents the predicted value, and n is the number of samples;

[0108] S45. After confirming that the quantization accuracy is normal through the error evaluation criteria, use the BMNETU tool provided by the BMNNSDK2 SDK, that is, through the bmnetu interface in the Qantization-Tool, a network model quantization tool independently developed by Bitmain, use the int8umdel model in step S43 as the input, compile it into the file required by BMRuntime, and during the compilation, compare the NPU model results of each layer with the CPU calculation results to obtain the int8bmodel model for pedestrian detection and pedestrian re-identification;

[0109] If the error is within the specified range, the final model can be generated. If the error is not within the specified range, check which layer has a large quantization error before and after, set the fpfwd_outputs parameter before and after this layer, and run these layers to maintain floating-point calculation without being quantized, without affecting the quantization accuracy of the network;

[0110] S46. Based on the int8bmodel model completed in quantization in step S45, input multiple monitoring video streams, create each video frame on the specified chip, and use the int8bmodel model for pedestrian detection completed in quantization in step S45 to perform pedestrian detection on each video frame in each video stream to obtain the pedestrian bounding box and confidence;

[0111] S47. Filter the generated pedestrian bounding boxes and confidences using the DIoU-NMS method. If the final pedestrian confidence is less than the confidence threshold, suppress it to obtain the filtered video surveillance pedestrian images. The formula is as follows:

[0112]

[0113] where ε is the NMS threshold, N i is the classification confidence, M is the detection box with the highest confidence, B i is the bounding box, and R DIoU is the center distance between two bounding boxes;

[0114] S48. Crop each frame of the pedestrian selected in step S47 in the video frame as a pedestrian library picture. When the set batch processing quantity is reached, send each batch of pedestrians in the library to be recognized into the int8bmodel model of the quantized pedestrian re-identification in step S45 to extract the pedestrian picture features, obtain the candidate set features, and input a pedestrian picture to be queried to obtain the query set features. For example, Figure 5 is the flowchart of the pedestrian search method according to the embodiment of the present invention. Input the pedestrian picture to be retrieved and each road of pedestrian video stream, and use the quantized pedestrian re-identification model to extract the features of the pedestrian picture to be retrieved to obtain the query set features.

[0115] S49. Calculate the Euclidean distance features between the candidate set features and the query set features in S48 to obtain the pedestrian similarity value, and determine whether the pedestrian similarity value is greater than the preset pedestrian threshold to obtain the re-identification result of the given target pedestrian.

[0116] Specifically, use the quantized pedestrian detection model to detect pedestrians in each road of video stream, and determine whether the detected pedestrian score is greater than the set threshold. If so, determine it as a pedestrian and crop and store the pedestrian from the corresponding video frame into the pedestrian library.

[0117] When each certain batch number is reached, use the pedestrian re-identification model to extract features from these pedestrian library pictures to obtain the candidate set features.

[0118] The query set features and the candidate set features use the cosine similarity to determine the pedestrian similarity, which is expressed as follows:

[0119]

[0120] x i ,y i are the feature vectors of two pictures respectively, and n is the vector dimension.

[0121] Determine whether the calculated pedestrian similarity is greater than the set pedestrian threshold. If so, the pedestrian to be recognized in the pedestrian database is a specific pedestrian, and the number of frames of the video stream corresponding to the pedestrian to be recognized is saved. If not, calculate the pedestrian similarity for the next pedestrian database picture, and repeat this step.

[0122] The beneficial effects of the present invention are as follows: The present invention constructs an improved Yolov5 pedestrian detection network, replaces the original CSP structure with a Ghost module, reduces the network calculation amount, and improves the detection efficiency. At the same time, DIoU-NMS is introduced in the inference stage to reduce missed detection cases and improve the detection accuracy. A residual network ResNet50 is constructed to extract global features of pedestrian image features, and combined with triplet loss, center loss, and Identity loss for training, effectively reducing the overfitting degree of the model and improving the generalization ability of the model. The channel-level sparsification pruning method is adopted to prune the convolutional network of the pedestrian re-identification network, reduce the size and number of parameters of the model, and reduce the memory occupancy and time consumption during the operation of the model. A pedestrian search system is built, and quantization of the model is realized by combining Sophon SC5 and a cloud AI computing acceleration card, further reducing the model size and accelerating the inference speed. Use SC5 + a specific interface to deploy the model, quickly and accurately retrieve specific pedestrians in the surveillance video, and optimize the efficiency of deep learning through a top-down method from the algorithm to the hardware.

[0123] Corresponding to the embodiments of the foregoing high-performance pedestrian retrieval and re-identification method, the present invention also provides embodiments of a high-performance pedestrian retrieval and re-identification device.

[0124] See Figure 6 , a high-performance pedestrian retrieval and re-identification device provided by an embodiment of the present invention includes a memory and one or more processors. An executable code is stored in the memory. When the processor executes the executable code, it is used to implement the high-performance pedestrian retrieval and re-identification method in the foregoing embodiments.

[0125] The embodiments of the high-performance pedestrian retrieval and re-identification device of the present invention can be applied to any device with data processing capabilities. The any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation. From the hardware level, as Figure 6 shown, it is a hardware structure diagram of any device with data processing capabilities where the high-performance pedestrian retrieval and re-identification device of the present invention is located. Except for Figure 6In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities where the device in the embodiment is located may generally include other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.

[0126] For the specific implementation process of the functions and roles of each unit in the above device, please refer to the implementation process of the corresponding steps in the above method, which will not be elaborated here.

[0127] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative work.

[0128] The embodiment of the present invention also provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, the high-performance pedestrian retrieval and re-identification method in the above embodiment is implemented.

[0129] The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities in the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device of any device with data processing capabilities, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both the internal storage unit and the external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store computer programs and other programs and data required by any device with data processing capabilities, and can also be used to temporarily store the data that has been output or will be output.

[0130] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity or device including the element.

[0131] The foregoing describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0132] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to limit one or more embodiments of this specification. The singular forms “a,” “an,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term “and / or” as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0133] It should be understood that although the terms first, second, third, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word “if” as used herein may be interpreted as “when” or “upon” or “in response to determining.”

[0134] The above are only the preferred embodiments of one or more embodiments of this specification and are not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.

Claims

1. A high-performance pedestrian retrieval and re-identification method, characterized in that, it includes the following steps: S1. Obtain single-view pedestrian data and multi-view pedestrian data in a real monitoring scenario respectively. Among them, use the single-view pedestrian data and part of the pedestrian data in the COCO dataset to jointly construct a pedestrian detection dataset, and use the multi-view pedestrian data to construct a pedestrian re-identification dataset; S2. Use the pedestrian detection dataset in S1 to train a pedestrian detection model based on the YOLOv5 pedestrian detection algorithm improved by the Ghost lightweight model; S3. Use the pedestrian re-identification dataset in S1 to train a pedestrian re-identification network after channel-level sparsification pruning to obtain a pedestrian re-identification model; S4. Use the trained pedestrian detection model in S2 and the pedestrian re-identification model in S3 to perform quantization deployment based on the Suanfeng SC5 and the cloud AI computing acceleration card, and build a pedestrian search system; specifically: S41. Select 3000 pictures from the pedestrian detection training set constructed in step S1 and 1500 pictures from the pedestrian re-identification training set constructed in step S1, and convert them into lmdb datasets respectively for subsequent calibration quantization; S42. Use the BMNNSDK2 SDK tool to convert the trained pedestrian detection model and pedestrian re-identification model in S2 and S3 into fp32umodel files and corresponding prototxt files, where fp32umodel is a format proprietary to the Bit platform; S43. Use the calibration_use_pb quantization tool to convert the fp32umodel converted in step S42 into an intermediate temporary model int8umodel proprietary to the Bit, and use the lmdb dataset in S41 as quantization calibration, where int8umodel is a network coefficient file in the int8 format generated by quantization; S44. Use the calibration visualization analysis tool to check the network error of the int8umodel converted in S43, and use the mean absolute percentage error and cosine function as the error evaluation criteria, which are defined as follows: Among them, Actual i represents the true value, Forecast i represents the predicted value, and n is the number of samples; S45. After confirming that the quantization accuracy is normal through the error evaluation criteria, use the BMNETU tool provided by the BMNNSDK2 SDK, use the int8umdel model in step S43 as the input, compile it into a file required by the BMRuntime, and during the compilation, compare the NPU model results of each layer with the CPU calculation results to obtain the int8bmodel models for pedestrian detection and pedestrian re-identification; S46. Based on the int8bmodel model completed by quantization in step S45, input multiple monitoring video streams, create each frame of video on the specified chip, and use the int8bmodel model for pedestrian detection completed by quantization in step S45 to perform pedestrian detection on each frame of video in each video stream to obtain pedestrian bounding boxes and confidence levels; S47. Use the generated pedestrian bounding boxes and confidences to filter through the DIoU-NMS method. If the final pedestrian confidence is less than the confidence threshold, suppress it to obtain the filtered video surveillance pedestrian images. The formula is as follows: where ε is the NMS threshold, N i is the classification confidence, M is the detection box with the highest confidence, B i is the bounding box, R DIoU is the center distance between two bounding boxes; S48. Crop each frame of the pedestrians screened in step S47 in the video frame as the pedestrian library pictures. When the set batch processing quantity is reached, send each batch of pedestrians to be recognized in the pedestrian library into the int8bmodel model of the quantized pedestrian re-identification in step S45 to extract the pedestrian picture features, obtain the candidate set features, input a pedestrian image to be queried, and obtain the query set features; S49. Calculate the Euclidean distance features between the candidate set features and the query set features in S48 to obtain the pedestrian similarity value, and determine whether the pedestrian similarity value is greater than the pre-set pedestrian threshold to obtain the re-identification result of the given target pedestrian.

2. The high-performance pedestrian retrieval and re-identification method according to claim 1, characterized in that, the pedestrian detection model in step S2 includes four modules, namely the input end module, the backbone network module, the neck network module, and the output end module, and the input is a picture in the pedestrian detection dataset; The picture is first input into the backbone network module through the input end module to extract the pedestrian feature image, and the pedestrian feature image is sent into the neck network module. The neck network module will enhance the detection of the pedestrian feature image of different scaling scales by the pedestrian detection model, and send the enhanced pedestrian feature image into the output end. The output end predicts the pedestrian feature image to generate the bounding box and the category in the predicted pedestrian feature image.

3. The high-performance pedestrian retrieval and re-identification method according to claim 2, characterized in that, the backbone network module includes three modules, namely the Focus sub-module, the CBL sub-module, and the GhostCSP sub-module; The Focus sub-module performs slicing operations on the input picture and downsampling operations at every other pixel. The CBL sub-module performs convolution operations on the input image. The GhostCSP sub-module is generated by replacing with the Ghost network. Among them, the Ghost network with a stride of 1 replaces the residual component in the CSP structure, and the Ghost network with a stride of 2 replaces the convolutional layer in the CSP structure to play the role of downsampling.

4. The high-performance pedestrian retrieval and re-identification method according to claim 3, characterized in that, the neck network module performs multiple feature extractions on the pedestrian feature image extracted by the backbone network module to generate pedestrian feature images of scales 8, 16, and 32. Calculate the loss based on the pedestrian feature images of scales 8, 16, and 32 to obtain the loss value. The pedestrian detection model is trained and updated according to the loss value to obtain the trained enhanced pedestrian detection model.

5. The high-performance pedestrian retrieval and re-identification method according to claim 1, characterized in that, the pedestrian re-identification network in step S3 includes a ResNet50 network and a BNNneck module, and the input is a picture in the pedestrian re-identification dataset; For the input image, it is randomly cropped into different sizes and aspect ratios, scaled to the same size, and randomly erased. A rectangular box filled with random values is used to occlude the image to obtain an enhanced image. The enhanced image is input into the ResNet50 network. The ResNet50 network is pre-trained with the ImageNet dataset to extract pedestrian image features. Global pooling is performed on the extracted features to obtain the pedestrian global feature F globsl ; The BNN Neck module separates the person re-identification loss into two different feature spaces for optimization to complete one learning.

6. The high-performance pedestrian retrieval and re-identification method according to claim 5, characterized in that the loss function Loss of the person re-identification network in step S3 is: Where: n is the number of samples, x i is the input image, y i is its class label, p(y i |x i ) represents the predicted probability that x i is recognized as y i after softmax classification; d p is the distance between the same-class image and the input image, d n is the distance between the different-class image and the input image, α and β are hyperparameters for balancing the loss, and max(·) is to take the maximum distance; represents the features before the fully connected layer, represents the feature center of the yi-th class, is the L2 norm.

7. The high-performance pedestrian retrieval and re-identification method according to claim 5, characterized in that the ResNet50 network is processed by a channel-level sparsification pruning method. A scaling factor α is introduced for each channel. First, the connectivity is learned through normal network training. During the training process, these scale factors are sparsified and regularized to automatically identify the importance of channels. Finally, the channels with lower scaling factors obtained from training are pruned.

8. The high-performance pedestrian retrieval and re-identification method according to claim 1, characterized in that the objective function of pruning the person re-identification model is as follows: where (x, y) are the training input and target. The first term as a whole represents the original loss function of the unpruned network. The second term is the penalty term on the scaling factor. A represents the trainable parameters in the network, α is the scaling factor, β is the balance factor of the two terms, and |·| is the L1 norm.

9. A high-performance pedestrian retrieval and re-identification device, including a memory and one or more processors. The memory stores executable code, characterized in that when the processor executes the executable code, it is used to implement the high-performance pedestrian retrieval and re-identification method according to any one of claims 1-8.

Citation Information

Patent Citations

  • A pedestrian rerecognition method based on multi-view image feature decomposition

    CN109543602A

  • Visible light-near infrared pedestrian re-identification method based on depth feature orthogonal decomposition

    CN111695470A