An unsupervised pedestrian re-identification method, system, device and storage medium
By performing feature extraction and feature clustering initialization of memory dictionary on pedestrian data, a joint memory learning model is built to train pedestrian re-identification network, which solves the problem that the unsupervised pedestrian re-identification model is affected by equipment, and achieves high-precision pedestrian recognition.
Patent Information
- Application Number
- CN202210401595.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-18
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-04-18
AI Technical Summary
The unsupervised pedestrian re-identification model is affected by the shooting equipment and is difficult, and the existing domain adaptation methods are cumbersome and have limited accuracy.
By performing feature extraction, feature classification and intra-class feature clustering on pedestrian data, initializing memory dictionary, adjusting feature spacing, building a joint memory learning model for pure unsupervised training, and generating a pedestrian re-identification network.
It improves the accuracy of pedestrian recognition and adaptability to various shooting scenes, and achieves high-precision recognition that is not affected by shooting equipment.
Smart Images

Figure CN114821139B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular, to an unsupervised person re-identification method, system, device, and storage medium. Background Art
[0002] Person re-identification (Re-ID) is a pedestrian retrieval technology across non-overlapping cameras. Briefly, the goal of person re-identification is to determine whether a pedestrian has been captured by another camera at another location at another time given a query object. Generally, the query object can be an image, a video sequence, or a text description. Due to differences in the shooting scenes of different cameras in terms of viewpoints, image resolutions, lighting changes, pedestrian postures, occlusions, and modalities, person re-identification has become a highly challenging task.
[0003] Traditional person re-identification can be divided into two directions: supervised and unsupervised, depending on whether the training samples are labeled. Among them, supervised person re-identification trains a person re-identification model on labeled samples, while unsupervised person re-identification trains a person re-identification model on unlabeled samples. Obviously, unsupervised person re-identification is more difficult than supervised person re-identification and is closer to the actual situation.
[0004] The model learning methods of unsupervised person re-identification mainly include domain adaptation methods and pure unsupervised methods. Among them, in domain adaptation methods, the model is trained on a labeled source domain dataset and then transferred to a new unlabeled dataset for further training; while pure unsupervised methods use clustering or graph matching methods to assign pseudo-labels to the target data and directly train the model on the target domain dataset. Since pure unsupervised methods are more difficult and the technology is not yet mature, domain adaptation methods are mostly used for model learning in practical applications. However, domain adaptation methods have the disadvantage of being cumbersome. And with the increasing requirement for model accuracy, the influence and limitation of domain adaptation methods by the source domain dataset are becoming more and more obvious. Summary of the Invention
[0005] An object of the present invention is to solve at least to some extent one of the technical problems existing in the prior art.
[0006] To this end, an object of an embodiment of the present invention is to provide an unsupervised person re-identification method, system, device, and storage medium to achieve high-precision and accurate unsupervised person re-identification.
[0007] To achieve the above technical object, the technical solutions adopted by the embodiments of the present invention include:
[0008] In a first aspect, an embodiment of the present invention provides an unsupervised person re-identification method, including the following steps:
[0009] Obtain first person data from the ImageNet dataset, where the first person data includes multiple first person pictures and corresponding first picture information, and the first picture information includes the shooting device information of the corresponding first person picture;
[0010] Extract features from the first person pictures to generate first person features;
[0011] Group the first person features according to the first picture information to generate a first feature group;
[0012] Initialize a first memory dictionary according to the first feature group;
[0013] Adjust the feature spacing of the first person features according to the first memory dictionary to generate second person features;
[0014] Initialize a second memory dictionary according to the second person features;
[0015] Construct a joint memory learning model according to the first memory dictionary and the second memory dictionary;
[0016] Train a network according to the joint memory learning model and second person data to obtain a person re-identification network, where the second person data is the input training set data;
[0017] Use the person re-identification network to identify the person data to be identified to obtain a person identification result.
[0018] An unsupervised person re-identification method according to an embodiment of the present invention initializes a first memory dictionary by performing feature extraction, feature classification, and intra-class feature clustering on person data, then initializes a second memory dictionary based on the feature spacing adjustment of the first memory dictionary, and then constructs a joint memory learning model according to the first memory dictionary and the second memory dictionary, and trains a person re-identification network with this, realizing pure unsupervised model learning and network training, making the person identification effect of the person re-identification network not affected by the shooting device, and improving the accuracy of person identification and the adaptability to various shooting scenarios.
[0019] In addition, an unsupervised person re-identification method according to the above embodiment of the present invention may further have the following additional technical features:
[0020] Further, in an unsupervised person re-identification method according to an embodiment of the present invention, the extracting features from the first person pictures to generate first person features includes:
[0021] Perform pedestrian detection on the first pedestrian image according to the pedestrian detection algorithm to generate a pedestrian calibration box;
[0022] Cut the first pedestrian image through the pedestrian calibration box to generate a second pedestrian image;
[0023] Extract features from the second pedestrian image through a feature extraction network to generate the first pedestrian feature.
[0024] Further, in an embodiment of the present invention, the initializing the first memory dictionary according to the first feature group includes:
[0025] Cluster the first pedestrian features in each first feature group to generate a first feature class, where the first feature class includes a third pedestrian feature, and the third pedestrian feature includes the first pedestrian feature and a corresponding pseudo label;
[0026] Take the average of the third pedestrian features in each first feature class to generate a fourth pedestrian feature, and the fourth pedestrian feature is the feature representation of the first feature class;
[0027] Construct the first memory dictionary according to the fourth pedestrian feature.
[0028] Further, in an embodiment of the present invention, the adjusting the feature spacing of the first pedestrian feature according to the first memory dictionary to generate a second pedestrian feature includes:
[0029] Calculate the similarity between the first pedestrian feature and the fourth pedestrian feature to obtain a first similarity;
[0030] Generate a similarity distribution according to the first similarity;
[0031] Adjust the feature spacing of the first pedestrian feature according to the similarity distribution to generate the second pedestrian feature.
[0032] Further, in an embodiment of the present invention, the initializing the second memory dictionary according to the second pedestrian feature includes:
[0033] Cluster the second pedestrian features to generate a second feature class, where the second feature class includes a fifth pedestrian feature, and the fifth pedestrian feature includes the second pedestrian feature and a corresponding pseudo label;
[0034] Take the average of the fifth pedestrian features in each second feature class to generate a sixth pedestrian feature, and the sixth pedestrian feature is the feature representation of the second feature class;
[0035] Construct the second memory dictionary according to the sixth pedestrian feature.
[0036] Further, in an embodiment of the present invention, the second pedestrian data includes multiple third pedestrian pictures and corresponding second picture information, and the second picture information includes the shooting device information of the corresponding third pedestrian picture;
[0037] The network training according to the joint memory learning model and the second pedestrian data to obtain a pedestrian re-identification network includes:
[0038] Feature extraction is performed on the third pedestrian picture to generate a seventh pedestrian feature;
[0039] The similarity between the seventh pedestrian feature and the fourth pedestrian feature and the sixth pedestrian feature is calculated respectively to obtain a second similarity;
[0040] An eighth pedestrian feature is obtained from the fourth pedestrian feature and the sixth pedestrian feature according to the maximum value of the second similarity;
[0041] Loss calculation is performed according to the fourth pedestrian feature, the sixth pedestrian feature, the seventh pedestrian feature and the eighth pedestrian feature, the first memory dictionary and the second memory dictionary are updated, and the step of performing feature extraction on the third pedestrian picture to generate a seventh pedestrian feature is returned.
[0042] Further, in an embodiment of the present invention, the pedestrian data to be recognized includes a fourth pedestrian picture and corresponding third picture information, the third picture information includes the shooting device information of the corresponding fourth pedestrian picture, and the pedestrian features after being recognized by the pedestrian re-identification network are stored in the server;
[0043] The recognition of the pedestrian data to be recognized by the pedestrian re-identification network to obtain a pedestrian recognition result includes:
[0044] A ninth pedestrian feature is extracted from the pedestrian data to be recognized;
[0045] The similarity between the ninth pedestrian feature and the pedestrian features in the server is calculated to obtain a third similarity;
[0046] A fourth similarity is obtained from the third similarity, and the fourth similarity is the maximum value in the third similarity;
[0047] It is judged whether the fourth similarity is greater than a preset threshold;
[0048] If so, the ninth pedestrian feature is matched with the first pedestrian label to generate and display the pedestrian recognition result, and the first pedestrian label is the pedestrian label corresponding to the tenth pedestrian feature, and the tenth pedestrian feature is the pedestrian feature in the server with the highest similarity to the ninth pedestrian feature;
[0049] Otherwise, match the ninth pedestrian feature with the second pedestrian label to generate and display the pedestrian recognition result, where the second pedestrian label is a newly set pedestrian label.
[0050] In a second aspect, an embodiment of the present invention provides an unsupervised pedestrian re-identification system, including:
[0051] A pedestrian data acquisition module, configured to acquire first pedestrian data from an ImageNet dataset, where the first pedestrian data includes multiple first pedestrian pictures and corresponding first picture information;
[0052] A feature extraction module, configured to extract features based on the first pedestrian pictures to generate first pedestrian features;
[0053] A feature grouping module, configured to group the first pedestrian features according to the first picture information to generate a first feature group;
[0054] A first memory dictionary initialization module, configured to initialize a first memory dictionary according to the first feature group;
[0055] A feature spacing adjustment module, configured to adjust the feature spacing of the first pedestrian features according to the first memory dictionary to generate second pedestrian features;
[0056] A second memory dictionary initialization module, configured to initialize a second memory dictionary according to the second pedestrian features;
[0057] A model construction module, configured to construct a joint memory learning model according to the first memory dictionary and the second memory dictionary;
[0058] A network training module, configured to perform network training according to the joint memory learning model and second pedestrian data to obtain a pedestrian re-identification network, where the second pedestrian data is input training set data;
[0059] A pedestrian recognition module, configured to use the pedestrian re-identification network to identify the pedestrian data to be recognized to obtain a pedestrian recognition result.
[0060] In a third aspect, an embodiment of the present invention provides an unsupervised pedestrian re-identification device, including:
[0061] At least one processor;
[0062] At least one memory, configured to store at least one program;
[0063] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned unsupervised pedestrian re-identification method.
[0064] In a fourth aspect, an embodiment of the present invention provides a storage medium storing a program executable by a processor. When the program executable by the processor is executed by the processor, it is used to implement the unsupervised person re-identification method described above.
[0065] Advantages and beneficial effects of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present application:
[0066] In the embodiment of the present invention, the first memory dictionary is initialized by performing feature extraction, feature classification, and intra-class feature clustering on pedestrian data, then the second memory dictionary is initialized based on the feature spacing of the first memory dictionary, and then a joint memory learning model is constructed according to the first memory dictionary and the second memory dictionary, and a person re-identification network is trained thereby, realizing pure unsupervised model learning and network training, making the person recognition effect of the person re-identification network not affected by the shooting device, and improving the accuracy of person recognition and the adaptability to various shooting scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the accompanying drawings of the relevant technical solutions in the embodiments of the present application or the prior art. It should be understood that the accompanying drawings in the following introduction are only for clearly presenting some embodiments of the technical solutions in the present application for the convenience of description. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0068] Figure 1 It is a schematic flowchart of a specific embodiment of an unsupervised person re-identification method of the present invention;
[0069] Figure 2 It is a schematic diagram of a joint memory learning model of a specific embodiment of an unsupervised person re-identification method of the present invention;
[0070] Figure 3 It is a schematic structural diagram of a specific embodiment of an unsupervised person re-identification system of the present invention;
[0071] Figure 4 It is a schematic structural diagram of a specific embodiment of an unsupervised person re-identification device of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0073] The terms "first", "second", "third", "fourth", etc. in the description and claims of the present invention and the accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices.
[0074] Referring to "embodiments" in the present invention means that specific features, structures or characteristics described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0075] Person re-identification is a pedestrian retrieval technology across non-overlapping cameras. Briefly speaking, the goal of person re-identification is to determine whether a pedestrian has been captured by another camera at another location at another time given a query object. Generally, the query object can be an image, a video sequence, or a text description. Since the captured images of different cameras vary in viewpoints, image resolutions, lighting changes, pedestrian postures, occlusions, and modalities, person re-identification has become a very challenging task.
[0076] Traditional person re-identification can be divided into two directions: supervised and unsupervised according to whether the training samples are labeled. Among them, supervised person re-identification trains a person re-identification model on labeled samples, while unsupervised person re-identification trains a person re-identification model on unlabeled samples. Obviously, unsupervised person re-identification is more difficult than supervised person re-identification and is also closer to the actual situation.
[0077] The model learning methods for unsupervised person re-identification mainly include domain adaptation methods and pure unsupervised methods. Among them, in the domain adaptation method, the model is trained on the labeled source domain dataset and then transferred to the new unlabeled dataset for further training; while the pure unsupervised method uses methods such as clustering or graph matching to assign pseudo-labels to the target data and directly trains the model on the target domain dataset. Since the pure unsupervised method is relatively difficult and the technology is not yet mature, the domain adaptation method is mostly used for model learning in practical applications. However, the domain adaptation method has the disadvantage of cumbersome steps. And with the improvement of the requirement for model accuracy, the influence and limitation of the domain adaptation method by the source domain dataset become more and more obvious.
[0078] Therefore, the present invention proposes an unsupervised person re-identification method and system. Different from the traditional person re-identification method, which has problems such as low model accuracy caused by the domain adaptation method and being affected and limited by the shooting device, the present invention initializes the first memory dictionary by extracting features, classifying features, and clustering features within the class of pedestrian data, then adjusts and initializes the second memory dictionary based on the feature distance of the first memory dictionary, and then constructs a joint memory learning model according to the first memory dictionary and the second memory dictionary, and trains the person re-identification network with this, realizing pure unsupervised model learning and network training, making the person recognition effect of the person re-identification network not affected by the shooting device, and improving the accuracy of person recognition and the adaptability to various shooting scenarios.
[0079] Next, a detailed description will be given with reference to the accompanying drawings of an unsupervised person re-identification method and system according to an embodiment of the present invention. First, an unsupervised person re-identification method according to an embodiment of the present invention will be described with reference to the accompanying drawings.
[0080] Refer to Figure 1 In an embodiment of the present invention, an unsupervised person re-identification method is provided. The unsupervised person re-identification method in the embodiment of the present invention can be applied to a terminal, can also be applied to a server, or can also be software running on a terminal or a server, etc. The terminal can be a tablet computer, a notebook computer, a desktop computer, etc., but is not limited thereto. The server can be an independent physical server, can also be a server cluster or a distributed system composed of multiple physical servers, or can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The unsupervised person re-identification method in the embodiment of the present invention mainly includes the following steps:
[0081] S101. Obtain first pedestrian data from the ImageNet dataset;
[0082] Among them, the first pedestrian data includes multiple first pedestrian pictures and corresponding first picture information, and the first picture information includes the shooting device information of the corresponding first pedestrian picture.
[0083] Specifically, before inputting the collected pedestrian pictures for person re-identification training, the first pedestrian data in the ImageNet dataset is first used for pre-training.
[0084] S102. Extract features from the first pedestrian picture to generate a first pedestrian feature;
[0085] Specifically, Resnet50 is used as the backbone network to extract the pedestrian feature corresponding to the first pedestrian picture in the first pedestrian data obtained in step S101. In an embodiment of the present invention, the weight file of the Resnet50 network pre-trained in the ImageNet dataset is directly downloaded for the configuration of Resnet50. Among them, the first pedestrian feature is a feature vector with a dimension of 2048.
[0086] The edge computing box adopted in the embodiment of the present invention has strong computing power. By deploying artificial intelligence algorithms through the edge computing box, hardware acceleration is achieved, meeting the requirements of real-time person re-identification.
[0087] S102 can be further divided into the following steps S1021 - S1023:
[0088] Step S1021. Perform pedestrian detection on the first pedestrian picture according to the pedestrian detection algorithm to generate a pedestrian calibration box;
[0089] Among them, in the embodiment of the present invention, the edge computing box is used to perform pedestrian detection on the first pedestrian picture.
[0090] Specifically, each single-frame pedestrian picture (the first pedestrian picture) is processed. The YOLOv5 pedestrian detection model is used to detect the number, position, and size of pedestrians in the picture, and a pedestrian calibration box is generated according to the number, position, and size of pedestrians to calibrate a single pedestrian in the picture.
[0091] Step S1022. Cut the first pedestrian picture through the pedestrian calibration box to generate a second pedestrian picture;
[0092] Specifically, obtain the coordinates, width, and height of the pedestrian calibration box in step S1021, and cut out the picture of a single pedestrian according to the coordinates, width, and height of the pedestrian calibration box to generate a second pedestrian picture. It can be understood that the second pedestrian picture only contains one pedestrian, and the pedestrian occupies most of the space in the second pedestrian picture.
[0093] Step S1023: Extract features from the second pedestrian image through a feature extraction network to generate the first pedestrian feature.
[0094] Specifically, in an embodiment of the present invention, first adjust the second pedestrian image to an RBG image with a width of 128 and a height of 256. Then input it into the feature extraction network (the trained Resnet50), and output a feature vector with a dimension of 2048, which is the first pedestrian feature.
[0095] S103: Group the first pedestrian features according to the first image information to generate a first feature group;
[0096] Specifically, refer to Figure 2 , group the first pedestrian features according to the shooting device information (the first image information) of the first pedestrian images corresponding to each first pedestrian feature to generate a first feature group.
[0097] S104: Initialize the first memory dictionary according to the first feature group;
[0098] Specifically, the first memory dictionary (memory-cam) in the embodiment of the present invention is constructed by classifying all input pedestrian images according to the shooting device information (such as camera ID), and then performing clustering within the class to extract the clustering level features.
[0099] S104 can be further divided into the following steps S1041 - S1043:
[0100] Step S1041: Cluster the first pedestrian features in each first feature group to generate a first feature class;
[0101] Among them, the first feature class includes a third pedestrian feature, and the third pedestrian feature includes the first pedestrian feature and the corresponding pseudo label.
[0102] Specifically, perform clustering on all the first pedestrian features in the first feature class using the DBSCAN clustering method, and assign a pseudo label to each first pedestrian feature.
[0103] Step S1042: Take the average of the third pedestrian features in each first feature class to generate a fourth pedestrian feature;
[0104] Among them, the fourth pedestrian feature is the feature representation of the first feature class.
[0105] Specifically, perform an averaging operation on the third pedestrian features in each first feature class to generate the clustering level feature of each first feature class, which is the fourth pedestrian feature.
[0106] Step S1043: Construct the first memory dictionary according to the fourth pedestrian feature.
[0107] Specifically, construct a first memory dictionary (memory-cam) according to the clustering level feature, that is, the fourth pedestrian feature, to complete the initialization of the first memory dictionary.
[0108] S105: Adjust the feature spacing of the first pedestrian feature according to the first memory dictionary to generate a second pedestrian feature;
[0109] In the embodiment of the present invention, feature distance recalculation is combined during the initialization processes of the first memory dictionary and the second memory dictionary.
[0110] S105 can be further divided into the following steps S1051 - S1053:
[0111] Step S1051: Calculate the similarity between the first pedestrian feature and the fourth pedestrian feature to obtain a first similarity;
[0112] Specifically, use the similarity calculation formula to calculate the similarity between each of the first pedestrian features and each of the fourth pedestrian features to obtain the first similarity:
[0113] sim(f m ,f n )=(f m ·f n ) / (‖f m ‖×‖f n ‖)+1
[0114] where f m represents the first pedestrian feature generated in step S1023, and f n represents the fourth pedestrian feature generated in step S1032.
[0115] Step S1052: Generate a similarity distribution according to the first similarity;
[0116] Specifically, according to step S1051, calculate the similarity between each of the first pedestrian features and each of the fourth pedestrian features through the similarity calculation formula to obtain a series of the first similarities. A similarity distribution is formed according to the series of the first similarities.
[0117] Step S1053: Adjust the feature spacing of the first pedestrian feature according to the similarity distribution to generate the second pedestrian feature.
[0118] Specifically, in the embodiments of the present invention, if the similarity distributions of two pedestrian features are close, it is considered that the probability that these two pedestrian features come from the same pedestrian is relatively high. The similarity of the similarity distributions of two pedestrian features is calculated by the following formula:
[0119] SIM(s m ,s n )=(s m ∩s n ) / (s m ∪s n )
[0120] Wherein, s m ∩s n represents the sum of the minimum values of two pedestrian features (corresponding elements of the vectors), and s m ∪s n represents the sum of the maximum values of two pedestrian features (corresponding elements of the vectors). The feature spacing of the first pedestrian feature is adjusted by using SIM(s m ,s n ) to generate the second pedestrian feature:
[0121] D(I m ,I n )=d(f m ,f n )+μ(1 - SIM(s m ,s n ))
[0122] Wherein, μ is a hyperparameter, and usually takes a value of 0.02.
[0123] S106. Initialize a second memory dictionary according to the second pedestrian feature;
[0124] Specifically, the second memory dictionary (memory - all) in the embodiments of the present invention is constructed by clustering all input pedestrian pictures to extract clustering - level features.
[0125] S106 can be further divided into the following steps S1061 - S1063:
[0126] Step S1061. Cluster the second pedestrian feature to generate a second feature class;
[0127] Wherein, the second feature class includes a fifth pedestrian feature, and the fifth pedestrian feature includes the second pedestrian feature and a corresponding pseudo - label.
[0128] Specifically, all the second pedestrian features in the second feature class are clustered by using the DBSCAN clustering method, and each of the second pedestrian features is assigned a pseudo - label.
[0129] Step S1062: Take the average of the fifth pedestrian features within each of the second feature classes to generate a sixth pedestrian feature;
[0130] Among them, the sixth pedestrian feature is the feature representation of the second feature class.
[0131] Specifically, perform an averaging operation on the fifth pedestrian features in each of the second feature classes to generate the clustering level feature of each of the second feature classes, that is, the sixth pedestrian feature.
[0132] Step S1063: Construct the second memory dictionary according to the sixth pedestrian feature.
[0133] Specifically, construct the second memory dictionary (memory-all) according to the clustering level feature, that is, the sixth pedestrian feature, to complete the initialization of the second memory dictionary.
[0134] S107: Construct a joint memory learning model according to the first memory dictionary and the second memory dictionary;
[0135] Specifically, construct a joint memory learning model according to the first memory dictionary (memory-cam) constructed in step S104 and the second memory dictionary (memory-all) constructed in step S106.
[0136] S108: Perform network training according to the joint memory learning model and the second pedestrian data to obtain a pedestrian re-identification network;
[0137] Among them, the second pedestrian data is the input training set data. It can be understood that the second pedestrian data includes multiple third pedestrian pictures and corresponding second picture information, and the second picture information includes the shooting device information of the corresponding third pedestrian pictures.
[0138] S108 can be further divided into the following steps S1081 - S1084:
[0139] Step S1081: Extract features from the third pedestrian pictures to generate a seventh pedestrian feature;
[0140] Specifically, in the embodiment of the present invention, sample the third pedestrian pictures, extract 64 third pedestrian pictures, and extract 64 seventh pedestrian features through a feature extraction network (trained Resnet50).
[0141] Step S1082: Calculate the similarity between the seventh pedestrian feature and the fourth pedestrian feature and the sixth pedestrian feature respectively to obtain a second similarity;
[0142] Specifically, referring to step S1041, the similarity between each of the seventh pedestrian features and each of the fourth pedestrian features and each of the sixth pedestrian features is calculated through a similarity calculation formula to obtain the second similarity.
[0143] Step S1083: Obtain an eighth pedestrian feature from the fourth pedestrian feature and the sixth pedestrian feature according to the maximum value of the second similarity;
[0144] Specifically, the maximum value is extracted from the second similarity, and the pedestrian feature corresponding to the maximum value of the second similarity is obtained, so as to obtain the pedestrian feature among the fourth pedestrian feature and the sixth pedestrian feature that is most similar to the seventh pedestrian feature, that is, the eighth pedestrian feature.
[0145] Step S1084: Calculate a loss according to the fourth pedestrian feature, the sixth pedestrian feature, the seventh pedestrian feature, and the eighth pedestrian feature, update the first memory dictionary and the second memory dictionary, and return to step S1081.
[0146] Specifically, the loss calculation:
[0147]
[0148] Among them, q represents the seventh pedestrian feature, and c + represents the eighth pedestrian feature, and c i represents the fourth pedestrian feature and the sixth pedestrian feature. In the embodiment of the present invention, the losses compared in the first memory dictionary and the second memory dictionary are added, and the network is updated through backpropagation. Specifically, the pedestrian features in the first memory dictionary and the second memory dictionary that are closest to the newly extracted pedestrian features are mainly updated, that is, the eighth pedestrian feature (the closest to the seventh pedestrian feature):
[0149] c + = m·c + +(1 - m)·q
[0150] Among them, m is a hyperparameter, usually taking a value of 0.1. In the embodiment of the present invention, the first memory dictionary compares a group of pedestrian features of the same shooting device stored, and the second memory dictionary compares all the pedestrian features stored.
[0151] After step S1084 ends, return to step S1081. Steps S1081 - S1084 in the embodiment of the present invention are one network training iteration process. For each epoch, 200 times of the said network iteration process are required. In the embodiment of the present invention, steps S102 - S108 are one epoch, steps S102 - S107 are the process of constructing a joint memory learning model (pretraining of the person re-identification network), and step S108 is the process of training the person re-identification network using the joint memory learning model and training data. The entire training process of the person re-identification network in the embodiment of the present invention needs to execute 100 epochs. Set the initial learning rate to 3.5e-4, and reduce the learning rate to 1 / 10 of the original every 40 epochs.
[0152] The person re-identification network in the embodiment of the present invention evaluates the person re-identification effect through mAP, Rank-1, Rank-5, and Rank-10 metrics.
[0153] S109. Use the said person re-identification network to identify the person data to be identified, and obtain the person identification result.
[0154] Among them, the person data to be identified includes the fourth person picture and the corresponding third picture information. The third picture information includes the shooting device information of the corresponding fourth person picture. The person features after being identified by the person re-identification network are saved in the server.
[0155] Referring to steps S101 - S108, S109 can be further divided into the following steps S1091 - S1096:
[0156] Step S1091. Extract the ninth person feature from the person data to be identified;
[0157] Specifically, use the edge computing box to process the person photos (the fourth person picture) taken by different shooting devices in various orientations and perspectives, use the YOLOv5 person detection model to detect the number, position, and size of the pedestrians appearing in the fourth person picture, generate a person calibration box according to the number, position, and size of the pedestrians, and calibrate a single pedestrian in the fourth person picture. Cut out the picture of a single pedestrian according to the coordinates, width, and height of the person calibration box, and extract the features of the cut pedestrian picture through ResNet50 to obtain the said ninth person feature.
[0158] Step S1092. Calculate the similarity between the ninth person feature and the person features in the server to obtain the third similarity;
[0159] Specifically, according to the similarity calculation formula, calculate the similarity between the ninth pedestrian feature and the pedestrian features recognized by the pedestrian re-identification network stored in the server to obtain the third similarity.
[0160] Step S1093: Obtain a fourth similarity from the third similarity, where the fourth similarity is the maximum value among the third similarities;
[0161] Step S1094: Determine whether the fourth similarity is greater than a preset threshold;
[0162] Step S1095: If so, match the ninth pedestrian feature with the first pedestrian label to generate and display the pedestrian recognition result;
[0163] Wherein, the first pedestrian label is the pedestrian label corresponding to the tenth pedestrian feature, and the tenth pedestrian feature is the pedestrian feature with the highest similarity to the ninth pedestrian feature in the server.
[0164] Step S1096: If not, match the ninth pedestrian feature with the second pedestrian label to generate and display the pedestrian recognition result.
[0165] Wherein, the second pedestrian label is a newly set pedestrian label.
[0166] Specifically, when the fourth similarity is less than the preset threshold, attach a new pedestrian label to the ninth pedestrian feature.
[0167] It can be understood that after completing the recognition of the pedestrian data to be recognized, store the pedestrian recognition result in the server for subsequent pedestrian re-identification, as well as the training and optimization of the joint memory learning model and the pedestrian re-identification network.
[0168] In the embodiment of the present invention, a separate computer is used to display through a network link to a processing server to achieve convenient remote monitoring. It has the following functions:
[0169] First, view the real-time images captured by any shooting device in the system and mark the pedestrian calibration frames in the real-time images.
[0170] Second, view the latest several single pedestrian pictures in real time and label the pedestrian labels under the pictures.
[0171] Third, retrieve the saved shooting videos.
[0172] Fourth, modify the labels of single pedestrian pictures, that is, modify the label of a certain pedestrian picture to another label, namely, perform manual annotation.
[0173] Fifth, uniformly modify the tag names of a certain group of pedestrians, and modify the image tags marked with the same pedestrian tag into a person's name or any other name that is convenient for users to identify the pedestrian's identity. It can be understood that if two pedestrians are modified to the same name, they are regarded as the same pedestrian when stored.
[0174] Sixth, use the pedestrian pictures stored in the server to train the pedestrian re-identification network on the GPU of the server through step S108. Further, multiple re-identification network weight files are stored in the system at the same time, and the user can decide which one to actually apply at any time.
[0175] According to steps S101 - S109, the present invention initializes the first memory dictionary by performing feature extraction, feature classification, and intra-class feature clustering on pedestrian data, then adjusts and initializes the second memory dictionary based on the feature spacing of the first memory dictionary, and then constructs a joint memory learning model according to the first memory dictionary and the second memory dictionary, and trains the pedestrian re-identification network with this, realizing pure unsupervised model learning and network training, making the pedestrian recognition effect of the pedestrian re-identification network not affected by the shooting device, and improving the accuracy of pedestrian recognition and the adaptability to various shooting scenarios.
[0176] Secondly, a kind of unsupervised pedestrian re-identification system proposed according to an embodiment of the present application is described with reference to the accompanying drawings.
[0177] Figure 3 It is a schematic structural diagram of an unsupervised pedestrian re-identification system according to an embodiment of the present application.
[0178] The system specifically includes:
[0179] A pedestrian data acquisition module 301, configured to acquire first pedestrian data from the ImageNet dataset, where the first pedestrian data includes multiple first pedestrian pictures and corresponding first picture information;
[0180] A feature extraction module 302, configured to perform feature extraction according to the first pedestrian pictures to generate first pedestrian features;
[0181] A feature grouping module 303, configured to group the first pedestrian features according to the first picture information to generate a first feature group;
[0182] A first memory dictionary initialization module 304, configured to initialize a first memory dictionary according to the first feature group;
[0183] A feature spacing adjustment module 305, configured to adjust the feature spacing of the first pedestrian features according to the first memory dictionary to generate second pedestrian features;
[0184] The second memory dictionary initialization module 306 is used to initialize the second memory dictionary according to the second pedestrian feature;
[0185] The model construction module 307 is used to construct a joint memory learning model according to the first memory dictionary and the second memory dictionary;
[0186] The network training module 308 is used to perform network training according to the joint memory learning model and the second pedestrian data to obtain a pedestrian re-identification network, where the second pedestrian data is the input training set data;
[0187] The pedestrian recognition module 309 is used to recognize the to-be-recognized pedestrian data by using the pedestrian re-identification network to obtain a pedestrian recognition result.
[0188] It can be seen that the content in the above method embodiments is applicable to the system embodiments. The functions specifically implemented in the system embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0189] Referring to Figure 4 , an unsupervised pedestrian re-identification device is provided in an embodiment of the present application, including:
[0190] At least one processor 401;
[0191] At least one memory 402, configured to store at least one program;
[0192] When the at least one program is executed by the at least one processor 401, the at least one processor 401 is caused to implement the above-mentioned unsupervised pedestrian re-identification method.
[0193] Similarly, the content in the above method embodiments is applicable to the device embodiments. The functions specifically implemented in the device embodiments are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.
[0194] In some alternative embodiments, the functions / operations mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present application are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.
[0195] In addition, although the present application has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present application. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Thus, those skilled in the art can implement the present application as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.
[0196] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several programs for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0197] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable programs for implementing logical functions and can be specifically implemented in any computer-readable medium for use by a program execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can retrieve and execute programs from the program execution system, apparatus, or device), or in conjunction with these program execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with a program execution system, apparatus, or device.
[0198] More specific examples (nonexhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.
[0199] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable program execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), and the like.
[0200] In the foregoing description of the present specification, descriptions with reference to the terms "one embodiment / embodiment example", "another embodiment / embodiment example", or "certain embodiments / embodiment examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0201] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the claims and their equivalents.
[0202] The above has specifically described the preferred embodiments of the present application, but the present application is not limited to the described embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
Claims
1. An unsupervised person re-identification method, characterized in that, Including the following steps: Obtain first pedestrian data from the ImageNet dataset. The first pedestrian data includes multiple first pedestrian pictures and corresponding first picture information, and the first picture information includes the shooting device information of the corresponding first pedestrian picture; Extract features from the first pedestrian pictures to generate first pedestrian features; Group the first pedestrian features according to the shooting device information of the first pedestrian pictures corresponding to the first pedestrian features to generate a first feature group; Initialize a first memory dictionary according to the first feature group; Adjust the feature spacing of the first pedestrian features according to the first memory dictionary to generate second pedestrian features; Initialize a second memory dictionary according to the second pedestrian features; Construct a joint memory learning model according to the first memory dictionary and the second memory dictionary; Perform network training according to the joint memory learning model and second pedestrian data to obtain a pedestrian re-identification network, where the second pedestrian data is the input training set data; Use the pedestrian re-identification network to identify the pedestrian data to be identified to obtain a pedestrian identification result; The initializing the first memory dictionary according to the first feature group includes: Cluster the first pedestrian features in each first feature group to generate first feature classes, and the first feature classes include third pedestrian features, and the third pedestrian features include the first pedestrian features and corresponding pseudo-labels; Take the average of the third pedestrian features in each first feature class to generate fourth pedestrian features, and the fourth pedestrian features are the feature representations of the first feature classes; Construct the first memory dictionary according to the fourth pedestrian features; The adjusting the feature spacing of the first pedestrian features according to the first memory dictionary to generate second pedestrian features includes: Calculate the similarity between the first pedestrian features and the fourth pedestrian features to obtain a first similarity; Generate a similarity distribution according to the first similarity; Adjust the feature spacing of the first pedestrian features according to the similarity distribution to generate the second pedestrian features; The initializing the second memory dictionary according to the second pedestrian features includes: Cluster the second pedestrian features to generate second feature classes, and the second feature classes include fifth pedestrian features, and the fifth pedestrian features include the second pedestrian features and corresponding pseudo-labels; Take the average of the fifth pedestrian features in each second feature class to generate sixth pedestrian features, and the sixth pedestrian features are the feature representations of the second feature classes; Construct the second memory dictionary according to the sixth pedestrian features.
2. The unsupervised person re-identification method according to claim 1, wherein The extracting features from the first pedestrian pictures to generate first pedestrian features includes: Perform pedestrian detection on the first pedestrian pictures according to a pedestrian detection algorithm to generate pedestrian calibration boxes; Cut the first pedestrian pictures through the pedestrian calibration boxes to generate second pedestrian pictures; Extract features from the second pedestrian pictures through a feature extraction network to generate the first pedestrian features.
3. An unsupervised pedestrian re-identification method according to claim 1, characterized in that The second pedestrian data includes multiple third pedestrian pictures and corresponding second picture information, and the second picture information includes the shooting device information of the corresponding third pedestrian picture; The network training based on the joint memory learning model and the second pedestrian data to obtain a person re-identification network includes: Performing feature extraction on the third pedestrian picture to generate a seventh pedestrian feature; Calculating the similarity between the seventh pedestrian feature and the fourth pedestrian feature and the sixth pedestrian feature respectively to obtain a second similarity; Obtaining an eighth pedestrian feature from the fourth pedestrian feature and the sixth pedestrian feature according to the maximum value of the second similarity; Calculating the loss according to the fourth pedestrian feature, the sixth pedestrian feature, the seventh pedestrian feature and the eighth pedestrian feature, updating the first memory dictionary and the second memory dictionary, and returning to the step of performing feature extraction on the third pedestrian picture to generate a seventh pedestrian feature.
4. An unsupervised person re-identification method according to claim 1, characterized in that, The pedestrian data to be recognized includes a fourth pedestrian picture and corresponding third picture information, and the third picture information includes the shooting device information of the corresponding fourth pedestrian picture. The pedestrian features after being recognized by the person re-identification network are stored in the server; The using the person re-identification network to recognize the pedestrian data to be recognized to obtain a pedestrian recognition result includes: Extracting a ninth pedestrian feature from the pedestrian data to be recognized; Calculating the similarity between the ninth pedestrian feature and the pedestrian features in the server to obtain a third similarity; Obtaining a fourth similarity from the third similarity, and the fourth similarity is the maximum value in the third similarity; Judging whether the fourth similarity is greater than a preset threshold; If so, matching the ninth pedestrian feature with the first pedestrian label to generate and display the pedestrian recognition result, where the first pedestrian label is the pedestrian label corresponding to the tenth pedestrian feature, and the tenth pedestrian feature is the pedestrian feature with the highest similarity to the ninth pedestrian feature in the server; If not, matching the ninth pedestrian feature with the second pedestrian label to generate and display the pedestrian recognition result, where the second pedestrian label is a newly set pedestrian label.
5. An unsupervised pedestrian re-identification system, characterized in that, Including: A pedestrian data acquisition module, configured to acquire first pedestrian data from the ImageNet dataset, where the first pedestrian data includes multiple first pedestrian pictures and corresponding first picture information; A feature extraction module, configured to perform feature extraction on the first pedestrian picture to generate a first pedestrian feature; A feature grouping module, configured to group the first pedestrian features according to the shooting device information of the first pedestrian pictures corresponding to the first pedestrian features to generate a first feature group; A first memory dictionary initialization module, configured to initialize a first memory dictionary according to the first feature group; A feature spacing adjustment module, configured to adjust the feature spacing of the first pedestrian features according to the first memory dictionary to generate second pedestrian features; A second memory dictionary initialization module, configured to initialize a second memory dictionary according to the second pedestrian features; A model construction module for constructing a joint memory learning model according to the first memory dictionary and the second memory dictionary; A network training module for training a network according to the joint memory learning model and second pedestrian data to obtain a pedestrian re-identification network, where the second pedestrian data is the input training set data; A pedestrian recognition module for recognizing the pedestrian data to be recognized by using the pedestrian re-identification network to obtain a pedestrian recognition result; The first memory dictionary initialization module is specifically used for: Clustering the first pedestrian features in each of the first feature groups to generate first feature classes, where the first feature classes include third pedestrian features, and the third pedestrian features include the first pedestrian features and corresponding pseudo-labels; Taking the average of the third pedestrian features in each of the first feature classes to generate fourth pedestrian features, where the fourth pedestrian features are the feature representations of the first feature classes; Constructing the first memory dictionary according to the fourth pedestrian features; The feature distance adjustment module is specifically used for: Calculating the similarity between the first pedestrian features and the fourth pedestrian features to obtain a first similarity; Generating a similarity distribution according to the first similarity; Adjusting the feature distance of the first pedestrian features according to the similarity distribution to generate the second pedestrian features; The second memory dictionary initialization module is specifically used for: Clustering the second pedestrian features to generate second feature classes, where the second feature classes include fifth pedestrian features, and the fifth pedestrian features include the second pedestrian features and corresponding pseudo-labels; Taking the average of the fifth pedestrian features in each of the second feature classes to generate sixth pedestrian features, where the sixth pedestrian features are the feature representations of the second feature classes; Constructing the second memory dictionary according to the sixth pedestrian features.
6. An unsupervised pedestrian re-identification device, characterized in that, Including: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements an unsupervised pedestrian re-identification method according to any one of claims 1-4.
7. A storage medium storing a program executable by a processor, characterized in that: The program executable by the processor, when executed by the processor, is used to implement an unsupervised pedestrian re-identification method according to any one of claims 1-4.
Citation Information
Patent Citations
An unsupervised image video pedestrian re-identification method and system based on a migration network
CN109948561A
Unsupervised pedestrian re-identification method and system based on deep clustering and sample learning
CN111401281A