A Pedestrian Re-identification Method Based on a Classification Weight Equi-Distant Distribution Loss Model
By adopting the equidistant distribution loss model of classifying weights in pedestrian recognition, and using the combination method of L2 normalization and equidistant distribution loss of classifying weights, the problem of poor constraint effect of pedestrian characteristics in the measurement space is solved, and the accuracy of pedestrian recognition and the robustness of the model are improved.
Patent Information
- Application Number
- CN202011247963.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-11-10
AI Technical Summary
In the prior art, the constraint effect on pedestrian characteristics in the measurement space is poor, which affects the performance of pedestrian re-identification. Especially when samples are unbalanced, the model has poor recognition performance for few-sample categories.
A pedestrian recognition method based on the equidistant distribution loss model of classification weights is proposed. Through the L2 normalization operation, the pedestrian features and the classification weight vector of the fully connected layer are normalized to the supersphere, and the classification weight equidistant distribution loss is used to constrain the distribution of the weight vector in the feature space, thereby guiding the uniform distribution of pedestrian features.
It effectively solves the performance degradation caused by sample imbalance in pedestrian recognition, and improves the robustness of the model and the accuracy of pedestrian recognition.
Smart Images

Figure CN112446305B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and deep learning, and relates to a pedestrian re-identification method based on a classification weight equidistant distribution loss model. Background Art
[0002] Pedestrian re-identification, also known as pedestrian re-ID, is a very popular research topic in the field of computer vision in recent years. Its goal is to be able to retrieve the pedestrian from other surveillance perspectives after a pedestrian image is taken from the perspective of a surveillance camera. Pedestrian re-ID technology can overcome the limitations of fixed perspectives and perform pedestrian retrieval in multiple unrelated surveillance domains. With the popularization of camera networks in public areas (such as schools, shopping malls, street intersections, etc.) and the improvement of the performance of high-definition cameras, pedestrian re-ID plays an important role in intelligent security, smart city monitoring, intelligent social security management, and criminal investigation.
[0003] With the rise of deep learning and the emergence of large-scale datasets, pedestrian re-identification technology based on deep learning has been widely studied and developed rapidly in recent years. However, due to the difficulty of large-scale data collection, the dataset often shows a sample imbalance phenomenon, that is, a small number of categories have a large number of samples while the number of samples of most categories is relatively small. This causes the trained model to show good recognition performance for categories with a large number of samples, but poor recognition performance for categories with a small number of samples, which reduces the generalization ability of the model and brings great difficulties to the task of pedestrian re-identification.
[0004] In recent years, many studies have proposed methods to solve the sample imbalance problem in pedestrian re-identification tasks. One of the methods is to constrain the distribution of pedestrian features in the metric space by combining classification loss and metric loss. This type of method first extracts high-dimensional pedestrian features from the input pedestrian image through a deep network model. Afterwards, the classification loss uses a classifier to guide the network model to learn from the perspective of pedestrian category information. It converts pedestrian features into probabilities related to the category and maximizes the probability of the category to which the feature belongs, thereby obtaining the classification loss; the metric loss constrains the features from the perspective of the distribution of pedestrian features in the metric space, so that similar features are close in the metric space and features of different categories are far away, thereby obtaining the metric loss. In the training stage, the classification loss and metric loss are combined and used to train the model together. Finally, the trained model is used to extract features from the query set pedestrian images and the candidate set pedestrian images in the test set, and the distance function is used to calculate the distance between the query set image features and all candidate set image features. The distance results are sorted from small to large, and the re-identification results are calculated and obtained. In the prior art, the constraint effect of pedestrian features in the metric space is poor, which affects the performance of pedestrian re-identification. Summary of the invention
[0005] The purpose of the present invention is to address the problems existing in the prior art and propose a pedestrian re-identification method based on a classification weighted equidistant distribution loss model. The method starts from the perspective of constraining the distribution of pedestrian features in the metric space, performs normalization constraints on the metric space where the features are located, regards the classification vectors in the classification layer as category centers, and performs equidistant distribution constraints on all classification centers in the feature space, guides the features to be evenly distributed, and helps the model learn more discriminative pedestrian features, thereby improving the accuracy of pedestrian re-identification.
[0006] The technical solution adopted by the present invention is: a pedestrian re-identification method based on a classification weight equidistant distribution loss model, which specifically includes the following steps:
[0007] Step 1: Preprocess the pedestrian images in the training set and input them into the convolutional neural network to obtain the high-dimensional feature expression vector of the image. , used for subsequent classification learning;
[0008] Step 2: Use a fully connected layer with the bias removed to classify pedestrian features; before classification, the classification weight vector of the fully connected layer is and the pedestrian feature vector used for training After L2 normalization, we can get and , at the same time, let , s is a constant. Then the normalized features are input into the fully connected layer for classification learning, the classification results are obtained, and the classification loss is calculated. ,in Indicates i The characteristics of the samples, represent The real label, Represents the classification layer j column vector, n is the number of samples for batch training, C is the total number of training categories;
[0009] Step 3: Normalize each classification weight vector in the classification layer in step 2 Considered as the center of each type of feature, the sum of the cosine distances between each vector and all other weight vectors is calculated to obtain the classification weight equidistant distribution loss ,in Represents two classification weight vectors and The angle between To avoid the loss value of equal distance classification being too large in the early stage of training;
[0010] Step 4: Combine the classification loss obtained in steps 2 and 3 with the classification weight equidistant distribution loss to obtain the total loss function, that is, ,in is the weight value of the equally distributed loss;
[0011] Step 5: Use the total loss function obtained in step 4 to train and optimize the convolutional neural network, and finally obtain a trained network model;
[0012] Step 6: Input the pedestrian images in the test set into the trained network model, obtain the corresponding feature vectors, and perform L2 normalization on these feature vectors to obtain the final feature vectors; perform Euclidean distance comparison on the normalized feature vectors of the pedestrian images in the query set with the normalized feature vectors of the pedestrian images in the candidate set one by one, sort them according to the distance results, and complete the pedestrian re-identification task.
[0013] Compared with the prior art, the present invention has the following advantages:
[0014] First, the present invention uses L2 normalization operation to normalize pedestrian features and classification weight vectors of the fully connected layer to a hypersphere, thereby being able to effectively measure the distances between features, between weight vectors, and between features and weight vectors, while solving the problem of inconsistent constraint spaces between measurement loss and classification loss, which is beneficial to model optimization and convergence;
[0015] Second, the present invention uses the classification weights of the fully connected classification layer as the feature centers of the training categories. Therefore, the classification weight equidistant distribution loss can be used to constrain the distribution of the weight vector in the feature space, thereby guiding the pedestrian features to be evenly distributed in the feature space, effectively solving the performance degradation caused by sample imbalance in pedestrian re-identification and improving the robustness of the model.
[0016] Third, the present invention uses a method of joint training of metric loss and classification loss. While utilizing equidistant distribution loss to evenly distribute features of different categories, it also uses classification loss to reduce the distance between features of the same category, so that the model can mine and learn the discernible features of pedestrians in the image, thereby improving the accuracy of pedestrian re-identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic diagram of the implementation process of the present invention.
[0018] Figure 2 It is a schematic diagram of the weighted equidistant distribution loss in the present invention.
[0019] Figure 3 It is a schematic diagram of the network model structure of the present invention. DETAILED DESCRIPTION
[0020] The present invention will be further described in detail below in conjunction with embodiments and drawings, but the embodiments of the present invention are not limited thereto.
[0021] A pedestrian re-identification method in this embodiment has the following process: Figure 1 The specific implementation includes the following steps:
[0022] Step 1: Read training data from pedestrian datasets (such as Market-1501, Duke-MTMC, CUHK03, etc.) and preprocess the training data. The preprocessing operations include adjusting the image size, horizontal random flipping, random erasing, etc.; use batch training to train the neural network. The specific batch collection operation is: randomly select from all the training data in the training set P training categories, and randomly select K training samples, and use these data to form a set of batch training data. Therefore, the number of samples in a set of batch training data is P x K , among which P= 16 ,K= 4.
[0023] Step 2: After preprocessing, the selected pedestrian samples are input into the convolutional neural network based on ResNet-50 to obtain a high-dimensional feature matrix, which is then passed through the following steps: Figure 2 The global average pooling layer, batch normalization layer, fully connected layer and batch normalization layer shown in the figure obtain the pedestrian feature vector ;
[0024] Step 3: After the output pedestrian feature vector passes through the L2 normalization layer, the normalized pedestrian feature vector is obtained. , and order ;
[0025] Step 4: Remove the bias term in the fully connected layer used for classification, and adjust the classification weight vector in the fully connected layer Perform initialization and L2 normalization operations to obtain ; The normalized pedestrian feature vector obtained in step 3 is used as the input of the classification layer to obtain the classification result, and then the softmax loss is used to calculate the classification loss. The classification loss can be expressed as ,in Indicates i The characteristics of the samples, represent The real label, Represents the classification layer j column vector, n is the number of samples for batch training, C is the total number of training categories;
[0026] Step 5: Normalize the weight vector in the classification layer Considered as the feature center of each class, the sum of the cosine distances between each weight vector and all other weight vectors is calculated respectively, and the sum of the sum of each cosine distance is averaged to obtain the classification weight equidistant distribution loss ,in Represents two classification weight vectors and The angle between To avoid the loss value of equal distance classification being too large in the early stage of training, the value is set to 10 -6 ;
[0027] Step 6: Combine the classification loss and the equidistant distribution loss to get the total loss function ,in is the weight value of the equally spaced distribution loss, which is 200;
[0028] Step 7: Calculate the gradient of each input training sample based on the total loss value obtained, and use the back propagation algorithm to update the parameters of the neural network, and train the network until the number of iterations required is met;
[0029] Step 8: After the network training is completed, remove the fully connected layer used for classification in the original network and the network structure after it, and the remaining network becomes a feature extractor; input the test image into the feature extractor to obtain the normalized feature vector , and calculate the Euclidean distance between each feature vector in the query set and all feature vectors in the candidate set, sort them according to the obtained distance results, and calculate the Rank1 and mAP results.
[0030] The above is only a preferred implementation method of the invention of the present invention, but the protection scope of the invention patent is not limited to this. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A pedestrian re-identification method based on a classification weight equidistant distribution loss model, characterized in that, it includes the following steps: Step 1: Input the training pictures into a convolutional neural network to obtain the high-dimensional feature expression vectors of the pictures for subsequent classification learning; Step 2: Use a fully connected layer to perform classification learning on pedestrian features; Before classification, normalize the classification weight vector of the fully connected layer and the pedestrian feature vector obtained in Step 1, and then input the normalized features into the fully connected layer to obtain the classification loss; Step 3: According to each normalized classification weight vector in Step 2, calculate the sum of the distances between each vector and all other weight vectors to obtain the equidistant distribution loss; Step 4: Combine the classification loss and the equidistant distribution loss obtained in Step 2 and Step 3 to obtain the total loss function; Step 5: Use the total loss function obtained in Step 4 to train and optimize the convolutional neural network, and finally obtain the trained network model; Step 6: Input the pedestrian pictures in the test set into the trained network model, and obtain the corresponding feature vectors. At the same time, normalize these feature vectors to obtain the final feature vectors; compare the feature vectors of the pedestrian pictures in the query set with the feature vectors of the pedestrian pictures in the candidate set one by one, and sort according to the distance results to complete the pedestrian re-identification task; In the classification loss required in Step 2, the bias term in the classification layer is removed, and the input feature vector and the classification weights both undergo L2 normalization operations, that is, we obtain and . Therefore, both the feature vector and the classification weights are projected onto the unit hypersphere; let , where s is a parameter for controlling the radius of the hypersphere; finally, the normalized classification loss is , where represents the feature of the i th sample, represents 's true label, represents the j th column vector in the classification layer, n is the number of samples used for batch training, C is the total number of training classes.
2. The method according to claim 1, characterized in that, In step 3, the classification weight vector in the fully connected layer for classification is regarded as the feature center of each type of feature, and the weight vector is subjected to L2 normalization processing; After the normalization operation, the Euclidean distance and the cosine distance between vectors are positively correlated. Therefore, the isometric distribution loss to be calculated is expressed by the cosine distance as , where represents the angle between two classification weight vectors and . is a parameter to avoid the isometric classification loss from being too large at the initial stage of training, C is the total number of training categories.
3. The method according to claim 1, characterized in that, The total loss function used to train the network model in Step 4 is the sum of the normalized classification loss and the equally spaced loss of the weighted classification weights, which is expressed by the mathematical formula as follows: , where is the weight value of the equally spaced loss.
4. The method according to claim 1, characterized in that, In Step 6, perform L2 normalization on the extracted pedestrian features in the test set, then calculate the Euclidean distance between the normalized features of the query set and the candidate set, and sort according to the distance results to complete the re-identification task.
Citation Information
Patent Citations
A method of kinship identification based on angle loss neural network
CN109344759A
Cross-domain pedestrian re-identification method based on unsupervised joint multi-loss model
CN111126360A