An Unsupervised Person Re-identification Method for Distinguishing Similar Pedestrians
By adopting unsupervised dual-stream contrast learning and camera style feature elimination methods based on parameter transfer in unsupervised pedestrian recognition, combined with the unsupervised attention mechanism based on the learnable matrix, the problem of similar pedestrian and camera style differences is solved, and more accurate and efficient pedestrian recognition is achieved.
Patent Information
- Application Number
- CN202211105764.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-09
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-09-09
AI Technical Summary
The existing unsupervised pedestrian re-identification method has difficulties in dealing with differences in styles between similar pedestrians and cameras, resulting in inconsistent feature spatial distribution and making it difficult to effectively distinguish similar pedestrians.
Unsupervised dual-stream comparison learning based on parameter transfer is adopted, and the clustering difference is highlighted by using weighted clustering maximum features and average features when initializing the memory, and the concept of camera style features is proposed to eliminate the impact of camera style offsets. At the same time, an unsupervised attention mechanism based on the learnable matrix is constructed.
Effectively distinguish similar pedestrians, enhance the accuracy of pedestrian feature clustering, improve the distinction ability of pedestrian re-identification models, and overcome the impact of camera style differences on pedestrian re-identification.
Smart Images

Figure CN115457597B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of deep learning and computer vision, and particularly relates to an unsupervised person re-identification method for distinguishing similar pedestrians. Background Art
[0002] Person re-identification aims to find an interested person from a network of multiple cameras. With the continuous growth of video surveillance requirements, the high cost of manual labeling increasingly restricts the application of person re-identification in real scenarios. In recent years, many excellent unsupervised person re-identification methods have been proposed to solve this problem. These methods are divided into two types: unsupervised domain adaptation person re-identification (UDA) and pure unsupervised person re-identification (USL). The person re-identification method based on UDA usually pre-trains the model on a labeled source dataset and fine-tunes the model on an unlabeled target dataset. Different from the two-stage strategy based on UDA, the method based on USL only trains the model on an unlabeled dataset without pre-training on other datasets. Therefore, the person re-identification method based on USL is more challenging. The present invention also belongs to the category of USL.
[0003] Advanced USL methods usually use pseudo-labels to supervise the learning process and adopt a contrastive learning framework with momentum to update the memory. However, these methods all have the following deficiencies:
[0004] 1) These methods use the average features of person clustering to represent the entire clustering, and this operation cannot avoid the similarity between clusters of pedestrians with similar appearances.
[0005] 2) Due to different installation scenarios of cameras, the collected data will have large differences in parameters such as illumination, resolution, and pose, and these differences will seriously affect the spatial distribution of person features in deep learning.
[0006] 3) Person re-identification methods based on the attention mechanism have taken the leading position in the field of supervised person re-identification. However, in unsupervised person re-identification, it is still very difficult to find a simple and effective attention mechanism.
[0007] The above problems are all problems that need to be urgently solved for unsupervised person re-identification. Summary of the Invention
[0008] To solve the above problems, the present invention proposes an unsupervised person re-identification method for distinguishing similar pedestrians. First, the method uses an unsupervised two-stream contrastive learning based on parameter transfer to effectively describe the similarity within clusters and the difference between clusters. From the perspective of specific implementation, the framework uses two different memories, and when initializing the memories, the weighted maximum features and average features of clustering are used respectively, thereby highlighting the differences in clustering and avoiding the mutual influence of pedestrians with similar features. Secondly, in order to eliminate the influence of the style difference between cameras on the distribution of pedestrian feature space, the present invention proposes the concept of camera style features, that is, features that can represent all the offset effects brought by camera differences, and enhances the accuracy of pedestrian feature clustering by eliminating camera style features. And, in view of the current situation that the unsupervised person re-identification lacks an attention mechanism, the present invention constructs a simple and effective unsupervised attention mechanism based on a learnable matrix to improve the attention of the person re-identification model to the key information of pedestrians. The technical solutions proposed by the present invention are as follows:
[0009] An unsupervised person re-identification method for distinguishing similar pedestrians. First, the method uses unsupervised two-stream contrastive learning based on parameter transfer to enhance the difference between clusters of similar pedestrian features. Secondly, in order to overcome the influence of camera differences on the performance of person re-identification, the method extracts camera style features that represent all the offset effects brought by camera differences, and eliminates the camera style features to weaken the influence brought by camera style changes. And, the method constructs a simple and effective unsupervised attention mechanism based on a learnable matrix to solve the problem that the unsupervised person re-identification lacks an attention mechanism. The specific steps include:
[0010] 1) Prepare an unlabeled person re-identification dataset;
[0011] 2) Initialize the unsupervised attention mechanism;
[0012] 3) Use learnable parameters λ a and λ m to weight average pooling and max pooling respectively;
[0013] 4) Integrate the clustering attention mechanism and the weighted max pooling and average pooling into ResNet50 to form a backbone network;
[0014] 5) Use the backbone network to extract pedestrian features;
[0015] 6) Eliminate the influence of camera style offset: Extract camera style features and eliminate the camera style features before the clustering process;
[0016] 7) Initialize the average feature memory M mean and the maximum feature memory M max, with dimensions of (N, C), where N represents the number of clusters, and the average feature memory M is initialized using the average features of pedestrian clusters mean , and the maximum feature memory M is initialized using the maximum features of pedestrian clusters max , when using the maximum feature to represent a cluster, the person re-identification model will focus on the prominent features of pedestrians in the cluster, and the clusters of different pedestrians will become more independent due to the huge differences in prominent features; at the same time, by extracting the weighted parameters λ a and λ m of average pooling and max pooling in the backbone network respectively, M mean and M max are weighted to balance the roles of the average feature memory M mean and the maximum feature memory M max in the training process. Most unsupervised person re-identification methods usually only use the average features of clusters to initialize the memory. This method can ensure the compactness of the pedestrian features within the cluster in the memory, but it cannot guarantee the differences between clusters in the memory. Therefore, the present invention adds the maximum feature memory M max to solve this problem. When using the maximum feature to represent a cluster, the person re-identification model will focus on the prominent features of pedestrians in the cluster, which is also an important basis for distinguishing different clusters. As the training process progresses, the clusters of different pedestrians will become more independent due to the huge differences in prominent features. And, the present invention extracts the weighted parameters λ a and λ m of average pooling and max pooling in the backbone network respectively, M mean and M max are weighted to balance the roles of the average feature memory M mean and the maximum feature memory M max in the training process;
[0017] 8) Generate a training set using pseudo-labels, and use data augmentation techniques to adjust the size, padding, randomly horizontally flip, randomly crop, and randomly erase all images;
[0018] 9) In each iteration, extract the features of the training set images and calculate the contrastive loss with the features in the average feature memory and the maximum feature memory respectively to obtain the losses L mean and L max , calculate the consistency loss L mean between L max and L con , and perform momentum update on the memory;
[0019] 10) During the inference process, use the pedestrian features that reduce the camera style features as the matching features during retrieval.
[0020] Specifically, to address the issue of the lack of an attention mechanism in unsupervised person re-identification, in step 2) of the present invention, a learnable matrix A is first defined to initialize a clustering attention mechanism adapted to unsupervised training. The size of A is the same as the number of channels of the features extracted by layer4 in ResNet50, both represented by C. During the training process, this matrix will be replicated to the size of the training batch, and after passing through the sigmoid activation function, it forms a heat value, and finally multiplies with the features extracted from each image to obtain the attention map of each pedestrian image.
[0021] Specifically, to eliminate the influence of camera differences on unsupervised person re-identification, in step 6) of the present invention, the average features of all images extracted under each camera are used to represent the camera style features: at the beginning of clustering, the features of each image are subtracted by the style features of the camera, so as to obtain pedestrian features without the influence of camera style offset. Finally, this feature is used to generate pseudo-labels through the DBSCAN clustering algorithm, which can ensure the consistency of pedestrian features across cameras.
[0022] Through the above technical solutions, an unsupervised person re-identification method for distinguishing similar pedestrians provided by the present invention achieves the following innovations:
[0023] First, unsupervised two-stream contrast learning based on parameter transfer is used to expand the clustering difference and distinguish similar pedestrians in a targeted manner. Most unsupervised person re-identification methods usually only use the clustering average features to initialize the memory. This method can ensure the compactness of pedestrian features within the clustering in the memory. However, it cannot ensure the difference between clusters in the memory. Therefore, the present invention adds a maximum feature memory M max to solve this problem. When using the maximum feature to represent a cluster, the person re-identification model will focus on the prominent features of pedestrians in the cluster, which is also an important basis for distinguishing different clusters. As the training process progresses, the clusters of different pedestrians will become more independent due to the huge differences in prominent features. And, the present invention extracts the weighted parameters λ a and λ m from the average pooling and maximum pooling in the backbone network respectively to weight M mean and M max respectively, to balance the roles of the average feature memory M mean and the maximum feature memory M max in the training process.
[0024] Second, by eliminating the camera style features, the influence of camera style deviation on pedestrian re-identification is weakened. Previous methods usually adopted multi-stage training and multi-stage losses for distinguishing within-camera and cross-camera, resulting in a complex and inefficient process of eliminating camera differences. To address this problem, the present invention proposes the concept of camera style features, that is, this feature represents the influence of all camera offsets brought about by camera differences, and eliminates the camera style features before the clustering process.
[0025] Third, a new unsupervised clustering attention mechanism is constructed, which effectively enhances the ability of the unsupervised pedestrian re-identification model to distinguish the features of pedestrians with different identities, and solves the problem of the lack of an attention mechanism for unsupervised pedestrian re-identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic flowchart of the present invention.
[0027] Figure 2 It is a model framework diagram proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0028] The method of the present invention will be described in detail below with reference to the drawings and embodiments.
[0029] Referring to Figure 1 and Figure 2 The present invention provides an unsupervised pedestrian re-identification method for distinguishing similar pedestrians, and the method includes the following steps:
[0030] S1: Collect pedestrian bounding boxes from video surveillance data without manual annotation as the dataset x = {x1, x2,..., x n} of the pedestrian re-identification model, where n represents the number of samples.
[0031] S2: Define a learnable matrix A with size C to initialize the clustering attention mechanism.
[0032] S3: Define learnable parameters λ a and λ m with size 1. Take λ a and λ m as the weighted parameters of average pooling and max pooling respectively.
[0033] S4: Integrate the clustering attention mechanism and weighted average pooling and max pooling into ResNet50 to form the backbone network f θ .
[0034] S5: Use the backbone network f θExtract pedestrian features. When the pedestrian image passes through layer 4 of ResNet50, the obtained feature is F, and its size is (B, C, H, W), where B, C, H, and W represent the batch size, the number of channels, the height, and the width, respectively. Then, the learnable matrix A is resized to (B, C, 1, 1) according to the size of B by replicating in the B dimension and adding dimensions in the H and W dimensions. Next, the sigmoid activation function is used to normalize and constrain A, and A and F are multiplied to obtain the attention. Finally, weighted average pooling with λ a and weighted max pooling with λ m are used for feature compression, and the two compressed features are added together to obtain the pedestrian features finally extracted by the backbone network.
[0035] S6: Eliminate the influence of camera style shift on feature clustering. The present invention takes the average feature of all images under the camera as the style feature of the camera. First, the backbone network is used to extract the features of each image under the camera, and the average value is taken to obtain the camera style feature S j , (j = 1, 2,..., N cam ), where N cam represents the number of cameras. To eliminate the influence of camera style shift, the present invention needs to remove the camera style feature from each image. Therefore, the present invention can obtain the feature f c = F ij - S j , where i represents the identity of the pedestrian and j represents the label of the camera. To supervise the training process using pseudo-labels, the present invention uses the clustering algorithm DBSCAN to cluster the pedestrian features. The pedestrian features used for clustering are stack(f c ), where stack represents the stacking operation. At this time, the size of this feature is (n, C), and the obtained pseudo-labels are P i , (i = 0, 1,..., N).
[0036] S7: Initialize the memory. Calculate the average feature c and the maximum feature c′ of the features in each pseudo-label set, and initialize the memories M mean and M max respectively. The sizes of M mean and M max are (N, C). Extract the parameters λ a and λ m of average pooling and max pooling in the backbone network, and weight M mean and M max respectively. c and c′ can be specifically expressed as:
[0037]
[0038] c′ = λ max max(f k )
[0039] where f k indicates that the feature belongs to the k-th cluster, and N fk represents the number of f k . The function of max() is to find the maximum value.
[0040] S8: Generate a training set using pseudo-labels, and use data augmentation techniques to resize, pad, randomly horizontally flip, randomly crop, and randomly erase all images;
[0041] S9: In each iteration, extract the features of the training set images and calculate the contrastive loss with the features in the average feature memory and the maximum feature memory respectively to obtain the losses L mean and L max . Calculate the consistency loss L mean between L max and L con , and perform momentum update on the memory. The specific form of the contrastive loss function is as follows:
[0042]
[0043]
[0044] where τ is the temperature hyperparameter. The process of updating the memory in the present invention is as follows:
[0045] c i ← αc i + (1 - α)f
[0046] c' i ← αc' i + (1 - α)f
[0047] where α represents the momentum update factor. The present invention adopts a smooth L1 norm loss function to calculate L con , and its form is as follows:
[0048] L con = H(s c , s c' )
[0049] where s c , s c' represent the predicted values generated by L mean and L max . Therefore, the final loss function can be expressed as:
[0050] L = λ1L mean + λ2Lmax +λ3L con
[0051] Among them, λ1, λ2, and λ3 represent the balance parameters of three sub-loss functions.
[0052] S10: During the inference process, the pedestrian feature f that reduces the camera style feature c is used as the matching feature during retrieval.
[0053] S11: To verify the accuracy and robustness of the present invention, the present invention is compared with existing algorithms on large-scale pedestrian re-identification datasets Market1501 and MSMT17. The mean average precision and cumulative matching characteristic curve are used to evaluate the model effect. Different from previous pedestrian re-identification algorithms based on PyTorch, the present invention uses Huawei's new generation full-scenario AI computing framework MindSpore to test the algorithm performance, and obtains the recognition result data shown in Table 1.
[0054] Table 1
[0055]
[0056]
[0057] The average precision and first-choice accuracy of the method of the present invention on Market1501 are 84.8% and 93.8% respectively, and the average precision and first-choice accuracy on MSMT17 are 45.4% and 73.7% respectively, exceeding most pedestrian re-identification algorithms.
Claims
1. An unsupervised person re-identification method for distinguishing similar pedestrians, characterized in that: This method first uses unsupervised two-stream contrastive learning based on parameter passing to enhance the differences between similar pedestrian feature clusters; Secondly, in order to overcome the impact of camera differences on the performance of pedestrian re-identification, this method extracts camera style features that represent all the offset effects brought by camera differences and eliminates the camera style features to weaken the impact brought by camera style changes; moreover, this method constructs a simple and effective unsupervised attention mechanism based on a learnable matrix to solve the problem of the lack of an attention mechanism in unsupervised pedestrian re-identification. The specific steps include: 1) Prepare an unlabeled pedestrian re-identification dataset; 2) Initialize the unsupervised attention mechanism; 3) Use the learnable parameters λ respectively a and λ m Weighting average pooling and max pooling; 4) Incorporate the clustering attention mechanism and weighted max pooling and average pooling into ResNet50 to form a backbone network; 5) Use the backbone network to extract pedestrian features; 6) Eliminate the impact of camera style offsets. The specific method is as follows: use the average feature of the features extracted from all images under each camera to represent the camera style feature; at the beginning of clustering, subtract the camera style feature from the feature of each image to obtain a pedestrian feature without the influence of camera style offsets. Finally, use this feature to generate pseudo-labels through the DBSCAN clustering algorithm, which can ensure the consistency of pedestrian features across cameras; 7) Initialize the average feature memory M mean and the maximum feature memory M max , both with dimensions (N, C), where N represents the number of clusters. Initialize the average feature memory M with the average features of pedestrian clusters mean , and initialize the maximum feature memory M with the maximum features of pedestrian clusters max . When using the maximum features to represent a cluster, the person re-identification model will focus on the prominent features of pedestrians in the cluster, and the clusters of different pedestrians will become more independent due to the huge differences in prominent features. At the same time, by extracting the weighted parameters λ a and λ m of average pooling and maximum pooling in the backbone network respectively, weight M mean and M max to balance the roles of the average feature memory M mean and the maximum feature memory M max in the training process; 8) Use the pseudo-labels to generate a training set, and use data augmentation techniques to adjust the size, padding, randomly horizontally flip, randomly crop, and randomly erase all images; 9) During each iteration, extract the features of the training set images and calculate the contrastive losses with the features in the average feature memory and the maximum feature memory respectively to obtain losses L mean and L max , calculate the consistency loss L mean between L max and L con , and perform momentum update on the memory; 10) During the inference process, use the pedestrian feature with the camera style feature reduced as the matching feature during retrieval.
2. The unsupervised person re-identification method for distinguishing similar pedestrians according to claim 1, characterized in that: In step 2), to address the problem of the lack of an attention mechanism in unsupervised pedestrian re-identification, a learnable matrix A is defined to initialize the clustering attention mechanism adapted to unsupervised training. The size of A is the same as the number of channels of the features extracted by layer4 in ResNet50, both denoted by C. This matrix will be replicated to the size of the training batch during training, form a heat value after passing through the sigmoid activation function, and finally multiply it with the feature extracted from each image to obtain the attention map of each pedestrian image.
Citation Information
Patent Citations
Unsupervised cross-domain pedestrian re-identification method based on clustering
CN111860678A
Pedestrian re-identification method based on unsupervised cross-camera clustering comparison
CN114863154A