An Unsupervised Person Re-identification Method Based on Invariance Constraint Contrastive Learning

By introducing a comparative learning method of invariance constraints in unsupervised pedestrian recognition, combining the center, example and camera invariance loss function, the accuracy problems caused by clustering noise and camera style changes are solved, and the accuracy of pedestrian recognition is significantly improved.

CN118691861BActive Publication Date: 2025-05-27WUXI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410807651.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-21
Publication Date
2025-05-27
Estimated Expiration
2044-06-21

AI Technical Summary

Technical Problem

The existing unsupervised pedestrian re-identification method has accuracy problems in clustering generation of noise labels and feature changes caused by different camera styles.

Method used

The comparison learning method based on invariance constraints is adopted to improve the discriminant performance of the ResNet network through central invariance, instance invariance and camera invariance loss functions, alleviate the negative impact of noise samples, and adapt to camera changes.

Benefits of technology

It significantly improves the accuracy of unsupervised pedestrian re-identification, is better than the most advanced methods before, and effectively solves the inaccurate identification problems caused by clustering noise and camera style changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118691861B_ABST
    Figure CN118691861B_ABST
Patent Text Reader

Abstract

The present invention discloses an unsupervised pedestrian re-identification method based on invariance-constrained contrastive learning, which obtains the representative features of the same-class features of pseudo-labels as centroid features. Based on the similarity matrix between the centroid features and query features, in the camera-invariance similarity matrix, samples captured by different cameras within the same cluster are used as positive samples; for the center-invariance similarity matrix and the instance-invariance similarity matrix; the contrastive loss of the pedestrian image is calculated by combining the center-invariance loss, the instance-invariance loss, and the camera-invariance loss; the total loss of the invariance-constrained contrast of the ResNet network is constructed, and the total loss function is used to train the ResNet network for multiple rounds. During each round of training, the unlabeled dataset is re-stratified. The center invariance and instance invariance can alleviate the negative impact of noisy samples, while the camera invariance improves the discriminability by utilizing the camera-aware classification strategy. The performance of learning representations from unlabeled data in the dataset is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of machine vision, and particularly relates to an unsupervised pedestrian re-identification method based on invariance constraint contrast learning. Background Art

[0002] In recent years, pedestrian re-identification has been widely studied in the field of computer vision. The goal is to retrieve a person in videos captured by several non-overlapping cameras given a pedestrian image to be retrieved and output it. Pedestrian re-identification (re-ID) aims to associate images of the same person through non-overlapping camera views. Although great progress has been made in fully supervised methods, in practical applications, label annotation still requires time and effort. To bypass the scarcity of annotations, unsupervised methods that learn discriminative features for identity retrieval from unlabeled data have received increasing attention in recent years. Generally, unsupervised Re-ID can be divided into two categories depending on whether additional labeled data is used, unsupervised domain adaptation (UDA) and fully unsupervised learning (USL). In UDA, data from a labeled source domain and an unlabeled target domain are obtained. The two domains have different distributions and are used to train a model with good generalization characteristics on the target domain. Fully unsupervised person Re-ID is more challenging because only unlabeled images are provided to train the deep model.

[0003] Although most existing methods show good accuracy, they are hindered by the following factors: 1) Noisy labels generated by clustering lead to poor optimization effects, thus hindering the accuracy of model recognition. 2) Feature changes caused by different camera styles result in inaccurate prediction of samples within the same category. Summary of the Invention

[0004] In view of the above-mentioned disadvantages of the prior art, the present invention provides an unsupervised pedestrian re-identification method based on invariance constraint contrast learning to improve the accuracy of pedestrian re-identification.

[0005] To achieve the above effects, the technical solution of the present invention is as follows:

[0006] The present invention provides an unsupervised pedestrian re-identification method based on invariance constraint contrast learning, including the following steps:

[0007] Step 1: Use a camera to perform camera perception classification on pedestrians to obtain pedestrian images; use a pre-trained ResNet network as the backbone encoder to obtain the feature vectors of the pedestrian images;

[0008] Step 2: Obtain the feature vector with the shortest clustering distance in the pedestrian image as the query feature;

[0009] Obtain the labeled dataset and the unlabeled dataset. Take the labeled dataset as one layer, and divide the unlabeled dataset into N layers,

[0010] and assign pseudo-labels to the unlabeled data of each layer respectively to form N layers of pseudo-labeled data, where N is a constant;

[0011] Step 3: Obtain the representative features of the same-class features of the pseudo-labels as the centroid features. Based on the similarity matrix between the centroid features and the query features, in the camera-invariance similarity matrix, use the samples captured by different cameras within the same cluster as positive samples; for the center-invariance similarity matrix and the instance-invariance similarity matrix; combine the center-invariance loss, the instance-invariance loss, and the camera-invariance loss to calculate the contrastive loss of the pedestrian image;

[0012] Step 4: Construct the total loss of the invariance-constrained contrast of the ResNet network, and perform multiple rounds of training on the ResNet network based on the total loss function. During each round of training, re-stratify the unlabeled dataset.

[0013] The present invention proposes center invariance, instance invariance for handling noisy samples, and camera invariance for adapting to camera changes. Center invariance and instance invariance can alleviate the negative impact of noisy samples, while camera invariance improves discriminability by using the camera-aware classification strategy. Extensive tests conducted on two public datasets show that the method proposed by the present invention is superior to the previous state-of-the-art methods and significantly improves the performance of unsupervised person re-identification.

[0014] Furthermore, the obtaining of the feature vector with the shortest clustering distance in the pedestrian image in step 2 includes:

[0015] Use the DBSCAN clustering method to cluster all the labeled data and unlabeled data, and apply the DBSCAN clustering method to the feature vector with the shortest clustering distance.

[0016] Furthermore, the way of assigning pseudo-labels to the unlabeled data of each layer in step 2 is:

[0017] For the nearest neighbor pseudo-labeled data, use the label of the labeled data with the smallest Euclidean distance to it as the pseudo-label of this nearest neighbor pseudo-labeled data;

[0018] For the clustering pseudo-labeled data, cluster all the labeled data and unlabeled data based on the extracted features, and use the label of the labeled data belonging to the same clustering type as the pseudo-label of the clustering pseudo-labeled data in this clustering type.

[0019] Furthermore, the center-invariance loss described in step 3 is specifically:

[0020] Construct an initialized cluster centroid matrix and dynamically update the cluster centroid matrix through the representative of each cluster. The center invariance loss formula is as follows:

[0021]

[0022] Among them, is the query encoding, is the centroid feature sharing the same label, is the first centroid of each cluster obtained according to the pseudo label. The update mechanism is as follows:

[0023]

[0024] Among them, is the momentum update factor, represents cluster sets; |.| represents the number of instances in each cluster, including all feature vectors clusters.

[0025] Furthermore, the instance invariance loss described in step 3 is specifically:

[0026] Obtain the distance between each feature vector and the second centroid , and regard the feature vector with the shortest clustering distance as the representative, and update the centroid to:

[0027]

[0028] Among them, is the initialization value of the first centroid;

[0029] Obtain the instance centroid φn through the above formula, and incorporate it into the contrast loss equation:

[0030]

[0031] Among them, φ + is the instance centroid feature sharing the same label.

[0032] Furthermore, the camera invariance described in step 3 is specifically:

[0033] The camera invariance similarity query function of the camera invariance loss is as follows:

[0034]

[0035] Among them, P represents the positive feature set set with samples captured by different cameras within the same cluster, and the camera cluster belongs to P, and Q represents the negative feature set; represents temperature; the camera invariance loss is expressed as follows:

[0036] .

[0037] Furthermore, the total loss of the invariance constraint contrast of the ResNet network in step 4 is:

[0038]

[0039] wherein, is a hyperparameter that controls the importance of camera invariance.

[0040] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:

[0041] The present invention proposes center invariance, instance invariance for processing noisy samples, and camera invariance loss for adapting to camera changes. By center invariance and instance invariance, the negative impact of noisy samples is alleviated. Camera invariance improves the discriminative performance of the ResNet network by using the camera-aware classification strategy, solves the problem of inaccurate models caused by noisy labels generated by clustering, and improves the accuracy of person re-identification; and solves the problem that the feature changes caused by different camera styles lead to inaccurate prediction of samples within the same category but captured by different cameras. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 is the ICCL framework diagram provided by this application.

[0043] Figure 2 is the t-SNE visualization of the feature distribution of 10 randomly selected labels on Market-1501 and MSMT17 provided by this application.

[0044] Figure 3 is the hyperparameter provided by this application effectiveness diagram.

[0045] Figure 4 is the effectiveness diagram of the training cycle provided by this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following will describe the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention, rather than for limiting the protection scope of the present invention.

[0047] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0048] Embodiment

[0049] This embodiment proposes an unsupervised pedestrian re-identification method based on invariance-constrained contrastive learning. Please refer to Figure 1 , including the following steps:

[0050] Step 1: Use a camera to perform camera perception classification on pedestrians to obtain pedestrian images; use a pre-trained ResNet network as the backbone encoder to obtain the feature vectors of the pedestrian images;

[0051] Step 2: Obtain the feature vector with the shortest clustering distance in the pedestrian image as the query feature;

[0052] Obtain a labeled dataset and an unlabeled dataset. Take the labeled dataset as one layer and divide the unlabeled dataset into N layers,

[0053] and assign pseudo-labels to the unlabeled data of each layer respectively to form N layers of pseudo-labeled data, where N is a constant;

[0054] Step 3: Obtain the representative features of the same-class features of the pseudo-labels as the centroid features. Based on the similarity matrix between the centroid features and the query features, in the camera invariance similarity matrix, use the samples captured by different cameras within the same cluster as positive samples; for the center invariance similarity matrix and the instance invariance similarity matrix, update the memory through different mechanisms during the training process; combine the center invariance loss, the instance invariance loss, and the camera invariance loss to calculate the contrast loss of the pedestrian image;

[0055] Step 4: Construct the total loss of the invariance-constrained contrast of the ResNet network, and perform multiple rounds of training on the ResNet network based on the total loss function. During each round of training, re-stratify the unlabeled dataset.

[0056] Among them, the ResNet network can use ResNet-50. The centroid feature refers to the characterization feature of the same type of features of the pseudo-label obtained by clustering; in the camera invariance similarity matrix, the same type of features are numbered according to the camera, and the feature samples of the same type but different cameras are used as positive samples; the purpose of this is to shorten the distance between the samples of the same type but different cameras to enhance the model's ability to distinguish camera differences; the query feature is the sample obtained from the ResNet network during training as the input value of the loss function. In deep learning, the similarity matrix is ​​a method that runs through the whole process. The similarity of each feature belonging to each class can be obtained through the similarity of the sample features and the centroid (characterization) features, that is, the probability value of belonging to each class; the specific method is to multiply the query sample feature dot product by the transposed matrix of the centroid feature.

[0057] As a preferred technical solution, in this embodiment, the step 2 of obtaining the feature vector with the shortest clustering distance in the pedestrian image includes:

[0058] The DBSCAN clustering method is used to cluster all labeled and unlabeled data, and the DBSCAN clustering method is applied to the feature vector with the shortest cluster distance.

[0059] As a preferred technical solution, in this embodiment, the method of assigning pseudo labels to each layer of unlabeled data in step 2 is:

[0060] For the nearest neighbor pseudo-label data, the label of the labeled data with the smallest Euclidean distance is used as the pseudo-label of this nearest neighbor pseudo-label data;

[0061] For clustered pseudo-label data, all labeled data and unlabeled data are clustered based on the extracted features, and the labels of labeled data belonging to the same cluster type are used as pseudo-labels for clustered pseudo-label data in the cluster type.

[0062] As a preferred technical solution, in this embodiment, the center invariance loss in step 3 is specifically:

[0063] The contrast method aims to minimize the distance between samples belonging to the same category and maximize the distance from different categories. To achieve this goal, an initialized cluster centroid matrix is ​​constructed and dynamically updated by the representative of each cluster. The center invariance loss formula is as follows:

[0064]

[0065] in, is the query code, are the centroid features that share the same label, The first centroid of each cluster obtained according to the pseudo-label is stored in the dynamically updated slots; the update mechanism is as follows:

[0066]

[0067] Among them, is the momentum update factor, denotes the sum of clusters; |.| represents the number of instances in each cluster, including all feature vectors cluster.

[0068] As a preferred technical solution, in this embodiment, the instance invariance loss described in step 3 is specifically:

[0069] The goal of obtaining the clustering centroid is to express all the features of the corresponding cluster to the greatest extent with the centroid; in this way, the multi-class scores of the query features can be calculated with better accuracy in the loss function; however, only using the center invariance loss will result in the loss of original information; if can be utilized, the diversity and accuracy of the representation can be improved; to achieve this goal, the present invention obtains the distance between each feature vector and the second centroid , and regards the feature vector with the shortest clustering distance as the representative, and updates the centroid to:

[0070]

[0071] Among them, is the initial value of the first centroid;

[0072] The instance centroid φn is obtained through the above formula, and is incorporated into the contrastive loss equation:

[0073]

[0074] Among them, φ+ is the instance centroid feature sharing the same label.

[0075] As a preferred technical solution, in this embodiment, the camera invariance described in step 3 is specifically:

[0076] In previous studies, most training schemes used pseudo-labels implemented by clustering methods to assign identities to each sample and compared query samples with the centroid features of each cluster; the performance of such models was relatively insufficient, and at the same time, the influence of camera shift was ignored, which is crucial for optimizing a more robust Re-ID model; the appearance of pedestrians under different cameras may be affected by environmental factors such as viewpoints and lighting, resulting in a large gap between intra-class features; without considering this phenomenon, the trained model may be sensitive to camera changes, which may reduce the clustering results and thus hinder model optimization; to solve this problem, the present invention proposes a camera-aware classification strategy to help the model learn camera-invariant representations.

[0077] The present invention uses camera information to classify pseudo-label data in terms of camera perception; by performing a cross-operation on camera and clustering information, camera clusters are obtained, and the centroid of the camera cluster is used as the initialized camera cluster. , the similarity representation between camera clusters and query instances is calculated by a computer; in the similarity matrix, samples that belong to the same cluster but different cameras are defined as positive samples; the camera-invariance loss has the following camera-invariance similarity query function:

[0078]

[0079] where P represents the positive feature set set with samples captured by different cameras within the same cluster, and the camera cluster belongs to P, and Q represents the negative feature set; represents the temperature; the camera-invariance loss is expressed as follows:

[0080] .

[0081] In this way, the camera-aware classification clustering features located within the same cluster but in different cameras are pulled together, while reducing the variance within a class caused by non-overlapping camera views.

[0082] As a preferred technical solution, in this embodiment, the total loss of the invariance-constrained contrast of the ResNet network in step 4 is:

[0083]

[0084] where is a hyperparameter that controls the importance of camera invariance. Central invariance learning and instance invariance learning effectively reduce the influence of noisy labels; camera invariance significantly reduces the within-class variance caused by non-overlapping camera views. Among them, the camera cluster centroid is finer than the pseudo-label cluster centroid, and noisy pseudo-labels may have a relatively greater adverse impact on the discriminability of the network. Therefore, the present invention adopts a mechanism that does not update the camera cluster centroid.

[0085] As a preferred technical solution, in this embodiment, during training, the data domain in the camera is shown as:

[0086] Cluster in the feature sets with multiple identical camera labels, directly assign pseudo-labels to each training sample feature using the DBSCAN clustering algorithm, calculate the centroid of each cluster according to the pseudo-labels, and initialize the memory storage units under each camera of the two modalities; then use the contrast loss function with distillation parameters to update the feature extractor and the momentum update strategy to update the memory storage units under each camera of the two modalities respectively.

[0087] It should be noted that in a preferred embodiment of the present invention, the unlabeled target domain dataset is clustered to generate pseudo-labels (that is, one data has only one corresponding label). Since using pseudo-labels will bring large noise and interference during the learning process, the predicted value based on ResNet is used as the soft label for learning, that is, its probability is used as the output, and the teacher model is used to mutually learn and supervise the pseudo-labels and soft labels.

[0088] Figure 2 It shows that the ICCL method performs better than the baseline method; Figure 3 It shows that when the parameter λ is selected as 0.3, the performance on the market1501 dataset is better than 0.7, while on MSMT17, 0.7 is better than 0.3; Figure 4 It shows that on the two datasets, as the training period increases, the performance index steadily improves and the function value tends to converge.

[0089] The present invention uses the ResNet network as the backbone of the feature extractor and pre-trains the model with parameter initialization on ImageNet. During the test process, the features of the global average pooling layer are used to calculate the distance. At the beginning of each Epoch, DBSCAN is used for clustering to generate pseudo-labels. The size of the input image is adjusted. For the training images, random horizontal flipping, 10-pixel padding, random cropping, and random erasing are performed. Each mini-batch contains 256 images of 16 pseudo-person identities (each person has 16 instances).

[0090] The present invention uses the Adam optimizer to train the Re-ID model with a weight decay of 5e -4 The initial learning rate is set to 3.5e -460 epochs were trained for the market - 1501 and 80 epochs were trained for MSMT17 respectively. Each epoch contains 200 iterations. The DBSCAN clustering method based on Jaccard distance and k - reciprocal encoding were used for clustering where the maximum distance between two samples is 0.6 and the minimum number of neighbors in the core points is 4. All tests were conducted on two NVIDIA TITAN V GPUs using the Pytorch platform. A large number of tests conducted on two public datasets show that the present invention improves the performance of unsupervised person re - identification.

[0091] Implementing the embodiments of the present invention has the following beneficial effects:

[0092] The present invention sets a loss function and uses the loss function to update a preset model, enabling the ResNet network to learn image features with center - invariance using training images of multiple training sample sets, and capture discriminative information related to image recognition, ignoring the noise of a specific training sample set, thereby obtaining a ResNet network with high generalization performance, which can ensure the accuracy of identifying the image to be recognized, and there is no need to perform additional fine - tuning on the ResNet network, which can avoid overfitting problems and improve the recognition accuracy of the ResNet network.

[0093] Obviously, the above - mentioned embodiments of the present invention are merely examples for clearly explaining the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, based on the above description, various changes or modifications in different forms can be made. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. An unsupervised person re-identification method based on invariance constrained contrastive learning, characterized in that: The following steps are involved: Step 1: Use the camera to perform camera perception classification on pedestrians and obtain pedestrian images; Use the pre-trained ResNet network as the backbone encoder to obtain the feature vector of the pedestrian image; Step 2: Obtain the feature vector with the shortest clustering distance in the pedestrian image as the query feature; Get the labeled dataset and the unlabeled dataset, take the labeled dataset as one layer, and divide the unlabeled dataset into N layers. And assign pseudo labels to the unlabeled data of each layer respectively, forming N layers of pseudo-labeled data, where N is a constant; Step 3: Obtain the representation features of the same type of features of the pseudo-label as the centroid features. Based on the similarity matrix between the centroid features and the query features, in the camera invariance similarity matrix, samples captured by different cameras in the same cluster are used as positive samples; for the center invariance similarity matrix and the instance invariance similarity matrix; combine the center invariance loss, instance invariance loss and camera invariance loss to calculate the contrast loss of the pedestrian image; The center invariance loss is specifically: Construct an initialized cluster centroid matrix and dynamically update the cluster centroid matrix through the representative of each cluster; the center invariance loss formula is as follows: Where q is the query code, ω + is the centroid feature that shares the same label, ω n is the first centroid of each cluster obtained according to the pseudo-label; the update mechanism is as follows: Where μ is the momentum update factor, represents the sum of the n-th clusters; |.| represents the number of instances in each cluster, including all feature vectors in the n-th cluster; The instance invariance loss is specifically: Get the distance b between each eigenvector k and the second centroid C(b k ), and the feature vector with the shortest cluster distance is regarded as the representative, and the centroid is updated as: in, is the initialization value of the first centroid; The instance centroid is obtained by the above formula And incorporate it into the contrast loss equation: in, are the centroid features of instances sharing the same label; The camera invariance is specifically: The camera invariance loss has the following camera invariance similarity query function: Among them, P represents the positive feature set set with samples in the same cluster but captured by different cameras, the camera cluster σ belongs to P, Q represents the negative feature set; τ represents temperature; the camera invariance loss is expressed as follows: Step 4: Construct the total loss of the invariance constraint contrast of the ResNet network, and perform multiple rounds of training on the ResNet network based on the total loss function. The unlabeled dataset is re-stratified during each round of training.

2. The method according to claim 1, characterized in that The step 2 of obtaining the feature vector with the shortest clustering distance in the pedestrian image includes: The DBSCAN clustering method is used to cluster all labeled and unlabeled data, and the DBSCAN clustering method is applied to the feature vector with the shortest cluster distance.

3. The method according to claim 1, characterized in that The method of assigning pseudo labels to the unlabeled data of each layer in step 2 is: For the nearest neighbor pseudo-label data, the label of the labeled data with the smallest Euclidean distance is used as the pseudo-label of this nearest neighbor pseudo-label data; For clustered pseudo-label data, all labeled data and unlabeled data are clustered based on the extracted features, and the labels of labeled data belonging to the same cluster type are used as pseudo-labels for clustered pseudo-label data in the cluster type.

4. The method according to claim 1, characterized in that: The total loss of the invariance constraint contrast of the ResNet network in step 4 is: Among them, λ cam is a hyperparameter that controls the importance of camera invariance.

Citation Information

Patent Citations

  • Unsupervised pedestrian re-identification method based on camera distribution difference alignment constraint

    CN113065409A

  • Unsupervised pedestrian re-identification method based on joint training strategy

    CN114187655A