Multi-cluster guide variable feature contrast learning method and system for unsupervised pedestrian re-identification
Through multi-clustering guided volatile feature comparison learning method, filtering and utilizing volatile features to initialize and update the memory library, the problem of noise label introduction in unsupervised pedestrian re-identification is solved, and the model's discrimination and generalization ability is improved.
Patent Information
- Application Number
- CN202510703480.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-29
AI Technical Summary
The existing unsupervised pedestrian re-identification method is prone to introducing noise labels during training, resulting in limited model generalization ability and retraining effect, especially because strong classes endanger weak classes, resulting in inter-class mergers and the introduction of noise information.
The multi-cluster-guided variable feature comparison learning method is used to filter out variable features by using two clustering algorithms with different parameters, and the initialization and update of the memory bank is used to guide the degree of contribution of the sample to the centroid, construct the variable feature memory bank, and optimize the model through backpropagation.
It effectively suppresses the introduction of noise during training, improves the model's learning ability of non-noise boundary samples, improves the model's discrimination ability and generalization ability, and significantly improves the performance of unsupervised pedestrian re-identification.
Smart Images

Figure CN120236242A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of unsupervised pedestrian re-identification, and more specifically to a multi-clustering guided variable feature comparison learning method and system for unsupervised pedestrian re-identification. Background Art
[0002] Unsupervised person re-identification aims to learn robust and discriminative features from unlabeled datasets to identify specific pedestrians. Currently, the most advanced methods mainly rely on contrastive learning strategies based on memory dictionaries, generate pseudo labels through clustering for model training, and build a memory library to calculate the loss. These methods usually use a one-parameter clustering algorithm to assign pseudo labels and use the average centroid of samples to initialize the memory library. Although they have achieved remarkable results, during the training process, strong classes gradually swallow up weak classes, resulting in inter-class merging and the introduction of a large number of noisy labels, which limits the generalization ability of the model and the retraining effect.
[0003] Unsupervised person re-identification mainly includes two categories: unsupervised domain adaptation (UDA) person re-identification and pure unsupervised (USL) person re-identification. UDA methods learn from an annotated source domain and transfer knowledge to an unlabeled target domain. They usually use a two-stage training strategy. The initial stage is to pre-train the model on a labeled source domain dataset, followed by fine-tuning on an unlabeled target domain dataset. However, due to the need for additional annotation labels, UDA methods are susceptible to the quality of knowledge learned from the source domain and the data differences between the source and target domains, which limits the performance of the model. In contrast, USL methods do not require additional annotation information and are therefore more suitable for real-world scenarios.
[0004] In recent years, the most advanced USL methods have made great progress by using pseudo labels generated by clustering algorithms for training. Such methods generally adopt the following training scheme: 1) Clustering, generating corresponding pseudo labels through clustering algorithms such as DBSCAN; 2) Training, establishing a memory library to store features, and calculating the contrast loss (such as InfoNCE or ClusterNCE) between the input instance (query) and the memory library in a "supervised" manner; 3) Update, in the next iteration, the cluster representation vector in the memory library needs to be updated, so that each iteration stage can prompt the network to learn more discriminative features. However, in the cycle of clustering and training, due to the model's blind trust in pseudo labels, it is often misled by some unreliable label information. In the training stage, the distance between classes gradually increases, and the distance within the class gradually decreases. The class with high density within the class will gradually swallow up the class with low density within the adjacent class, that is, there is a process of strong class swallowing weak classes. As the number of iterations advances, the strong class gradually swallows up the weak class, thereby introducing a large amount of noise information in the strong class, which limits the retraining of the model. Although existing methods have refined pseudo-labels for noisy information, often using auxiliary information such as camera IDs and body part predictions, they have not paid attention to the interactions between classes of different densities during training.
[0005] In USL, the principle of assigning pseudo labels to unlabeled samples through one clustering is usually followed. However, during the training process, there is a specific boundary feature that directly induces the merging phenomenon between classes, which is called "volatile feature". Unlike noise information, the volatile feature itself is not noise information, but a feature that may lead to the introduction of potential noise in subsequent training. Such features often exist in weak classes with low density, while weak classes contain rich clustering boundary information and are close to features of other classes. Due to the low intra-class density of the class to which it belongs, the volatile feature is easily attracted by the strong class with high intra-class density, thereby introducing a lot of noise in the strong class. Figure 1 (a) shows this specific process. It is observed that the mutable features have significant sensitivity to the clustering parameters, such as Figure 1 As shown in (b), using a variety of different clustering parameters can effectively affect the distribution trend of volatile features and play a screening role to a certain extent. Since volatile features embed boundary sample information that is easily confused with other classes, they are highly sensitive to clustering parameters. As a non-noise boundary information, volatile features can guide the model to form more reliable clusters.
[0006] Therefore, how to propose a multi-cluster-guided volatile feature contrast learning method and system for unsupervised person re-identification, cluster and screen out volatile features, use volatile features to guide the initialization of the memory bank, prompt the model to learn a more fine-grained feature distribution during the update process of the memory bank, and to a certain extent solve the respective disadvantages of using the hardest samples and average samples, propose a volatile feature mining loss, taking into account both the cluster level and the instance level, and guide the model to more fully exploit the information in the volatile features is an urgent problem to be solved by those skilled in the art. Summary of the Invention
[0007] In view of this, the present invention provides a multi-cluster-guided volatile feature contrast learning method and system for unsupervised person re-identification, guiding the model to more fully exploit the information in the volatile features. To achieve the above object, the present invention adopts the following technical solutions: A multi-cluster-guided volatile feature contrast learning method for unsupervised person re-identification, comprising: Collect pedestrian image data, and construct an image data set; Extract features from the image data set through an unsupervised person re-identification model to obtain samples to be processed; Use two clustering algorithms with different parameters to assign labels to the samples to be processed, and screen out volatile features by comparing the two clusters to guide the initialization of the memory bank; Process the initialized memory bank through a multi-cluster centroid scheduler to obtain a volatile feature memory bank; Utilize the context information in the mini-batch samples to be processed, dynamically assign different weights to each sample, and construct a dynamic centroid to update the volatile feature memory bank; Perform backpropagation optimization on the unsupervised person re-identification model through the volatile feature memory bank to obtain an optimal unsupervised person re-identification model.
[0008] Optionally, the use of two clustering algorithms with different parameters includes: ; Wherein, , are the clustering results and labels of the baseline clustering respectively, , are the clustering results and labels of the extended clustering respectively, , are the hyperparameters of the control clustering respectively, and are the expansion factors of the control hyperparameters.
[0009] Optionally, it further includes adding the tags of the corresponding features in the reference tag and the extended tag to obtain a comparison tag. The features at the intersection of classes in the reference clustering will have a situation where the number of intra-class features in the comparison clustering is less than the hyperparameter of the reference clustering. In this case, these features are marked as clustering-sensitive features: ; ; ; ; Among them, and are respectively the clustering result and the tag of the comparison clustering. represents the mapping from the tag to the clustering result, num represents the set of the number of features in each class in the comparison clustering, represents the number of features in the i-th class among them, is the j-th feature in the i-th class of the comparison clustering, is the clustering hyperparameter, is the clustering-sensitive feature.
[0010] Optionally, construct an uncertainty factor to screen the clustering-sensitive features to obtain volatile samples: ; Among them, represents the j-th feature in the i-th class, is the clustering-sensitive feature, represents the number of features in the i-th class. Set a threshold Q to control the selection of volatile features. For the clustering-sensitive features, when the value of p is greater than Q, perform clustering-sensitive feature filtering to obtain the final volatile features .
[0011] Optionally, it further includes using the final volatile features to guide the formation of the centroid. The centroid of the k-th class obtained is as follows: ; Among them, is the weight of the volatile feature, is the volatile feature, M is the total number of features in the k-th class, , by increasing the weight of the volatile features in the process of forming the clustering centroid, a centroid containing more non-noise boundary information is obtained, enabling the model to learn the boundary of each clustering, and using the centroids regulated by weights to initialize the memory bank.
[0012] Optionally, the construction of a dynamic centroid to update the volatile feature memory bank by using the context information in a small batch of samples to be processed and dynamically assigning different weights to each sample includes: According to the distance between the sample and the centroid, use the Softmax function to assign different weights to each query sample, reflecting the overall distribution of features: ; where N represents the number of instances of the i-th class in the small batch, is the centroid of the i-th class, and are sample instances, is the linear scheduler.
[0013] Optionally, the linear scheduler includes: ; where is the initial linear scheduler, and the weights are dynamically changed throughout the training process: The dynamic centroid of the i-th class is obtained as follows: ; Update the memory bank using the dynamic centroid: ; where is the hyperparameter that controls momentum update.
[0014] Optionally, the backpropagation optimization of the unsupervised person re-identification model through the volatile feature memory bank includes: Based on all volatile features, select representative features from them to construct a loss function. By calculating the inner product of each query instance and the volatile feature and selecting the volatile representative positive sample vector and the volatile representative negative sample vector , capture the information of the volatile features: ; where and are the volatile features of the same class and different classes as the query instance q respectively, is the volatile feature with the farthest distance in the feature space among the volatile features of the same class as the query instance q, while is the volatile feature with the closest distance in the feature space among the volatile features of different classes from the query instance q.
[0015] Optionally, it also includes obtaining the volatile positive sample and volatile negative sample with the largest boundary information through the screening of vectors, and constructing a volatile feature loss: ; Among them, is the temperature hyperparameter, is the expected value, and the total loss of volatile feature mining consists of two parts: ClusterNCE loss and volatile feature loss: ; Among them, is the ClusterNCE loss, is the weight coefficient used to balance these two losses.
[0016] Optionally, a multi-cluster-guided volatile feature contrast learning system for unsupervised person re-identification includes: Acquisition module: used to acquire pedestrian captured image data and construct an image data set; Feature extraction module: used to extract features from the image data set through an unsupervised person re-identification model to obtain samples to be processed; Initialization module: used to assign labels to the samples to be processed using two clustering algorithms with different parameters, and screen out volatile features by comparing the two clusters to guide the initialization of the memory bank; Volatile feature memory bank construction module: used to process the initialized memory bank through a multi-cluster centroid scheduler to obtain a volatile feature memory bank; Update module: used to utilize the context information in the mini-batch samples to be processed, dynamically assign different weights to each sample, and construct a dynamic centroid to update the volatile feature memory bank; Backpropagation optimization module: used to perform backpropagation optimization on the unsupervised person re-identification model through the volatile feature memory bank to obtain an optimal unsupervised person re-identification model.
[0017] It can be seen from the above technical solutions that compared with the prior art, the present invention discloses a multi-cluster-guided volatile feature contrast learning method and system for unsupervised person re-identification, which has the following beneficial effects: The present invention proposes a multi-cluster guided variable feature contrast learning method for unsupervised person re-identification, including: collecting pedestrian image data, constructing an image dataset; extracting features from the image dataset through an unsupervised person re-identification model to obtain samples to be processed; using two clustering algorithms with different parameters to assign labels to the samples to be processed, and screening out variable features by comparing the two clusters to guide the initialization of the memory bank; processing the initialized memory bank through a multi-cluster centroid scheduler to obtain a variable feature memory bank; using the context information in the mini-batch samples to be processed, dynamically assigning different weights to each sample, and constructing a dynamic centroid to update the variable feature memory bank; optimizing the unsupervised person re-identification model through backpropagation using the variable feature memory bank to obtain an optimal unsupervised person re-identification model. The present invention proposes a multi-cluster guided variable feature contrast learning method (MGVF), designs a multi-cluster centroid regulator (MCM), screens out variable features through multiple clustering algorithms with different parameters, and dynamically adjusts the contribution degree of each sample in the memory bank to its corresponding centroid, strengthening the model's learning of non-noise boundary samples, thereby effectively suppressing the introduction of noise during the training process. In addition, a dynamic global feature update strategy (DGF) is also proposed, which couples the update process of the feature bank with different stages of model training and the context relationship of samples. Finally, in order to further improve the discriminative ability and generalization ability of the model, a variable feature mining loss (VFE) is constructed to enhance the model's adaptability to complex scenarios. A large number of experiments show that MGVF is effective and achieves advanced performance in the field of unsupervised person re-identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0019] Figure 1 It is a schematic diagram of the principle of traditional pure unsupervised (USL) person re-identification provided by the present invention.
[0020] Figure 2 It is a schematic diagram of the principle of the process of screening variable features and initializing the variable feature memory bank provided by the present invention.
[0021] Figure 3 It is a schematic diagram of the principle of a multi-cluster guided variable feature contrast learning method for unsupervised person re-identification provided by the present invention.
[0022] Figure 4 It is a schematic diagram of the principle of the formation process of the dynamic centroid provided by the present invention.
[0023] Figure 5(a) is a schematic diagram showing the relationship between the Φ value and the number of training rounds under the action of the logarithmic scheduler provided by the present invention.
[0024] Figure 5(b) is a schematic diagram showing the relationship between the Φ value and the number of training rounds under the action of the linear scheduler provided by the present invention.
[0025] Figure 5(c) is a schematic diagram showing the relationship between the Φ value and the number of training rounds under the action of the exponential scheduler provided by the present invention.
[0026] Figure 5(d) is a graph showing the relationship between the model performance (mAP) and the Φ value on the Market-1501 dataset when there is no scheduler action (i.e., when Φ is a constant) provided by the present invention.
[0027] Figure 6 Figure is a graph showing the TSNE visualization analysis results of the baseline model and the MGVF model provided by the present invention on the Market-1501 dataset.
[0028] Figure 7(a) is a graph showing the visualization analysis results of the intra-class and inter-class distances of the baseline model on the Market-1501 dataset provided by the present invention.
[0029] Figure 7(b) is a graph showing the visualization analysis results of the intra-class and inter-class distances of the MGVF model on the Market-1501 dataset provided by the present invention.
[0030] Figure 8(a) is a graph showing the visualization analysis results of the Rank-list of the baseline model provided by the present invention.
[0031] Figure 8(b) is a graph showing the visualization analysis results of the Rank-list of the MGVF model provided by the present invention.
[0032] Figure 9(a) shows the analysis results of the parameters on the Market-1501 dataset provided by the present invention.
[0033] Figure 9(b) shows the analysis results of the parameters on the Market-1501 dataset provided by the present invention. Detailed implementation manners
[0034] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0035] An embodiment of the present invention discloses a multi-cluster-guided variable feature contrast learning method for unsupervised person re-identification, including: Collect pedestrian image data and construct an image dataset; Extract features from the image dataset through an unsupervised person re-identification model to obtain samples to be processed; Use two clustering algorithms with different parameters to assign labels to the samples to be processed, and screen out variable features by comparing the two clusters to guide the initialization of the memory bank; Process the initialized memory bank through a multi-cluster centroid scheduler to obtain a variable feature memory bank; Utilize the context information in the mini-batch samples to be processed, dynamically assign different weights to each sample, and construct a dynamic centroid to update the variable feature memory bank; Optimize the unsupervised person re-identification model through the variable feature memory bank by backpropagation to obtain an optimal unsupervised person re-identification model.
[0036] Furthermore, the use of two clustering algorithms with different parameters includes: ; Among them, , are the clustering results and labels of the baseline clustering respectively, , are the clustering results and labels of the extended clustering respectively, , are the hyperparameters of the control clustering respectively, and is the extension factor of the control hyperparameter.
[0037] Furthermore, it also includes adding the labels of the corresponding features in the baseline label and the extended label to obtain a comparison label. For the features at the intersection of classes in the baseline clustering, the number of intra-class features in the comparison clustering will be less than the hyperparameter of the baseline clustering. Mark these features as clustering-sensitive features: ; ; ; ; Among them, and are the clustering results and labels of the comparison clustering respectively, represents the mapping from the label to the clustering result, num represents the set of the number of features in each class in the comparison clustering, represents the number of features in the i-th class among them, For the j-th feature of the i-th class in the contrast clustering, is the clustering hyperparameter, and is the clustering sensitive feature.
[0038] Furthermore, construct an uncertainty factor for the clustering sensitive feature to screen and obtain volatile samples: ; wherein, represents the j-th feature in the i-th class, is the clustering sensitive feature, represents the number of features in the i-th class, set a threshold Q to control the selection of volatile features. For the clustering sensitive feature, when the value of p is greater than Q, perform clustering sensitive feature filtering to obtain the final volatile features .
[0039] Furthermore, it also includes using the final volatile features to guide the formation of the centroid, and the centroid of the k-th class obtained is as follows: ; wherein, is the weight of the volatile feature, is the volatile feature, M is the total number of features in the k-th class, , by increasing the weight of the volatile feature in the process of forming the clustering centroid, obtain a centroid containing more non-noise boundary information, let the model learn the boundary of each cluster, and use the centroid regulated by the weight to initialize the memory bank.
[0040] Furthermore, using the context information in the mini-batch of samples to be processed, dynamically assign different weights to each sample, and constructing a dynamic centroid to update the volatile feature memory bank includes: According to the distance between the sample and the centroid, use the Softmax function to assign different weights to each query sample, reflecting the overall distribution of the features: ; wherein, N represents the number of instances of the i-th class in the mini-batch, is the centroid of the i-th class, and are sample instances, is the linear scheduler.
[0041] Furthermore, the linear scheduler includes: ; wherein, is the initial linear scheduler, and the weight changes dynamically during the entire training process: The dynamic centroid of the $i$-th class is obtained as follows: ; Update the memory bank using the dynamic centroid: ; where, is the hyperparameter that controls the momentum update.
[0042] Optionally, the backpropagation optimization of the unsupervised person re-identification model through the volatile feature memory bank includes: Based on all volatile features, select representative features from them to construct a loss function, and calculate the inner product of each query instance and the volatile features and select the volatile representative positive sample vector and the volatile representative negative sample vector to capture the information of the volatile features: ; where, and are the volatile features of the same class and different classes as the query instance $q$ respectively, is the volatile feature that is farthest in the feature space among the volatile features of the same class as the query instance $q$, while is the volatile feature that is closest in the feature space among the volatile features of different classes from the query instance $q$.
[0043] Optionally, it also includes obtaining the volatile positive sample and volatile negative sample with the largest boundary information through the screening of vectors, and constructing the volatile feature loss: ; where, is the temperature hyperparameter, is the expected value, and the total volatile feature mining loss is composed of two parts: the ClusterNCE loss and the volatile feature loss: ; where, is the ClusterNCE loss, is the weight coefficient used to balance these two losses.
[0044] In a specific implementation manner, a multi-cluster guided volatile feature contrast learning system for unsupervised person re-identification includes: Acquisition module: used to acquire pedestrian captured image data and construct an image data set; Feature extraction module: used to extract features from the image data set through an unsupervised person re-identification model to obtain samples to be processed; Initialization module: used to assign labels to samples to be processed using clustering algorithms with two different sets of parameters, and to screen out volatile features by comparing the two clusters to guide the initialization of the memory bank; Volatile feature memory bank construction module: used to obtain a volatile feature memory bank by processing the initialized memory bank through a multi-cluster centroid scheduler; Update module: used to utilize the context information in a small batch of samples to be processed, dynamically assign different weights to each sample, and construct a dynamic centroid to update the volatile feature memory bank; Backpropagation optimization module: used to perform backpropagation optimization on an unsupervised person re-identification model through the volatile feature memory bank to obtain an optimal unsupervised person re-identification model.
[0045] In the specific implementation, for Re-id methods such as USL, the goal is to train a robust deep neural network to maximize the distance between the features of different individuals and minimize the difference between the features of the same individual. In Cluster Contrast learning, the dataset is represented as , where represents an unlabeled image, and there are a total of N images. For each sample in the dataset, the Re-id model will generate a feature for it. At the beginning of each epoch, DBSCAN is used to cluster these features and assign corresponding pseudo-labels to them. At the same time, before each round of training, according to the clustering results, the average feature vector in each cluster of samples is calculated, and the memory bank is initialized using these feature vectors and the corresponding pseudo-labels. In each clustering result, the clustering representation vector of the k-th class is calculated as follows: ; where, represents the set of features of the samples falling into the k-th class, represents the number of features in the set, and the network is optimized using the obtained pseudo-labels. The objective function is ClusterNCE, and its formula is: ; where, is the clustering representation vector corresponding to the feature q, C is the total number of clusters in the pseudo-labels, is the temperature factor, is the expected value. In each epoch, the memory bank provides rich and diverse samples for loss calculation by maintaining a long-term feature storage. Therefore, the model can effectively utilize the information of historical samples, thereby enhancing the understanding of new samples and improving its adaptability in different perspectives and environments. To ensure that the model can access diverse data, learn more comprehensive features, and reduce the risk of overfitting, the memory bank is updated in a momentum manner during the training iteration: ; where, is the momentum update factor, is the clustering representation vector of the i-th class in the current mini-batch.
[0046] In the specific implementation manner, the embodiment of the present invention follows the baseline model of clustering contrast learning, and on this basis, proposes a contrast learning algorithm (MGVF) guided by volatile features. Its framework is as shown in Figure 3 . In the clustering algorithm, the purpose of using bi-clustering is mainly to screen volatile features (star image markers). The initial feature library is formed by assigning pseudo-labels through solid-line clustering in the figure. After obtaining the volatile features, the MCM is used to guide the formation of the initial memory bank into a volatile feature memory bank for subsequent loss calculation. It also includes other modules such as a backbone network and a clustering algorithm. There are mainly three differences from previous work: 1) Memory bank. Instead of using the average centroid of samples as the final memory bank, at the beginning of clustering, two clustering algorithms with different parameters are used to assign labels to samples, and a multi-clustering centroid scheduler is used to provide a better clustering representation vector for feature learning. 2) Update of the memory bank. Instead of using the average centroid in the mini-batch as the update of the memory bank, different weights are assigned to each sample. On this basis, the weight of each sample will change with the change of the epoch. 3) Loss function. Combining ClusterNCE, a volatile feature mining loss is constructed to further mine volatile features, forming a new total loss function.
[0047] In the specific implementation manner, the multi-clustering centroid scheduler specifically includes: S11: In methods such as USLRe-id, clustering algorithms such as DBSCAN are often used to assign a pseudo-label to samples for subsequent training. However, the hyperparameters of clustering directly control the classification results, and some clustering boundary features are often affected by this single label assignment method and are classified into categories that do not belong to them, which is particularly obvious in the early stage of training. To solve the above problems, two clusterings with different parameters are used, and volatile features are screened by comparing the two clusterings to guide the initialization of the memory bank. The clustering formula is as follows: ; Among them, , are the clustering results and labels of the baseline clustering respectively, , are the clustering results and labels of the extended clustering respectively. , are the hyperparameters of the control clustering respectively, and is the extension factor for controlling the hyperparameters. Among them, is too large, is too small will introduce a large amount of noise, is too small, is too large will exclude a large number of valid samples.
[0048] S12: Add the labels corresponding to the features in the baseline label and the extended label to obtain the comparison label. Affected by the volatile features, the features at the boundary between classes in the baseline clustering will have a situation where the number of intra-class features in the comparison clustering is less than that in the baseline clustering of this situation, mark these features as clustering-sensitive features: Among them, and are the clustering results and labels of the comparison clustering respectively, represents the mapping from the label to the clustering result, num represents the set of the number of features in each class in the comparison clustering, represents the number of features in the i-th class among them. is the j-th feature in the i-th class in the comparison clustering. is the clustering hyperparameter.
[0049] S13: The clustering-sensitive features contain a large amount of indistinguishable edge information, which is one of the reasons for the confusion between features. However, these features are not completely equivalent to the volatile features. Especially in the early stage of model training, a large number of features will be marked as clustering-sensitive features, but a considerable part of these marked features are noise information generated due to insufficient discriminative ability of the model. Therefore, in order to further screen out the volatile samples, an uncertainty factor is proposed: ; Among them, represents the j-th feature in the i-th class, is the class-sensitive feature, Represents the number of features in the \(i\)-th class. The higher the uncertainty factor, the lower the confidence of the class. If the volatile features are still marked as volatile classes, it will mislead the model to learn incorrect noise information. Therefore, a threshold is set to control the selection of volatile features. For clustering-sensitive features, when the value of \(p\) is greater than 0.5, it means that there is more noise information in them. Filter out these clustering-sensitive features to obtain the final volatile features: Volatile features are rich in a large amount of clustering boundary information, which is the main reason for the merger between classes.
[0050] S14: Different from the average centroid obtained by averaging the samples, using volatile features to guide the formation of the centroid, the \(k\)-th class centroid obtained is as follows: ; where is the weight of the volatile features, \(M\) is the total number of features in the \(k\)-th class. It should be noted that . Figure 2 Shows the process of screening volatile features and initializing the volatile feature memory bank. For the clustering in the first figure on the left, the hexagons represent clustering-sensitive features. When the uncertainty factor is greater than 0.5, the class has a lower confidence, and the potential noise information it contains is larger. Therefore, the marking of this part of the clustering-sensitive features should be discarded. After screening by the uncertainty factor, volatile features are obtained, and the volatile features are used to guide the formation of the volatile memory bank. For the four clusters in the second figure on the left, the solid arrows represent embedding more volatile feature information into the centroid of the ordinary memory bank, and the thickness represents the degree of embedding, while the dashed arrows do not perform any processing. Specifically, by increasing the weight of volatile features in the process of forming the clustering centroid, a centroid containing more non-noise boundary information can be obtained, thereby guiding the model to better learn the boundaries of each cluster, which greatly improves the separability between classes. In addition, this method effectively inhibits the merger of strong classes and weak classes, preventing further introduction of noise information during the training process. Finally, the memory bank is initialized using the centroids with weight regulation.
[0051] In the specific implementation manner, the specific steps of the dynamic global update strategy include: S21: Although the average strategy adopted in SpCL and CC has achieved remarkable results in memory bank updates, this strategy fails to effectively reflect the overall distribution of features, and the model performs poorly when learning more challenging samples. In HHCL, ICE, and HDCRL, to further improve the generalization ability of the model, a method of updating the memory bank with hard samples is adopted. However, at the initial stage of training, the discriminative ability of the model is insufficient, and this method may introduce a large number of incorrect samples, thus misleading the learning of the model. To address the deficiencies of these two update methods and enable the model to make more full use of context information, a method of updating the memory bank based on the centroid of dynamic weights is proposed. Specifically, according to the distance between the sample and the centroid, the Softmax function is used to assign different weights to each query sample, so as to more accurately reflect the overall distribution of features: ; where N represents the number of instances of the i-th class in the mini-batch, is the centroid of the i-th class, and are sample instances. The greater the distance between the sample and the centroid in the feature space, the greater the assigned weight, and thus the richer the embedded feature information. At the initial stage of training, due to more noise information, the model should carefully mine hard samples to avoid introducing too much noise information. As the number of epochs increases, the intra-class features become more compact, the inter-class distance gradually increases, and the model should gradually emphasize the mining of hard samples and increase the proportion of difficult samples in the clustering representation vector. Therefore, a linear scheduler : ; Specifically, the larger the value of, the more emphasis is placed on the importance of boundary samples.
[0052] S22: The weights are not only dynamically assigned within the same epoch but also change dynamically throughout the training process. The dynamic centroid of the i-th class is as follows: ; These dynamic centroids are used to update the memory bank: ; where, is a hyperparameter that controls momentum update. Figure 4Shows the formation process of the dynamic centroid. The DGF consists of two parts, dynamic weight allocation and linear scheduler. The static average weight vector is obtained from the samples in the mini-batch, and the dynamic weight vector is obtained after passing through the dynamic weight allocation and linear scheduler. Samples in the mini-batch that are farther away from the centroid in the feature space will be assigned higher weights, and this assignment ratio is also affected by the linear scheduler during the training loop. The dynamic centroid not only incorporates rich context information but is also more sensitive to the distribution of features. In addition, as the number of epochs increases, the dynamic centroid can be adaptively adjusted to better fit the changing trends of intra-class and inter-class distances.
[0053] In the specific implementation, the specific steps of the volatile feature mining loss include: The multi-cluster centroid scheduler reconstructs the proportion of each feature in the centroid, providing a more reliable basis for the initialization of the memory bank. However, the formation of the centroid is still limited to the level of each class and does not take into account the information of each volatile feature itself, which may affect the adaptability of the model in various complex scenarios. Therefore, a volatile feature mining loss is proposed, aiming to deeply mine the volatile feature information and make full use of the potential of volatile samples. Volatile features are often challenging, but not all volatile features are worth mining. Some volatile features are located close to the centroid of the class and may be marked due to insufficient discriminative ability of the model. Therefore, based on all volatile features, the most representative features are selected to construct the loss function. Calculate the inner product of each query instance and the volatile feature and select the volatile representative positive sample vector and the volatile representative negative sample vector . In this way, the information of volatile features can be captured more effectively, thereby improving the adaptability of the model: ; where, and are volatile features of the same class and different classes as the query instance q respectively, is the volatile feature with the farthest distance in the feature space among the volatile features of the same class as the query instance q, while is the volatile feature with the closest distance in the feature space among the volatile features of different classes from the query instance q. For ordinary classes, since there are no corresponding volatile features to mine, the average centroid is used for substitution. By screening the vectors, the volatile positive samples and volatile negative samples with the largest boundary information can be obtained, thereby constructing the following volatile feature loss: ; where, is the temperature hyperparameter. The total volatile feature mining loss consists of two parts, the ClusterNCE loss and the volatile feature loss: ; wherein, is the weight coefficient for balancing these two losses.
[0054] In the specific implementation manner, the experiments on the multi-cluster-guided variable feature contrast learning method for unsupervised person re-identification include: S31: Dataset and evaluation metrics The proposed method was verified on the Market1501, DukeMTMC-reID, and MSMT17 datasets. The Market150 dataset consists of 32,668 images covering 1501 identities, which were captured by 6 cameras. The training set includes 12,936 images involving 751 pedestrian identities, while the test set contains 19,732 images involving 750 pedestrian identities.
[0055] The DukeMTMC-reID dataset consists of 36,411 images of 1812 pedestrians captured by 8 different cameras. Among them, the training set includes 16,522 images of 702 pedestrian IDs, and the test set contains 17,661 images of 702 pedestrians and 408 interfering pedestrians. The query set contains 702 pedestrians in the test set, and one image is randomly selected for each of the 702 pedestrians in each camera, for a total of 2228 images. MSMT17 is composed of 126,441 images of 4101 pedestrian IDs captured by 15 cameras. The training set contains 32,621 images of 1041 pedestrian IDs, and the test set contains 93,820 images from 3060 pedestrian IDs.
[0056] All experiments used the Rank-1, Rank-5, Rank-10 accuracies and mean average precision (mAP) of cumulative matching characteristics (CMC). No post-processing methods such as Re-ranking were used during the testing process.
[0057] Table 1 Comparison results with the current state-of-the-art unsupervised Re-ID methods on the Market-1501, DukeMTMC-reID, and MSMT17 datasets The best results are marked in bold. † indicates the use of additional camera information.
[0058] S32: Experimental details Based on the work of others in the previous field, ResNet-50 pre-trained on ImageNet was used as the backbone network to endow it with the most basic discriminative ability. All modules after the fourth layer were removed, and a Generalized Mean Pooling (GeM) layer was added, followed by a batch normalization layer and a normalization layer, and finally 2048-dimensional features were generated as the output for each image.
[0059] Table 2 Effectiveness results of each component in the proposed Multi-Clustering Guided Variable Feature Contrastive Learning (MGVF) Among them, MGVF includes a Multi-Clustering Centroid Modulator (MCM), a Dynamic Global Feature (DGF), and a Variable Feature Mining Loss (VFE).
[0060] Table 3 Comparison results of different scheduler strategies on the Market1501 and DukeMTMC-reID datasets Training was carried out on two RTX3090s. At the beginning of the training phase, the size of the input images was adjusted to 256×128. Random inversion, cropping, and erasing were adopted as data augmentation. Each mini-batch contained 64 images from four pedestrian IDs, and the batch size was 256. These images were randomly selected from the training set. The Adam optimizer was used, and its weight decay coefficient was . The initial learning rate was 0.00035, and a warmup strategy was adopted in the first 10 epochs. and were set to 0.1 and 1 respectively. For the hyperparameters of DBSCAN, the maximum distance between two samples on Market-1501 and DukeMTMC-reID was set to 0.6, and the minimum number of neighbors of the core points was set to 4. On MSMT17, they were set to 0.45 and 4 respectively. According to the experience of previous work, the momentum update hyperparameter was set to 0.2. Training was carried out for 70 epochs on Market-150 and DukeMTMC-reID, and 80 epochs on MSMT17.
[0061] S33: Comparison with the performance of the state-of-the-art models Comparative experiments were conducted on the proposed method (MGVF) and the current state-of-the-art pure unsupervised person re-identification (USL) methods on three datasets, Market-1501, DukeMTMC-reID, and MSMT17. The experimental results are shown in Table 1. As can be seen from Table 1, the proposed method achieved 86.7% mAP and 94.5% Rank-1 on the Market-1501 dataset, 76.7% mAP and 86.9% Rank-1 on the DukeMTMC-reID dataset, and 41.8% mAP and 72.6% Rank-1 on the MSMT17 dataset. The overall performance is better than BUC, SSL, MMCL, JVCT+, HCT, SpCL, GCL, IICS, ClusterContrast, PPLR, ISE, HDCRL, AdaMG, RTMem, and DHCCN, demonstrating significant advantages and strong competitiveness. It is worth mentioning that different from CAP, ICE, and IIDS, the present invention does not use any additional camera information, indicating that MGVF can achieve better performance under more limited information conditions. In addition, it also performs well compared with some well-known supervised methods.
[0062] S34: Ablation Experiments To verify the effectiveness of the proposed method, detailed ablation experiments were conducted on the Market1501 and DukeMTMC-reID datasets. Compared with the baseline model (Cluster Contrast), three key components were deeply verified: Multicluster Centroid Modulator, Dynamic Global Feature Update, and Variable Feature Exploration Loss. The relevant results are shown in Table 2.
[0063] Effectiveness of the Multicluster Centroid Modulator. By comparing #1 and #2, #3 and #5, #4 and #6, #7 and #8, the effectiveness of MCM can be illustrated. For #1 and #2, the baseline model improved by 1.3% and 1.1% in mAP and 0.4% and 0.6% in Rank-1 on the Market1501 and DukeMTMC-reID datasets respectively through MCM. Compared with the strategy of using the average centroid to initialize the memory bank in the baseline model, MCM effectively guides the model to push the distance between different classes farther by embedding more reliable marginal sample information into the cluster centroids, avoiding the merging of classes during the training process.
[0064] Effectiveness of Dynamic Global Feature Update. By comparing #1 and #3, #2 and #5, #4 and #7, #6 and #8, the effectiveness of DGF can be illustrated. Dynamic Global Feature Update assigns different weights according to the distance between the feature and the mini-batch center, enabling more context information to be embedded in the cluster representation vector. In addition, to avoid introducing incorrect samples in the early stage of training, a linear scheduler is used to adjust the proportion of hard samples during training, while overcoming the deficiencies of using the average centroid and the hardest sample update in the minibatch. For #1 and #3, the baseline model improved the mAP by 1.0% and 0.9% and Rank-1 by 0.2% and 0.5% on the datasets Market1501 and DukeMTMC-reID respectively through MCM. To further illustrate the effectiveness of each part inside the dynamic global feature update and compare with other strategies that can gradually increase the proportion of difficult samples, detailed comparative experiments were conducted. As shown in Figure 5 (a), Figure 5 (b), Figure 5 (c) and Figure 5 (d), three different strategies of logarithmic growth, linear growth, and exponential growth were compared, and the value range was explored. When takes 0.05, the mAP reaches the maximum value of 83.4%. Therefore, the value range of the scheduler is selected between 0.01 - 0.08. In addition, to explore the impact of the three schedulers on the performance of DGF, a comparative experiment was designed, and the experimental results are shown in Table 3. It can be seen from Table 3 that the linear function best conforms to the change trend between classes during training and has the best performance indicators.
[0065] Effectiveness of Volatile Feature Mining Loss. By comparing #1 and #4, #2 and #6, #3 and #7, #5 and #8, the effectiveness of VFE can be illustrated. The present invention conducts a more in-depth mining of volatile samples through the loss, making up for the deficiency that MCM only focuses on the cluster level. In particular, the combination of MCM and VFE is more conducive to the separation between different classes both at the cluster level and at the instance level, and learning more challenging volatile features, further improving the generalization ability of the model. It can be seen from #1 and #6 that under the combined action of MCM and VFE, the baseline model has an mAP as high as 86.0% and 74.9% and Rank-1 as high as 94.4% and 86.3% on the datasets Market1501 and DukeMTMC-reID respectively. By comparing #4 and #6, it can be concluded that compared with the case of using VFE alone in #4, #6 has an mAP improvement of 2.6% and 1.7% and a Rank-1 improvement of 1.8% and 1.2% on the two datasets respectively.
[0066] Visualization analysis, the features of some test samples were dimensionally reduced and visualized using the T-SNE method to compare the baseline model (baseline) with the MGVF method, and the results are asFigure 6 As shown. Each point represents a sample, and a typical representative is marked by an elliptical dotted line. From the samples marked by the elliptical dotted line, it can be seen that the features obtained by this method have a tighter intra-class structure and better inter-class separability. In addition, all the test set samples were extracted, and the intra-class and inter-class distance distributions were visually displayed. The results are shown in Figures 7(a) and 7(b), where d represents the peak distance between the intra-class and inter-class distances. Experiments show that compared with the baseline model, the intra-class distance of the overall samples is smaller, the inter-class distance is larger, and the peak distance between the two is also farther (d1 < d2). To more intuitively demonstrate the superiority of the algorithm, the Rank-list visualization of MGVF and the baseline model was also carried out. The results are shown in Figures 8(b) and 8(a). The leftmost column is the query image, and then from left to right are Rank-1 to Rank-10 in turn. The images marked with a thick frame are the misrecognized images. The experimental results show that compared with the baseline, the algorithm of the present invention significantly improves the matching hit rate of the overall query images in the gallery. Especially in the top 5 rankings, there are many misrecognition cases in the baseline model, while MGVF has a very significant improvement in the hit rate, indicating that the algorithm has stronger applicability.
[0067] S35: Parameter Analysis The hyperparameters and were analyzed in detail. Among them, controls the contribution degree of the volatile features to the centroid. The larger is, the more information of the volatile features is embedded in the centroid, and the model will focus more on learning the clustering boundary information, but the global information reflecting the feature distribution in the centroid will become less. On the contrary, the smaller the value of , the model will fall into a local optimal solution and cannot obtain higher performance. Therefore, in order to find the best value of , experiments were carried out on the dataset Market-1501 within the range of 1-11 for its value. The experimental results are shown in Figure 9(a). The results show that as increases, the performance of the model gradually improves and reaches the peak when the value of is 6, and then slowly decreases. The hyperparameter controls the ratio between the ClusterNCE loss and the volatile feature loss. The balance between the two is crucial for the performance. Experiments were carried out on the dataset Market-1501 within the range of 0.0-1.0 for its value. The experimental results are shown in Figure 9(b). When takes 0.4, the model obtains the best performance. The larger the value taken, the slower the model convergence speed, but it can more strongly emphasize the diversity of samples and improve the generalization ability of the model. It is worth mentioning that If the value of is greater than 0.8, the model may even experience negative optimization and its performance will drop sharply. This may be due to the fact that the model is unable to learn more generalized information.
[0068] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0069] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multi-cluster-guided variable feature contrastive learning method for unsupervised person re-identification, characterized in that, Including: Collect pedestrian captured image data and construct an image dataset; Extract features from the image dataset through an unsupervised person re-identification model to obtain samples to be processed; Use two clustering algorithms with different parameters to assign labels to the samples to be processed, and screen out volatile features by comparing the two clusterings to guide the initialization of the memory bank; Process the initialized memory bank through a multi-cluster centroid scheduler to obtain a volatile feature memory bank; Utilize the context information in the mini-batch samples to be processed, dynamically assign different weights to each sample, and construct a dynamic centroid to update the volatile feature memory bank; Perform backpropagation optimization on the unsupervised person re-identification model through the volatile feature memory bank to obtain an optimal unsupervised person re-identification model.
2. The multi-cluster-guided variable feature contrastive learning method for unsupervised person re-identification according to claim 1, wherein, The use of two clustering algorithms with different parameters includes: ; Among them, , are the clustering results and labels of the baseline clustering respectively, , are the clustering results and labels of the extended clustering respectively, , are the hyperparameters of the control clustering respectively, and is the extension factor for controlling the hyperparameters.
3. The multi-cluster-guided variable feature contrastive learning method for unsupervised person re-identification according to claim 2, wherein It also includes adding the tags of the corresponding features in the reference tag and the extended tag to obtain a comparison tag. For the features at the boundary between classes in the reference clustering, the number of intra-class features in the comparison clustering may be less than the hyperparameter of the reference clustering in which case, these features are marked as clustering-sensitive features: ; ; ; ; Among them, and are the clustering results and labels of the contrast clustering respectively, represents the mapping from labels to clustering results, num represents the set of the number of features of each class in the contrast clustering, represents the number of features in the i-th class among them, is the j-th feature of the i-th class in the contrast clustering, is the clustering hyperparameter, is the clustering sensitive feature.
4. A multi-cluster-guided volatile feature contrastive learning method for unsupervised person re-identification according to claim 3, characterized in that Construct an uncertainty factor for the clustering-sensitive features Perform screening to obtain volatile samples: ; Among them, represents the j-th feature in the i-th class, is a clustering-sensitive feature, represents the number of features in the i-th class. A threshold Q is set to control the selection of volatile features. For clustering-sensitive features, when the value of p is greater than Q, clustering-sensitive feature filtering is performed to obtain the final volatile features .
5. A multi-cluster-guided volatile feature contrastive learning method for unsupervised person re-identification according to claim 4, characterized in that It also includes using the final volatile features to guide the formation of centroids, and the k-th class centroid obtained is as follows: ; Among them, is the weight of the mutable feature, is the mutable feature, M is the total number of features in the k-th class, , by increasing the weight of the mutable feature in the formation process of the clustering centroid, a centroid containing more boundary information of non-noise is obtained, enabling the model to learn the boundary of each cluster, and using the centroids regulated by weights to initialize the memory bank.
6. The multi-cluster guided volatile feature contrastive learning method for unsupervised person re-identification according to claim 1, characterized in that, The utilization of the context information in the mini-batch samples to be processed, dynamically assign different weights to each sample, and construct a dynamic centroid to update the volatile feature memory bank includes: According to the distance between the sample and the centroid, use the Softmax function to assign different weights to each query sample, reflecting the overall distribution of features: ; where N represents the number of instances of the i-th class in the mini-batch, is the centroid of the i-th class, and is a sample instance, is a linear scheduler.
7. A multi-cluster-guided variable feature contrastive learning method for unsupervised person re-identification according to claim 6, characterized in that, The linear scheduler comprises: ; Among them, is the initial linear scheduler that dynamically changes the weights throughout the training process: The dynamic centroid of the i-th class obtained is as follows: ; Update the memory bank using the dynamic centroid: ; Among them, is a hyperparameter for controlling momentum update.
8. A multi-cluster-guided volatile feature contrastive learning method for unsupervised person re-identification according to claim 1, characterized in that The performing of backpropagation optimization on the unsupervised person re-identification model through the volatile feature memory bank includes: Based on all mutable features, representative features are selected from them to construct a loss function. By calculating the inner product of each query instance with the mutable features and selecting the mutable representative positive sample vector and the mutable representative negative sample vector , the information of the mutable features is captured: ; Among them, and are mutable features of the same class and different classes as the query instance q respectively, is the mutable feature with the farthest distance in the feature space among the features of the same class as the query instance q, while is the mutable feature with the closest distance in the feature space among the features of different classes from the query instance q.
9. A multi-cluster-guided volatile feature contrastive learning method for unsupervised person re-identification according to claim 8, characterized in that It also includes, through the screening of vectors, obtaining the volatile positive sample and volatile negative sample with the largest boundary information, and constructing a volatile feature loss: ; Among them, is the temperature hyperparameter, is the expected value, and the total loss of volatile feature mining consists of two parts: ClusterNCE loss and volatile feature loss: ; Among them, is the ClusterNCE loss, is the weight coefficient used to balance these two losses.
10. A multi-cluster guided volatile feature contrastive learning system for unsupervised person re-identification, characterized in that, Including: Collection module: used to collect pedestrian captured image data and construct an image dataset; Feature extraction module: used to extract features from the image dataset through an unsupervised person re-identification model to obtain samples to be processed; Initialization module: used to use two clustering algorithms with different parameters to assign labels to the samples to be processed, and screen out volatile features by comparing the two clusterings to guide the initialization of the memory bank; Volatile feature memory bank construction module: used to process the initialized memory bank through a multi-cluster centroid scheduler to obtain a volatile feature memory bank; Update module: used to utilize the context information in the mini-batch samples to be processed, dynamically assign different weights to each sample, and construct a dynamic centroid to update the volatile feature memory bank; Backpropagation optimization module: used to perform backpropagation optimization on the unsupervised person re-identification model through the volatile feature memory bank to obtain an optimal unsupervised person re-identification model.
Citation Information
Patent Citations
Unsupervised pedestrian re-identification method based on joint training strategy
CN114187655A
Unsupervised cross-domain target re-identification method based on comparative learning
CN115205570A
Unsupervised cross-domain pedestrian re-identification method based on clustering and multi-scale learning
CN115641613A
Pedestrian re-identification method and device based on unsupervised learning
CN116030502A
Unsupervised pedestrian re-identification method based on comparative learning
CN116524534A
Cited By
Pedestrian re-recognition method based on real-time memory update
CN121354204A