Low-Cost Pedestrian Re-Identification Method Based on Deep Active Learning

Through the combination of deep active learning and unsupervised clustering, pseudo-label generation and sample selection are optimized, which solves the performance improvement problem of the pedestrian re-identification model when the labeling costs are high, and achieves efficient pedestrian re-identification under low-cost conditions.

CN114187610BActive Publication Date: 2025-07-04NANJING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111461351.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-02
Publication Date
2025-07-04
Estimated Expiration
2041-12-02

AI Technical Summary

Technical Problem

In the case of high labeling costs, existing pedestrian re-identification technology is difficult to effectively improve model performance, especially in the new environment, which is a barrier to deployment to real-life scenarios.

Method used

The deep active learning method is adopted, combined with the unsupervised clustering model and the idea of ​​active learning, and manually annotate by generating pseudo-labels and selecting hard-negative and hard-positive samples, optimize the class cluster splitting and merging, and gradually improve the model performance.

Benefits of technology

The performance of the pedestrian re-identification model is significantly improved under low annotation cost, making the clustering results closer to the real label, reducing the cost of manual annotation, and adapting to different cluster structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114187610B_ABST
    Figure CN114187610B_ABST
Patent Text Reader

Abstract

The present invention discloses a low-cost pedestrian re-identification method based on deep active learning, which integrates dynamic sample selection and clustering prediction of pseudo-labels into a unified framework. The method includes: obtaining a pedestrian dataset, inputting it into an unsupervised clustering model to generate pseudo-labels; based on the idea of active learning, selecting hard-negative and hard-positive pedestrian sample pairs for manual labeling, splitting and merging clusters according to the labeling results to optimize the clustering structure and obtain more reliable pseudo-labels; the pedestrian samples with pseudo-labels are used for model training, so as to learn more discriminative feature expressions, and this cycle continues until the model is stable. The present invention can solve the problem of the decline in pedestrian re-identification accuracy when the labor cost of labeled samples is limited and a large amount of training data lacks labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and pattern recognition, and particularly to a low-cost pedestrian re-identification method based on deep active learning. Background Art

[0002] In recent years, with the acceleration and promotion of urbanization construction in China, the population density has been increasing, and the urban operation system has become increasingly complex. Therefore, the country has vigorously carried out the construction of the social surveillance system, installed a large number of video surveillance devices at all levels, and comprehensively covered the key parts of public security. However, the large-scale deployment of surveillance points and the sharp increase in the collected data have made the labor cost of traditional processing and analysis of surveillance data extremely high. To solve this practical problem, automated and intelligent surveillance management systems have gradually come into the public eye and become a major construction hotspot in the field of public security. As an important part of the intelligent surveillance system, pedestrian re-identification has received extensive attention from researchers.

[0003] Pedestrian re-identification is an important research direction in the field of computer vision. Pedestrian re-identification refers to the process of identifying the same pedestrian through a set of cameras with non-overlapping perspectives in a multi-camera surveillance video system in different geographical regions. Through pedestrian re-identification technology, the features of pedestrians are extracted to perform cross-camera retrieval on pedestrians whose faces cannot be clearly photographed, enhancing the spatio-temporal continuity of data.

[0004] Although the research on supervised learning and same-domain methods for pedestrian re-identification has been very mature and good retrieval results have been achieved on existing datasets, there are still many problems and challenges to be faced in solving the problem of pedestrian re-identification. The most obvious one is that data annotation is relatively difficult. Pedestrian re-identification requires a large amount of annotation work, and the annotation cost of large-scale datasets is extremely high both in terms of time and money. The contradiction between the need for a large number of sample labels for model training and the difficulty of data annotation in the new environment is a major obstacle to the deployment of pedestrian re-identification models into real scenarios. Therefore, more and more researchers are turning their attention to low-cost pedestrian re-identification technology, studying how to improve the performance of pedestrian re-identification models when the labeling cost is less. Summary of the Invention

[0005] The purpose of the present invention is to provide a low-cost pedestrian re-identification method based on deep active learning to improve the performance of pedestrian re-identification models when the annotation cost is limited.

[0006] The technical solution for achieving the purpose of the present invention is: a low-cost pedestrian re-identification method based on deep active learning, including the following steps:

[0007] Step A1: Read the pedestrian dataset and input it into an unsupervised clustering model to extract pedestrian features and generate a clustering result, and assign pseudo-labels to each pedestrian sample.

[0008] Step A2: Based on the idea of active learning, for each cluster, select several hard-negative pedestrian sample pairs for manual labeling, and use the labeling results to guide cluster splitting.

[0009] Step A3: Based on the idea of active learning, for each cluster, select several hard-positive pedestrian sample pairs for manual labeling, and use the labeling results to guide cluster merging.

[0010] Step A4: Use the sample pseudo-labels obtained in Step A3 to train the unsupervised clustering model.

[0011] Step A5: Iterate Steps A1 to A4 until the model converges.

[0012] Preferably, the unsupervised clustering model in Step A1 consists of two parts: a deep feature extraction network and a clustering algorithm; the deep feature extraction network uses ResNet-50 as the basic network and integrates and adds a one-dimensional BatchNorm and an L2 normalization layer; input pedestrian samples are used to extract features, and pseudo-labels are output through the DBSCAN clustering algorithm.

[0013] Preferably, Step A2 is specifically as follows:

[0014] Step A201: For each cluster in the clustering result of Step A1, select the number of K-means clustering centers K. According to the K-means clustering result, split the original cluster into K smaller clusters and obtain the central features corresponding to the clusters.

[0015] Step A202: Select the corresponding pedestrian samples for the K central features in the result of Step A201 and form sample pairs pairwise. Have experts label whether the sample pairs match. If they are not the same person, assign a new pseudo-label to the cluster corresponding to the sample. If they are the same person, retain the original clustering pseudo-label.

[0016] Preferably, the method for selecting K in Step A201 is to limit the maximum value of the number of centers to k max , traverse each integer k in the range [2, k max as the number of centers in the current round, perform K-means clustering on the data, and calculate the reliability score c_score k , after the traversal is completed, select the number of clustering centers corresponding to the round with the highest c_score k as the final value of K, and the calculation formula is as follows:

[0017]

[0018] Among them, N represents the number of clustering samples.

[0019] Preferably, the clustering reliability score c_score k is obtained by multiplying the compactness index comp of each cluster and the cluster independence index indep. A more compact cluster requires the similarity between samples within the cluster to be as high as possible. When the minimum similarity between samples is relatively high, it indicates that the cluster is sufficiently compact. Therefore, the minimum sample similarity within the cluster is used to measure the compactness index. A more independent cluster requires the similarity between clusters to be sufficiently small. The similarity between the central features of two clusters is used to represent the cluster similarity. If the maximum similarity between a certain cluster and the remaining clusters is relatively small, it indicates that the cluster has high independence. Thus, the maximum similarity between clusters is used to measure the independence index. Given the set of clusters G after clustering, the size of the set is k, and G j represents the j-th cluster among them, and c_score k has the following specific calculation form:

[0020] comp j = minSim intra (G j ) ∈ [0, 1]

[0021] indep j = 1 - maxSim inter (G j , G) ∈ [0, 1]

[0022]

[0023] Among them, Sim intra represents the set of similarities between all samples within the cluster G j , and Sim inter represents the set of similarities between the central point corresponding to the cluster G j and the central points of the remaining clusters within the cluster set G.

[0024] Preferably, step A3 is specifically as follows:

[0025] Step A301: According to the clustering result of step A2, obtain the set C = {C1, C2,... C n} composed of n clusters. For each cluster C j within the set, calculate the average feature of all samples within the cluster, and the feature with the closest distance is the central feature of the cluster. Sort the central feature similarities between all the remaining clusters and the cluster C j from large to small to obtain the similarity sequence [s1, s2,... s nCalculate the difference between the previous term and the next term in the sequential calculation sequence to obtain the similarity difference sequence d. The result after linearly normalizing the data in d is denoted as Select the truncation threshold δ. The clusters corresponding to the values greater than δ within form the candidate cluster set. Since the difference sequence can reflect the change in similarity between clusters, the truncation position of δ is at the position where the similarity trend starts to flatten, meaning that the similarity after the truncation is very small and there is no obvious change, thus narrowing the range of candidate clusters. For each cluster in the candidate cluster set, obtain the pedestrian sample corresponding to the central feature, and at the same time obtain the central sample of cluster C j The matching relationship between the samples is labeled by an expert, and the calculation form is:

[0026] d = {d l = s l - s l+1 | l ∈ [1, l max - 1]}

[0027]

[0028] Step A302: According to the labeling result of Step A301, if the two central samples are of the same person, assign the same label to the two clusters corresponding to the central sample; if not, retain the original clustering pseudo-label.

[0029] Preferably, in Step A4, the unsupervised clustering model training uses the contrastive loss as the loss function. For each pedestrian sample in the training set, its loss calculation form is:

[0030]

[0031] where x T c k represents the similarity between the pedestrian feature x and the k-th cluster center c k , n represents the number of clusters, and τ is the temperature parameter.

[0032] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the above-mentioned low-cost pedestrian re-identification method based on deep active learning.

[0033] A computer-readable storage medium stores a computer program, and when the program is executed by a processor, it implements the above-mentioned low-cost pedestrian re-identification method based on deep active learning.

[0034] Compared with the prior art, the beneficial effects of the present invention are:

[0035] (1) The present invention combines active learning and unsupervised clustering, greatly improving the performance of the pedestrian re-identification model on the premise of low labeling cost;

[0036] (2) The present invention mines hard-negative and hard-positive sample pairs, and guides the splitting and merging of clusters according to expert annotations, thereby optimizing the clustering pseudo-labels and making the clustering results closer to the true sample labels;

[0037] (3) The present invention can adaptively adjust the number of manual annotations according to different cluster structures, effectively reducing the annotation cost of traditional active learning algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is the overall framework diagram of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0039] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0040] According to an embodiment of the present invention, a low-cost pedestrian re-identification method based on deep active learning is proposed to improve the performance of the pedestrian re-identification model when the labeling cost is limited. Combining Figure 1 as shown, the implementation of this method generally includes the following 5 steps:

[0041] Step A1, read the pedestrian dataset and input it into the unsupervised clustering model, extract pedestrian features and generate clustering results, and assign pseudo-labels to each pedestrian sample;

[0042] Step A3, based on the active learning idea, according to the clustering results of Step A2, select several hard-positive pedestrian sample pairs for each cluster for manual labeling, and use the labeling results to guide the cluster merging;

[0043] Step A4, based on the active learning idea, use the sample pseudo-labels obtained in Step A3 to train the unsupervised clustering model;

[0044] Step A5, iterate Steps A1 to A4 until the model converges.

[0045] Preferably, the unsupervised clustering model in step A1 consists of two parts: a feature extraction network and a clustering algorithm. The feature extraction network is based on ResNet-50, and a one-dimensional BatchNorm and an L2 normalization layer are integrated and added. The model is set to use the Adam optimizer, with a weight decay of 0.0005 and a learning rate of 3.5*10 -4 . The pedestrian training samples are input to extract features, and pseudo-labels are output through DBSCAN clustering.

[0046] Preferably, step A2 is specifically as follows:

[0047] Step A201: For each cluster in the clustering result of step A1, select the number of K-means clustering centers K. According to the K-means clustering result, split the original cluster into K smaller clusters, and obtain the central features corresponding to the clusters;

[0048] Step A202: Select the corresponding pedestrian samples from the K central features in the result of step A201, and form sample pairs in pairs. An expert labels whether the sample pairs match. If they are not the same person, new pseudo-labels are assigned to the pedestrian samples within the corresponding cluster. If they are the same person, the original clustering pseudo-labels are retained.

[0049] Preferably, the method for selecting K in step A201 is to limit the maximum value of the number of centers to k max , and traverse each integer k in the range [2, k max as the number of centers in the current round, perform K-means clustering on the data, and calculate the reliability score c_score of this clustering result k . After the traversal is completed, select the number of clustering centers corresponding to the round with the highest c_score k as the final value of K. The calculation formula is as follows:

[0050]

[0051] Among them, N represents the number of clustering samples.

[0052] Preferably, the clustering reliability score c_score k is obtained by multiplying the values of the cluster compactness index comp and the cluster independence index indep. The specific calculation formula of c_score k is as follows:

[0053] comp j = minSim intra (G j ) ∈ [0, 1]

[0054] indep j = 1 - maxSim inter (Gj , G) ∈ [0, 1]

[0055]

[0056] Among them, Sim intra represents the set of similarities between all samples within cluster G j , and Sim inter represents the set of similarities between the center point corresponding to cluster G j and the center points of the remaining clusters within the cluster set G.

[0057] Preferably, step A3 is specifically as follows:

[0058] Step A301: According to the clustering result of step A2, obtain a set C = {C1, C2,... C n} composed of n clusters. For each cluster C j within the set, calculate the average feature of all samples within the cluster, and set the feature closest to the center feature of the cluster. Sort the similarity of the center features between all the remaining clusters and cluster C j from largest to smallest to obtain a similarity sequence [s1, s2,... s n . Sequentially calculate the difference between the previous term and the next term of the sequence to obtain a similarity difference sequence d. Denote the result after linearly normalizing the data in d as Select a truncation threshold δ, and the clusters corresponding to the values greater than δ within

[0059] d = {d l = s l - s l+1 |l ∈ [1, l max - 1]}

[0060]

[0061] For each cluster in the candidate cluster set, obtain the pedestrian sample corresponding to the center feature of the cluster, and at the same time obtain the center sample of cluster C j . Let an expert annotate the matching relationship between the samples.

[0062] Step A302: According to the annotation result of step A301, when the two center samples are the same person, assign the same pseudo-label to the corresponding clusters.

[0063] Preferably, in step A4, the unsupervised clustering model training uses contrastive loss as the loss function. For each pedestrian sample in the training set, its loss calculation form is:

[0064]

[0065] where x T c k represents the similarity between the pedestrian feature x and the k-th cluster center c k , n represents the number of clusters, and τ is the temperature parameter.

[0066] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A low-cost pedestrian re-identification method based on deep active learning, characterized in that, Including the following steps: Step 1: Obtain a pedestrian dataset, input it into an unsupervised clustering model, extract pedestrian features and generate a clustering result, and assign pseudo-labels to each pedestrian sample; Step 2: Based on the idea of active learning, for each cluster, select several hard-negative pedestrian sample pairs for manual annotation, and use the annotation results to guide cluster splitting; specifically: Step 201: For each cluster in the clustering result of Step 1, select the number of K-means clustering centers K. According to the K-means clustering result, split the original cluster into K smaller clusters. The method for selecting the number of K-means centers K is to limit the maximum value of the center number to k max , traverse each integer k in the range [2, k max as the number of clustering centers in the current round, perform K-means clustering on the data, and calculate the reliability score c_score of this clustering result k . After the traversal is completed, select the number of clustering centers corresponding to the highest score as the final value of K. The calculation formula is as follows: where N represents the number of clustering samples; The reliability score c_score k is composed of the cluster compactness index comp and the cluster independence index indep; given the cluster set G, G j represents the j-th cluster among them, c_score k has the specific calculation form as follows: comp j = minSim intra (G j ) ∈ [0, 1] indep j = 1 - maxSim inter (G j , G) ∈ [0, 1] Among them, Sim intra represents the similarity set among all samples within the cluster G j ; Sim inter represents the similarity set between the corresponding center point of the cluster G j and the center points of the remaining clusters within the cluster set G; Step 202: Select the central pedestrian samples corresponding to K clusters in the result of Step 201, form sample pairs pairwise, and have experts annotate whether the sample pairs match. If they are not the same person, assign a new pseudo-label to the cluster where the sample is located. If they are the same person, retain the original clustering pseudo-label; Step 3: Based on the idea of active learning, for each cluster, select several hard-positive pedestrian sample pairs for manual annotation, and use the annotation results to guide cluster merging; specifically: Step 301: According to the clustering results in Step 2, obtain a set C = {C1, C2,... C n} composed of n clusters. For each cluster C j in the set, calculate the average feature of all samples within the cluster, and the feature with the closest distance is the central feature of the cluster; sort the similarity of the central features between all the remaining clusters and cluster C j from large to small to obtain a similarity sequence [s1, s2,... s n ; sequentially calculate the difference between the previous item and the next item in the sequence to obtain a similarity difference sequence d, and denote the result after linearly normalizing the data in d as Select a truncation threshold δ, and the clusters corresponding to the values greater than δ within d = {d l = s l - s l+1 | l ∈ [1, l max - 1]} For each cluster in the candidate cluster set, obtain the pedestrian sample corresponding to its central feature, and at the same time obtain the central sample of cluster C j ; the matching relationship between the samples is labeled by experts Step 302: According to the annotation results of Step 301, if the two central samples are the same person, assign the same label to the two clusters corresponding to the central samples. If they are not the same person, retain the original clustering pseudo-label; Step 4: Use the sample pseudo-labels obtained in Step 3 to train the unsupervised clustering model; for the training of the unsupervised clustering model, use the contrastive loss as the loss function. For each pedestrian sample in the training set, the form of its loss calculation is: where x T c k represents the similarity between the pedestrian feature x and the k-th cluster center c k , n represents the number of clusters, and τ is the temperature parameter; Step 5: Iterate Steps 1 to 4 until the model converges.

2. The low-cost pedestrian re-identification method based on deep active learning according to claim 1, wherein The unsupervised clustering model in Step 1 consists of two parts: a deep feature extraction network and a clustering algorithm; the deep feature extraction network uses ResNet-50 as the basic network and integrates and adds a one-dimensional BatchNorm and an L2 normalization layer; the clustering algorithm selects the DBSCAN algorithm.

3. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the low-cost pedestrian re-identification method based on deep active learning as described in any one of Claims 1-2.

4. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the low-cost pedestrian re-identification method based on deep active learning as described in any one of Claims 1-2.