A post-reranking method for pedestrian re-identification based on human-in-the-loop

By adopting a human-in-the-loop person re-identification and post-ranking method, using user feedback to optimize the initial ranking list, and adopting a semi-supervised label graph propagation and k-fusion algorithm, the problem of large fluctuations in the ranking list in the existing technology is solved, the accuracy and stability of pedestrian re-identification are improved, and the generalization ability of the algorithm is enhanced.

CN116778525BActive Publication Date: 2025-09-09ZHEJIANG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310403963.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-17
Publication Date
2025-09-09
Estimated Expiration
2043-04-17

AI Technical Summary

Technical Problem

Existing person re-identification technologies suffer from large fluctuations in ranking lists when processing incremental feedback and are difficult to adapt to the differences between human cognition and model calculations, resulting in insufficient accuracy and stability.

Method used

By introducing a human-in-the-loop based person re-identification and re-ranking method, user feedback is used to optimize the initial ranking list, and semi-supervised label graph propagation and k-fusion algorithms are adopted to reduce sample jitter and improve ranking accuracy and stability.

Benefits of technology

It improves the accuracy and stability of pedestrian re-identification, enhances the generalization ability of the algorithm, reduces sample jitter in the incremental feedback process, and improves the efficiency of user verification and retrieval in ranking lists.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116778525B_ABST
    Figure CN116778525B_ABST
Patent Text Reader

Abstract

A human-in-the-loop (HIL) person re-identification and post-ranking method comprises: (1) using an unsupervised HIL model to obtain an initial ranking of all images in a gallery for a search target; (2) a human performs iterative feedback on the ranking results; (3) using a semi-supervised label graph propagation method to optimize the HIL ranking scores using user feedback. (4) using a k-fusion algorithm to eliminate sample "jitter" that may occur during step 3, and obtaining a ranking optimized after feedback. The present invention can improve the accuracy and stability of HIL.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target detection, and in particular to a pedestrian re-identification and rearrangement method based on a person-in-the-loop. Background Art

[0002] Person re-identification (Re-ID) is a computer vision technology whose primary purpose is to identify and track people across multiple cameras. Using image matching, Re-ID identifies images of the same person captured by different cameras at different times and locations. This helps users quickly retrieve and track people of interest within massive amounts of video data. In simple terms, Re-ID technology primarily aims to identify and track the same person captured by different cameras.

[0003] The implementation of pedestrian re-identification technology usually includes the following steps: first, the appearance characteristics of pedestrians are characterized by extracting feature vectors from pedestrian images; then, a similarity calculation method is used to measure the similarity between two pedestrian images to determine whether they are from the same pedestrian; finally, the matching results are displayed to the user through a ranked list or other means, so that pedestrians of interest can be quickly found and tracked.

[0004] Pedestrian re-identification technology has a wide range of applications in public safety, traffic management, commercial security, and other fields. For example, in public safety, Re-ID technology can be used to track suspects and missing persons, helping police locate targets more quickly. In traffic management, Re-ID technology can be used to monitor traffic violations and illegal parking, helping to improve road safety and efficiency. In commercial security, Re-ID technology can be used to identify suspected thieves and conduct customer counts, helping businesses maintain safety and optimize operations.

[0005] However, existing person re-identification technologies have some limitations and challenges. For example, existing algorithms have difficulty coping with various complex situations in the real world, such as different lighting conditions, viewing angles, pedestrian occlusion, etc. In addition, existing algorithms have limited processing capabilities for large-scale data, and also have difficulty handling changes in pedestrian appearance features. Therefore, a new person re-identification technology is needed that can overcome these limitations and challenges, improve the accuracy and robustness of the algorithm, and adapt to the ever-changing real-world environment. Therefore, in recent years, some re-ranking methods have emerged, which can further optimize and correct the initial ranking list through human feedback, thereby improving the accuracy and stability of pedestrian re-identification. The basic idea of ​​the re-ranking method is to optimize the order of the ranking list by re-sorting the ranking list, so that images that are more likely to match are ranked higher.

[0006] However, existing re-ranking methods still have some problems. They can only handle one-time feedback and cannot adapt well to incremental feedback. In incremental user feedback, as feedback accumulates, the generated ranking list may fluctuate significantly. This fluctuation may cause positive samples that were originally far apart in the initial ranking to become closer. Due to the difference between model calculation and human cognition, the constant change in ranking is prone to sample jitter. Jitter is an inherent problem in existing graph propagation algorithms because they only propagate labels based on sample feature similarity. When there is a disagreement between human cognition and model calculation, existing algorithms only accept label information based on human feedback, but not pairwise relationship information. This may lead to ranking results that are inconsistent with human cognition. Therefore, traditional post-re-ranking algorithms cannot handle the problem of incremental feedback well, and further research and improvement of re-ranking methods are needed to improve their application value in the field of pedestrian re-identification. Summary of the Invention

[0007] The present invention aims to provide a post-reranking method for person re-identification based on human-in-the-loop, which is used to solve the problems existing in existing pedestrian re-identification technologies. The generalization of current pedestrian re-identification technologies in practical applications is very limited. After running the automatic algorithm, the user will obtain a suboptimal ranking list consisting of thousands of candidates, and each candidate needs to be manually verified, which is an extremely time-consuming and error-prone task. In addition, the target may show different patterns in different cameras, and users often do not have enough clues when browsing and retrieving the ranking list. The innovation of the present invention is to propose a post-reranking method for pedestrian re-identification based on human-in-the-loop. The algorithm can further optimize and correct the initial ranking list by using human feedback through human-computer interaction, thereby improving the accuracy and stability of pedestrian re-identification.

[0008] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0009] A method for pedestrian re-identification and re-ranking based on people in the loop, the method comprising the following steps:

[0010] Step 1: Use the unsupervised person re-identification model to get the initial ranking of all images in the gallery for the retrieval target. Given a retrieval object q and a gallery containing n images, As the input of the unsupervised pedestrian re-identification model, the features of each image are obtained and the initial ranking score S of each image in the gallery is calculated. init .

[0011] Step 1.1, data preprocessing: given a retrieval object q and a library containing n images As the input of the unsupervised person re-identification model, the original image is scaled to 256x128 and randomly horizontally flipped and randomly cropped to increase the diversity and robustness of the data.

[0012] Step 1.2, Feature Extraction: Use the ImageNet pre-trained ResNet-50 as the feature extractor, remove the last fully connected layer, and only retain the previous convolutional and pooling layers. Use this ResNet-50 to convert the 256x128 image into a 2048-dimensional feature vector and perform L2 normalization.

[0013] Step 1.3, attention mechanism: Introduce a learnable weight matrix, perform weighted summation on the feature vectors, obtain an attention vector, and add it to the original feature vector to enhance the identity-related parts, reduce the interference of irrelevant information, and improve recognition accuracy. Specifically, it includes the following steps:

[0014] First, the input image is passed through the feature extractor to produce a feature map of size C x H x W, where C is the number of channels, H and W are the height and width.

[0015] The feature map is then passed through an adaptive average pooling layer to obtain an output feature map of size C x 1 x 1. This step can compress the spatial dimension to 1x1, thereby reducing the amount of computation and the number of parameters.

[0016] Next, the output feature map is passed through a fully connected layer to obtain a weight vector of size C. This step maps the 1x1 value corresponding to each channel to a real value, indicating the contribution of the channel to the final feature vector.

[0017] Finally, the weight vector is multiplied by the original feature map and summed along the channel dimension to obtain the final 2048-dimensional feature vector. This step allows each channel to be weighted averaged according to its weight value and retains the useful information in the original feature map.

[0018] Step 1.4, Fully Connected Layer and Batch Normalization Layer: Use a fully connected layer to map the feature vector to a low-dimensional space. The size of the weight matrix W of the fully connected layer is 128x2048, and the size of the bias vector b is 128. Then use the batch normalization layer to perform a standardization operation and normalize again to reduce the dimension and variance of the feature vector, improving computational efficiency and stability.

[0019] Step 1.5, Category Prediction: Using the representative vectors stored in the memory bank, we use cosine similarity to calculate the distance between the feature vector and each category. We then find the nearest neighbor category as the predicted label to guide the model in learning to better distinguish between different categories. The memory bank is a matrix that stores the representative vector for each category in a 128-dimensional space. Distance is calculated using the Euclidean distance, which is the L2 norm between two vectors.

[0020] Step 1.6, Instance Retrieval: Use the instance feature vectors stored in the memory bank to calculate the similarity between the feature vectors and each instance, find the most similar k instances as retrieval targets, and sort them in descending order of similarity to obtain the initial ranking score S of each image in the gallery corresponding to the retrieval target. init .

[0021] Step 2: The user performs iterative feedback operations on the ranking results. This step requires the user to provide at least one positive sample and one negative sample for each feedback. If the user cannot find a positive sample, then a negative sample similar to the query image q is allowed to be selected as a positive sample. We denote the set of all feedback operations performed as L, and use these feedback operations to further optimize and correct the initial ranking list to improve the accuracy and stability of pedestrian re-identification. The user provides at least one positive label and one negative label. It should be noted that if the user cannot find a positive sample, then a negative sample similar to the query q can be selected as a positive sample. We denote the set of samples marked as positive samples by the user as The negative samples are represented as The remaining unlabeled instances G in the gallery are denoted as unlabeled samples U = G\L.

[0022] Step 3: A semi-supervised label graph propagation method is used to leverage user feedback to optimize the person re-ID ranking score. The first step is to construct a kNN graph to characterize the similarity between images. The second step is to perform label propagation on this graph, transferring the labels of positive and negative samples to the remaining unlabeled samples. When constructing the kNN graph, a high-dimensional vector is used to represent the features of each sample, and the similarity between samples is calculated, ultimately constructing a kNN graph. During label propagation, an optimization function is constructed using the positive and negative labels provided by the user, and the function is solved using the Newton solver. Finally, the resulting new ranking score is used to update the similarity score of each sample in the image library with the detection sample.

[0023] Step 3.1: Construct a kNN proximity graph to characterize the similarity between images. First, feature extraction is performed. For each image in G, the high-dimensional vector extracted by the person re-identification algorithm is used to represent each image.

[0024] Step 3.2: Image similarity calculation: After obtaining the high-dimensional vector of each image, the cosine similarity method is used to calculate the similarity score between samples.

[0025]

[0026] Step 3.3. Select the parameter k. Based on the similarity score, select the k most similar neighbors of each data point and connect them to the data point. Here, use the priority queue data structure to store the distance between samples and the nearest neighbor information.

[0027] Step 3.4: Construct a proximity graph. Add the lines between each data point (i.e., sample) and its neighbors to the graph to obtain an undirected graph consisting of nodes and edges, namely the kNN proximity graph.

[0028] Step 3.5: Label propagation on the graph. Based on the constructed proximity graph, propagate the labeled sample L to the remaining unlabeled samples U. The user feedback l input to the model is expressed as follows:

[0029]

[0030] Step 3.6, Initialization: Set the label value of the sample that the user has fed back to the formula shown above, and the label value of the sample that has not fed back to 0.

[0031] Step 3.7, iterative update: For each unlabeled node i, update its label value according to the label value of its neighbor node j. The specific formula is:

[0032]

[0033] Among them, y i represents the label value of node i, w i,j represents the edge weight between node i and node j, N i Represents the set of neighbor nodes of node i.

[0034] Step 3.8: To ensure the smoothness and continuity of label propagation, we use a propagation optimization function based on the Laplace support vector machine. The formula is as follows:

[0035]

[0036] The optimization function consists of three parts. The first is the hinge loss, a commonly used loss function based on user feedback, used to measure the error of a classifier. In the optimization function, hinge loss is used to measure the variance between samples, encouraging the classifier to correctly classify samples and penalizing incorrectly classified samples. Specifically, in the optimization function, hinge loss can help make the ranking scores more accurate and reliable. The second part is the norm of the function in the Reproducing Kernel Hilbert Space (RKHS), which is used to ensure the smoothness of the model. In the optimization function, smoothness refers to the norm of the function in the RKHS, which reflects the rate of change of the function value. Therefore, a smooth function has a smaller norm in the RKHS. The RKHS norm in the optimization function is used to reduce the volatility of the ranking scores, making them smoother and more consistent, thereby improving the accuracy of the ranking. The third part is the manifold regularizer, which ensures the continuity of the propagated label information on the graph. In the optimization function, this part constrains the continuity of label information on the graph to prevent large jumps during the propagation process, making the label information smoother. This part is very important for the task of label propagation on the graph because it can ensure that the label information is more reliable and accurate during the propagation process. A is a parameter that controls the smoothness of the manifold regularizer, λ I It is a parameter used to ensure the relationship between label propagation and kNN graph. When using this optimization function for label propagation, it is necessary to adjust the smoothness parameter λ of the manifold regularizer according to the actual situation. A and the parameter λ that ensures the relationship between label propagation and the kNN graph I Finally, the optimization function is solved by Newton's method. New ranking score S for the retrieval object pr Updated to

[0037] Step 4: Use the k-fusion algorithm to eliminate sample jitter that may have occurred during step 3. "Jitter" refers to the significant fluctuations in the generated ranking list as incremental user feedback accumulates. This fluctuation can cause positive samples that were initially far apart to move closer together. Due to the discrepancy between model computation and human perception, constantly changing rankings are prone to jitter. Jitter is an inherent problem in existing graph propagation algorithms, as they propagate labels solely based on sample feature similarity. When there is a discrepancy between human perception and model computation, existing algorithms only accept label information based on human feedback and not pairwise relationship information. This can lead to ranking results that contradict human perception. First, all samples with jitter are detected using rules, and then they are updated using a progressive approach. Since jitter does not exist in single-shot feedback, samples with jitter can be identified by comparing the ranking of each sample across branches. Based on the detected samples, the accumulated feedback is divided into single-shot feedback and a branch-ranked list based on each single-shot feedback is calculated using existing algorithms. Finally, the highest score in each branch-ranked list is used to update the sample with jitter.

[0038] Step 4.1: Input the search object q, the search image set G, the user feedback information L and the new ranking score S calculated in step 3. pr .

[0039] Step 4.2, remove L + Each of and L - As a branch feedback input f pr Calculate the ranking score S of each branch ranking list i and branch ranking list r i .

[0040] Step 4.3: Merge the ranking lists of all branches into a ranking list set R.

[0041] Step 4.4: For the top k samples in each ranking list, check their ranking changes in other ranking lists. If the ranking change of a sample in other ranking lists exceeds N places, update the score of the sample in the current ranking list. Here, we set N to 500 and k to 20.

[0042] Step 4.5: Use the maximum value of each sample score in all branches to update the ranking score of the sample.

[0043] Step 4.6: Output the updated ranking score S pr* As the output of the k-fusion algorithm. The updated S pr*This is the result obtained by our entire algorithm after one feedback. Users can provide multiple feedbacks to continuously optimize the ranking.

[0044] Step 5: Testing phase: We apply our person-in-the-loop-based post-permutation method to the MSMT17 dataset and compare it with the existing post-permutation method (POP).

[0045] In step 5.1, during the testing phase, an unsupervised person re-identification algorithm was applied to the MSMT17 dataset. For each query target q in the query set, the person re-identification algorithm generated a ranked list of results from the image gallery. For each ranked list, samples with the same person ID and camera ID as q were removed, and the top 1000 results were retained. 200 query targets q and their ranked lists were then selected for subsequent manual feedback.

[0046] In step 5.2, three participants conducted an evaluation experiment with manual feedback on the 200 samples. Each user was randomly assigned a sample and one of the two algorithms (our human-in-the-loop person re-identification post-ranking method and POP). The user provided multiple rounds of feedback on the ranked list according to the requirements of step 2 on an interactive web page. In each round of feedback, the user prioritized the top 10 samples, labeled images that were considered to have the same pedestrian ID as p as positive samples, and labeled images that were least similar to p as negative samples. If no positive sample was found in the top 10, the user could search backwards until a positive sample was found. Each sample received five rounds of feedback.

[0047] Step 5.3. After feedback from both methods on all samples, evaluate the accuracy of the method using mAP, Rank-1, and Rank-5 evaluation indicators.

[0048] The technical concept of the present invention is as follows: The present invention provides a post-reranking method for pedestrian re-identification based on human-in-the-loop, which aims to solve the problems existing in existing pedestrian re-identification technologies. The algorithm uses user feedback information to optimize and correct the initial ranking list, thereby improving the accuracy and stability of pedestrian re-identification. The method includes the steps of constructing a kNN graph, performing label propagation on the constructed graph, solving the optimization function using the Newton solver, and further optimizing the results by using a k-fusion method that reduces sample jitter. The innovation of the present invention is that it proposes a reranking method for pedestrian re-identification based on human feedback, which allows users to verify and retrieve the ranking list more efficiently.

[0049] The advantages of the present invention are: 1) Improving the accuracy and stability of pedestrian re-identification: The present invention proposes a post-reranking method for pedestrian re-identification based on a human in the loop, which can further optimize and correct the initial ranking list through user feedback, thereby improving the accuracy and stability of pedestrian re-identification. 2) Enhancing the generalization of the algorithm: The generalization of traditional pedestrian re-identification technology is very limited in practical applications. The post-reranking method for pedestrian re-identification based on a human in the loop proposed by the present invention can optimize and correct the initial ranking list through human feedback, thereby enhancing the generalization of the algorithm. 3) Reducing sample jitter during incremental feedback: The present invention provides a k-fusion method for reducing sample jitter during incremental feedback, which can effectively reduce the impact of sample jitter during the feedback process and improve the stability and accuracy of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0051] Figure 1 Schematic diagram of the post-reordering algorithm for pedestrian re-identification based on people in the loop of the present invention;

[0052] Figure 2 This is a schematic diagram of the human feedback interface of the present invention;

[0053] Figure 3 Schematic diagram of data processing of the present invention;

[0054] Figure 4 Schematic diagram of the human-in-the-loop feedback mechanism of the present invention;

[0055] Figure 5 Schematic diagram of a user retrieval and feedback mechanism according to an embodiment of the present invention. Specific implementation plan

[0056] Various exemplary embodiments of the present invention will now be described in detail. This detailed description should not be considered as limiting the present invention, but rather as a more detailed description of certain aspects, features, and embodiments of the present invention.

[0057] Example 1

[0058] See also Figures 1 to 4 , Figure 1The complete process of this method is presented. In order to be able to compare with other post-reranking algorithms, the MSMT17 dataset is selected. An unsupervised pedestrian re-identification model is used to obtain the initial ranking of all images in the MSMT17 dataset for the retrieval target; the user performs iterative feedback operations on the ranking results in the feedback interface; a semi-supervised label graph propagation method is used to optimize the pedestrian re-identification ranking score with user feedback; the k-fusion algorithm is used to eliminate sample jitter that may occur in step 3, and the optimized ranking is obtained after one feedback; the user feedback interface and ranking list are updated to the latest status.

[0059] A pedestrian re-identification and re-ranking method based on people in the loop specifically comprises the following steps:

[0060] Step 1: Figure 2 , the unsupervised person re-identification model is used on the MSMT17 dataset to obtain the initial ranking of all images in the gallery for the retrieval target. Given a retrieval object q and a gallery containing n images As the input of the unsupervised pedestrian re-identification model, the features of each image are obtained and the initial ranking score S of each image in the gallery is calculated. init .

[0061] Step 1.1, data preprocessing: given a retrieval object q and a library containing n images As the input of the unsupervised person re-identification model, the original image is scaled to 256x128 and randomly horizontally flipped and randomly cropped to increase the diversity and robustness of the data.

[0062] Step 1.2, Feature Extraction: Use the ImageNet pre-trained ResNet-50 as the feature extractor, remove the last fully connected layer, and only retain the previous convolutional and pooling layers. Use this ResNet-50 to convert the 256x128 image into a 2048-dimensional feature vector and perform L2 normalization.

[0063] Step 1.3, attention mechanism: Introduce a learnable weight matrix, perform weighted summation on the feature vectors, obtain an attention vector, and add it to the original feature vector to enhance the identity-related parts, reduce the interference of irrelevant information, and improve recognition accuracy. Specifically, it includes the following steps:

[0064] First, the input image is passed through the feature extractor to produce a feature map of size C x H x W, where C is the number of channels, H and W are the height and width.

[0065] The feature map is then passed through an adaptive average pooling layer to obtain an output feature map of size C x 1 x 1. This step can compress the spatial dimension to 1x1, thereby reducing the amount of computation and the number of parameters.

[0066] Next, the output feature map is passed through a fully connected layer to obtain a weight vector of size C. This step maps the 1x1 value corresponding to each channel to a real value, indicating the contribution of the channel to the final feature vector.

[0067] Finally, the weight vector is multiplied by the original feature map and summed along the channel dimension to obtain the final 2048-dimensional feature vector. This step allows each channel to be weighted averaged according to its weight value and retains the useful information in the original feature map.

[0068] Step 1.4, Fully Connected Layer and Batch Normalization Layer: Use a fully connected layer to map the feature vector to a low-dimensional space. The size of the weight matrix W of the fully connected layer is 128x2048, and the size of the bias vector b is 128. Then use the batch normalization layer to perform a standardization operation and normalize again to reduce the dimension and variance of the feature vector, improving computational efficiency and stability.

[0069] Step 1.5, Category Prediction: Using the representative vectors stored in the memory bank, we use cosine similarity to calculate the distance between the feature vector and each category. We then find the nearest neighbor category as the predicted label to guide the model in learning to better distinguish between different categories. The memory bank is a matrix that stores the representative vector for each category in a 128-dimensional space. Distance is calculated using the Euclidean distance, which is the L2 norm between two vectors.

[0070] Step 1.6, Instance Retrieval: Use the instance feature vectors stored in the memory bank to calculate the similarity between the feature vectors and each instance, find the most similar k instances as retrieval targets, and sort them in descending order of similarity to obtain the initial ranking score S of each image in the gallery corresponding to the retrieval target. init .

[0071] Step 2: The user performs iterative feedback on the ranking results. This step requires the user to provide at least one positive sample and one negative sample for each feedback. Figure 3, users can label images with positive labels, negative labels, and undetermined labels (this label will not affect the model results). If the user cannot find a positive sample, they are allowed to select a negative sample similar to the query image q as a positive sample. We denote the set of all feedback operations performed as L, and use these feedback operations to further optimize and correct the initial ranking list to improve the accuracy and stability of pedestrian re-identification. The user provides at least one positive label and one negative label. It should be noted that if the user cannot find a positive sample, they can select a negative sample similar to the query q as a positive sample. We denote the set of samples marked as positive samples by the user as The negative samples are represented as The remaining unlabeled instances G in the gallery are denoted as unlabeled samples U = G\L.

[0072] Step 3: Figure 4 A semi-supervised label graph propagation method leverages user feedback to optimize person re-ID ranking scores. The first step is to construct a kNN graph to characterize the similarity between images. The second step is to perform label propagation on this graph, transferring the labels of positive and negative samples to the remaining unlabeled samples. When constructing the kNN graph, a high-dimensional vector is used to represent the features of each sample, and the similarity between samples is calculated, ultimately constructing a kNN graph. During label propagation, an optimization function is constructed using the positive and negative labels provided by the user, and the function is solved using the Newton solver. Ultimately, the resulting new ranking score is used to update the similarity score of each sample in the image library with the detection sample.

[0073] Step 3.1: Construct a kNN proximity graph to characterize the similarity between images. First, feature extraction is performed. For each image in G, the high-dimensional vector extracted by the person re-identification algorithm is used to represent each image.

[0074] Step 3.2: Image similarity calculation: After obtaining the high-dimensional vector of each image, the cosine similarity method is used to calculate the similarity score between samples.

[0075]

[0076] Step 3.3. Select the parameter k. Based on the similarity score, select the k most similar neighbors of each data point and connect them to the data point. Here, a priority queue data structure is used to store the distance between samples and the nearest neighbor information.

[0077] Step 3.4: Construct a proximity graph. Add the lines between each data point (i.e., sample) and its neighbors to the graph to obtain an undirected graph consisting of nodes and edges, namely the kNN proximity graph.

[0078] Step 3.5: Label propagation on the graph. Based on the constructed proximity graph, propagate the labeled sample L to the remaining unlabeled samples U. The user feedback l input to the model is expressed as follows:

[0079]

[0080] Step 3.6, Initialization: Set the label value of the sample that the user has fed back to the formula shown above, and the label value of the sample that has not fed back to 0.

[0081] Step 3.7, iterative update: For each unlabeled node i, update its label value according to the label value of its neighbor node j. The specific formula is:

[0082]

[0083] Among them, y i Indicates the label value of node y, w i,j represents the edge weight between node i and node j, N i Represents the set of neighbor nodes of node i.

[0084] Step 3.8: To ensure the smoothness and continuity of label propagation, we use a propagation optimization function based on the Laplace support vector machine. The formula is as follows:

[0085]

[0086] The optimization function consists of three parts. The first is the hinge loss, a commonly used loss function based on user feedback, used to measure the error of a classifier. In the optimization function, hinge loss is used to measure the variance between samples, encouraging the classifier to correctly classify samples and penalizing incorrectly classified samples. Specifically, in the optimization function, hinge loss can help make the ranking scores more accurate and reliable. The second part is the norm of the function in the Reproducing Kernel Hilbert Space (RKHS), which is used to ensure the smoothness of the model. In the optimization function, smoothness refers to the norm of the function in the RKHS, which reflects the rate of change of the function value. Therefore, a smooth function has a smaller norm in the RKHS. The RKHS norm in the optimization function is used to reduce the volatility of the ranking scores, making them smoother and more consistent, thereby improving the accuracy of the ranking. The third part is the manifold regularizer, which ensures the continuity of the propagated label information on the graph. In the optimization function, this part constrains the continuity of label information on the graph to prevent large jumps during the propagation process, making the label information smoother. This part is very important for the task of label propagation on the graph because it can ensure that the label information is more reliable and accurate during the propagation process. A is a parameter that controls the smoothness of the manifold regularizer, λ I It is a parameter used to ensure the relationship between label propagation and kNN graph. When using this optimization function for label propagation, it is necessary to adjust the smoothness parameter λ of the manifold regularizer according to the actual situation. A and the parameter λ that ensures the relationship between label propagation and the kNN graph I Finally, the optimization function is solved by Newton's method. New ranking score S for the retrieval object pr Updated to

[0087] Step 4: Use the k-fusion algorithm to eliminate sample jitter that may have occurred during step 3. "Jitter" refers to the significant fluctuations in the generated ranking list as incremental user feedback accumulates. This fluctuation can cause positive samples that were initially far apart to move closer together. Due to the discrepancy between model computation and human perception, constantly changing rankings are prone to jitter. Jitter is an inherent problem in existing graph propagation algorithms, as they propagate labels solely based on sample feature similarity. When there is a discrepancy between human perception and model computation, existing algorithms only accept label information based on human feedback and not pairwise relationship information. This can lead to ranking results that contradict human perception. First, all samples with jitter are detected using rules, and then they are updated using a progressive approach. Since jitter does not exist in single-shot feedback, samples with jitter can be identified by comparing the ranking of each sample across branches. Based on the detected samples, the accumulated feedback is divided into single-shot feedback and a branch-ranked list based on each single-shot feedback is calculated using existing algorithms. Finally, the highest score in each branch-ranked list is used to update the sample with jitter.

[0088] Step 4.1: Input the search object q, the search image set G, the user feedback information L and the new ranking score S calculated in step 3. pr .

[0089] Step 4.2, remove L + Each of and L - As a branch feedback input f pr Calculate the ranking score S of each branch ranking list i and branch ranking list r i .

[0090] Step 4.3: Merge the ranking lists of all branches into a ranking list set R.

[0091] Step 4.4: For the top k samples in each ranking list, check their ranking changes in other ranking lists. If the ranking change of a sample in other ranking lists exceeds N places, update the score of the sample in the current ranking list. Here, we set N to 500 and k to 20.

[0092] Step 4.5: Use the maximum value of each sample score in all branches to update the ranking score of the sample.

[0093] Step 4.6: Output the updated ranking score S pr* As the output of the k-fusion algorithm. The updated S pr*This is the result obtained by our entire algorithm after one feedback. Users can provide multiple feedbacks to continuously optimize the ranking.

[0094] Step 5: Testing phase: We apply our person-in-the-loop-based post-permutation method to the MSMT17 dataset and compare it with the existing post-permutation method (POP).

[0095] In step 5.1, during the testing phase, an unsupervised person re-identification algorithm was applied to the MSMT17 dataset. For each query target q in the query set, the person re-identification algorithm generated a ranked list of results from the image gallery. For each ranked list, samples with the same person ID and camera ID as q were removed, and the top 1000 results were retained. 200 query targets q and their ranked lists were then selected for subsequent manual feedback.

[0096] In step 5.2, three participants conducted an evaluation experiment with manual feedback on the 200 samples. Each user was randomly assigned a sample and one of the two algorithms (our human-in-the-loop person re-identification post-ranking method and POP). The user provided multiple rounds of feedback on the ranked list according to the requirements of step 2 on an interactive web page. In each round of feedback, the user prioritized the top 10 samples, labeled images that were considered to have the same pedestrian ID as p as positive samples, and labeled images that were least similar to p as negative samples. If no positive sample was found in the top 10, the user could search backwards until a positive sample was found. Each sample received five rounds of feedback.

[0097] Step 5.3. After providing feedback to both methods on all samples, we evaluate the accuracy of our method using mAP, Rank-1, and Rank-5 evaluation indicators. The results show that our method outperforms the POP method in terms of mAP, Rank-1, and Rank-5 indicators in all five rounds of feedback.

[0098] Example 2

[0099] Reference Figure 5 The present invention provides a method for tracking people based on a pedestrian re-identification and rearrangement method in a loop. Specifically, the method includes the following steps:

[0100] Step 1: Data acquisition and processing. This method acquires video stream data across cameras. After acquiring the video stream, the target detection method is first used to obtain the detection frame of all people in the video stream to obtain a library containing n images. The user specifies the search target q( Figure 5 A) As the tracked person in this method, the retrieval object q and the image gallery G are used as the input of the unsupervised pedestrian re-identification model, the features of each image are obtained and the initial ranking score S of each image in the gallery is calculated. init .

[0101] Step 1.1, data preprocessing: For a given retrieval object q and a library containing n images As the input of the unsupervised person re-identification model, the original image is scaled to 256x128 and randomly horizontally flipped and randomly cropped to increase the diversity and robustness of the data.

[0102] Step 1.2, Feature Extraction: Use the ImageNet pre-trained ResNet-50 as the feature extractor, remove the last fully connected layer, and only retain the previous convolutional and pooling layers. Use this ResNet-50 to convert the 256x128 image into a 2048-dimensional feature vector and perform L2 normalization.

[0103] Step 1.3, attention mechanism: Introduce a learnable weight matrix, perform weighted summation on the feature vectors, obtain an attention vector, and add it to the original feature vector to enhance the identity-related parts, reduce the interference of irrelevant information, and improve recognition accuracy. Specifically, it includes the following steps:

[0104] First, the input image is passed through the feature extractor to produce a feature map of size C x H x W, where C is the number of channels, H and W are the height and width.

[0105] The feature map is then passed through an adaptive average pooling layer to obtain an output feature map of size C x 1 x 1. This step can compress the spatial dimension to 1x1, thereby reducing the amount of computation and the number of parameters.

[0106] Next, the output feature map is passed through a fully connected layer to obtain a weight vector of size C. This step maps the 1x1 value corresponding to each channel to a real value, indicating the contribution of the channel to the final feature vector.

[0107] Finally, the weight vector is multiplied by the original feature map and summed along the channel dimension to obtain the final 2048-dimensional feature vector. This step allows each channel to be weighted averaged according to its weight value and retains the useful information in the original feature map.

[0108] Step 1.4, Fully Connected Layer and Batch Normalization Layer: Use a fully connected layer to map the feature vector to a low-dimensional space. The size of the weight matrix W of the fully connected layer is 128x2048, and the size of the bias vector b is 128. Then use the batch normalization layer to perform a standardization operation and normalize again to reduce the dimension and variance of the feature vector, improving computational efficiency and stability.

[0109] Step 1.5, Category Prediction: Using the representative vectors stored in the memory bank, we use cosine similarity to calculate the distance between the feature vector and each category. We then find the nearest neighbor category as the predicted label to guide the model in learning to better distinguish between different categories. The memory bank is a matrix that stores the representative vector for each category in a 128-dimensional space. Distance is calculated using the Euclidean distance, which is the L2 norm between two vectors.

[0110] Step 1.6, Instance Retrieval: Use the instance feature vectors stored in the memory bank to calculate the similarity between the feature vectors and each instance, find the most similar k instances as retrieval targets, and sort them in descending order of similarity to obtain the initial ranking score S of each image in the gallery corresponding to the retrieval target. init .

[0111] Step 2: Users can view the top ten detection results for the tracked personnel ( Figure 5 Ba), perform iterative feedback operation on the ranking results. This step requires the user to provide at least one positive sample and one negative sample for each feedback. Figure 3 , users can label images with positive labels, negative labels, and undetermined labels (this label will not affect the model results). If the user cannot find a positive sample, they are allowed to select a negative sample similar to the query image q as a positive sample. Figure 5 B, labels the a1 sample positively and the a2 sample negatively. We denote the set of all feedback operations performed as L, which can be used to further optimize and correct the initial ranking list, improving the accuracy and stability of pedestrian re-identification. The user provides at least one positive label and one negative label. It should be noted that if the user cannot find a positive sample, a negative sample similar to the query q can be selected as a positive sample. We denote the set of samples marked as positive samples by the user as The negative samples are represented as The remaining unlabeled instances G in the gallery are denoted as unlabeled samples U = G\L.

[0112] Step 3: Figure 4 A semi-supervised label graph propagation method leverages user feedback to optimize person re-ID ranking scores. The first step is to construct a kNN graph to characterize the similarity between images. The second step is to perform label propagation on this graph, transferring the labels of positive and negative samples to the remaining unlabeled samples. When constructing the kNN graph, a high-dimensional vector is used to represent the features of each sample, and the similarity between samples is calculated, ultimately constructing a kNN graph. During label propagation, an optimization function is constructed using the positive and negative labels provided by the user, and the function is solved using the Newton solver. Ultimately, the resulting new ranking score is used to update the similarity score of each sample in the image library with the detection sample.

[0113] Step 3.1: Construct a kNN proximity graph to characterize the similarity between images. First, feature extraction is performed. For each image in G, the high-dimensional vector extracted by the person re-identification algorithm is used to represent each image.

[0114] Step 3.2: Image similarity calculation: After obtaining the high-dimensional vector of each image, the cosine similarity method is used to calculate the similarity score between samples.

[0115]

[0116] Step 3.3. Select the parameter k. Based on the similarity score, select the k most similar neighbors of each data point and connect them to the data point. Here, use the priority queue data structure to store the distance between samples and the nearest neighbor information.

[0117] Step 3.4: Construct a proximity graph. Add the lines between each data point (i.e., sample) and its neighbors to the graph to obtain an undirected graph consisting of nodes and edges, namely the kNN proximity graph.

[0118] Step 3.5: Label propagation on the graph. Based on the constructed proximity graph, propagate the labeled sample L to the remaining unlabeled samples U. The user feedback l input to the model is expressed as follows:

[0119]

[0120] Step 3.6, Initialization: Set the label value of the sample that the user has fed back to the formula shown above, and the label value of the sample that has not fed back to 0.

[0121] Step 3.7, iterative update: For each unlabeled node i, update its label value according to the label value of its neighbor node j. The specific formula is:

[0122]

[0123] Among them, y i represents the label value of node i, w i,j represents the edge weight between node i and node j, N i Represents the set of neighbor nodes of node i.

[0124] Step 3.8: To ensure the smoothness and continuity of label propagation, we use a propagation optimization function based on the Laplace support vector machine. The formula is as follows:

[0125]

[0126] The optimization function consists of three parts. The first is the hinge loss, a commonly used loss function based on user feedback, used to measure the error of a classifier. In the optimization function, hinge loss is used to measure the variance between samples, encouraging the classifier to correctly classify samples and penalizing incorrectly classified samples. Specifically, in the optimization function, hinge loss can help make the ranking scores more accurate and reliable. The second part is the norm of the function in the Reproducing Kernel Hilbert Space (RKHS), which is used to ensure the smoothness of the model. In the optimization function, smoothness refers to the norm of the function in the RKHS, which reflects the rate of change of the function value. Therefore, a smooth function has a smaller norm in the RKHS. The RKHS norm in the optimization function is used to reduce the volatility of the ranking scores, making them smoother and more consistent, thereby improving the accuracy of the ranking. The third part is the manifold regularizer, which ensures the continuity of the propagated label information on the graph. In the optimization function, this part constrains the continuity of label information on the graph to prevent large jumps during the propagation process, making the label information smoother. This part is very important for the task of label propagation on the graph because it can ensure that the label information is more reliable and accurate during the propagation process. A is a parameter that controls the smoothness of the manifold regularizer, λ I It is a parameter used to ensure the relationship between label propagation and kNN graph. When using this optimization function for label propagation, it is necessary to adjust the smoothness parameter λ of the manifold regularizer according to the actual situation. A and the parameter λ that ensures the relationship between label propagation and the kNN graph I Finally, the optimization function is solved by Newton's method. New ranking score S for the retrieval object pr Updated to

[0127] Step 4: Use the k-fusion algorithm to eliminate sample jitter that may have occurred during step 3. "Jitter" refers to the significant fluctuations in the generated ranking list as incremental user feedback accumulates. This fluctuation can cause positive samples that were initially far apart to move closer together. Due to the discrepancy between model computation and human perception, constantly changing rankings are prone to jitter. Jitter is an inherent problem in existing graph propagation algorithms, as they propagate labels solely based on sample feature similarity. When there is a discrepancy between human perception and model computation, existing algorithms only accept label information based on human feedback and not pairwise relationship information. This can lead to ranking results that contradict human perception. First, all samples with jitter are detected using rules, and then they are updated using a progressive approach. Since jitter does not exist in single-shot feedback, samples with jitter can be identified by comparing the ranking of each sample across branches. Based on the detected samples, the accumulated feedback is divided into single-shot feedback and a branch-ranked list based on each single-shot feedback is calculated using existing algorithms. Finally, the highest score in each branch-ranked list is used to update the sample with jitter.

[0128] Step 4.1: Input the search object q, the search image set G, the user feedback information L and the new ranking score S calculated in step 3. pr .

[0129] Step 4.2, remove L + Each of and L - As a branch feedback input f pr Calculate the ranking score S of each branch ranking list i and branch ranking list r i .

[0130] Step 4.3: Merge the ranking lists of all branches into a ranking list set R.

[0131] Step 4.4: For the top k samples in each ranking list, check their ranking changes in other ranking lists. If the ranking change of a sample in other ranking lists exceeds N places, update the score of the sample in the current ranking list. Here, we set N to 500 and k to 20.

[0132] Step 4.5: Use the maximum value of each sample score in all branches to update the ranking score of the sample.

[0133] Step 4.6: Output the updated ranking score S pr* As the output of the k-fusion algorithm. The updated S pr*This is the result obtained by our entire algorithm after one feedback. Users can view the updated ranking results of the tracked persons.

[0134] Step 5: After multiple rounds of feedback updates, accurate and reliable tracking results for the tracked person are obtained, including the camera location, time, and photos of the tracked person. Once the user has sufficient information, they can stop searching and conduct more in-depth analysis.

Claims

1. A pedestrian re-identification and re-ranking method based on people in the loop, characterized in that: The following steps are involved: Step 1: Use the unsupervised person re-identification model to obtain the initial ranking of all images in the gallery for the retrieval target. Given a retrieval object , and contains Gallery of images As the input of the unsupervised pedestrian re-identification model, the features of each image are obtained and the initial ranking score of each image in the gallery is calculated. ; Step 2: The user performs iterative feedback on the ranking results; the user is required to provide at least one positive sample and one negative sample for each feedback; if the user cannot find a positive sample, he is allowed to select one that matches the query image. Similar negative samples are used as positive samples; The set of all executed feedback operations is denoted as , using these feedback operations can further optimize and correct the initial ranking list, improve the accuracy and stability of pedestrian re-identification; the user provides at least one positive label and one negative label; it should be noted that if the user cannot find a positive sample, he can choose one that matches the query Similar negative samples are used as positive samples; The sample set marked by the user as positive samples is represented as , and the negative samples are represented as ; The remaining unlabeled instances in the gallery Represented as unlabeled samples ; Step 3: Use a semi-supervised label graph propagation method to utilize user feedback to optimize the pedestrian re-identification ranking score; the first step is to construct a kNN graph to characterize the similarity between images; The second step is to perform label propagation on the constructed graph, transferring the labels of positive and negative samples to the remaining unlabeled samples. When constructing the kNN graph, a high-dimensional vector is used to represent the characteristics of each sample, and the similarity between samples is calculated to ultimately construct a kNN graph. When performing label propagation, the positive and negative labels provided by the user are used to construct a propagation optimization function based on the Laplace support vector machine and solve it using the Newton solver. Finally, the new ranking score is used to update the similarity score of each sample in the image library with the detection sample. Step 4. Use the k-fusion algorithm to eliminate sample jitter that may be generated during step 3. First, detect all samples with jitter through rules, and then update them using a progressive method. Since there is no jitter in a single feedback, the samples with jitter can be determined by comparing the ranking of each sample between branches. Based on the detected samples, the cumulative feedback is divided into single feedbacks, and the existing algorithm is used to calculate the branch ranking list based on each single feedback. Finally, the sample with jitter is updated using the highest score in each branch ranking list.

2. The method for pedestrian re-identification and rearrangement based on human-in-the-loop according to claim 1, characterized in that: The step 1 specifically includes the following steps: Step 1.1, Data preprocessing: Given a retrieval object , and contains Gallery of images As the input of the unsupervised person re-identification model, the original image is scaled to 256x128, and randomly horizontally flipped and randomly cropped to increase the diversity and robustness of the data; Step 1.2, Feature Extraction: Use the ImageNet pre-trained ResNet-50 as the feature extractor, remove the last fully connected layer, and only keep the previous convolutional and pooling layers. Use this ResNet-50 to convert the 256x128 image into a 2048-dimensional feature vector and perform L2 normalization. Step 1.3, attention mechanism: Introduce a learnable weight matrix, perform weighted summation on the feature vectors, obtain an attention vector, and add it to the original feature vector to enhance the identity-related parts, reduce the interference of irrelevant information, and improve recognition accuracy; specifically, it includes: First, the input image is passed through the feature extractor to obtain a feature map of size C x H x W, where C is the number of channels, H and W are the height and width; Then, the feature map is passed through an adaptive average pooling layer to obtain an output feature map of size C x 1 x 1; this step can compress the spatial dimension to 1x1, thereby reducing the amount of computation and the number of parameters; Next, the output feature map is passed through a fully connected layer to obtain a weight vector of size C. This step maps the 1x1 value corresponding to each channel to a real value, indicating the contribution of the channel to the final feature vector. Finally, the weight vector is multiplied by the original feature map and summed along the channel dimension to obtain the final 2048-dimensional feature vector; this step allows each channel to be weighted averaged according to its weight value and retains the useful information in the original feature map; Step 1.4, Fully connected layer and batch normalization layer: Use the fully connected layer to map the feature vector to a low-dimensional space, and the weight matrix of the fully connected layer The size of is 128x2048, and the size of the bias vector b is 128; then the batch normalization layer is used for standardization and normalization again to reduce the dimension and variance of the feature vector and improve computational efficiency and stability; Step 1.5, Category Prediction: Using the representative vectors stored in the memory bank, we use cosine similarity to calculate the distance between the feature vector and each category, and find the nearest neighbor category as the predicted label to guide the model to learn to better distinguish different categories. The memory bank is a matrix that stores the representative vector of each category in a 128-dimensional space. Distance is calculated using the Euclidean distance, which is the L2 norm between two vectors. Step 1.6, Instance Retrieval: Using the instance feature vectors stored in the memory bank, calculate the similarity between the feature vectors and each instance, and find the most similar Instances are used as retrieval targets and sorted in descending order of similarity to obtain the initial ranking score of each image in the gallery corresponding to the retrieval target. .

3. The method for pedestrian re-identification and rearrangement based on human-in-the-loop according to claim 1, characterized in that: The step 3 specifically includes the following steps: Step 3.1, construct a kNN neighbor graph to describe the similarity between images; the first step is feature extraction, for For each picture in , use the high-dimensional vector of each picture extracted by the pedestrian re-identification algorithm to represent each picture; Step 3.2: Image similarity calculation: After obtaining the high-dimensional vector of each image, use the cosine similarity method to calculate the similarity score between samples; (1) Step 3.

3. Select parameters ; Based on the similarity score, select the most similar one for each data point neighbors and connect them to the data point; here, the priority queue data structure is used to store the distance between samples and the nearest neighbor information; Step 3.4: Construct a proximity graph. Add the lines between each data point (i.e., the sample) and its neighbors to the graph to obtain an undirected graph consisting of nodes and edges, i.e., the kNN proximity graph. Step 3.5: Label propagation on the graph; Based on the constructed proximity graph, the labeled samples Propagate to the remaining unlabeled samples ; User feedback input to the model It is expressed as follows: (2) Step 3.6, Initialization: Set the label value of the sample that the user has fed back to the formula shown above, and the label value of the sample that has not fed back to 0; Step 3.7, iterative update: For each unlabeled node , according to its neighbor nodes The label value of the tag is updated by the tag value of the tag. The specific formula is: (3) in, Representation node The label value of Representation node and nodes The edge weights between Representation node The set of neighbor nodes of Step 3.8: To ensure the smoothness and continuity of label propagation, a propagation optimization function based on the Laplace support vector machine is used. The formula is as follows: (4) The optimization function consists of three parts. The first part is the hinge loss based on user feedback. Hinge loss is a commonly used loss function used to measure the error of the classifier. In the optimization function, hinge loss is used to measure the difference between samples because it can encourage the classifier to correctly classify samples and penalize incorrectly classified samples. Specifically, in the optimization function, hinge loss can help make the ranking score more accurate and reliable. The second part is the function in the Reproducing Kernel Hilbert The norm in Space, referred to as RKHS norm, is used to ensure the smoothness of the model; in the optimization function, smoothness refers to the norm of the function in RKHS, which reflects the rate of change of the function value; the RKHS norm in the optimization function is used to reduce the volatility of the ranking score, making the ranking score smoother and more consistent, thereby improving the accuracy of the ranking; the third part is the manifold regularizer, which is used to ensure the continuity of the propagated label information on the graph; in the optimization function, this part constrains the continuity of the label information on the graph; this part is very important for the task of label propagation on the graph, because it can ensure that the label information is more reliable and accurate during the propagation process; among them, is a parameter that controls the smoothness of the manifold regularizer, It is a parameter used to ensure the relationship between label propagation and kNN graph; when using this optimization function for label propagation, it is necessary to adjust the smoothness parameter of the manifold regularizer according to the actual situation. and parameters that ensure the relationship between label propagation and kNN graph Finally, the optimization function is solved by Newton's method; Medium sample New ranking score for the search object Updated to .

4. The method for pedestrian re-identification and rearrangement based on human-in-the-loop according to claim 1, characterized in that: The step 4 specifically includes the following steps: Step 4.1: Enter the search object , Retrieve image collection , user feedback information and the new ranking score calculated in step 3 ; Step 4.2, remove Each of and As a branch feedback input Calculate the ranking score of each branch ranking list and branch ranking list ; Step 4.3: Merge the ranking lists of all branches into a ranking list set R; Step 4.4: For each top ranking list samples, check their ranking changes in other ranking lists; if the ranking change of a sample in other ranking lists exceeds If the sample is in the current ranking list, the score of the sample is updated; Step 4.5: Use the maximum value of each sample score in all branches to update the ranking score of the sample; Step 4.6: Output the updated ranking score As the output of the k-fusion algorithm; the updated This is the result obtained by the entire algorithm after one feedback. Users can provide multiple feedbacks to continuously optimize the ranking.

5. The method for pedestrian re-identification and rearrangement based on human-in-the-loop according to claim 4, characterized in that: In step 4.4, Set to 500, Set to 20.