A Pedestrian Re-identification Method and System Based on Momentum Network and Contrastive Learning

By performing labeling and feature extraction under a single camera, combining momentum network and comparison learning methods, the problem of insufficient utilization of data labels in pedestrian recognition is solved, and the pedestrian recognition performance under cross-camera is improved while reducing costs.

CN114724075BActive Publication Date: 2025-07-08SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210214375.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-04
Publication Date
2025-07-08
Estimated Expiration
2042-03-04

AI Technical Summary

Technical Problem

The prior art does not fully utilize pedestrian picture data labels under a single camera in pedestrian re-identification, and the pedestrian picture correlation constraints across cameras are insufficient, resulting in poor model performance and high manual labeling costs.

Method used

Using a method based on momentum network and contrast learning, by performing labeling and feature extraction under a single camera, combining hierarchical clustering and cross-camera comparison loss calculation, the pedestrian re-identification model is optimized, the cost of manual labeling is reduced and the recognition effect is improved.

Benefits of technology

While reducing the cost of manual labeling, training noise interference is reduced through momentum networks, and more discriminant pedestrian picture features are obtained using contrast loss and hierarchical clustering, which improves the pedestrian re-identification effect under the camera.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114724075B_ABST
    Figure CN114724075B_ABST
Patent Text Reader

Abstract

The present invention discloses a pedestrian re-identification method and system based on a momentum network and contrastive learning. The method includes: obtaining pedestrian pictures and performing annotation processing; initializing a momentum network and a backbone network; extracting pedestrian picture features; calculating contrastive loss for the average feature set of pedestrian pictures and updating the backbone network and the momentum network; performing hierarchical clustering processing on the average feature set of pedestrian pictures and assigning pseudo-labels; calculating contrastive loss for the clustering results and updating the backbone network and the momentum network; updating the average feature set of pedestrian pictures; iteratively optimizing the pedestrian re-identification model multiple times and outputting the optimized pedestrian re-identification model. The present invention reduces the interference of noise in the data training process by designing a momentum network, and can improve the pedestrian re-identification effect while reducing the manual annotation cost. As a pedestrian re-identification method and system based on a momentum network and contrastive learning, the present invention can be widely applied to the field of computer vision technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and particularly to a pedestrian re-identification method and system based on momentum network and contrast learning. Background Art

[0002] With the development of deep learning, neural network technology has been applied in more and more scenarios. As a popular research direction in the field of computer vision, pedestrian re-identification has also received increasing attention. Pedestrian re-identification (Re-ID) mainly solves the problem of retrieving and matching the same pedestrian under different cameras. By comparing and sorting the features of the pedestrian image to be queried with the features of the pedestrian images in the retrieval library, the final retrieval result is obtained. Pedestrian re-identification has wide application scenarios and great practical significance in the fields of smart city, security, etc. In the process of annotating pedestrian pictures, it is relatively easy to annotate pedestrians under the same camera. Therefore, in recent years, some scholars have proposed a weakly supervised pedestrian re-identification setting with in-camera supervision and cross-camera unsupervised, hoping to achieve better model performance while reducing the pedestrian annotation cost. However, the existing technologies in this direction do not make full use of the data labels of pedestrian pictures under a single camera, and do not consider the association constraints well when associating pedestrian pictures under different cameras, resulting in poor model performance. Summary of the Invention

[0003] To solve the above technical problems, the purpose of the present invention is to provide a pedestrian re-identification method and system based on momentum network and contrast learning, which can improve the pedestrian re-identification effect while reducing the manual annotation cost.

[0004] The first technical solution adopted by the present invention is: A pedestrian re-identification method based on momentum network and contrast learning, comprising the following steps:

[0005] Obtain pedestrian pictures and perform annotation processing to obtain a pedestrian picture data set with labels under a single camera;

[0006] Based on using a pre-trained deep convolutional neural network model to initialize the momentum network and the backbone network in the pedestrian re-identification model respectively, obtain the initialized momentum network and the initialized backbone network;

[0007] Based on the initialized momentum network, perform feature extraction processing on the pedestrian picture data set with labels under a single camera to obtain a set of average features of pedestrian pictures;

[0008] Based on the initialized backbone network, calculate the contrast loss under a single camera for the set of average features of pedestrian pictures and update the initialized backbone network and the initialized momentum network in sequence to obtain the first updated backbone network and the first updated momentum network;

[0009] Perform hierarchical clustering on the average feature set of pedestrian images based on the first updated momentum network and assign pseudo-labels to obtain a clustering result;

[0010] Calculate the contrast loss under cross-cameras for the clustering result based on the first updated backbone network and perform backpropagation to sequentially update the first updated backbone network and the first updated momentum network to obtain a second updated backbone network and a second updated momentum network;

[0011] Update the average feature set of pedestrian images based on the second updated momentum network to obtain an updated average feature set of pedestrian images;

[0012] Loop through the steps of the hierarchical clustering algorithm and the steps of calculating the contrast loss under cross-cameras to perform multiple iterative optimizations on the pedestrian re-identification model until the number of iterations reaches a preset value, and output the optimized pedestrian re-identification model;

[0013] Perform cross-camera pedestrian re-identification applications based on the optimized pedestrian re-identification model.

[0014] Furthermore, the step of obtaining pedestrian images and performing annotation processing to obtain a pedestrian image dataset with single-camera labels specifically includes:

[0015] Obtain pedestrian images based on multiple cameras and construct a pedestrian image dataset;

[0016] Perform annotation on the pedestrian images under each camera based on the pedestrian image dataset to obtain a pedestrian image dataset with single-camera labels.

[0017] Furthermore, the step of performing feature extraction processing on the pedestrian image dataset with single-camera labels based on the initialized momentum network to obtain an average feature set of pedestrian images specifically includes:

[0018] Perform feature extraction processing on the pedestrian image dataset with single-camera labels based on the initialized momentum network to obtain pedestrian image features;

[0019] Take the average of the pedestrian image features of the same pedestrian under the same camera and integrate to obtain an average feature set of pedestrian images.

[0020] Furthermore, the contrast loss calculation under single-camera uses contrast loss calculation, and the formula is as follows:

[0021]

[0022] In the above formula, L represents the first loss value, c represents the index value of the camera, N c represents the total number of pedestrian labels under camera c, D cDenote the average feature set of pedestrian images under camera c. j and k represent the index values of the average feature set of pedestrian images under a single camera, and f(x i ) represents the feature of sample x i . Denote the average feature set of all pedestrian images with label j corresponding to sample x under camera c. τ represents the adjustment factor. i

[0023] Furthermore, the step of performing hierarchical clustering on the average feature set of pedestrian images and assigning pseudo-labels based on the first updated momentum network to obtain the clustering result specifically includes:

[0024] Calculate the feature fusion weight of the average feature set of pedestrians under a single camera based on the updated momentum network to obtain the feature fusion weight value;

[0025] Process the average feature set of pedestrians according to the feature fusion weight value to obtain the fused average feature set of pedestrians;

[0026] According to the fused average feature set of pedestrians, calculate the jaccard distance between the features of different pedestrian images under different cameras in the fused average feature set of pedestrians and construct a distance set;

[0027] Perform traversal sorting on the distance set to obtain the clustering result;

[0028] Assign the same pseudo-label to the features of the same pedestrian image under different cameras in the clustering result.

[0029] Furthermore, the decoupled contrast loss calculation is adopted for the cross-camera contrast loss calculation, and the formula is as follows:

[0030]

[0031] In the above formula, L represents the second loss value, N represents the total number of samples with pseudo-labels, f(x i ) represents the feature of sample x i , m j represents the average feature set of pedestrian images under a single camera with the same pseudo-label as sample x i , P represents the number of positive sample pedestrians from different cameras included in the cluster to which sample x i belongs, N neg represents the sampling number of negative sample pedestrians, m k represents the average feature set of pedestrian images under a single camera of negative samples belonging to other clusters that are relatively close to sample x i , τ represents the adjustment factor, and [·] + means taking the value of 0 when the value is negative.

[0032] Furthermore, the step of updating the average feature set of pedestrian images based on the second updated momentum network to obtain the updated average feature set of pedestrian images specifically includes:

[0033] Extract features from the pedestrian image dataset with single-camera labels based on the second updated momentum network and take the average to obtain the average feature set of the extracted pedestrian images;

[0034] Update the average feature set of pedestrian images based on the average feature set of the extracted pedestrian images to obtain the updated average feature set of pedestrian images.

[0035] Furthermore, the step of returning the hierarchical clustering algorithm steps and the cross-camera contrast loss calculation steps to iteratively optimize the pedestrian re-identification model until the number of iterations reaches a preset value and output the optimized pedestrian re-identification model specifically includes:

[0036] Judge the number of iterative optimizations of the pedestrian re-identification model;

[0037] If it is judged that the number of iterative optimizations of the pedestrian re-identification model has not reached the preset value, repeat the hierarchical clustering algorithm and the cross-camera contrast loss calculation for the iterative optimization of the pedestrian re-identification model;

[0038] If it is judged that the number of iterative optimizations of the pedestrian re-identification model has reached the preset value, output the optimized pedestrian re-identification model.

[0039] The second technical solution adopted by the present invention is: A system for pedestrian re-identification based on a momentum network and contrast learning, including:

[0040] A data annotation processing module for annotating the pedestrian image dataset to obtain a pedestrian image dataset with single-camera labels;

[0041] A data feature extraction module for extracting pedestrian image features under different cameras to obtain an average feature set of pedestrian images;

[0042] A first contrast loss calculation module for calculating the loss value between the average feature sets of pedestrians under each camera;

[0043] A hierarchical clustering module for performing hierarchical clustering processing on the average feature set of pedestrian images to obtain a clustering result;

[0044] A second contrast loss calculation module for calculating the loss value between the same pedestrian image features under different cameras in the clustering result;

[0045] A data feature update module for updating the average feature set of pedestrian images to obtain the updated average feature set of pedestrian images;

[0046] A model optimization module for iteratively optimizing a person re-identification model to obtain an optimized person re-identification model.

[0047] The beneficial effects of the method and system of the present invention are as follows: By annotating pedestrian pictures under a single camera, the present invention reduces the cost of manual annotation, mitigates the interference of noise during the data training process by designing a momentum network, and obtains more discriminative pedestrian picture features through contrastive loss calculation and hierarchical clustering, thereby improving the effect of person re-identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is a flowchart of the steps of a person re-identification method based on a momentum network and contrastive learning according to the present invention;

[0049] Figure 2 is a structural block diagram of a person re-identification system based on a momentum network and contrastive learning according to the present invention;

[0050] Figure 3 is a process framework diagram of a person re-identification method based on a momentum network and contrastive learning according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0051] The following further elaborates on the present invention in detail in conjunction with the drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of elaboration and description, and no limitation is made on the order between steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0052] The present invention mitigates the interference of noise during the data training process by designing a momentum network, and can improve the person re-identification effect while reducing the manual annotation cost.

[0053] Referring to Figure 1 and Figure 3 , the present invention provides a person re-identification method based on a momentum network and contrastive learning, and the method includes the following steps:

[0054] S1. Obtain pedestrian pictures and perform annotation processing to obtain a pedestrian picture dataset with labels under a single camera;

[0055] S11. Based on multiple cameras, obtain pedestrian pictures and construct a pedestrian picture dataset;

[0056] S12. Based on the pedestrian picture dataset, annotate the pedestrian pictures under each camera to obtain a pedestrian picture dataset with labels under a single camera.

[0057] Specifically, by selecting a pedestrian image with multiple cameras, the pedestrian images under each camera are separately labeled. Among them, for the same pedestrian, the images under different cameras may have different single-camera labeling tags, and a pedestrian image dataset with tags under a single camera but without global cross-camera tags is constructed.

[0058] S2. Initialize the momentum network and the backbone network in the pedestrian re-identification model respectively using a pre-trained deep convolutional neural network model to obtain the initialized momentum network and the initialized backbone network.

[0059] Specifically, both the backbone network and the momentum network adopt the ResNet-50 neural network. Remove the last fully connected layer, add a batch normalization layer, and initialize the network using the parameters pre-trained on the ImageNet dataset.

[0060] S3. Perform feature extraction processing on the pedestrian image dataset with tags under a single camera based on the initialized momentum network to obtain the pedestrian image average feature set.

[0061] S31. Perform feature extraction processing on the pedestrian image dataset with tags under a single camera based on the initialized momentum network to obtain the pedestrian image features.

[0062] S32. Take the average of the pedestrian image features of the same pedestrian under the same camera and integrate them to obtain the pedestrian image average feature set.

[0063] Specifically, use the initialized momentum network to perform feature extraction processing on the pedestrian image dataset with tags under a single camera, take the average of the image features of the same pedestrian to obtain the pedestrian average feature set under a single camera, and integrate them to obtain the pedestrian image average feature set, which can be represented by where c represents the camera number, N c represents the total number of pedestrian tags under camera c, represents the average feature of the image features of the pedestrian with tag j under camera c.

[0064] S4. Calculate the contrast loss under a single camera for the pedestrian image average feature set based on the initialized backbone network and update the initialized backbone network and the initialized momentum network in turn to obtain the first updated backbone network and the first updated momentum network.

[0065] Specifically, based on the initialized backbone network, the contrast loss calculation is performed on the average feature set of pedestrian images under a single camera, and the classification difference calculation is performed on the centers of multiple types of pedestrians under each camera. The average feature set of each batch of pedestrian images is composed of P randomly selected from all single-camera pedestrian labels, and each pedestrian selected is randomly selected from its picture K, a total of P × K pictures, and the backbone network parameters are updated, and the momentum network parameters are updated at the same time to obtain the first updated backbone network and the first updated momentum network. The contrast loss calculation under a single camera adopts contrast loss calculation, and the formula is as follows:

[0066]

[0067] In the above formula, L represents the first loss value, c represents the index value of the camera, and N c represents the total number of pedestrian labels under camera c, D c represents the average feature set of pedestrian images under camera c, j and k represent the index values ​​of the average feature set of pedestrian images under a single camera, and f(x i ) represents the sample x i Features, Represents the sample x under camera c i The corresponding average feature set of all pedestrian images with label j, τ represents the adjustment factor;

[0068] The momentum network update formula after the initialization is as follows:

[0069]

[0070] In the above formula, represents the parameters of the initialized backbone network, represents the momentum network parameters after initialization, represents the first updated momentum network parameter, and μ represents the update rate.

[0071] S5. Based on the first updated momentum network, hierarchical clustering is performed on the average feature set of the pedestrian image and pseudo labels are assigned to obtain a clustering result.

[0072] S51, calculating the feature fusion weight of the average feature set of pedestrians under a single camera based on the updated momentum network, and obtaining the feature fusion weight value;

[0073] Specifically, the feature fusion weight of the average feature set of pedestrians under a single camera is calculated based on the updated momentum network, and the feature fusion weight value is calculated for the same pedestrian picture with the same camera. The feature fusion weight w is ij The calculation formula is as follows:

[0074]

[0075] In the above formula, w ij represents the feature fusion weight value, M represents the number of pictures of the pedestrian under this camera, and τ represents the adjustment factor used to adjust the feature fusion weight.

[0076] S52. Process the pedestrian average feature set according to the feature fusion weight value to obtain the fused pedestrian average feature set;

[0077] Specifically, the feature fusion formula is expressed as follows:

[0078]

[0079] S53. According to the fused pedestrian average feature set, calculate the jaccard distance between the features of different pedestrian pictures under different cameras in the fused pedestrian average feature set, and construct a distance set;

[0080] S54. Perform traversal and sorting processing on the distance set to obtain a clustering result;

[0081] Specifically, calculate the jaccard distance between the features of pedestrian pictures after feature fusion. For two cross-camera pedestrians a and b from different cameras c a and c b , define their distance as D(a, b) = mind(x, y), where x represents the image of pedestrian a from camera c a , and y represents the image of pedestrian b from camera c b , and d(x, y) represents the jaccard distance between the features of images x and y; record all pairs of pedestrians from different cameras with D(a, b) < T, and form a set E = {(a, b, D(a, b))}, where T represents the distance threshold used to adjust the clustering result. Sort the elements in set E according to the value of D(a, b) from small to large, and define the clustering result set as Traverse the elements (a, b, D(a, b)) in the sorted set E. If neither pedestrian a nor pedestrian b belongs to any element g in the clustering result set G, then define a new set g = {a, b}, and the corresponding camera set C g = {c a , c b}, and add g to G; if one of a and b belongs to a certain element g in the clustering result set G, taking a ∈ g as an example, when the camera where b is located then add b to g, otherwise skip the current element (a, b, D(a, b)); when a and b each belong to two different elements in G, that is, a ∈ g a , b ∈ g b , when then put ga and g b Merge them, otherwise skip the current element (a, b, D(a, b)).

[0082] S55. Assign the same pseudo-label to the image features of the same pedestrian under different cameras in the clustering results.

[0083] Specifically, after completing the traversal, obtain the clustering result set G, and assign the same pseudo-label to all the images of pedestrians from different cameras belonging to the same element g in G.

[0084] S6. Calculate the contrastive loss across cameras for the clustering results based on the first updated backbone network and perform backpropagation to sequentially update the first updated backbone network and the first updated momentum network to obtain the second updated backbone network and the second updated momentum network;

[0085] Specifically, use the cross-camera pedestrian contrastive loss to calculate the clustering results and backpropagate to update the parameters of the first updated backbone network, and at the same time update the parameters of the second updated momentum network. The cross-camera contrastive loss calculation uses decoupled contrastive learning loss, which is used to calculate the loss value between the feature of each training sample and the average feature of pedestrians under a single camera. Its calculation formula is as follows:

[0086]

[0087] In the above formula, N represents the total number of samples with pseudo-labels, f(x i ) represents the feature of sample x i , m j represents the average feature of pedestrians under a single camera with the same pseudo-label as sample x i , P represents the number of positive sample pedestrians from different cameras included in the cluster to which sample x i belongs, N neg represents the sampling number of negative sample pedestrians, m k represents the average feature of negative sample pedestrians under a single camera that belongs to other clusters and is relatively close to sample x i , τ represents a regulation factor used to adjust the compactness between the sample and the average feature of pedestrians under a single camera, [·] + means taking the value of 0 when the value is negative;

[0088] The update formula for the momentum network is as follows:

[0089]

[0090] In the above formula, represents the parameters of the first updated backbone network, represents the parameters of the first updated momentum network, represents the parameters of the second updated momentum network, and μ is the update rate.

[0091] For the contrast loss calculation of the clustering results across cameras, each batch of training data is composed of randomly selecting P pseudo-labels from all the pseudo-labels first, and then randomly selecting K images from the images assigned with the selected pseudo-labels, with a total of P×K images.

[0092] S7. Update the pedestrian image average feature set based on the second updated momentum network to obtain the updated pedestrian image average feature set;

[0093] S71. Extract features from the pedestrian image dataset with single-camera labels based on the second updated momentum network and take the average to obtain the extracted pedestrian image average feature set;

[0094] S72. Update the pedestrian image average feature set based on the extracted pedestrian image average feature set to obtain the updated pedestrian image average feature set.

[0095] S8. Return the hierarchical clustering algorithm steps and the contrast loss calculation steps across cameras to perform multiple iterations of optimization on the pedestrian re-identification model until the number of iterations reaches the preset value, and output the optimized pedestrian re-identification model;

[0096] S81. Judge the number of iterations of the pedestrian re-identification model for optimization;

[0097] S82. If it is judged that the number of iterations of the pedestrian re-identification model for optimization has not reached the preset value, repeat the hierarchical clustering algorithm and the contrast loss calculation across cameras for the iterative optimization of the pedestrian re-identification model;

[0098] S83. If it is judged that the number of iterations of the pedestrian re-identification model for optimization has reached the preset value, output the optimized pedestrian re-identification model.

[0099] Specifically, update the pedestrian image average feature set based on the extracted pedestrian image features, extract the updated pedestrian average features through the second updated momentum network, repeat the hierarchical clustering and cross-camera pedestrian contrast learning until the number of iterations reaches the maximum number of iterations, and output the optimized pedestrian re-identification model.

[0100] S9. Perform cross-camera pedestrian re-identification applications according to the optimized pedestrian re-identification model.

[0101] Specifically, the person re-identification model first inputs a dataset of person images with labels under a single camera but without global cross-camera labels. The person labels between the average feature sets of person images are compared through a contrast loss function to calculate the contrast loss value between the average feature sets of person images under a single camera. Further, the loss value between the average feature sets of person images under cross-cameras is calculated, the average feature sets of person images are updated, and finally, the average feature set of person images is output through continuous iterative optimization of the dataset and applied to person re-identification.

[0102] Referring to Figure 2 , a person re-identification system based on a momentum network and contrast learning includes:

[0103] A data annotation processing module for performing annotation processing on a dataset of person images to obtain a dataset of person images with labels under a single camera;

[0104] A data feature extraction module for extracting the features of person images under different cameras to obtain an average feature set of person images;

[0105] A first contrast loss calculation module for calculating the loss value between the average feature sets of persons under each camera;

[0106] A hierarchical clustering module for performing hierarchical clustering processing on the average feature set of person images to obtain a clustering result;

[0107] A second contrast loss calculation module for calculating the loss value between the features of the same person image under different cameras in the clustering result;

[0108] A data feature update module for updating the average feature set of person images to obtain an updated average feature set of person images;

[0109] A model optimization module for performing iterative optimization processing on the person re-identification model to obtain an optimized person re-identification model.

[0110] The content in the above method embodiments is applicable to the system embodiments of the present invention. The functions specifically implemented by the system embodiments of the present invention are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those of the above method embodiments.

[0111] The above is a specific description of the preferred embodiments of the present invention. However, the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A pedestrian re-identification method based on a momentum network and contrastive learning, characterized in that It includes the following steps: Obtain pedestrian pictures and perform annotation processing to obtain a pedestrian picture dataset with labels under a single camera; Initialize the momentum network and the backbone network in the pedestrian re-identification model respectively based on the use of a pre-trained deep convolutional neural network model to obtain the initialized momentum network and the initialized backbone network; Perform feature extraction processing on the pedestrian picture dataset with labels under a single camera based on the initialized momentum network to obtain a set of average features of pedestrian pictures; Calculate the contrast loss under a single camera for the set of average features of pedestrian pictures based on the initialized backbone network and update the initialized backbone network and the initialized momentum network in sequence to obtain the first updated backbone network and the first updated momentum network; Perform hierarchical clustering processing on the set of average features of pedestrian pictures based on the first updated momentum network and assign pseudo-labels to obtain a clustering result; Calculate the contrast loss across cameras for the clustering result based on the first updated backbone network and perform backpropagation, and update the first updated backbone network and the first updated momentum network in sequence to obtain the second updated backbone network and the second updated momentum network; Update the set of average features of pedestrian pictures based on the second updated momentum network to obtain an updated set of average features of pedestrian pictures; Loop through the steps of the hierarchical clustering algorithm and the steps of calculating the contrast loss across cameras to perform multiple iterative optimizations on the pedestrian re-identification model until the number of iterations reaches a preset value, and output the optimized pedestrian re-identification model; Perform cross-camera pedestrian re-identification applications based on the optimized pedestrian re-identification model; The contrast loss calculation under a single camera uses contrast loss calculation, and the formula is as follows: In the above formula, L represents the first loss value, c represents the index value of the camera, and N c represents the total number of pedestrian labels under camera c, D c represents the average feature set of pedestrian images under camera c, j and k represent the index values ​​of the average feature set of pedestrian images under a single camera, and f(x i ) represents the sample x i Features, Represents the sample x under camera c i The corresponding average feature set of all pedestrian images with label j, τ represents the adjustment factor; The contrast loss calculation across cameras uses decoupled contrast loss calculation, and the formula is as follows: In the above formula, L represents the second loss value, N represents the total number of samples with pseudo-labels, f(x i ) represents the feature of sample x i , m j represents the average feature set of pedestrian pictures under a single camera with the same pseudo-label as sample x i , P represents the number of positive sample pedestrians from different cameras included in the cluster to which sample x i belongs, N neg represents the sampling number of negative sample pedestrians, m k represents the average feature set of pedestrian pictures under a single camera of negative samples from other clusters that are relatively close to sample x i , τ represents the adjustment factor, [·] + represents taking the value of 0 when the value is negative.

2. The pedestrian re-identification method based on a momentum network and contrastive learning according to claim 1, wherein, The step of obtaining pedestrian pictures and performing annotation processing to obtain a pedestrian picture dataset with labels under a single camera specifically includes: Obtain pedestrian pictures based on multiple cameras and construct a pedestrian picture dataset; Annotate the pedestrian pictures under each camera based on the pedestrian picture dataset to obtain a pedestrian picture dataset with labels under a single camera.

3. The pedestrian re-identification method based on a momentum network and contrastive learning according to claim 2, wherein The step of performing feature extraction processing on the pedestrian picture dataset with labels under a single camera based on the initialized momentum network to obtain a set of average features of pedestrian pictures specifically includes: Perform feature extraction processing on the pedestrian picture dataset with labels under a single camera based on the initialized momentum network to obtain pedestrian picture features; Take the average of the pedestrian picture features of the same pedestrian under the same camera and integrate them to obtain a set of average features of pedestrian pictures.

4. The pedestrian re-identification method based on a momentum network and contrastive learning according to claim 3, characterized in that, The step of performing hierarchical clustering processing on the set of average features of pedestrian pictures based on the first updated momentum network and assigning pseudo-labels to obtain a clustering result specifically includes: Calculate the feature fusion weights of the set of average features of pedestrians under a single camera based on the updated momentum network to obtain feature fusion weight values; Process the set of average features of pedestrians according to the feature fusion weight values to obtain a fused set of average features of pedestrians. Calculate the Jaccard distances between the features of different pedestrian images under different cameras in the fused pedestrian average feature set, and construct a distance set based on the fused pedestrian average feature set; Perform traversal and sorting on the distance set to obtain the clustering results; Assign the same pseudo-label to the features of the same pedestrian image under different cameras in the clustering results.

5. The pedestrian re-identification method based on momentum network and contrastive learning according to claim 4, wherein, The step of updating the pedestrian image average feature set based on the second updated momentum network to obtain the updated pedestrian image average feature set specifically includes: Extract and average the features of the pedestrian image dataset with labels under a single camera based on the second updated momentum network to obtain the extracted pedestrian image average feature set; Update the pedestrian image average feature set based on the extracted pedestrian image average feature set to obtain the updated pedestrian image average feature set.

6. The pedestrian re-identification method based on a momentum network and contrastive learning according to claim 5, wherein The step of iteratively optimizing the person re-identification model using the loop hierarchical clustering algorithm steps and the cross-camera contrast loss calculation steps until the number of iterations reaches a preset value and outputting the optimized person re-identification model specifically includes: Judge the number of iterative optimizations of the person re-identification model; When it is judged that the number of iterative optimizations of the person re-identification model has not reached the preset value, repeat the hierarchical clustering algorithm and the cross-camera contrast loss calculation for the iterative optimization of the person re-identification model; When it is judged that the number of iterative optimizations of the person re-identification model reaches the preset value, output the optimized person re-identification model.

7. A pedestrian re-identification system based on momentum network and contrastive learning, characterized in that, A method for performing person re-identification based on a momentum network and contrast learning as described in claim 1, comprising the following modules: A data annotation processing module for annotating the pedestrian image dataset to obtain a pedestrian image dataset with labels under a single camera; A data feature extraction module for extracting the features of pedestrian images under different cameras to obtain a pedestrian image average feature set; A first contrast loss calculation module for calculating the loss value between the pedestrian average feature sets under each camera; A hierarchical clustering module for performing hierarchical clustering on the pedestrian image average feature set to obtain clustering results; A second contrast loss calculation module for calculating the loss value between the features of the same pedestrian image under different cameras in the clustering results; A data feature update module for updating the pedestrian image average feature set to obtain the updated pedestrian image average feature set; A model optimization module for iteratively optimizing the person re-identification model to obtain the optimized person re-identification model.

Citation Information

Patent Citations

  • Cross-domain pedestrian re-identification algorithm based on momentum network guidance

    CN113326731A

  • Unsupervised domain adaptive pedestrian re-recognition algorithm based on pseudo label optimization

    CN113378632A