An Unsupervised Pedestrian Re-Identification Method Based on Cross-Camera Channel Information Interaction and Spatiotemporal Semantic Perception
Through the cross-camera joint learning framework and the multi-scale channel interactive attention module, combined with the space-time label punishment mechanism, the pedestrian re-identification accuracy problem caused by cross-camera differences is solved, and the robustness and stability of the model are improved.
Patent Information
- Application Number
- CN202510190011.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-02-20
AI Technical Summary
The existing unsupervised pedestrian re-identification method has poor accuracy when dealing with cross-camera differences, ignoring the viewing angle, lighting and background changes between cameras, resulting in performance degradation.
A cross-camera joint learning framework is adopted, combining local and global features, and through multi-scale channel interactive attention module and spatiotemporal label punishment mechanism, pseudo-label generation and feature learning are optimized to enhance the robustness of the model.
It improves the accuracy of pedestrian re-identification in complex scenarios, improves the robustness and stability of the model, enhances the distinction between pedestrian identity characteristics and suppresses background interference.
Smart Images

Figure CN120032392B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of pedestrian re-identification, and specifically relates to an unsupervised pedestrian re-identification method based on cross-camera channel information interaction and spatio-temporal semantic perception. Background Art
[0002] As an indispensable part of the social public security management system, the intelligent video surveillance system plays an indispensable role in maintaining social security and stability. Due to the exponential growth of the data volume generated by surveillance videos, the traditional method relying on manual monitoring cannot meet the real-time processing requirements of massive video data. Therefore, intelligent analysis technologies have emerged and are widely used in fields such as public security, traffic management, and retail analysis. With the continuous growth of the data volume, how to efficiently store these huge amounts of data and quickly retrieve the required information has become an urgent problem to be solved.
[0003] The purpose of pedestrian re-identification is to find images of people with a given identity in a large-scale image library and determine whether pedestrians in different perspectives, different cameras, and different video segments are the same person. The acquisition of labeled data incurs a large amount of labor and computational costs. Therefore, supervised pedestrian re-identification cannot be extended to large-scale data sets. To solve the scalability problem of supervised pedestrian re-identification, in recent years, researchers have begun to focus on unsupervised pedestrian re-identification. Unsupervised pedestrian re-identification aims to achieve accurate re-identification by using the latent structure and patterns between data without the need for labeled data.
[0004] Existing unsupervised pedestrian re-identification methods directly perform a clustering algorithm on the target features and assign pseudo-labels to the images. Then, based on the clustering, the cross-camera feature similarity is measured, and the images are roughly divided into classes. However, they ignore the feature distribution differences caused by the camera domain gap, resulting in an inevitable performance decline.
[0005] In pedestrian re-identification, the image contains not only the features of the pedestrian but also background, illumination, and noise information. Using the attention mechanism in unsupervised pedestrian re-identification can enhance the extraction of identity-related features while suppressing background interference. Luo et al. designed a novel cross-channel attention module that focuses on the max-pooling features and average-pooling features and generates channel weights through a convolutional layer. The feature enhancement module in PGAL improves the feature representation of the non-key point regions of the human body through multi-head self-attention. However, these methods only focus on the single-scale channel features or spatial features of the pedestrian image separately and do not pay attention to these features simultaneously.
[0006] The confidence-based sample filtering method uses the similarity between a sample and the center of its pseudo-label as the confidence score, and eliminates samples with low confidence or assigns them lower weights. Miao et al. use confidence to guide pseudo-labels, which enables samples to approach not only the initially assigned centroids but also other centroids that may embed their identity information. Wang et al. propose self-consistency constraints to soften label noise in a single model, reducing the computational cost and the cost of network parameters. Modeling the relationships between samples can better capture the similarities between samples of the same class and the discriminability between samples of different classes, thereby optimizing the pseudo-label generation and feature learning processes. These methods have alleviated the uncertainty problem of pseudo-labels to a certain extent, but they often ignore the potential constraint of time information on the quality of pseudo-labels.
[0007] Therefore, in the field of person re-identification, although deep learning-based research has made significant progress in improving the accuracy of models, current research often ignores the camera differences caused by perspective, illumination, and background changes in practical applications, increasing the difficulty of person re-identification, and the accuracy of existing cross-camera unsupervised person re-identification models is relatively poor. Therefore, there is an urgent need for an unsupervised person re-identification method focusing on cross-camera to solve the above problems. Summary of the Invention
[0008] The purpose of the present invention is to provide an unsupervised person re-identification method based on cross-camera channel information interaction and spatio-temporal semantic perception, including the following steps:
[0009] S1: Obtain an unsupervised person re-identification dataset and input it into a cross-camera joint learning framework;
[0010] S2: In the cross-camera joint learning framework, simultaneously extract local features within the camera and global features across cameras, generate a feature representation that is consistent inside and outside the camera, combine local and global features, capture the dependency relationship of the cross-camera feature distribution, and generate pseudo-labels;
[0011] S3: In the multi-scale channel interaction attention module, optimize the feature expression by modeling the relationships between channels, focus on the discriminability of identity-related features, and suppress background interference;
[0012] The multi-scale channel interaction attention module includes channel grouping, channel shuffling, multi-scale channel interaction branches, and multi-scale spatial branches; the multi-scale channel interaction branches use multi-scale convolutions for channel information interaction between multi-scale branches, and each channel combines information from other channels; the multi-scale spatial branches use multi-scale convolutions combined with spatial attention mechanisms to effectively integrate spatial information with different receptive fields, enhancing the model's perception ability of local key regions and global feature distributions of pedestrians; finally, channel shuffling is used to model and interact the relationships between all channels;
[0013] S4: Use the spatio-temporal label penalty mechanism to endow samples with a more detailed label distribution by combining time information;
[0014] The spatio-temporal label penalty mechanism includes in-camera time constraints and inter-camera time and number constraints, and generates the final matching confidence by combining the two constraints; Use the in-camera time constraint, for pedestrians belonging to the same camera, adjust their matching confidence according to the proximity of the pseudo-timestamps; Use the inter-camera time and number constraints, for pedestrians from different cameras, adjust their matching confidence according to the time interval and camera number information; Integrate the in-camera and inter-camera constraints to generate the final matching confidence, and incorporate the matching confidence into the pseudo-label generation process to optimize the label distribution;
[0015] S5: Iteratively adjust the network parameters by designing a loss function to optimize the model performance;
[0016] S6: In each training stage, compare the similarity between the queried pedestrian images and the pedestrian images in the image library to find images of the same pedestrian.
[0017] Furthermore, the overall framework includes a cross-camera joint learning framework; a backbone network and a multi-scale channel interaction attention module; a spatio-temporal label penalty mechanism.
[0018] Furthermore, the cross-camera joint learning framework includes an in-camera local feature extraction branch and an inter-camera global feature extraction branch; The inter-camera global feature extraction branch extracts features from the input images and segmented images of all cameras to generate global features; The in-camera local feature extraction branch extracts features from the pedestrian images in each camera; Cluster the samples according to the extracted features to generate pseudo-labels.
[0019] Furthermore, the multi-scale channel interaction branch encodes the channel features respectively through multi-scale convolution operations of 1x1 convolution and 3x3 convolution and a channel interaction module to obtain the relationship between channels.
[0020] Furthermore, the channel interaction module includes segmentation, average pooling, GN, Sigmoid, matrix multiplication, concatenation, and element-wise addition.
[0021] Furthermore, the multi-scale spatial branch enhances the model's perception ability of the pedestrian's local key regions and global feature distribution respectively through multi-scale convolution operations of 1x1 convolution and 3x3 convolution and a spatial attention module, including GN, ReLU, Sigmoid, and element-wise addition.
[0022] Furthermore, the loss function is calculated by using cross-entropy loss and triplet loss, and the obtained results are used for the training of pseudo-label generation.
[0023] Compared with the prior art, the present invention has the following beneficial effects:
[0024] The present invention proposes a novel cross-camera joint learning framework, which can simultaneously combine local features within the camera and global features between cameras to capture the global dependencies of the cross-camera feature distribution; the present invention proposes a novel multi-scale channel interaction attention module, which can model the relationships between channels, optimize the feature representation, and improve the robustness of the model in complex scenarios; the present invention proposes a novel spatio-temporal label penalty mechanism, which improves the stability and accuracy of the training process by refining the label distribution and semantic relationship modeling; the present invention verifies the effectiveness and superiority of the method of the present invention on two unsupervised person re-identification datasets, Market1501 and MSMT17. Comprehensive evaluation metrics are used to evaluate the model accuracy: including mAP, Rank-1, Rank-5, and Rank-10. The experimental results under the four metrics fully prove the effectiveness of the method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings:
[0026] Figure 1 It is a flowchart of the steps of an unsupervised person re-identification method based on cross-camera channel information interaction and spatio-temporal semantic perception provided by the present invention;
[0027] Figure 2 It is an overall framework diagram of an unsupervised person re-identification method based on cross-camera channel information interaction and spatio-temporal semantic perception provided by the present invention;
[0028] Figure 3 It is a schematic structural diagram of a cross-camera joint learning framework of a preferred embodiment provided by the present invention;
[0029] Figure 4 It is a schematic structural diagram of a multi-scale channel interaction attention module of a preferred embodiment provided by the invention;
[0030] Figure 5 It is a flowchart of a spatio-temporal label penalty mechanism of a preferred embodiment provided by the invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] In order to make those skilled in the art understand the present invention more clearly, the following will be described in detail with specific examples. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0032] Such as Figure 1As shown in the figure, it is a step flowchart of an unsupervised pedestrian re-identification method based on cross-camera channel information interaction and spatio-temporal semantic perception provided by the present invention, including:
[0033] S1: Obtain an unsupervised pedestrian re-identification dataset and input it into the cross-camera joint learning framework;
[0034] S2: In the cross-camera joint learning framework, simultaneously extract local features within the camera and global features across cameras to generate a feature representation that is consistent both inside and outside the camera. Combine local and global features to capture the dependency relationship of the cross-camera feature distribution and generate pseudo-labels;
[0035] S3: In the multi-scale channel interaction attention module, optimize the feature expression by modeling the relationship between channels, focus on the discriminability of identity-related features, and suppress background interference;
[0036] The multi-scale channel interaction attention module includes channel grouping, channel shuffling, multi-scale channel interaction branches, and multi-scale spatial branches; the multi-scale channel interaction branches use multi-scale convolutions for channel information interaction between multi-scale branches, and each channel combines information from other channels; the multi-scale spatial branches use multi-scale convolutions combined with spatial attention mechanisms to effectively integrate spatial information with different receptive fields and enhance the model's perception ability of pedestrian local key regions and global feature distributions; finally, use channel shuffling to model and interact the relationships between all channels;
[0037] S4: Use the spatio-temporal label penalty mechanism to endow samples with a more detailed label distribution by combining time information;
[0038] The spatio-temporal label penalty mechanism includes intra-camera time constraints and inter-camera time and number constraints, and generates the final matching confidence by combining the two constraints; use intra-camera time constraints, for pedestrians belonging to the same camera, adjust their matching confidence according to the proximity of the pseudo-timestamps; use inter-camera time and number constraints, for pedestrians from different cameras, adjust their matching confidence according to the time interval and camera number information; integrate intra-camera and inter-camera constraints to generate the final matching confidence, and incorporate the matching confidence into the pseudo-label generation process to optimize the label distribution;
[0039] S5: Iteratively adjust the network parameters by designing a loss function to optimize the model performance;
[0040] S6: In each training stage, compare the similarity between the queried pedestrian image and the pedestrian images in the image gallery to find images of the same pedestrian.
[0041] Such as Figure 2As shown in the figure, it is the overall framework diagram of an unsupervised pedestrian re-identification method based on cross-camera channel information interaction and spatio-temporal semantic perception provided by the present invention. It mainly includes three modules, namely, a cross-camera joint learning framework, including an in-camera local feature extraction branch and a cross-camera global feature extraction branch, and a multi-scale channel interaction attention module, including a backbone network, a multi-scale channel interaction attention module, and a spatio-temporal label penalty mechanism.
[0042] The present invention provides a preferred embodiment to execute S2, simultaneously extracting local features within the camera and global features across cameras, generating a feature representation that is consistent inside and outside the camera, combining local and global features, capturing the dependency relationship of the cross-camera feature distribution, and generating pseudo-labels. As Figure 3 shown, it includes the following steps:
[0043] S2-1: Based on the feature map, it is segmented into multiple local regions and vertically divided into n blocks. Applying multi-scale channel interaction attention to each region separately can be expressed as:
[0044]
[0045] where F gi represents the i-th region after segmentation.
[0046] Generate optimized local features, and this process can be expressed as:
[0047]
[0048] Pool and splice the local features to generate an overall local feature representation:
[0049]
[0050] S2-2: Pass all input images X through the backbone network ResNet to extract the feature map F g , which is expressed as follows:
[0051] F g = ResNet(X),
[0052] where C is the number of channels, and H, W are the spatial dimensions.
[0053] Apply multi-scale channel interaction attention to the feature map to model the inter-channel and spatial relationships and generate the optimized global feature F' g , and the formula is as follows:
[0054] F' g = MSCI(F g ),
[0055] Among them, MSCI represents multi-scale channel interaction attention.
[0056] S2-3: For the pedestrian image X in each camera i After passing through the backbone network ResNet, the feature map F is extracted ci , which is expressed as follows:
[0057]
[0058] Among them, C is the number of channels, H, W are the spatial dimensions, and i is the camera number.
[0059] Apply multi-scale channel interaction attention on the feature map to model the inter-channel and spatial relationships, and generate the optimized in-camera local feature F’ ci , and the formula is as follows:
[0060]
[0061] Among them, MSCI represents multi-scale channel interaction attention.
[0062] S2-4: Generate pseudo-labels by clustering the samples of each camera in the in-camera local feature extraction branch; generate pseudo-labels by clustering all samples of all cameras in the cross-camera global feature extraction branch.
[0063] The present invention provides a preferred embodiment to execute S3, which optimizes the feature expression by modeling the inter-channel relationship, improves the distinguishability of identity-related features and suppresses background interference. As Figure 4 shown, it includes the following steps:
[0064] S3-1: First, group each channel of the feature map; secondly, perform multi-scale channel interaction on the sub-features through two branches.
[0065] In the multi-scale channel interaction branch, 1x1 convolution and 3x3 convolution are respectively used for multi-scale operations. In the 1x1 convolution, one-dimensional global average pooling operation is used to encode the channels along two spatial directions of width and height respectively, and the spatial distribution features of each channel are obtained for channel interaction. At the same time, combined with 3x3 convolution, through the interaction between multi-scale branches, each channel can combine information from other channels. The formula is as follows:
[0066] f1 = S(GN(Conv 1×1 (Concat(HAP(F),WAP(F))))),
[0067] f3 = Conv 3×3 (F),
[0068] F1 = mat(AP(f1), Softmax(f3)),
[0069] F2 = mat(AP(f3), Softmax(f1)),
[0070] F’ c = F1 + F2,
[0071] F C = F · F’, c .
[0072] S3-3: In the multi-scale spatial branch, the multi-scale convolution is combined with the spatial attention mechanism to effectively integrate the information of different receptive fields and enhance the model's perception ability of the local key regions and global feature distributions of pedestrians. The formula is as follows:
[0073] f s3 = SA(Conv 3×3 (F)),
[0074] f s = SA(F),
[0075] F S = F · (f s3 + f s ),
[0076] S3-4: Finally, perform a channel shuffle operation on the obtained feature map as follows:
[0077] F' = Shuffle(Concat(F C + F S ))
[0078] The present invention provides an embodiment to execute S4, which assigns a more detailed label distribution to the samples by combining time information. As Figure 5 shown, it includes the following steps:
[0079] S4-1: Camera internal time constraint: Assume that two images i, j belong to the same camera, that is, camID i = camID j , then according to the proximity of the pseudo time stamps t i and t j , adjust their matching confidence:
[0080]
[0081] where γ > 0 controls the speed of time decay. Image pairs with close times, that is, |t i - t j | is small, will obtain a higher matching confidence.
[0082] S4-2: Camera - to - Camera Time and ID Constraints: Assume two images i and j are from different cameras, i.e., camID i ≠camID j , and their matching confidence is affected by the time interval and the difference in camera IDs:
[0083]
[0084] Among them, δ > 0: the cross - camera time - interval decay coefficient. κ: the threshold of allowable camera - ID difference. When the difference in camera IDs is large, i.e., camID i -camID j > κ, directly set the matching confidence to zero to suppress possible false matches.
[0085] S4-3: Integrate the intra - camera and inter - camera constraints to generate the final matching confidence:
[0086]
[0087] Integrate the matching confidence S(i, j) into the pseudo - label generation process to optimize the label distribution. Combine the initial pseudo - labels and the confidence S(i, j), update the pseudo - labels, and decide whether to accept the pseudo - labels according to the confidence. It can be expressed as:
[0088]
[0089] Among them, the neighborhood set of image i, ΙΙ(·) is the indicator function.
[0090] The present invention provides an embodiment to execute S5. In this embodiment, a loss function including cross - entropy loss and triplet loss is used for calculation, and the results obtained are used for the training of pseudo - label generation, including the following steps:
[0091] S5-1: Construct the intra - camera loss function. This loss function is used to optimize the local features within the camera and is implemented using cross - entropy loss. It can be expressed as:
[0092]
[0093] Among them, p i is the pedestrian prediction score value of the local feature, y i is the pseudo - label value of the pedestrian, and n represents the number of samples.
[0094] S5-2: Construct the inter - camera loss function. This loss function is used to optimize the global features across cameras and is implemented using triplet loss. It can be expressed as:
[0095]
[0096] Among them, d is the distance function between features, and ɑ is the margin parameter of the triplet loss.
[0097] S5-3: Construct the penalty loss function. This loss function is used to reduce the uncertainty of pseudo-labels and is implemented using cross-entropy loss. It can be expressed as:
[0098]
[0099] Among them, p i is the pedestrian prediction score value of the global feature, is the pseudo-label value of the pedestrian, and n represents the number of samples.
[0100] S5-4: Construct the classification loss function. This loss function is used to optimize the final classification result and is implemented using cross-entropy loss. It can be expressed as:
[0101]
[0102] Among them, p i is the final pedestrian prediction score value, y i is the pseudo-label value of the pedestrian, and n represents the number of samples.
[0103] S5-5: Construct the joint loss function. This loss function is the sum of the above loss functions. It can be expressed as:
[0104] L Joint = L Intra + L Inter + L penalty + L Classifier .
[0105] In this embodiment, the proposed method is compared with other advanced methods in terms of performance on two unsupervised pedestrian re-identification datasets, Market1501 and MSMT17, and common evaluation metrics are selected: mAP, Rank-1, Rank-5, and Rank-10. The experimental results are shown in Table 1.
[0106] Table 1 Comparison of the proposed method on Market1501 and MSMT17 datasets
[0107]
[0108] As can be seen from Table 1, this embodiment leads the existing methods in multiple metrics on the two datasets, proving the effectiveness of the method in this embodiment.
[0109] The above is the specific description of the specific embodiments of the present invention. However, the present invention is not limited to the above specific embodiments. Those skilled in the art can make various deformations or modifications to the above embodiments without affecting the essence of the present invention, and all should be within the protection scope of the present invention.
Claims
1. An unsupervised pedestrian re-identification method based on cross-camera channel information interaction and spatio-temporal semantic perception, characterized in that It includes the following steps: S1: Obtain an unsupervised person re-identification dataset and input it into the constructed model; S2: In the cross-camera joint learning framework, simultaneously extract local features within the camera and global features across cameras to generate a feature representation that is consistent inside and outside the camera. Combine local and global features to capture the dependency relationship of the cross-camera feature distribution and generate pseudo-labels; S3: In the multi-scale channel interaction attention module, optimize the feature representation by modeling the relationship between channels, focus on the distinctiveness of identity-related features, and suppress background interference; The multi-scale channel interaction attention module includes channel grouping, channel shuffling, multi-scale channel interaction branches, and multi-scale spatial branches; The multi-scale channel interaction branches use multi-scale convolutions for channel information interaction between multi-scale branches, and each channel combines information from other channels; The multi-scale spatial branches use a combination of multi-scale convolutions and spatial attention mechanisms to effectively integrate spatial information with different receptive fields, enhance the model's perception ability of local key regions and global feature distributions of pedestrians; finally, use channel shuffling to model and interact the relationships between all channels; S4: Use a spatio-temporal label penalty mechanism to endow samples with a more detailed label distribution by combining time information; The spatio-temporal label penalty mechanism includes intra-camera time constraints and inter-camera time and number constraints, and generates the final matching confidence by combining the two constraints; use intra-camera time constraints, for pedestrians belonging to the same camera, adjust their matching confidence according to the proximity of the pseudo-timestamps; use inter-camera time and number constraints, for pedestrians from different cameras, adjust their matching confidence according to the time interval and camera number information; combine intra-camera and inter-camera constraints to generate the final matching confidence, and incorporate the matching confidence into the pseudo-label generation process to optimize the label distribution; S5: Iteratively adjust network parameters by designing a loss function to optimize the model performance; S6: In each training stage, compare the similarity between the queried pedestrian image and the pedestrian images in the image gallery to find images of the same pedestrian.
2. The unsupervised pedestrian re-identification method based on cross-camera channel information interaction and spatio-temporal semantic perception according to claim 1, wherein The overall structure includes a cross-camera joint learning framework; a backbone network and a multi-scale channel interaction attention module; a spatio-temporal label penalty mechanism.
3. The unsupervised pedestrian re-identification method based on cross-camera channel information interaction and spatio-temporal semantic perception according to claim 1, characterized in that, The cross-camera joint learning framework includes an intra-camera local feature extraction branch and an inter-camera global feature extraction branch; The inter-camera global feature extraction branch extracts features from the input images and segmented images of all cameras to generate global features; the intra-camera local feature extraction branch extracts features from the pedestrian images in each camera; cluster the samples according to the extracted features to generate pseudo-labels.
4. The unsupervised pedestrian re-identification method based on cross-camera channel information interaction and spatio-temporal semantic perception according to claim 1, characterized in that The multi-scale channel interaction branches encode the channel features through multi-scale convolution operations of 1x1 convolution and 3x3 convolution and a channel interaction module to obtain the relationship between channels.
5. The unsupervised pedestrian re-identification method based on cross-camera channel information interaction and spatio-temporal semantic perception according to claim 4, characterized in that The channel interaction module includes segmentation, average pooling, GN, Sigmoid, matrix multiplication, concatenation, and element-wise addition.
6. The unsupervised pedestrian re-identification method based on cross-camera channel information interaction and spatio-temporal semantic perception according to claim 1, characterized in that The multi-scale spatial branches enhance the model's perception ability of local key regions and global feature distributions of pedestrians through multi-scale convolution operations of 1x1 convolution and 3x3 convolution and spatial attention modules, including GN, ReLU, Sigmoid, and element-wise addition.
7. The unsupervised pedestrian re-identification method based on cross-camera channel information interaction and spatio-temporal semantic perception according to claim 1, characterized in that The loss function is calculated by using cross-entropy loss and triplet loss, and the results obtained are used for the training of pseudo-label generation.
Citation Information
Patent Citations
Unsupervised pedestrian re-identification method and system and computer readable medium
CN113158815A
Unsupervised pedestrian re-identification method based on pseudo label improvement of local features
CN117011776A