A Semi-Supervised Cross-Modal Person Re-Identification Method Based on Dual Self-Supervised Learning

By constructing a semi-supervised cross-modal pedestrian re-identification method with dual self-supervised learning, and using labelless data for self-supervised learning, the problem of insufficient label samples is solved, and the accuracy and robustness of cross-modal pedestrian re-identification is improved.

CN116052212BActive Publication Date: 2025-07-18HENAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310027835.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-09
Publication Date
2025-07-18
Estimated Expiration
2043-01-09

AI Technical Summary

Technical Problem

The existing cross-modal pedestrian re-identification method is limited by insufficient label training samples in practical applications, making it difficult to effectively utilize labelless data, and cross-modal image differences and in-modal differences lead to insufficient recognition accuracy.

Method used

The semi-supervised cross-modal pedestrian re-identification method based on dual self-supervised learning is adopted. By building a backbone network, a context-based rotating self-supervised network and a self-supervised network based on contrast learning, self-supervised learning is used to obtain a more comprehensive pedestrian feature representation.

Benefits of technology

It improves the accuracy and robustness of cross-modal pedestrian re-identification, realizes efficient recognition under unsupervised conditions, simplifies the training process, and improves the generalization ability and recognition accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116052212B_ABST
    Figure CN116052212B_ABST
Patent Text Reader

Abstract

The present invention discloses a semi-supervised cross-modal pedestrian re-identification method based on dual self-supervised learning, comprising the following steps: A: constructing a cross-modal pedestrian re-identification data set; B: performing data augmentation processing on pedestrian images in the cross-modal pedestrian re-identification data set; C: constructing a backbone network for semi-supervised cross-modal pedestrian re-identification based on dual self-supervised learning, a context-based rotation self-supervised network, and a contrastive learning-based self-supervised network; D: obtaining final pedestrian image features, a first probability matrix, and a second probability matrix through the constructed network model; E: performing a self-supervised pedestrian re-identification task using the obtained final pedestrian image features, first probability matrix, and second probability matrix, and outputting a final recognition result. The present invention can utilize a large amount of unlabeled data, learn the consistency information of images in different modalities, obtain a more comprehensive pedestrian feature representation, and thus more accurately achieve cross-modal pedestrian re-identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a pedestrian image recognition method, and in particular to a semi-supervised cross-modal pedestrian re-identification method based on dual self-supervised learning. Background Art

[0002] Person re-identification, also known as pedestrian re-identification, is a technology that uses computer vision technology to determine whether a specific pedestrian exists in an image or video sequence. It is widely considered to be a sub-problem of image retrieval, that is, given a monitored pedestrian image, the pedestrian image is retrieved across devices. In existing pedestrian re-identification research, the datasets used for training and testing are often single-modal RGB images. However, in real-world applications, pedestrian images captured and described by infrared mode cameras, depth cameras, and eyewitness statements are very common. Therefore, how to re-identify pedestrians across visible light and infrared modalities is one of the urgent problems to be solved. Cross-modal pedestrian re-identification mainly studies the problem of retrieving and matching images belonging to the same individual in the image library under the two modalities given a visible light image or infrared image of a specific individual.

[0003] At present, the cross-modal pedestrian re-identification problem mainly faces the following challenges:

[0004] (1) There are significant differences between the images captured in the two modes. RGB images have three channels, including the visible light color information of red, green and blue, while infrared images have only one channel, including the intensity information of near-infrared light. In addition, from the perspective of imaging principles, the wavelength ranges of the two are also different. Different clarity and lighting conditions will produce very different effects on the two types of images.

[0005] (2) The intra-modal differences that exist in traditional person re-ID, such as low resolution, occlusion, and viewpoint changes, also exist in cross-modal person re-ID.

[0006] In addition, although existing methods have made some progress in cross-modal person re-identification under extreme degradation conditions, there is still much room for improvement in performance. Most of the existing methods are trained based on supervised frameworks, and their performance depends heavily on a large number of labeled training samples. However, labeling enough training samples requires a lot of manpower and material resources. Therefore, the lack of labeled training data severely limits the practical application of supervised models. Summary of the invention

[0007] The object of the present invention is to provide a semi-supervised cross-modal pedestrian re-identification method based on dual self-supervised learning, which can more effectively utilize a large amount of unlabeled data, learn the consistency information of images in different modalities, obtain a more comprehensive pedestrian feature representation, and thus more accurately achieve cross-modal pedestrian re-identification.

[0008] The present invention adopts the following technical solutions:

[0009] A semi-supervised cross-modal pedestrian re-identification method based on dual self-supervised learning, comprising the following steps:

[0010] A: Construct a cross-modal pedestrian re-identification data set, and preprocess the pedestrian images in the cross-modal pedestrian re-identification data set to obtain input images for supervised training;

[0011] B: Perform data augmentation on the pedestrian images in the cross-modal pedestrian re-identification data set to obtain the pedestrian images after data augmentation, and the pedestrian images after data augmentation include context-based rotation self-supervised images and contrast-based self-supervised images;

[0012] C: Construct a backbone network and a self-supervised training network for semi-supervised cross-modal pedestrian re-identification based on dual self-supervised learning; wherein, the self-supervised training network includes a context-based rotation self-supervised network and a contrast learning-based self-supervised network; the backbone network, the context-based rotation self-supervised network and the contrast learning-based self-supervised network are arranged in parallel and share network weights;

[0013] Among them, the backbone network is used to perform supervised learning on the input images for supervised training to obtain the final pedestrian image features; the context-based rotation self-supervised network is used to perform self-supervised learning on the context-based rotation self-supervised images to obtain a first probability matrix for rotation angle prediction; the contrast learning-based self-supervised network is used to perform self-supervised learning on the contrast-based self-supervised images to obtain a second probability matrix for contrast self-supervised learning;

[0014] D: Use the pedestrian images after enhancement processing in step B to construct a training set, and the training set includes labeled samples and unlabeled samples. Use the labeled samples to learn the final pedestrian image features for supervised training through the backbone network, and use the unlabeled samples to pass through the context-based rotation self-supervised network and the contrast learning-based self-supervised network respectively to obtain a first probability matrix for rotation angle prediction and a second probability matrix for contrast self-supervised learning;

[0015] E: Using the final pedestrian image features obtained in step D, the first probability matrix for rotation angle prediction, and the second probability matrix for contrastive self-supervised learning, perform a self-supervised pedestrian re-identification task through the backbone network and self-supervised training network of semi-supervised cross-modal pedestrian re-identification based on dual self-supervised learning, and output the final recognition result.

[0016] The said step A includes the following specific steps:

[0017] A1: Construct a cross-modal pedestrian re-identification dataset, obtain the pedestrian images in the training set of the cross-modal pedestrian re-identification dataset, and set the total number of images input to the network model of the semi-supervised cross-modal pedestrian re-identification method based on dual self-supervised learning.

[0018] A2: Adjust the size of the pedestrian images in the cross-modal pedestrian re-identification dataset, and adjust both the width and height of the pedestrian images to the same size.

[0019] A3: Randomly horizontally flip the pedestrian images after size adjustment in step A2.

[0020] A4: Perform pixel padding on the pedestrian images after random horizontal flipping in step A3.

[0021] A5: Randomly crop the padded pedestrian images in step A4.

[0022] A6: Normalize the randomly cropped pedestrian images in step A5.

[0023] A7: Perform channel random erasing on the normalized pedestrian images in step A6 to obtain the input images for supervised training.

[0024] The said step B includes the following specific steps:

[0025] B1: For each of the size-adjusted pedestrian images in turn, randomly select an angle from the set of rotation angles {0, 90, 180, 270} for rotation, and generate a pseudo-label for each rotated pedestrian image respectively to obtain context-based rotation self-supervised images.

[0026] B2: Perform channel random erasing on the obtained input images for supervised training.

[0027] B3: Use channel swapping on the pedestrian images after completing channel random erasing to obtain contrast-based self-supervised images.

[0028] In step C described above, the backbone network sequentially includes a first convolutional layer, a first pooling layer, first to third residual layers, a first modality attention layer, a fourth residual layer, a second modality attention layer, and a part alignment attention layer; the first convolutional layer and the first to fourth residual layers perform feature extraction on the downsampled pedestrian image features layer by layer to learn the shallow features of the pedestrian image; the first modality attention layer is used to learn the deep features of the pedestrian image in two modalities; the part alignment attention layer is used to discover the small differences between the visible light modality and the infrared modality to obtain the final pedestrian image features.

[0029] The first modality attention layer and the second modality attention layer have the same structure, both consisting of two second convolutional layers with a kernel size of 1, a ReLU activation function, and a Sigmod activation function; the calculation formulas of the first modality attention layer and the second modality attention layer are as follows:

[0030]

[0031] where Z represents the depth features of the obtained pedestrian image, represents the matrix after Z is instance-normalized, and m C is the channel mask, representing the channels related to identity, and the calculation formula of m C is m C = σ(W2δ(W1g(Z))); g(·) represents the global average pooling layer, δ(·) represents the ReLU activation function, σ(·) represents the Sigmod activation function, and W1 and W2 respectively represent the two fully connected layers in the modality attention layer, and the two fully connected layers are located after the ReLU activation function and the Sigmod activation function.

[0032] The context-based rotation self-supervised network sequentially includes a third convolutional layer, a second pooling layer, a third modality attention layer, a fifth residual layer, a fourth modality attention layer, a global average pooling layer, a BN layer, and a first fully connected layer. The third modality attention layer and the fourth modality attention layer have the same structure as the first modality attention layer.

[0033] The structure of the contrastive learning-based self-supervised network is to add a second fully connected layer on the basis of the backbone network, and the second fully connected layer is located after the part alignment attention layer.

[0034] In step E described above, for the final pedestrian image features of the input image in supervised training, the weights of the backbone network are updated by backpropagation using the set first cross-entropy loss function and center loss function;

[0035] Cross-entropy loss The calculation formula is:

[0036]

[0037] Among them, n and m respectively represent the numbers of visible light modality and infrared modality images in the current batch, f v , f r respectively represent the pedestrian image features of the visible light modality and the pedestrian image features of the infrared modality, and respectively represent f v , f r corresponding image labels, C(f v ) and C(f r ) respectively represent the probability matrices obtained by passing the pedestrian image features of the two modalities through two classifiers with parameters θ, and P(·) is the softmax function;

[0038] Center loss function The calculation formula is:

[0039]

[0040] Among them, f i represents the pedestrian image features, represents the mean of the features with the current batch label y i , represents the mean of the features with the current batch label y k , represents the mean of the features with the current batch label y j , T is the number of pedestrians in the current batch, and ρ refers to the minimum distance between all centers.

[0041] In the described step E, the context-based rotation self-supervised network judges the rotation angle through the second cross-entropy loss function, and finally the backbone network outputs the average precision result of pedestrian detection, and the average precision result is used to evaluate the accuracy of pedestrian re-identification;

[0042] Second cross-entropy loss function The calculation formula is:

[0043]

[0044] Among them, represents the image after random rotation, is the label generated by the random rotation angle of the image, and R represents the total number of image samples in a batch.

[0045] In the described step E, the second probability matrix output by the contrast learning-based self-supervised network uses the KL divergence as the consistency constraint loss function,

[0046] Among them, p(x i) represents the probability matrix obtained by the color image classifier in supervised learning, q(x i ) represents the probability matrix obtained after the contrast-based self-supervised features pass through the classifier.

[0047] The present invention adopts a semi-supervised cross-modal person re-identification method based on dual self-supervised learning, effectively utilizes a large amount of unlabeled data, obtains more comprehensive person image features therefrom, and enhances the feature extraction ability and generalization ability of the backbone network based on the context rotation self-supervised network and the contrast learning-based self-supervised network, so as to more accurately realize cross-modal person re-identification. Secondly, the semi-supervised algorithm achieves the state-of-the-art prediction performance of supervised and unsupervised image classification without introducing additional hyperparameters for optimization. At the same time, the semi-supervised algorithm does not require a separate pre-training step, but is trained end-to-end in parallel to achieve simplicity, efficiency and practicality. Brief Description of the Drawings

[0048] Figure 1 is a schematic flow chart of the present invention. Detailed Embodiment

[0049] The following describes the present invention in detail with reference to the drawings and embodiments:

[0050] As Figure 1 shown, the semi-supervised cross-modal person re-identification method based on dual self-supervised learning according to the present invention includes the following steps:

[0051] A: Construct a cross-modal person re-identification dataset, and preprocess the person images in the cross-modal person re-identification dataset to obtain the input images for supervised training;

[0052] In the present invention, the cross-modal person re-identification dataset includes the datasets SYSU and RegDB, both of which are publicly available person re-identification datasets. The dataset SYSU is a large-scale dataset collected by four visible light cameras and two near-infrared cameras, including indoor and outdoor environments. The training set in the dataset SYSU contains 22,258 visible images and 11,909 infrared images, involving 395 identities, while the query set and the gallery set contain 3,803 infrared images and 3,010 randomly sampled visible images. The dataset RegDB consists of a pair of aligned cameras (one visible camera and one thermal camera). The dataset RegDB contains 8,240 images of 412 identities, with 10 images from the visible camera and 10 images from the thermal imaging camera in each image.

[0053] In the present invention, step A includes the following specific steps:

[0054] A1: Construct a cross-modal pedestrian re-identification dataset, obtain pedestrian images in the training set of the cross-modal pedestrian re-identification dataset, and set the total number of images input to the network model of the semi-supervised cross-modal pedestrian re-identification method based on dual self-supervised learning to 2*p*k; where p is the number of pedestrians input to the network model of the semi-supervised cross-modal pedestrian re-identification method based on dual self-supervised learning in each batch, and k is the number of randomly sampled images in each single modality of each pedestrian.

[0055] In this embodiment, the Python programming language can be used to read the pedestrian images in the training set of the cross-modal pedestrian re-identification dataset into the memory. The hardware device of the experimental environment of the present invention has a CPU of Intel(R) Core(TM) i9-10900K CPU@3.70GHz, a memory size of 32GB, a GPU model of NVIDIA Geforce RTX3090, a Python version of 3.83 for the software platform, a CUDA version of 11.1, and a deep learning framework with a PyTorch version of 17.0 is used to build the model structure. Set the number of pedestrian images input to the network model of the semi-supervised cross-modal pedestrian re-identification method based on dual self-supervised learning in each batch to p, and randomly select k images from the pedestrian images in each batch, that is, the total number of images input to the network model at one time is 2*p*k.

[0056] A2: Resize the pedestrian images in the cross-modal pedestrian re-identification dataset, and adjust both the width and height of the pedestrian images to 224 pixels.

[0057] Since in the rotation self-supervised module, if the height and width of the pedestrian are still set to 256 pixels and 128 pixels respectively in the traditional way, the shape of this similar rectangular pedestrian image will change after rotation and will be easily recognized by the model. In this embodiment, both the width and height of all input pedestrian images are set to 224 pixels, and the aspect ratio of the image is 1:1. After rotation, the external features of the pedestrian image hardly change, which can effectively increase the difficulty of the training task, prompt the network model to pay more attention to extracting the pedestrian detail features in the pedestrian image, and thus increase the generalization ability and robustness of the model.

[0058] In addition, the training of the backbone network of the semi-supervised cross-modal pedestrian re-identification based on dual self-supervised learning also uses pedestrian images with both width and height set to 224 pixels, which can enable the shallow features of the pedestrian image learned by the rotation self-supervised module from the 224*224-sized pedestrian image to effectively act on the supervised training of the backbone network, thereby prompting the backbone network to converge faster and improving the training accuracy.

[0059] A3: Randomly horizontally flip the pedestrian image after size adjustment in step A2 to enhance the model's generalization ability and alleviate overfitting.

[0060] A4: Pad the pedestrian image after random horizontal flipping in step A3 with 10 pixels using the torchvision.transforms.pad() function, and the pixel value of each padded pixel is 127;

[0061] A5: Randomly crop the padded pedestrian image in step A4;

[0062] In this embodiment, random cropping means randomly selecting a rectangular area from the pedestrian image to cause different degrees of occlusion of the pedestrian image, and correcting the misalignment in the pedestrian image through the spatial transformation network layer in the affine estimation branch. By cropping the part with a large background and filling the missing part of the pedestrian image, the phenomenon of network overfitting is reduced, the network generalization ability is improved, and no additional parameter learning or more memory consumption is required.

[0063] A6: Normalize the randomly cropped pedestrian image in step A5 so that the preprocessed pedestrian image data is limited within a set range, thereby eliminating the adverse effects caused by singular sample data. After data normalization, the speed of gradient descent to find the optimal solution can be accelerated, thereby improving the accuracy.

[0064] A7: Randomly erase channels of the normalized pedestrian image in step A6 with a probability of 0.5 to finally obtain the input image for supervised training.

[0065] B: Perform data augmentation on the pedestrian images in the cross-modal pedestrian re-identification dataset to obtain the pedestrian images after data augmentation, and the pedestrian images after data augmentation include context-based rotation self-supervised images and contrast-based self-supervised images;

[0066] The said step B includes the following specific steps:

[0067] B1: For each pedestrian image after size adjustment in step A2 in sequence, randomly select an angle from the set of rotation angles {0, 90, 180, 270} for rotation, and generate a pseudo-label for each rotated pedestrian image to obtain context-based rotation self-supervised images;

[0068] In this embodiment, by using the obtained context-based rotation self-supervised image and cooperating with the context-based rotation self-supervised network in step C, it is possible to focus on the beneficial attributes of pedestrian image features while ignoring the background features of training images, effectively learn the meaningful features in the semantics including the rotation-related part and the unrelated part, and incorporate rotation invariance into the self-supervised network learning framework. The context-based rotation self-supervised network in step C learns a segmentation representation including a rotation-related part and an unrelated part, and trains the neural network by jointly predicting image rotation and distinguishing individual instances. In the present invention, decoupling rotation recognition from instance recognition can improve rotation prediction by reducing the influence of rotation label noise, and recognizing instances regardless of image rotation, ensuring that the obtained features have better generalization ability.

[0069] B2: Perform channel random erasure with a probability of 0.5 on the supervised training input image obtained in step A7;

[0070] B3: Use channel swapping for the pedestrian image after channel random erasure in step B2 to finally obtain a contrast-based self-supervised image.

[0071] In this embodiment, color-independent images are uniformly generated by channel swapping (i.e., randomly swapping color channels). The three-channel color visible light image contains rich pedestrian feature information, and the color information in the pedestrian feature information is beneficial to visible light-infrared matching, thus continuously improving the robustness to color changes. Combining the channel random erasure strategy in step B2 and the random cropping strategy in step A5 can further enrich the diversity to obtain stronger distinguishability.

[0072] C: Construct a backbone network and a self-supervised training network for semi-supervised cross-modal pedestrian re-identification based on dual self-supervised learning; among them, the self-supervised training network includes a context-based rotation self-supervised network and a contrast learning-based self-supervised network; the backbone network, the context-based rotation self-supervised network, and the contrast learning-based self-supervised network are arranged in parallel and share network weights;

[0073] The backbone network is used to perform supervised learning on the supervised training input image obtained in step A to obtain the final pedestrian image features with robustness;

[0074] The backbone network sequentially includes a first convolutional layer, a first pooling layer, first to third residual layers, a first modality attention layer, a fourth residual layer, a second modality attention layer, and a part alignment attention layer; the first convolutional layer and the first to fourth residual layers perform feature extraction layer by layer on the downsampled pedestrian image features (including the shallow features and deep features of the pedestrian image); the first convolutional layer and the first to fourth residual layers are used to learn the shallow features of the pedestrian image; the first modality attention layer is used to learn the deep features of the pedestrian image in two modalities. The first modality attention layer and the second modality attention layer have the same structure, both consisting of two second convolutional layers with a kernel size of 1, a ReLU activation function, and a Sigmod activation function; the calculation formulas of the first modality attention layer and the second modality attention layer are:

[0075]

[0076] where Z represents the depth features of the obtained pedestrian image, represents the matrix after Z is instance-normalized, and m C is the channel mask, representing the channels related to identity, and the calculation formula of m C is m C = σ(W2δ(W1g(Z))); g(·) represents the global average pooling layer, δ(·) represents the ReLU activation function, σ(·) represents the Sigmod activation function, and W1 and W2 respectively represent the two fully connected layers in the modality attention layer, and the two fully connected layers are located after the ReLU activation function and the Sigmod activation function.

[0077] The part alignment attention layer is used to discover the small differences between the two modalities, divides the global pedestrian image features into six blocks, and obtains the final pedestrian image features after processing the feature vector obtained by combining the global pedestrian image features and the local pedestrian image features through an adaptive average pooling layer;

[0078] The context-based rotation self-supervised network is used to perform self-supervised learning on the context-based rotation self-supervised image obtained in step B1, and finally obtain the first probability matrix for predicting the rotation angle; the context-based rotation self-supervised network sequentially includes a third convolutional layer, a second pooling layer, a third modality attention layer, a fifth residual layer, a fourth modality attention layer, a global average pooling layer, a BN layer, and a first fully connected layer. The third modality attention layer and the fourth modality attention layer have the same structure as the first modality attention layer;

[0079] In the present invention, a context-based rotation self-supervised network applies a set of random geometric transformations to randomly rotate the input context-based rotation self-supervised image. Each randomly rotated rotation self-supervised image corresponds to a pseudo-label. The context-based rotation self-supervised network is used to identify the rotation angle of the rotation self-supervised image after random rotation. If the context-based rotation self-supervised network fails to capture the deep features of the pedestrian image in the rotated rotation self-supervised image, it cannot identify the rotation angle of the rotation self-supervised image. Therefore, in the present invention, the context-based rotation self-supervised network and the contrast learning-based self-supervised network share the weights of the backbone network, and the context-based rotation self-supervised network is also provided with a separate output. Through the global average pooling layer, the BN layer and the fully connected layer, the deep features of the pedestrian image in the rotated rotation self-supervised image are flattened into a vector with a dimension of 4 for classification, and finally a first probability matrix for rotation angle prediction is obtained to accurately identify the rotation angle of the rotation self-supervised image; the dimension of 4 represents the labels corresponding to four rotation angles.

[0080] The contrast learning-based self-supervised network is used to perform self-supervised learning on the contrast-based self-supervised image obtained in step B3, and finally obtain a second probability matrix for contrast self-supervised learning; the structure of the contrast learning-based self-supervised network is to add a second fully connected layer for color image classification on the basis of the backbone network, and the second fully connected layer is located after the part alignment attention layer.

[0081] In the present invention, the contrast learning-based self-supervised network is used to perform data augmentation on a given visible light image to obtain images of the same person with different color effects. The data augmentation is to perform channel random erasure and channel swapping on the R, G, and B channels of the visible light image. Subsequently, after extracting the pedestrian image features through the backbone network with shared weight parameters, a second probability matrix for contrast self-supervised learning is obtained through the second fully connected layer.

[0082] In the supervised task, only the consistency constraint is considered for the image classifiers (i.e., fully connected layers) of different modalities, and the backbone network only learns the shallow features of the pedestrian images between different modalities, but the shallow features of the pedestrian images between the same modalities are not considered. Because in the present invention, through the contrast learning-based self-supervised network, the shallow features of the pedestrian images between different modalities and between the same modalities are learned simultaneously, and the invariance between the input images of the supervised training and the pedestrian images after enhancement processing can be well learned.

[0083] D: Construct a training set using the enhanced pedestrian images in step B. The training set includes labeled samples and unlabeled samples. Use the labeled samples to learn the final pedestrian image features for supervised training through the backbone network. Use the unlabeled samples to obtain the first probability matrix for rotation angle prediction and the second probability matrix for contrastive self-supervised learning through the context-based rotation self-supervised network and the contrastive learning-based self-supervised network respectively;

[0084] E: Use the final pedestrian image features for supervised training obtained in step D, as well as the first probability matrix for rotation angle prediction and the second probability matrix for contrastive self-supervised learning, to perform the self-supervised pedestrian re-identification task through the backbone network and the self-supervised training network of semi-supervised cross-modal pedestrian re-identification based on dual self-supervised learning, and output the final recognition result;

[0085] In step E, for the final pedestrian image features of the input images for supervised training, the weights of the backbone network are updated through the set first cross-entropy loss function and center loss function using backpropagation;

[0086] For the final pedestrian image features of the input images for supervised training, supervised training is performed through the first cross-entropy loss function and the center loss function. The cross-entropy loss The calculation formula is:

[0087]

[0088] where, n and m respectively represent the number of visible light modality and infrared modality images in the current batch, f v , f r respectively represent the pedestrian image features of the visible light modality and the infrared modality, and respectively represent the image labels corresponding to f v , f r . C(f v ) and C(f r ) respectively represent the probability matrices obtained by the pedestrian image features of the two modalities through two classifiers with parameter θ. P(·) is the softmax function (normalized exponential function);

[0089] The center loss function The calculation formula is:

[0090]

[0091] where, f i represents the pedestrian image features, represents the mean of the features with the current batch label y i , Indicates that the label of the current batch is y k The mean of the features, Indicates that the label of the current batch is y j The mean of the features, T is the number of pedestrians in the current batch, and ρ refers to the minimum distance between all centers;

[0092] The context-based rotation self-supervised network determines the rotation angle through the second cross-entropy loss function, and finally the backbone network outputs the average precision result of pedestrian detection, which is used to evaluate the accuracy of pedestrian re-identification;

[0093] In the present invention, the first probability matrix output by the context-based rotation self-supervised network is calculated through the second cross-entropy loss function The calculation formula is:

[0094]

[0095] Wherein, Indicates the image after random rotation, Is the label generated by the random rotation angle of the image, and R represents the total number of image samples in a batch;

[0096] The second probability matrix output by the contrastive learning-based self-supervised network uses the KL divergence as the consistency constraint loss function:

[0097]

[0098] Wherein, p(x i ) represents the probability matrix obtained by the color image classifier in supervised learning, and q(x i ) represents the probability matrix obtained by the contrastive self-supervised features passing through the classifier.

[0099] In the present invention, the average precision result and Rank-1 (average correct rate of the first match) obtained by using the semi-supervised cross-modal pedestrian re-identification method based on dual self-supervised learning of the present invention are respectively increased by 5.9% (85.97% - 80.07%) and 8.27% (91.07% - 82.8%) on the RegDB dataset. In the indoor scenario of the SYSU dataset, the average precision result and Rank-1 are respectively increased by 1.35% (82.3% - 80.95%) and 1.96% (78.7% - 76.74%). The present invention not only successfully applies unsupervised learning to the field of pedestrian re-identification, but also a large number of experiments on several datasets show that the semi-supervised cross-modal pedestrian re-identification method based on dual self-supervised learning of the present invention can enhance the robustness of recognition and effectively improve the accuracy of pedestrian re-identification.

Claims

1. A semi-supervised cross-modal pedestrian re-identification method based on dual self-supervised learning, characterized in that, It includes the following steps: A: Construct a cross-modal pedestrian re-identification dataset, and preprocess the pedestrian images in the cross-modal pedestrian re-identification dataset to obtain the input images for supervised training; B: Perform data augmentation on the pedestrian images in the cross-modal pedestrian re-identification dataset to obtain the pedestrian images after data augmentation. The pedestrian images after data augmentation include context-based rotation self-supervised images and contrast-based self-supervised images; C: Construct a backbone network and a self-supervised training network for semi-supervised cross-modal pedestrian re-identification based on dual self-supervised learning. Among them, the self-supervised training network includes a context-based rotation self-supervised network and a contrast learning-based self-supervised network. The backbone network, the context-based rotation self-supervised network, and the contrast learning-based self-supervised network are set in parallel and share network weights; Among them, the backbone network is used to perform supervised learning on the input images for supervised training to obtain the final pedestrian image features. The context-based rotation self-supervised network is used to perform self-supervised learning on the context-based rotation self-supervised images to obtain the first probability matrix for rotation angle prediction. The contrast learning-based self-supervised network is used to perform self-supervised learning on the contrast-based self-supervised images to obtain the second probability matrix for contrast self-supervised learning; D: Use the pedestrian images after enhancement in step B to construct a training set. The training set includes labeled samples and unlabeled samples. Use the labeled samples to learn the final pedestrian image features for supervised training through the backbone network. Use the unlabeled samples to obtain the first probability matrix for rotation angle prediction and the second probability matrix for contrast self-supervised learning through the context-based rotation self-supervised network and the contrast learning-based self-supervised network respectively; E: Use the final pedestrian image features for supervised training obtained in step D, as well as the first probability matrix for rotation angle prediction and the second probability matrix for contrast self-supervised learning, to perform a self-supervised pedestrian re-identification task through the backbone network and the self-supervised training network for semi-supervised cross-modal pedestrian re-identification based on dual self-supervised learning, and output the final recognition result.

2. The semi-supervised cross-modal person re-identification method based on dual self-supervised learning according to claim 1, wherein, The step A includes the following specific steps: A1: Construct a cross-modal pedestrian re-identification dataset, obtain the pedestrian images in the training set of the cross-modal pedestrian re-identification dataset, and set the total number of images input to the network model of the semi-supervised cross-modal pedestrian re-identification method based on dual self-supervised learning; A2: Adjust the size of the pedestrian images in the cross-modal pedestrian re-identification dataset, and adjust the width and height of the pedestrian images to the same size; A3: Randomly horizontally flip the pedestrian images after size adjustment in step A2; A4: Perform pixel filling on the pedestrian images after random horizontal flipping in step A3; A5: Randomly crop the pedestrian images after filling in step A4; A6: Normalize the pedestrian images after random cropping in step A5; A7: Randomly erase channels of the pedestrian images after normalization in step A6 to obtain the input images for supervised training.

3. The semi-supervised cross-modal person re-identification method based on dual self-supervised learning according to claim 2, wherein The step B includes the following specific steps: B1: For each pedestrian image after size adjustment, randomly select an angle from the set of rotation angles {0, 90, 180, 270} for rotation, and generate a pseudo-label for each rotated pedestrian image respectively to obtain context-based rotation self-supervised images; B2: Perform channel random erasure on the obtained input images for supervised training; B3: Use channel swapping on the pedestrian images after completing channel random erasure to obtain contrast-based self-supervised images.

4. The semi-supervised cross-modal person re-identification method based on dual self-supervised learning according to claim 1, characterized in that: In the step C described above, the backbone network sequentially includes a first convolutional layer, a first pooling layer, a first to third residual layer, a first modality attention layer, a fourth residual layer, a second modality attention layer, and a part alignment attention layer; the first convolutional layer and the first to fourth residual layers perform feature extraction on the pedestrian image features after dimensionality reduction layer by layer to learn the shallow features of the pedestrian image; the first modality attention layer is used to learn the deep features of the pedestrian images in two modalities; The part alignment attention layer is used to discover the small differences between the visible light modality and the infrared modality to obtain the final pedestrian image features.

5. The semi-supervised cross-modal person re-identification method based on dual self-supervised learning according to claim 4, characterized in that: The first modality attention layer and the second modality attention layer have the same structure, both consisting of two second convolutional layers with a kernel size of 1, a ReLU activation function, and a Sigmod activation function; the calculation formulas of the first modality attention layer and the second modality attention layer are: Among them, Z represents the depth feature of the obtained pedestrian image, represents the matrix after Z undergoes instance normalization, m C is the channel mask, representing the channels related to identity, m C The calculation formula of is m C =σ(W2δ(W1g(Z))); g(·) represents the global average pooling layer, δ(·) represents the ReLU activation function, σ(·) represents the Sigmod activation function, and W1 and W2 respectively represent the two fully connected layers in the modality attention layer, and the two fully connected layers are located after the ReLU activation function and the Sigmod activation function.

6. The semi-supervised cross-modal person re-identification method based on dual self-supervised learning according to claim 1, characterized in that: The context-based rotation self-supervised network sequentially includes a third convolutional layer, a second pooling layer, a third modality attention layer, a fifth residual layer, a fourth modality attention layer, a global average pooling layer, a BN layer, and a first fully connected layer. The third modality attention layer and the fourth modality attention layer have the same structure as the first modality attention layer.

7. The semi-supervised cross-modal person re-identification method based on dual self-supervised learning according to claim 1, wherein: The structure of the contrast learning-based self-supervised network is to add a second fully connected layer on the basis of the backbone network, and the second fully connected layer is located after the part alignment attention layer.

8. The semi-supervised cross-modal person re-identification method based on dual self-supervised learning according to claim 1, characterized in that: In the step E described above, the final pedestrian image features of the input images for supervised training are used to update the weights of the backbone network through the set first cross-entropy loss function and center loss function using backpropagation; Cross-entropy loss The calculation formula is as follows: where n and m respectively represent the numbers of visible-light modality and infrared modality images in the current batch, f v , f r respectively represent the pedestrian image features of the visible-light modality and the pedestrian image features of the infrared modality, and respectively represent the image labels corresponding to f v , f r , C(f v ) and C(f r ) respectively represent the probability matrices obtained by passing the pedestrian image features of the two modalities through two classifiers with parameters θ, and P(·) is the softmax function; Center loss function The calculation formula is as follows: Among them, f i represents the pedestrian image feature, represents the mean of the features with the current batch label being y i and represents the mean of the features with the current batch label being y k and represents the mean of the features with the current batch label being y j . T is the number of pedestrians in the current batch, and ρ refers to the minimum distance between all centers.

9. The semi-supervised cross-modal person re-identification method based on dual self-supervised learning according to claim 1, wherein: In the step E described above, the context-based rotation self-supervised network judges the rotation angle through the second cross-entropy loss function, and finally the backbone network outputs the average precision result of pedestrian detection. The average precision result is used to evaluate the accuracy of pedestrian re-identification; The second cross-entropy loss function The calculation formula is as follows: Among them, represents the image after random rotation, is the label generated by the random rotation angle of the image, and R represents the total number of image samples in a batch.

10. The semi-supervised cross-modal person re-identification method based on dual self-supervised learning according to claim 1, characterized in that: In the step E described above, the second probability matrix output by the contrast learning-based self-supervised network uses the KL divergence as the consistency constraint loss function; Among them, p(x i ) represents the probability matrix obtained by the color image classifier in supervised learning, and q(x i ) represents the probability matrix obtained after the contrast-based self-supervised features pass through the classifier.

Citation Information

Patent Citations

  • Pedestrian re-identification method based on unsupervised cross-modal

    CN114495004A

  • Cross-modal pedestrian re-identification method using local supervision

    CN115565204A