A Transformer-based Camera Domain Adaptation Person Re-identification Method

Through the Transformer-based adversarial learning framework, image features are extracted using the cross-patch encoder and the Transformer encoder, the problem of image style changes in pedestrian recognition under multi-camera shooting is solved, the recognition accuracy is improved, and the limitations of convolutional neural networks are overcome.

CN114155554BActive Publication Date: 2025-07-22SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111463655.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-02
Publication Date
2025-07-22
Estimated Expiration
2041-12-02

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of image style changes caused by differences in lighting, background and resolution in pedestrian recognition under multi-camera shooting, and data enhancement methods may introduce errors.

Method used

Adopting an adversarial learning framework based on Transformer, image features are extracted through cross-patch encoder and Transformer encoder, the network is optimized by combining identity and camera classification loss functions, and the discriminator is used to judge camera categories, and robust features are directly learned from the original data.

Benefits of technology

It improves the accuracy of pedestrian re-identification, overcomes the limitations of convolutional neural networks, realizes adaptive feature extraction for camera style differences, and avoids the error introduced by data augmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114155554B_ABST
    Figure CN114155554B_ABST
Patent Text Reader

Abstract

The present invention discloses a Transformer-based camera domain adaptation pedestrian re-identification method, which includes: encoding an input image into a vector sequence using a cross-patch encoder; inputting the vector sequence into a Transformer encoder to learn image features, and constructing an identity information loss using the image features to optimize the network; regarding the cross-patch encoder and the Transformer encoder together as a feature generator, inputting the features generated by the generator into a discriminator to judge the camera category, and on this basis, constructing a camera classification loss and a camera domain adaptation loss to optimize the discriminator and the generator respectively; extracting the feature vectors of pedestrian images using the generator, calculating the Euclidean distance between the feature vectors of the query image and the feature vectors of each image, sorting them from small to large in terms of distance, and selecting the pedestrian identity of the image with the topmost ranking as the recognition result. The method of the present invention has a high accuracy rate and can effectively solve the problem of image style differences brought about by images collected by multiple cameras in the pedestrian re-identification task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision and image retrieval, and particularly relates to a Transformer-based camera domain adaptation person re-identification method. Background Art

[0002] Person re-identification is a technology for retrieving specific persons from a large range of image sets. This technology has important practical significance in the fields of intelligent video surveillance, intelligent security, etc. In recent years, person re-identification technology has developed rapidly, but the uncontrolled appearance changes of images between multiple cameras make person re-identification still a challenging task. In actual re-identification scenarios, images captured by different cameras often have differences in illumination, background, and resolution, and these differences will have an adverse impact on the performance of re-identification. At present, a considerable number of generative models have attempted to handle camera style differences, and the adaptation of these methods to camera style differences is mainly reflected in data augmentation. The process of generating images of different camera styles is relatively complicated, and the data augmentation process is relatively independent of feature representation learning, so the data augmentation process may introduce unnecessary errors. Therefore, the present invention designs an adversarial loss to solve the problem of camera style differences from the perspective of metric learning.

[0003] In recent years, studies have shown that the ability of convolutional neural networks to retain fine-grained information and learn long-range dependencies is not ideal, but the vast majority of re-identification methods still choose to use deep convolutional neural networks to extract image features. Recently, as a network structure that does not rely on convolutional operations at all, Transformer has become increasingly popular in the field of computer vision, so it is very meaningful to explore a person re-identification method based on the Transformer structure. Summary of the Invention

[0004] In view of the above problems, the present invention designs an adversarial learning framework based on Transformer from the perspective of metric learning to solve the problem of differences in person images between multiple cameras, thereby effectively improving the accuracy of person re-identification.

[0005] To achieve the above object, the technical solution of the present invention is as follows:

[0006] A Transformer-based camera domain adaptation person re-identification method, comprising the following steps:

[0007] (1) Decompose the input person image into image patches with a fixed resolution, and the image patches and the corresponding cross-image blocks are encoded by a cross-patch encoder to obtain a vector sequence;

[0008] (2) Input the vector sequence into the Transformer encoder to learn the feature vector of the image, and use the learned image features to construct the identity classification loss and the triplet loss to optimize the cross-patch encoder and the Transformer encoder;

[0009] (3) Regard the cross-patch encoder and the Transformer encoder together as a feature generator, input the image features generated by the generator into the discriminator to judge the camera category of this feature, and on this basis, construct the camera classification loss and the camera domain adaptation loss to alternately optimize the discriminator and the generator respectively;

[0010] (4) Use the trained generator to extract the feature vector of the pedestrian image, calculate the Euclidean distance between the feature vector of the image to be queried and the feature vector of each image, sort them in ascending order of the distance, and select the pedestrian identity of the image with the top-ranked distance as the recognition result.

[0011] The framework proposed by the present invention consists of a cross-patch encoder, a Transformer encoder, and a discriminator. The cross-patch encoder encodes the input pedestrian image into a vector sequence, the Transformer encoder learns the feature representation from the vector sequence, and the discriminator is used to judge the camera category to which the feature belongs. During the training process, the cross-patch encoder and the Transformer encoder are connected in series as a feature generator G, and the feature generator and the discriminator are alternately updated until the model converges.

[0012] In step (1), use a linear transformation to map the image patches with a fixed resolution into vectors with a fixed dimension At the same time, use depthwise separable convolution to map the cross-image blocks corresponding to the image patches into vectors with the same dimension. Finally, the vector e i generated by the encoder is:

[0013]

[0014] where i represents the serial number of the pedestrian image, j represents the serial number of the image patch, and represent the vectors mapped by the horizontal and vertical image blocks respectively, and p i is the position vector containing position information.

[0015] In step (2), the identity information loss function used to optimize the cross-patch encoder and the Transformer encoder is:

[0016]

[0017] Represents the identity classification loss function, and the formula is as follows:

[0018]

[0019] Where p(y i |x i ) represents the predicted probability that the input image x i belongs to the identity class y i . At the same time, in order to strengthen intra-class aggregation and inter-class separation, a triplet loss function is introduced during training The formula is as follows:

[0020]

[0021] Where m represents the margin, G(·) represents the image features output by the Transformer encoder, d represents the distance between two features, x p , x n are the positive and negative samples of the reference sample x i respectively.

[0022] In step (3), the discriminator is used to identify the camera class of the pedestrian features, while the generator tries to generate pedestrian features that are difficult to be identified by the discriminator. The camera classification loss function for optimizing the discriminator is:

[0023]

[0024]

[0025] Where q i represents the correct camera class of the pedestrian image x i , p(q i |x i ) represents the probability that the pedestrian image x i belongs to the camera class q i , G(x i ) represents the image features extracted by the generator, D(G(x i ))[j] represents the predicted score of the discriminator output for the camera class j, and K represents the total number of camera classes. The camera domain adaptation loss function for optimizing the generator is:

[0026]

[0027] Where, p(g|x i ) represents the pedestrian image x iThe probability of belonging to camera category g, where δ(·) represents the Dirac δ function. During the training process of the generator and the discriminator, the parameters of one party are fixed, and the parameters of the other party are updated, iterating alternately until the model converges. The specific training process can be expressed as:

[0028]

[0029]

[0030] where, θ G and θ D represent the parameter variables of the generator and the discriminator respectively, and represent the fixed network parameters, and λ represents the hyperparameter that adjusts the contributions of the two loss functions.

[0031] The beneficial effects of the present invention are as follows:

[0032] (1) The present invention uses the Transformer as the backbone network to extract effective features of pedestrian images. The entire backbone network does not use pooling and convolution operations, enabling the method of the present invention to overcome the limitations of the method based on convolutional neural networks.

[0033] (2) The present invention designs a novel cross-patch encoder, which obtains a more effective vector sequence from pedestrian images at a lower computational cost.

[0034] (3) The method of the present invention does not rely on any data augmentation techniques and can directly learn pedestrian features that are robust to camera style changes from the original dataset. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a schematic flow chart of a method for camera domain adaptation pedestrian re-identification based on Transformer according to the present invention;

[0036] Figure 2 is a schematic structural diagram of a cross-patch encoder;

[0037] Figure 3 is a schematic framework diagram of a system for camera domain adaptation pedestrian re-identification based on Transformer according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0038] The following further clarifies the present invention in conjunction with the accompanying drawings and specific embodiments. It should be noted that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0039] As Figure 1As shown in the figure, a Transformer-based camera domain adaptation pedestrian re-identification method of the present invention includes the following steps:

[0040] Step 1: Encode the input image into a sequence of vectors using a cross-patch encoder;

[0041] The structure of the cross-patch encoder in the present invention is as Figure 2 shown.

[0042] Given a training dataset where N1 represents the total number of pedestrian images in the dataset, y i and q i represent the identity label and camera type label of the pedestrian image x i respectively. First, all pedestrian images are resized to a fixed size of H×W, and then the pedestrian images are divided into N2 non-overlapping image patches of size I×I where represents the j-th patch of the i-th pedestrian image, and these image patches are mapped into M-dimensional vectors The formula is as follows:

[0043]

[0044] where F P represents a linear mapping, i represents the serial number of the pedestrian image, and j represents the serial number of the patch. On this basis, the cross-patch encoder maps the cross-image blocks corresponding to the image patches into vectors of the same dimension as

[0045]

[0046] where represents a horizontal image block of size I×W, represents a vertical image block of size H×I, represents a horizontal vector, represents a vertical vector, F h and F v represent depthwise separable convolutions applied to the horizontal and vertical image blocks respectively. Finally, the position vector p i is added to the output vector of the cross-patch encoder, and is expressed by the formula:

[0047]

[0048] In this embodiment, the fixed size of the input image is 256×128, the size of the image patch is 16×16, and M is set to 768.

[0049] Step 2: Input the vector sequence into the Transformer encoder to learn the feature vectors of the image, and use the learned image features to construct the identity classification loss and the triplet loss to optimize the cross-patch encoder and the Transformer encoder;

[0050] As Figure 3 shown, before the vector sequence is input into the Transformer encoder, a trainable classification vector is appended to the vector sequence, so the Transformer encoder processes the input (N2 + 1) vectors. The structure of the Transformer encoder enables information to propagate among the vectors, and finally only the image features corresponding to the classification vector are used to construct the identity classification loss and the triplet loss. Among them, the identity information loss function used to optimize the cross-patch encoder and the Transformer encoder is:

[0051]

[0052] denotes the identity classification loss function, and the formula is as follows:

[0053]

[0054] where p(y i |x i ) represents the predicted probability that the input image x i belongs to the identity class y i , and the predicted probability is obtained through the classifier connected after the feature vector. At the same time, in order to strengthen the intra-class aggregation and inter-class separation, the triplet loss function is introduced during the training process, and the formula is as follows:

[0055]

[0056] where m represents the margin, G(·) represents the image features output by the Transformer encoder, d represents the distance between two features, x p , x n respectively represent the positive sample and the negative sample of the reference sample x i in a batch of training samples.

[0057] In this embodiment, ViT-Base is selected as the Transformer encoder to extract the pedestrian feature vectors. Before starting the training, ViT-Base is pre-trained on two datasets, ImageNet-21K and ImageNet-1K.

[0058] Step 3: Regard the cross-patch encoder and the Transformer encoder together as a feature generator, input the image features generated by the generator into the discriminator to judge the camera category of this feature, and construct a camera classification loss and a camera domain adaptation loss on this basis to alternately optimize the discriminator and the generator;

[0059] As Figure 3 shown, the discriminator is used to discriminate the camera category of pedestrian features, while the generator tries to generate pedestrian features that are difficult to be discriminated by the discriminator. The camera classification loss function for optimizing the discriminator can be expressed as:

[0060]

[0061]

[0062] where q i represents the correct camera category of the pedestrian image x i , p(q i |x i ) represents the probability that the pedestrian image x i belongs to the camera category q i , G(x i ) represents the image features extracted by the generator, and D(G(x i ))[j] represents the predicted score of the discriminator output for the camera category j, and K represents the total number of camera categories. The camera domain adaptation loss function for optimizing the generator can be expressed as:

[0063]

[0064] where p(g|x i ) represents the probability that the pedestrian image x i belongs to the camera category g, and δ(·) represents the Dirac δ function. During the training process of the generator and the discriminator, fix the parameters of one party and update the parameters of the other party, and iterate alternately until the model converges. The specific training process can be expressed as:

[0065]

[0066]

[0067] where θ G and θ D represent the parameter variables of the generator and the discriminator respectively, and represent the fixed network parameters, and λ represents the hyperparameter for adjusting the contributions of the two loss functions.

[0068] In this embodiment, the discriminator is a shallow fully-connected network. The number of camera categories K is 15. The SGD optimizer with a learning rate of 0.008, a momentum coefficient of 0.9, and a weight decay of 0.0001 is applied to the generator, and the Adam optimizer with a learning rate of 0.0003 is applied to the discriminator.

[0069] Step 4: Use the trained generator to extract the feature vectors of pedestrian images, calculate the Euclidean distances between the feature vectors of the query image and those of each image, sort them in ascending order of distance, and select the pedestrian identity of the image with the smallest distance as the recognition result.

[0070] To verify the effectiveness of the present invention, experiments are conducted on the MSMT17 dataset. The MSMT17 dataset consists of 126,441 images of 4,101 pedestrians captured by 15 cameras, of which 32,621 pedestrian images are used for training and the other 93,820 pedestrian images are used for testing.

[0071] In the test phase, the cumulative matching characteristic metric (CMC) and the mean average precision (mAP) are used to quantitatively evaluate the performance of the model. Finally, the method of the present invention achieves a Rank-1 accuracy of 62.9% and a mean average precision of 83.4% on the MSMT17 dataset.

[0072] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.

Claims

1. A Transformer-based camera domain adaptation pedestrian re-identification method, characterized in that Including the following steps: (1) Decompose the input pedestrian image into image patches with a fixed resolution. The image patches and their corresponding cross-image blocks are encoded by a cross-patch encoder to obtain a vector sequence; (2) Input the vector sequence into a Transformer encoder to learn the feature vectors of the image. Use the learned image features to construct an identity classification loss and a triplet loss to optimize the cross-patch encoder and the Transformer encoder; (3) Regard the patch encoder and the Transformer encoder together as a feature generator. Input the image features generated by the generator into a discriminator to judge the camera category of this feature. On this basis, construct a camera classification loss and a camera domain adaptation loss to alternately optimize the discriminator and the generator respectively; Camera classification loss function for optimizing discriminator is as follows: Among them, q i represents the correct camera category of the pedestrian image x i , p(q i |x i ) represents the probability that the pedestrian image x i belongs to the camera category q i , G(x i ) represents the image features extracted by the generator, D(G(x i ))[j] represents the prediction score of the discriminator output for the camera category j, and K represents the total number of camera categories; the camera domain adaptation loss function used to optimize the generator is as follows: where p(g|x i ) represents the probability that the pedestrian image x i belongs to the camera category g, and δ(·) represents the Dirac δ function; during the training process of the generator and the discriminator, the parameters of one party are fixed, and the parameters of the other party are updated, and the iteration is alternated until the model converges; the specific training process is as follows: Among them, θ G and θ D represent the parameter variables of the generator and the discriminator respectively, and represent the fixed network parameters, and λ represents the hyperparameter for adjusting the contributions of the two loss functions; is the identity information loss function for optimizing the cross-patch encoder and the Transformer encoder; (4) Use the trained generator to extract the feature vectors of the pedestrian image. Calculate the Euclidean distance between the feature vector of the query image and the feature vectors of each image. Sort them in ascending order of distance, and select the pedestrian identity of the image with the smallest distance as the recognition result.

2. The method for camera domain adaptation pedestrian re-identification based on Transformer according to claim 1, wherein In step (1), the image patch with a fixed resolution is mapped into a vector with a fixed dimension by a linear transformation Meanwhile, the cross-image patch corresponding to the image patch is mapped into a vector with the same dimension by a depthwise separable convolution, and finally the vector e generated by the encoder i is as follows: where \(i\) represents the serial number of the pedestrian image, and \(j\) represents the serial number of the image patch, and respectively represent the vectors corresponding to the horizontal and vertical image patch mappings, and \(\mathbf{p}\) i is the position vector containing position information.

3. A Transformer-based camera domain adaptation pedestrian re-identification method according to claim 1, characterized in that, In step (2), the identity information loss function for optimizing the patch encoder and the Transformer encoder is as follows: Among them, represents the identity classification loss function, represents the triplet loss function.

Citation Information

Patent Citations

  • Pedestrian re-recognition method based on multi-task learning

    CN112149538A

  • Video pedestrian re-recognition method based on Transform space-time modeling

    CN113627266A