Cross-modal Person Re-identification Method Based on Spatiotemporal Features and Hetero-center Loss
By connecting the spatial feature extraction module in the cross-modal pedestrian re-identification model and using heterocenter loss, the problem of in-modal differential constraints is solved, and the accuracy of cross-modal pedestrian re-identification is achieved, and the recognition and compactness of feature representations are enhanced.
Patent Information
- Application Number
- CN202211169495.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-22
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-09-22
AI Technical Summary
The existing cross-modal pedestrian re-identification method uses modal differences in the modal differences, and the existing modal alignment modules have key detail features that are not fully utilized at the spatial level.
A cross-modal pedestrian re-identification method based on spatiotemporal features and heterocenter loss is adopted. By connecting the spatial feature extraction module and modal alignment module, combining heterocenter sample loss, identity loss and central cluster loss, network parameters are optimized, global, spatial and channel features are extracted, and loss functions are constructed to improve feature compactness and recognition.
The accuracy of cross-modal pedestrian re-identification is improved. By simultaneously utilizing channel and spatial characteristics, the identification of pedestrian representation and the compactness of feature distribution are enhanced, and the matching accuracy of cross-modal pedestrian re-identification is improved.
Smart Images

Figure CN115497121B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image retrieval algorithms, and relates to the technology of pedestrian re-identification based on cross-modal images. Specifically, it is a cross-modal pedestrian re-identification method based on spatio-temporal features and heterocentric loss. Background Art
[0002] As an important branch in the field of image retrieval, pedestrian re-identification has a wide range of applications in ensuring social public safety and other aspects. With the emergence of more and more monitoring devices integrated with infrared acquisition devices, cross-modal pedestrian re-identification using both visible light images and infrared images has become an important research direction.
[0003] Compared with traditional single-modal pedestrian re-identification, cross-modal pedestrian re-identification is more challenging. In addition to factors such as view changes, occlusions, and pose changes, cross-modal pedestrian re-identification also faces the challenge of huge modal differences. To reduce the impact brought by modal differences, some methods adopt the method of modal conversion to convert infrared images and visible light images into the same modality. Other methods adopt the method of representation learning to align the image features of the two modalities in a unified feature space.
[0004] There is a existing cross-modal pedestrian re-identification method that uses a single-stream network followed by a mode alignment module to extract image detail features, and then uses cross-entropy loss and center cluster loss to train network parameters to achieve the purpose of cross-modal pedestrian re-identification. However, although the center cluster loss handles the differences between modalities well, the constraint on intra-modal differences is not tight enough; the mode alignment module extracts detail features from the channel level, but there are also some key detail features that can be utilized at the spatial level. If these information can be utilized simultaneously, the performance of the cross-modal pedestrian re-identification model can be further improved. Summary of the Invention
[0005] In order to overcome the deficiencies existing in the prior art, the present invention provides a cross-modal pedestrian re-identification method based on spatio-temporal features and heterocentric loss, which simultaneously extracts channel features and spatial features, making the representation of pedestrians more discriminative, and uses heterocentric sample loss to make the feature distribution of the same pedestrian feature more compact.
[0006] The technical solution adopted by the present invention to solve its technical problems is: a cross-modal pedestrian re-identification method based on spatio-temporal features and heterocentric loss. The cross-modal pedestrian re-identification model extracts pedestrian image features from the spatial, channel, and global dimensions, and uses a loss function composed of identity loss, heterocentric sample loss, center cluster loss, and total feature loss to train the cross-modal pedestrian re-identification model to achieve cross-modal pedestrian re-identification of pedestrian images.
[0007] As a further embodiment of the present invention, it specifically includes the following steps:
[0008] S1: Construct a cross-modal person re-identification model, and the cross-modal person re-identification model includes a parallel modal alignment module and a spatial feature extraction module;
[0009] S2: Perform data augmentation on the person images input into the cross-modal person re-identification model;
[0010] S3: The cross-modal person re-identification model extracts the global features, spatial local features, and channel local features of the person images;
[0011] S4: Calculate the heterocentric sample loss and identity loss of the extracted spatial local features, and calculate the central cluster loss and identity loss of the extracted channel local features;
[0012] S5: Concatenate the global features, spatial local features, and channel local features as the total features, and calculate the total feature loss;
[0013] S6: Add all the losses in steps S4 and S5 to form the loss function of the overall cross-modal person re-identification model, and train and optimize the parameters in the cross-modal person re-identification model according to this loss function;
[0014] S7: After the cross-modal person re-identification model is trained, input the person images to be queried and the images in the test set into the cross-modal person re-identification model, calculate the similarity between them, and return the M values with the highest similarity, which are the results of cross-modal person re-identification.
[0015] As a further embodiment of the present invention, in step S2, the data augmentation of the person images is specifically: mixing visible light images and infrared images to form several batches.
[0016] As a further embodiment of the present invention, in step S4, the calculation formula for the heterocentric sample loss of the extracted spatial local features is as follows:
[0017]
[0018] Among them, L HCS represents the heterocentric sample loss, ρ is the margin parameter, δ is a balance coefficient, [x] + = max(x, 0) represents the standard hinge loss, ||x a - x b ||2 represents the L2 norm between x a and x b ; P represents the total number of different classes in a mini-batch, and respectively represent the feature representations of the j-th visible light image and the j-th infrared image of class i, and respectively represent the central features of class i in the visible light modality and the infrared modality in a mini-batch;
[0019] and are obtained by taking the mean of all samples of class i in their respective modalities, and their calculation formulas are as follows:
[0020]
[0021] where K represents that the number of visible light images and infrared images in a mini-batch is both K.
[0022] The beneficial effects of the present invention include: On the basis of the original modality alignment module, a spatial feature extraction module is connected in parallel, and both channel features and spatial features are extracted, making the representation of pedestrians more discriminative; at the same time, the heterogeneous center sample loss is used to handle both cross-modal and intra-modal variations, making the feature distribution of the same pedestrian feature more compact. It has reference significance for exploring the feature representation of pedestrians and improving the accuracy of pedestrian re-identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is the model framework diagram of the present invention;
[0024] Figure 2 is the diagram for explaining the heterogeneous center sample loss of the present invention;
[0025] Figure 3 is the t-SNE diagram for comparing the effects of the heterogeneous center sample loss. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0027] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0028] Embodiment 1
[0029] A cross-modal pedestrian re-identification method based on spatio-temporal features and heterogeneous center loss. On the basis of the original cross-modal pedestrian re-identification model, a spatial feature extraction structure is connected in parallel to construct a new cross-modal pedestrian re-identification model, Figure 1This is the model framework diagram of the present invention.
[0030] First, perform data augmentation on the infrared images and visible light images in the dataset. Then send them into a single-stream network composed of Resnet50 and MPM to obtain an embedded feature map X. Next, use three branches to extract the spatial features, channel features, and global feature representations of the embedded feature map X respectively. Finally, use the identity loss and the relevant cross-modal triplet loss for these three feature representations respectively to optimize the parameters of the network and obtain a cross-modal person re-identification model.
[0031] It specifically includes the following technical steps:
[0032] 1. Construct the network model and initialize the network parameters.
[0033] The backbone network of the present invention is a single-stream network composed of Resnet50 and MAM, and its main purpose is to extract a coarse-grained pedestrian feature Among them, C, H, and W respectively represent the number of channels, length, and width of the feature map. Resnet50 initializes the network parameters with the pre-trained parameters on ImageNet, and removes the last downsampling layer, that is, the stride of the last convolutional block in layer 4 is changed to 1. MAM is essentially an instance normalization structure (specifically refer to: Qiong Wu, Pingyang Dai, Jie Chen, Chia-Wen Lin, Yongjian Wu, Feiyue Huang, Bineng Zhong, and Rongrong Ji. Discover cross-modality nuances for visible-infrared person re-identification. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pages 4330–4339, 2021.), and embed it behind layer 3 and layer 4 of Resnet50. Then there are three branches behind the backbone network, one is the spatial feature extraction module branch, one is the channel feature extraction module branch, and the last one is the global feature extraction branch.
[0034] For the spatial feature branch, the coarse-grained pedestrian feature X is evenly divided into p small blocks, and each small block uses operations such as pooling, 1*1 convolution, and reshaping to extract local spatial features. It can be specifically expressed as:
[0035] S i= Reshape(Conv(pool(X i )))
[0036] Among them, represents the feature map of each small block. Finally, all the spatial local features are connected as the final feature representation of the spatial feature extraction module.
[0037] The channel feature branch is mainly composed of a pattern alignment module and a downsampling layer with 1*1 convolution.
[0038] The global feature extraction branch performs pooling and reshaping operations on the coarse-grained pedestrian feature X. Finally, the pedestrian features of the three branches are connected as the representation of the final pedestrian feature.
[0039] 2. Image preprocessing
[0040] The images in the dataset are randomly cropped into images of size 288*144, and then randomly horizontal flipping and random grayscaling are used to increase the diversity of the data. For visible light images, an additional local random channel enhancement is used, that is, a region in the visible light image is randomly selected, and the pixel values in this region are randomly replaced with one of the gray value, the pixel values of the R channel, the G channel, and the B channel. This is done to force the model to reduce its sensitivity to color changes and promote its learning of color-independent features.
[0041] 3. Loss function
[0042] In the process of training the network model of the present invention, multiple loss functions are used. Specifically, they include off-center sample loss, center cluster loss, identity loss, and other losses.
[0043] The off-center sample loss is an improvement of the traditional triplet loss, which simultaneously optimizes the distance between each class center and the distance between the sample and the class center. As Figure 2 shown, the squares and circles represent different classes, the hollow ones represent each sample feature, and the solid ones represent the center features calculated from the samples. In a mini-batch, the off-center sample loss, on the one hand, makes the center features of the same class and different modalities approach each other, and the center features of different classes move away from each other; on the other hand, it makes the distance of each sample from the class center as close as possible. Its formula is as follows:
[0044]
[0045] Among them, ρ is the margin parameter, and δ is a balance coefficient. [x]+ = max(x, 0) represents the standard hinge loss, ||x a - x b ||2 represents x a and x bThe second norm between them. P represents the total number of different classes in a mini-batch. and respectively represent the feature representations of the j-th visible light image and the j-th infrared image of class i. and respectively represent the central features of class i in the visible light modality and the infrared modality in a mini-batch. They are obtained by taking the mean of all samples of class i in their respective modalities, and their calculation formulas are as follows:
[0046]
[0047] where K represents that the number of visible light pictures and infrared pictures in a mini-batch is both K. The different-center sample loss makes the sample distribution more compact, and its comparison effect with the traditional triplet loss is as Figure 3 shown.
[0048] The identity loss is essentially a cross-entropy loss. Regarding cross-modal pedestrian re-identification as a multi-classification problem, each identity of a pedestrian is regarded as a class.
[0049] The central cluster loss and other losses are all the losses used in the above-mentioned literature. The total loss used, after removing the central cluster loss, is denoted as other losses.
[0050] In the present invention, for the spatial feature extraction module branch, the identity loss and the different-center sample loss are used; for the channel feature extraction module branch, the identity loss and the central cluster loss are used; for the global feature extraction branch, other losses are used.
[0051] In response to the requirements of the cross-modal pedestrian re-identification task, the present invention extracts the local features of pedestrians from the spatial perspective, the channel perspective, and the global perspective, enhances the feature representation of pedestrians, narrows the distance between different-modal features, and improves the matching accuracy of the cross-modal pedestrian re-identification task.
[0052] Obviously, the above embodiments are only examples given for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A cross-modal pedestrian re-identification method based on spatio-temporal features and off-center loss, characterized in that The cross-modal pedestrian re-identification model extracts pedestrian image features from the parallel spatial, channel, and global dimensions, and uses a loss function composed of identity loss, off-center sample loss, center cluster loss, and total feature loss to train the cross-modal pedestrian re-identification model to achieve cross-modal pedestrian re-identification of pedestrian images; The spatial dimension evenly divides the coarse-grained pedestrian feature X into p small blocks. Each small block uses pooling, 1*1 convolution, and reshaping operations to extract spatial local features, which are specifically expressed as: S i = Reshape(Conv(pool(X i ))) Among them, represents the feature map of each small block; finally, all the spatial local features are connected as the final feature representation of the spatial feature extraction module; The channel dimension consists of a pattern alignment module and a downsampling layer with 1*1 convolution; The global dimension performs pooling and reshaping operations on the coarse-grained pedestrian feature X, and finally concatenates the pedestrian image features of the three dimensions as the representation of the final pedestrian image features; The formula for calculating the off-center sample loss is as follows: Among them, L HCS represents the off-center sample loss, ρ is the margin parameter, δ is a balance coefficient, and [x] + = max(x, 0) represents the standard hinge loss, and ||x a - x b ||2 represents the L2 norm between x a and x b ; P represents the total number of different classes in a mini-batch, and respectively represent the feature representations of the j-th visible light image and the j-th infrared image of class i, and respectively represent the central features of class i in the visible light modality and the infrared modality in a mini-batch; and is obtained by averaging the samples of all classes i in their respective modalities, and its calculation formula is as follows: where K represents that the number of visible light images and infrared images in a mini-batch is both K.
2. The cross-modal pedestrian re-identification method based on spatio-temporal features and off-center loss according to claim 1, wherein Specifically, it includes the following steps: S1: Construct a cross-modal pedestrian re-identification model, which includes a parallel modal alignment module and a spatial feature extraction module; S2: Perform data augmentation on the pedestrian images input into the cross-modal pedestrian re-identification model; S3: The cross-modal pedestrian re-identification model extracts the global features, spatial local features, and channel local features of the pedestrian images; S4: Calculate the off-center sample loss and identity loss of the extracted spatial local features, and calculate the center cluster loss and identity loss of the extracted channel local features; S5: Concatenate the global features, spatial local features, and channel local features as the total feature, and calculate the total feature loss; S6: Add all the losses in steps S4 and S5 to form the loss function of the overall cross-modal pedestrian re-identification model, and train and optimize the parameters in the cross-modal pedestrian re-identification model according to this loss function; S7: After the cross-modal pedestrian re-identification model is trained, input the pedestrian images to be queried and the images in the test set into the cross-modal pedestrian re-identification model, calculate the similarity between them, and return the M values with the highest similarity, which are the results of cross-modal pedestrian re-identification.
3. The cross-modal pedestrian re-identification method based on spatio-temporal features and off-center loss according to claim 2, wherein In step S2, the data augmentation of the pedestrian images is specifically: mixing the visible light images and infrared images to form several batches.