A cross-appearance person re-identification method based on multimodal information
By introducing multimodal information, including pedestrian edge and component semantic prior information, in the pedestrian re-identification model, the problem of cross-appearance pedestrian matching is solved, and the recognition ability and retrieval performance of the model are improved.
Patent Information
- Application Number
- CN202210820445.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-13
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-07-13
AI Technical Summary
The existing pedestrian re-identification technology is difficult to effectively deal with cross-appearance pedestrian matching problems across time, across cameras, and across scenes, and the model recognition ability is not ideal.
The cross-appearance pedestrian re-identification method based on multimodal information is adopted. By introducing the pedestrian edge and component semantic prior information extracted by the pre-trained network, the dependence on traditional features is reduced, and the detailed information in the visual image and high-level semantic information that is robust to the appearance is fused.
It improves the recognizability of pedestrians across appearances, improves the search performance of pedestrian re-identification model, and enhances the feature extraction ability that is robust to appearance changes.
Smart Images

Figure CN115376159B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural networks, and in particular relates to a cross-appearance pedestrian re-identification method based on multimodal information. Background Art
[0002] Person re-identification, also known as person retrieval, aims to solve the problem of pedestrian matching across time, cameras, and scenes. Given a pedestrian target of interest, an ideal person re-identification system should be able to identify the target pedestrian when it appears again at different times, locations, and devices. Existing person re-identification tasks mainly focus on the re-identification of pedestrians with the same appearance in a short period of time. There is a serious lack of methods for long-term and cross-appearance pedestrian re-identification with changes in appearance such as clothing and accessories. In fact, the application of cross-appearance pedestrian re-identification is extremely common: comparative identification of long-term missing persons, analysis of customer business behavior, etc.
[0003] At present, the public datasets of cross-appearance pedestrian re-identification collected in monitoring environments mainly include NKUP+ and PRCC, which contain 40217 and 33698 pedestrian images respectively. As for the research on cross-appearance pedestrian re-identification, some of the work focuses on studying the association between different parts in pedestrian images, such as face, top, pants, etc., and forms robust cross-appearance features by adjusting the feature fusion of local features and global features of different parts. Typical methods include CCAN, 2S-IDE, 3APF, etc. Another part of the work attempts to introduce prior information such as contour and posture that is robust to appearance changes into the network. Typical methods include SPT and FSAM. For example, the SPT algorithm samples the pedestrian contour map from the Cartesian coordinate system and converts it to the polar coordinate system with the center of the human body as the origin to obtain more refined contour features. Finally, the ASE attention mechanism is added to obtain a relatively complete and robust pedestrian identity feature. Existing pedestrian re-identification models often focus on pedestrian appearance information such as clothing color and texture, and the recognition ability of the model is not ideal. Summary of the invention
[0004] In response to the technical problems existing in the prior art, the present invention provides a cross-appearance pedestrian re-identification method based on multimodal information, which improves the recognizability of cross-appearance pedestrians by reducing the model's dependence on traditional features, and introduces pedestrian edge and component semantic prior information extracted by a pre-trained network into the network. The information of three different modalities enables the model to comprehensively learn the detail information in the visual image and the high-level semantic information that is robust to the appearance, effectively alleviating the problem of the network focusing too much on the appearance information of pedestrians, and improving the retrieval performance of the cross-appearance pedestrian re-identification model.
[0005] The technical solution adopted by the present invention is: a cross-appearance pedestrian re-identification method based on multimodal information, comprising the following steps:
[0006] Step 1: Use data augmentation strategies to preprocess the cross-appearance person re-identification dataset; the data augmentation strategies include: scaling, random horizontal flipping, padding, random cropping, mean and variance subtraction, and random erasing.
[0007] Step 2: Use the contour recognition network and semantic segmentation network pre-trained with public datasets to obtain the pedestrian contour image and part semantic image from the preprocessed image respectively.
[0008] The pre-trained contour recognition network and semantic segmentation network are used to extract contour images and part semantic images from the pre-processed pedestrian visual images, respectively. The images of three different modalities are all represented by RGB color images.
[0009] Step 3: Use three non-shared weights of contour feature extraction network models, visual feature extraction network models and semantic feature extraction network models to extract pedestrians' high-dimensional contour feature matrix, high-dimensional visual feature matrix and high-dimensional semantic feature matrix from contour images, visual images and component semantic images, respectively. This is performed by inputting data into the feature extraction network model to obtain the feature map output before the classification layer of the network model.
[0010] Step 4: Concatenate the high-dimensional contour feature matrix, high-dimensional visual feature matrix, and high-dimensional semantic feature matrix into a fused feature matrix. The concatenation method is used to fuse the features of different modal information. Without adding additional parameters and training time required by methods such as attention mechanisms, the retrieval characteristics of different modal features in different focus directions can be integrated to comprehensively improve the cross-appearance retrieval capability of the model.
[0011] The fused feature matrix incorporates a variety of prior information that is robust to appearance changes. For long-term, cross-appearance pedestrian re-identification problems, cross-appearance pedestrian matching often fails due to excessive appearance-sensitive information such as clothing and accessories in the visual image. The pedestrian's contour information is actually mainly manifested as the pedestrian's edge information. Since the pedestrian's posture generally does not change drastically, it has a certain degree of robustness. At the same time, the semantic information of human body parts can obtain fine-grained pedestrian area information to avoid the influence of color and problems on the extraction of cross-appearance pedestrian features. The present invention comprehensively considers the prior knowledge such as contours and component semantics that are robust to changes in pedestrian appearance in the image, and improves the problem of using only a single visual modality information in the previous network, so that the network learns the correlation between three different modality features end-to-end, and improves the cross-appearance pedestrian retrieval effect.
[0012] Step 5: Perform pooling and downsampling on the high-dimensional contour feature matrix, high-dimensional visual feature matrix, high-dimensional semantic feature matrix and fusion feature matrix to obtain high-dimensional contour features, high-dimensional visual features, high-dimensional semantic features and fusion features respectively; use generalized mean pooling to downsample different modalities and their fusion features, which combines the advantages of maximum pooling and average pooling, so that the model can focus on significant features in different modal images and improve the retrieval effect of the model.
[0013] Step 6: For high-dimensional contour features, high-dimensional visual features, high-dimensional semantic features, and fusion features, batch normalization and fully connected layers are used to obtain high-dimensional contour classification features, high-dimensional visual classification features, high-dimensional semantic classification features, and fusion classification features, respectively.
[0014] Step 7: Calculate the most difficult ternary loss of high-dimensional contour features, high-dimensional visual features, high-dimensional semantic features, and fusion features respectively, and then calculate the identity classification loss of high-dimensional contour classification features, high-dimensional visual classification features, high-dimensional semantic classification features, and fusion classification features respectively, and then perform weighted summation to obtain the total loss.
[0015] Among them, the most difficult three losses are:
[0016]
[0017] Among them, α represents the interval parameter, D represents the distance metric, represents the kth image of the pth person in the batch High-dimensional features, 1≤p≤P, 1≤k≤K, p′ is the p′th person, k′ is the k′th image;
[0018] Identity Classification Loss:
[0019]
[0020] where x i ,y i denotes the image and its identity category respectively, p(y I |x i ) represents the image x i Recognized by the model as identity category y i The probability is 1≤i≤N.
[0021] The multimodal network model calculates the losses of each branch of vision, contour, component semantics, and fusion features end-to-end, where each branch calculates the most difficult triple loss and identity classification loss. Branch loss:
[0022] L=λ1L HardTri +λ2L ID
[0023] Among them, λ1 and λ2 represent the weight parameters of the hardest triplet loss and identity classification loss, respectively; λ1 and λ2 are both 1.0.
[0024] The total loss is the sum of the four branch losses of contour, vision, part semantics and fusion features.
[0025] The pedestrian identity classification loss and metric learning loss are calculated for the pedestrian's high-dimensional visual, contour, component semantic features and fusion features, thereby strengthening the loss function's guidance on the learning of different branch features, so that each branch feature has a certain representation ability, and ultimately improves the robust retrieval effect of the fusion feature.
[0026] Step 8: Loss layer gradient back propagation, update the weight parameters of the three non-weight-sharing contour feature extraction network models, visual feature extraction network models, and semantic feature extraction network models and their fully connected layers. The contour recognition network and semantic segmentation network do not participate in the weight update.
[0027] Step 9: Repeat steps 2-8 until the contour feature extraction network model, the visual feature extraction network model, and the semantic feature extraction network model converge, or the maximum number of iterations is reached, and the model training is completed.
[0028] Step 10: Input the query image and the gallery image into the trained model, and use the fused inference feature as the pedestrian feature representation for retrieval. The fused inference feature is obtained by batch normalization of the fused feature. Complete the evaluation and visualization of pedestrian re-identification, and calculate the top 1, 5, 10 hit rates (Rank1, Rank5, Rank10) and the average retrieval precision mAP to prove the role of multimodal information in promoting pedestrian retrieval.
[0029] Compared with the prior art, the present invention has the following beneficial effects: the strategy of fusing multimodal prior information proposed in the present invention can reduce the weight of appearance-sensitive information in the features of a single visual RGB image, and the fused two modal information that are relatively robust to appearance changes can promote the network to learn pedestrian features that are robust to appearance, and ultimately promote the model's pedestrian retrieval performance in cross-appearance scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a flow chart of an embodiment of the present invention;
[0031] Figure 2 A network structure diagram of the fusion branch loss according to an embodiment of the present invention;
[0032] Figure 3 This is a flow chart of a test of an embodiment of the present invention;
[0033] Figure 4Schematic diagram showing images of three different modalities used in the embodiments of the present invention;
[0034] Figure 5 This is a schematic diagram of the top ten search results of some pedestrians on NKUP+ using the benchmark network according to an embodiment of the present invention;
[0035] Figure 6 It is a schematic diagram of the top ten search results of some pedestrians on NKUP+ according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0037] The embodiment of the present invention provides a method for cross-appearance person re-identification based on multimodal information, such as Figure 1 As shown, it includes the following steps:
[0038] Step 1: Preprocess the cross-appearance person re-identification dataset. The images in the training set need to be processed by the data augmentation strategy and normalized before being used as the network input. The preprocessing order is as follows: 1) adjust the image size to the network input size (256*128); 2) randomly flip the image horizontally with a probability of 50%; 3) fill 10 pixels with a value of 0 around the image; 4) randomly crop an image of the network input size (256*128) from the image;
[0039] 5) The image is normalized by subtracting the mean and dividing the variance, using the mean (0.485, 0.456, 0.406) and variance (0.229, 0.224, 0.225) of the images in ImageNet; 6) The area of 2% to 40% of the area in the image is randomly erased with a probability of 50%. When testing the model, only the above operations 1) and 5) are used to process the images of the model set.
[0040] The cross-appearance person re-identification datasets mainly include NKUP+ and PRCC, which contain 40217 and 33698 pedestrian images respectively.
[0041] Table 1 Attribute statistics of NKUP+ dataset
[0042]
[0043] Table 2 PRCC dataset attribute statistics
[0044]
[0045] Step 2: Use the contour recognition network R (RCF Net) and semantic segmentation network P (PSP Net) trained on the public contour recognition dataset (BSDS500) and pedestrian semantic segmentation dataset (LIP) to extract the contour recognition network R (RCF Net) and semantic segmentation network P (PSP Net) from the visual image X of the pedestrian RGB Extract the contour image X C and component semantic image X P , the images of the three different modalities are all represented by RGB color images, so they have the same dimension. The example images of different modalities are as follows Figure 4 shown.
[0046] X C =R(X RGB ), X P =P(X RGB )
[0047] Step 3: Use three non-shared weights trained on a public dataset (ImageNet) to extract the feature of the Densenet121 network model: Contour feature extraction network model N C , Visual feature extraction network model N RGB And semantic feature extraction network model N P Extract high-dimensional feature matrices of pedestrian vision, contour and component semantics from contour images, visual images and component semantic images respectively: high-dimensional contour feature matrix High-dimensional visual feature matrix and a high-dimensional semantic feature matrix
[0048]
[0049] Step 4: Concatenate the high-dimensional feature matrices of three different modal information, namely pedestrian vision, contour and component semantics, into a fusion feature matrix
[0050]
[0051] Step 5: Based on Generalized Mean Pooling (GeM Pooling), the high-dimensional profile feature matrix High-dimensional visual feature matrix High-dimensional semantic feature matrix And the fusion feature matrix Downsample to the corresponding high-dimensional features: high-dimensional contour features High-dimensional visual features High-dimensional semantic features and fusion features
[0052]
[0053]
[0054] Step 6: High-dimensional contour features of pedestrians High-dimensional visual features High-dimensional semantic features and fusion features First, batch normalization (BN) is used to obtain inference features:
[0055] High-dimensional profile inference features High-dimensional visual reasoning features High-dimensional semantic reasoning features and fusion reasoning features Then use the fully connected layer (FC) to obtain identity classification features: high-dimensional profile classification features High-dimensional visual classification features High-dimensional semantic classification features and fusion classification features
[0056]
[0057]
[0058] Step 7: Calculate the overall branch loss L for each of the visual, contour, component semantic, and fusion features RGB , L C , L P , L F , and then sum the losses of different branches to get the final total loss L All .
[0059]
[0060]
[0061]
[0062]
[0063] L All =L RGB +L C +L P +L F
[0064] Among them, λ1 and λ2 represent the weight parameters of the hardest triplet loss and identity classification loss, respectively; λ1 and λ2 are both 1.0.
[0065] The most difficult three-way loss:
[0066]
[0067] Among them, α represents the interval parameter, D represents the distance metric, represents the kth image of the pth person in the batch High-dimensional features, 1≤p≤P, 1≤k≤K, p′ is the p′th person, k′ is the k′th image;
[0068] Identity Classification Loss:
[0069]
[0070] where x i ,y i denotes the image and its identity category respectively, p(yi|x i ) represents the image x i Recognized by the model as identity category y i The probability is 1≤i≤N.
[0071] The network structure of the fusion branch loss is as follows Figure 2 As shown in Figure 2, the network structure of the branch loss for vision, contour, and part semantics is similar.
[0072] Step 8: Back propagate the gradient of the loss layer and update the contour feature extraction network model N C , Visual feature extraction network model N RGB And semantic feature extraction network model N P , and the weight parameters of its corresponding fully connected layer.
[0073] Step 9: The multimodal model is optimized and trained on the person re-identification dataset for 120 rounds, and the initial learning rate of the network is 3.5×10 -6 In the first 10 epochs, the network learning rate will increase linearly to 3.5×10 -4 , then, the learning rate will decay to 0.1 times of the current value in rounds 31, 61, and 91 respectively to fine-tune the network weights. The model training is completed and a trained multimodal model is obtained.
[0074] Step 10: For the network testing process, Figure 3 All query images and gallery images in the test set are input into the multimodal model for forward propagation, and the normalized inference features of the fused features are used. As the final pedestrian feature vector representation. Assume that the feature representation of the query image is fq , the feature representation of the candidate image is f g , use the Euclidean distance to calculate the distance d between the two q,g =||F Q -F g ||2, if the distance is smaller, the similarity between the image pairs is higher, otherwise it is lower. Calculate the distance between each query image and all candidate images and sort them from large to small according to the similarity to obtain a sorted list, and finally calculate the top k hit rate Rank-k and average retrieval accuracy mAP. Comparative experiments are conducted on the NKUP+ and PRCC datasets to prove the robustness of multimodal fusion features.
[0075] Figure 5 and Figure 6 The results of some pedestrian re-identification of the baseline network model Densenet121 and the multimodal model M2Net in the NKUP+ cross-appearance subset are shown. The top ten search results of a pedestrian to be searched are shown in each row. The leftmost one is the search image, and the query images are arranged from high to low according to the similarity. The black and gray bounding boxes represent the correct and incorrect search results respectively. As can be seen from the figure, the appearance information such as clothing and backpacks of pedestrians in the search results of the baseline network model (Densenet121) greatly affects the search results. After adopting the multimodal model M2Net, some images with obvious changes in the appearance of pedestrians are also retrieved, which confirms that multimodal information can improve the performance of the cross-appearance pedestrian re-identification model.
[0076] Tables 3 and 4 quantify the experimental Rank-k and mAP indicators, which are two important evaluation criteria in the field of person re-identification. In the PRCC dataset with a relatively small number of images and little appearance change, the features extracted by the multimodal model M2Net improved the Rank1 value of 0.7% / 7.5% and the mAP accuracy of 1.7% / 6.1% on the same / cross-appearance subsets respectively; while in the NKUP+ dataset with a large number of images and obvious appearance changes, the multimodal network M2Net improved the Rank1 value of 1.6% and the mAP of 0.7% on the cross-appearance subset while keeping the same appearance retrieval ability basically unchanged, proving the retrieval ability of multimodal features for cross-appearance pedestrians.
[0077] Table 3 Comparison of retrieval indicators of each feature extraction network in PRCC dataset
[0078]
[0079] Table 4. Comparison of retrieval indicators of each feature extraction network in NKUP+ dataset
[0080]
[0081] The present invention has been described in detail above through embodiments, but the contents described are only exemplary embodiments of the present invention and cannot be considered to limit the scope of implementation of the present invention. The protection scope of the present invention is defined by the claims. Anyone who utilizes the technical solution described in the present invention, or a technician in the field, inspired by the technical solution of the present invention, designs a similar technical solution within the essence and protection scope of the present invention to achieve the above technical effects, or makes equal changes and improvements to the scope of application, etc., shall still fall within the scope of protection covered by the patent of the present invention.
Claims
1. A cross-appearance person re-identification method based on multimodal information, characterized by: The following steps are involved: Step 1: Preprocess the cross-appearance person re-identification dataset using data augmentation strategies; Step 2: Use the pre-trained contour recognition network and semantic segmentation network to obtain the pedestrian contour image and component semantic image from the pre-processed image respectively; Step 3: A contour feature extraction network model, a visual feature extraction network model, and a semantic feature extraction network model with non-shared weights are used to extract a high-dimensional contour feature matrix, a high-dimensional visual feature matrix, and a high-dimensional semantic feature matrix of pedestrians from the contour image, the visual image, and the component semantic image, respectively; Step 4: Concatenate the high-dimensional contour feature matrix, the high-dimensional visual feature matrix, and the high-dimensional semantic feature matrix into a fused feature matrix; Step 5: Perform pooling downsampling on the high-dimensional contour feature matrix, high-dimensional visual feature matrix, high-dimensional semantic feature matrix and fusion feature matrix to obtain high-dimensional contour features, high-dimensional visual features, high-dimensional semantic features and fusion features respectively; Step 6: For high-dimensional contour features, high-dimensional visual features, high-dimensional semantic features, and fusion features, batch normalization and fully connected layers are used to obtain high-dimensional contour classification features, high-dimensional visual classification features, high-dimensional semantic classification features, and fusion classification features, respectively; Step 7: Calculate the most difficult ternary loss of high-dimensional contour features, high-dimensional visual features, high-dimensional semantic features, and fusion features respectively, and then calculate the identity classification loss of high-dimensional contour classification features, high-dimensional visual classification features, high-dimensional semantic classification features, and fusion classification features respectively, and then perform weighted summation to obtain the total loss; Step 8: Back-propagate the gradient of the loss layer to update the weight parameters of the contour feature extraction network model, the visual feature extraction network model, the semantic feature extraction network model and their fully connected layers; Step 9: Repeat steps 2-8 until the contour feature extraction network model, the visual feature extraction network model, and the semantic feature extraction network model converge, or the maximum number of iterations is reached, and the model training is completed; Step 10: The query image and the gallery image are input into the trained model, and the fused inference features are used as the pedestrian feature representation for retrieval. The fused inference features are obtained by fusion features using batch normalization.
2. The cross-appearance person re-identification method based on multimodal information according to claim 1, characterized in that: In step 1, the data augmentation strategies include: scaling, random horizontal flipping, padding, random cropping, mean and variance reduction, and random erasing.
3. The cross-appearance person re-identification method based on multimodal information according to claim 1, characterized in that: In step 2, the pre-trained contour recognition network and semantic segmentation network are used to extract contour images and component semantic images from the pre-processed pedestrian visual images, respectively. The images of three different modalities are all represented by RGB color images.
4. The cross-appearance person re-identification method based on multimodal information according to claim 1, characterized in that: In step 7, the hardest ternary loss is: Among them, α represents the interval parameter, D represents the distance metric, represents the kth image of the pth person in the batch High-dimensional features, 1≤p≤P, 1≤k≤K, p′ is the p′th person, k′ is the k′th image; Identity Classification Loss: where x i ,y i denotes the image and its identity category respectively, p(y i |x i ) represents the image x i Recognized by the model as identity category y i The probability is 1≤i≤N.
5. The cross-appearance person re-identification method based on multimodal information according to claim 4, characterized in that: Branch loss: L=λ1L HardTri +λ2L ID Among them, λ1 and λ2 represent the weight parameters of the hardest triplet loss and identity classification loss respectively; The total loss is the sum of the four branch losses of contour, vision, part semantics and fusion features.
6. The cross-appearance person re-identification method based on multimodal information according to claim 5, characterized in that: Both λ1 and λ2 are 1.0.
Citation Information
Patent Citations
Cross-modal pedestrian re-identification method based on multi-modal image style conversion
CN111539255A
Cross-modal pedestrian re-recognition method based on dual attribute information
CN112001279A