Method and device for re-identifying a dressing pedestrian guided by a black clothing and a head image

Through the black-clothing and head image guidance method, a re-identification network of pedestrians with three network branches was constructed, which solved the problems of long image generation time and impact of clothing color in the existing methods, and achieved more robust feature extraction and higher recognition accuracy.

CN115620338BActive Publication Date: 2025-08-05HENAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211258905.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-14
Publication Date
2025-08-05
Estimated Expiration
2042-10-14

AI Technical Summary

Technical Problem

The existing method of re-identification of pedestrians with changing clothes requires a lot of image generation. It is affected by the different colors of the clothes, making it difficult to extract robust features that are independent of the clothes, and ignores the influence of head features. Traditional networks lead to feature loss.

Method used

Using the black clothes and head image guidance method, the clothes of all pedestrians tend to be consistent through the cloaking strategy. Using the improved Transformer network, a re-identification network of dress changers from three network branches was constructed, and the original, black clothes and head characteristics were learned respectively, and feature extraction was optimized through knowledge distillation and triple loss.

Benefits of technology

It effectively reduces the image generation time, improves the robustness and robustness of the model, can extract more discriminant fine-grained features, and improves the accuracy of re-identification of pedestrians with changing clothes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115620338B_ABST
    Figure CN115620338B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for re-identifying dressing pedestrians guided by black clothes and head images. The method includes: First, the original image is processed using a newly designed method for occluding clothes to obtain the corresponding black clothes image, and the obtained black clothes image is placed in the black clothes branch for pre-training; then, joint learning is performed on the framework, the original pedestrian image is placed in the original branch, and the pre-trained black clothes branch is used to guide the original branch to learn; at the same time, the pedestrian head image is placed in the head branch to obtain more fine-grained pedestrian features. The present invention occludes the clothes of all pedestrians to obtain black clothes images, making the clothes colors of pedestrians unified, so that the model pays more attention to parts other than the clothes colors, thereby improving the robustness of the model; and it can effectively utilize the information in the original image, effectively reduce the information loss in the process of obtaining black clothes images, and improve the robustness of features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pedestrian re-identification, and particularly to a method and device for re-identifying undressing pedestrians guided by black clothes and head images. Background Art

[0002] The purpose of pedestrian re-identification is to solve the problem of pedestrian retrieval under different conditions, such as different cameras, different lighting, or different observation angles. There are various methods for pedestrian re-identification research, such as sub-fields like lightweight networks, domain generalization, unsupervised learning, etc., which have achieved good results in recent years. However, these methods generally assume that a person's clothes remain consistent for a long time.

[0003] However, in the real world, people's clothes do not remain the same. For example, people always wear different clothes for a long time, and some suspects may change their clothes in a short time to avoid being tracked. Therefore, a different version of the pedestrian re-identification problem has been proposed, which is called long-term undressing pedestrian re-identification and has become a hot issue today. In the long run, undressing pedestrian re-identification is a hot issue today. The core of solving undressing pedestrian re-identification is to extract discriminative relevant features that are only related to identity. To remove the interference of clothes, researchers usually adopt two general strategies.

[0004] The first is the data strategy. A common method is to construct a large-scale dataset, in which each person should have multiple pictures with a large number of different clothes, and then force the model to learn features unrelated to clothes from these pictures. However, it is very difficult and almost impossible to construct such an undressing dataset purely by manpower. Therefore, some researchers use GAN or other methods to expand the original dataset.

[0005] The second is the feature separation strategy. A common operation is to separate clothing features from other identity features. By doing so, other features except clothes can be used for identity judgment. For example, Yang et al. used the silhouette of pedestrians as queries and galleries, and used polar coordinates to better obtain the silhouette features of pedestrians. However, although learning from the silhouette can obtain features unrelated to clothes, it also abandons a large part of features unrelated to clothes (such as the head). In addition, Hong et al. proposed using an appearance branch and a shape branch to extract fine-grained features. However, this method is often affected by different colors of clothes and cannot extract more robust features unrelated to clothes.

[0006] Main problems existing in the prior art:

[0007] 1. Existing methods for re-identifying undressing pedestrians require a large amount of image generation work and a long training time.

[0008] 2. The existing methods for re-identifying pedestrians with changing clothes are often affected by different colors of clothes and cannot extract more robust features that are independent of clothes.

[0009] 3. Most of the existing methods for re-identifying pedestrians with changing clothes adopt traditional convolutional neural networks as the training network. Due to downsampling and pooling, traditional convolutional neural networks bring certain losses.

[0010] 4. Most of the existing methods for re-identifying pedestrians with changing clothes ignore the influence of head features on overall judgment. Summary of the Invention

[0011] In view of the above problems in the background technology, the present invention proposes a method and device for re-identifying pedestrians with changing clothes guided by black clothes and head images. A non-GAN method is used to extract features independent of clothes in the image. A new strategy for occluding clothes is proposed to make the clothes of all pedestrians tend to be the same, forcing the model to learn robust features independent of clothes. An improved Transformer is adopted as the training network, and a head branch is designed separately to obtain the fine-grained head features of the original image.

[0012] To achieve the above object, the present invention adopts the following technical solutions:

[0013] On the one hand, the present invention proposes a method for re-identifying pedestrians with changing clothes guided by black clothes and head images, including:

[0014] Step 1: Delete the clothes features from the original pedestrian image to obtain a black-clothed pedestrian image.

[0015] Step 2: Use the pre-trained HRNet to process the original pedestrian image to obtain the pedestrian head image in the original pedestrian image.

[0016] Step 3: Construct a network for re-identifying pedestrians with changing clothes, which consists of three network branches, namely the original branch, the black-clothed branch, and the head branch, for learning the features of the original pedestrian image, the black-clothed pedestrian image, and the pedestrian head image respectively; the main network structures of the three network branches are the same, but the parameters are not shared.

[0017] Step 4: Input the black-clothed pedestrian image obtained in Step 1 into the black-clothed branch for training to obtain a pre-trained black-clothed branch.

[0018] Step 5: Train the original branch under the guidance of the pre-trained black-clothed branch to obtain features related to the pedestrian but independent of clothes in the original pedestrian image.

[0019] Step 6: Input the pedestrian head image obtained in Step 2 into the head branch for learning, and combine the learned pedestrian head image features with the features related to the pedestrian but not related to the clothes obtained in Step 5 to obtain the overall features related to the pedestrian but not related to the clothes, thus completing the training of the clothing-changing pedestrian re-identification network;

[0020] Step 7: Perform clothing-changing pedestrian re-identification based on the trained clothing-changing pedestrian re-identification network.

[0021] Further, Step 1 includes:

[0022] Use a pre-trained human parsing model to obtain the images of each body part of the pedestrian in the original pedestrian image, and recombine the obtained images of each body part to get six parts: background, head, upper garment, trousers, arms, and legs. Extract the pixels of the upper garment and trousers images to form the clothing area; set all the pixels in the clothing area to zero to obtain a black clothing pedestrian image.

[0023] Further, the backbone network of the three network branches is the imViT network. In each imViT network, there are two outputs for extracting global features and local features respectively. The backbone network is optimized by triplet loss and ID loss on global features and local features respectively, where the ID loss is cross-entropy loss without label smoothing.

[0024] Further, Step 5 includes:

[0025] According to the knowledge distillation algorithm, train the branch of the original image under the guidance of the pre-trained black clothing branch, and use the mean squared error loss to regularize the training of the original branch to train more features related to identity but not related to clothes.

[0026] On the other hand, the present invention proposes a clothing-changing pedestrian re-identification device guided by black clothing and head images, including:

[0027] Black clothing image obtaining module, used to delete the clothing features from the original pedestrian image to obtain a black clothing pedestrian image;

[0028] Head image obtaining module, used to process the original pedestrian image with a pre-trained HRNet to obtain the pedestrian head image in the original pedestrian image;

[0029] Clothing-changing pedestrian re-identification network construction module, used to construct a clothing-changing pedestrian re-identification network, which consists of three network branches, namely the original branch, the black clothing branch, and the head branch, used to learn the features of the original pedestrian image, the black clothing pedestrian image, and the pedestrian head image respectively; the backbone network structures of the three network branches are the same but do not share parameters;

[0030] The black-clothes branch training module is used to input the black-clothed pedestrian images obtained by the black-clothes image obtaining module into the black-clothes branch for training, and obtain a pre-trained black-clothes branch;

[0031] The original branch training module is used to train the original branch under the guidance of the pre-trained black-clothes branch, and obtain features related to the pedestrian but not related to the clothes in the original pedestrian image;

[0032] The head branch training module is used to input the pedestrian head images obtained by the head image obtaining module into the head branch for learning, and combine the learned pedestrian head image features with the features related to the pedestrian but not related to the clothes obtained by the original branch training module, to obtain overall features related to the pedestrian but not related to the clothes, and complete the training of the clothing-changing pedestrian re-identification network;

[0033] The clothing-changing pedestrian re-identification module is used to perform clothing-changing pedestrian re-identification based on the trained clothing-changing pedestrian re-identification network.

[0034] Furthermore, the black-clothes image obtaining module specifically is used for:

[0035] Adopt a pre-trained human parsing model to obtain the images of each body part of the pedestrian in the original pedestrian image, and recombine the obtained images of each body part to obtain six parts: background, head, upper garment, trousers, arms and legs, extract the pixels of the upper garment and trousers images from them to form a clothing area; set all the pixels in the clothing area to zero to obtain a black-clothed pedestrian image.

[0036] Furthermore, the backbone networks of the three network branches are all imViT networks. In each imViT network, there are two outputs respectively used to extract global features and local features. The backbone network is optimized by triplet loss and ID loss on the global features and local features respectively, where the ID loss is cross-entropy loss without label smoothing.

[0037] Furthermore, the original branch training module specifically is used for:

[0038] According to the knowledge distillation algorithm, train the branch of the original image under the guidance of the pre-trained black-clothes branch, and use the mean square error loss to regularize the training of the original branch, so as to train more features related to the identity but not related to the clothes.

[0039] Compared with the prior art, the beneficial effects of the present invention are:

[0040] 1. Compared with the method of using GAN to generate a large number of clothing-changing images to expand the dataset, the present invention uses a non-GAN method. By using a human semantic parsing model to obtain the upper and lower body parts of a human body and applying the method of the present invention to occlude them, this method enables the model to more intensively learn features unrelated to clothing without expanding the dataset, saving both space and time.

[0041] 2. Compared with other methods that separate clothing features and identity features, the present invention occludes the clothing of all pedestrians to obtain black-clothed images, thus making the clothing colors of pedestrians uniform, enabling the model to pay more attention to parts other than clothing colors, and improving the robustness of the model.

[0042] 3. Compared with directly learning features from images after excluding clothing, the present invention uses a black-clothed branch to guide the original branch to directly learn identity features unrelated to clothing from the original RGB images, which can effectively utilize the information in the original images, effectively reduce information loss during the acquisition of black-clothed images, and improve the robustness of the features.

[0043] 4. Compared with the method of directly extracting clothing-unrelated features from the original images, the present invention can extract more discriminative fine-grained features by adding specialized processing of head image patches, which has a good complementary effect on global features.

[0044] 5. The test results on the PRCC dataset show that the method of the present invention achieves excellent results in re-identifying pedestrians with clothing changes. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 FIG. is a basic flowchart of a method for re-identifying pedestrians with clothing changes guided by black clothing and head images according to an embodiment of the present invention;

[0046] Figure 2 FIG. is a schematic diagram of the network architecture for re-identifying pedestrians with clothing changes constructed according to an embodiment of the present invention;

[0047] Figure 3 FIG. is a comparison example of feature activation maps obtained according to an embodiment of the present invention;

[0048] Figure 4 FIG. is a schematic diagram of the structure of a device for re-identifying pedestrians with clothing changes guided by black clothing and head images according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] The following further explains the present invention in conjunction with the drawings and specific embodiments:

[0050] As Figure 1 [[ID=4,0]]shown, a method for re-identifying pedestrians with clothing changes guided by black clothing and head images includes:

[0051] Step 1: Remove the clothing features from the original pedestrian image to obtain a black-clothed pedestrian image (abbreviated as the black-clothed image).

[0052] Step 2: Process the original pedestrian image using the pre-trained HRNet to obtain the pedestrian head image in the original pedestrian image.

[0053] Step 3: Construct a cross-dressing pedestrian re-identification network, which consists of three network branches, namely the original branch, the black-clothed branch, and the head branch, as Figure 2 shown, which are used to learn the features of the original pedestrian image, the black-clothed pedestrian image, and the pedestrian head image respectively; the backbone network structures of the three network branches are the same, but they do not share parameters to better utilize the discriminative feature space.

[0054] Step 4: Input the black-clothed pedestrian image obtained in Step 1 into the black-clothed branch for training to obtain the pre-trained black-clothed branch.

[0055] Step 5: Train the original branch under the guidance of the pre-trained black-clothed branch to obtain the features related to the pedestrian but not related to the clothing in the original pedestrian image.

[0056] Step 6: Input the pedestrian head image obtained in Step 2 into the head branch for learning, and combine the learned pedestrian head image features with the features related to the pedestrian but not related to the clothing obtained in Step 5 to obtain the overall features related to the pedestrian but not related to the clothing, and complete the training of the cross-dressing pedestrian re-identification network.

[0057] Step 7: Perform cross-dressing pedestrian re-identification based on the trained cross-dressing pedestrian re-identification network.

[0058] Furthermore, Step 1 includes:

[0059] Adopt a pre-trained human parsing model to obtain the images of each body part of the pedestrian in the original pedestrian image, and recombine the obtained images of each body part to get six parts: background, head, upper body, pants, arms, and legs. Extract the pixels of the upper body and pants images to form the clothing area; set all the pixels in the clothing area to zero to obtain a black-clothed pedestrian image.

[0060] Furthermore, the backbone network of the three network branches is the imViT network. In each imViT network, there are two outputs respectively used to extract global features and local features. The backbone network is optimized by the triplet loss and the ID loss on the global features and local features respectively, where the ID loss is the cross-entropy loss without label smoothing.

[0061] Furthermore, Step 5 includes:

[0062] According to the knowledge distillation algorithm, the branch of the original image is trained under the guidance of the pre-trained black-clothes branch, and the mean squared error loss is used to regularize the training of the original branch to train more identity-related but clothing-unrelated features.

[0063] Specifically, the method includes:

[0064] 1. Obtain black-clothes images

[0065] To eliminate the influence of clothes in feature extraction, we removed the clothing features from the original images to obtain black-clothed pedestrian images. First, we used HRNet, a pre-trained human parsing model, to obtain images of each body part. The prediction results of this model divide the body into 20 parts. Since this model divides the body into 20 parts, we recombined them to obtain six parts: background, head, upper body, pants, arms, and legs. We extracted the upper body and pants to form the clothing region. We only used two of them (upper body, pants).

[0066] First, for a batch of samples x i [i = 1....B], where B is the batch size, x i is actually an image, x i ∈R H×W×C , where H, W, and C represent its height, width, and number of channels respectively. First, we represent the semantic map obtained by parsing x i through the human parsing model as s i [i = 1....B], s i ∈R 1×H×W . The pixel values of s i are defined as s i ∈{0,1,2,3,4,5} representing six parts of the body respectively. Each pixel of s i is defined as 0, 1, 2, 3, 4, or 5, representing six parts of the body respectively.

[0067] Secondly, obtain the pixels of the upper body and pants. Each pixel of x i is defined as v j , with c values. Each input sample x i has a total of W×H pixel vectors. And we represent all the pixel vectors of the upper body and pants in x i as:

[0068] B(clothes and pants) = {v j |v j = x i [s i == 2 || s i==3], i ∈ [1, B], j ∈ [1, N]}(1)

[0069] where N is the total number of trousers and coats in each x i and is different in each x i ; j represents the j th pixel vector, s i represents the semantic segmentation map, 2 represents the index of the coat, 3 represents the index of the trousers, x i [s i ==2 || s i ==3] represents the coats and trousers of the x i pixel vector. Therefore, the original image x i can be regarded as x i =[v1, v2... v c1 ,..., v cn ,... v last , where [v c1 ... v cn belongs to the pixels of the trousers and coats obtained from formula (1).

[0070] Finally, the black-clothed image is obtained by setting the pixels of the trousers and coats to zero. Specifically, we set [v c1 ... v cn =0, and then we can get x i′ =[v1, v2... 0,... v last , which is the corresponding black-clothed image of x i .

[0071] 2. Backbone Networks of Each Branch

[0072] We select imViT (see specifically [He, Shuting, et al. "Transreid: Transformer-based object re-identifification." Proceedings of the IEEE / CVF International Conference on Computer Vision. 2021. 1, 2, 4]) as the backbone network for each of our branches. In each imViT, there are two outputs for extracting global and local features. For the local part, we can choose how many local sub-features the local features in the network consist of, and 4 local sub-features are selected in the experiments of this invention. Therefore, for the i-th (i ∈ {0, 1, 2}) branch, we can obtain the local feature F li and the global feature F gi , where F li =[Fli1 , F li2 , F li3 , F li4 , i ∈ {0, 1, 2}。

[0073] The backbone network is optimized by triplet loss and ID loss on global and local features respectively. The ID loss is the cross-entropy loss without label smoothing. As for the triplet loss L tri is as follows:

[0074] L tri = log(1 + exp(||f a - f p ||² - ||f a - f n ||²))

[0075] where f a represents the anchor, f p represents the positive sample, and f n represents the negative sample. Therefore, each branch of our framework has two sets of loss functions, namely L tri-gi , L id-gi , and L tri-li , L id-li , where L tri-gi , L id-gi are the triplet loss and ID loss corresponding to the global feature F gi respectively, and L tri-li , L id-li are the triplet loss and ID loss corresponding to the local feature F li respectively.

[0076] 3. The original branch guided by the black clothing branch

[0077] The black clothing branch can learn features irrelevant to the clothing. However, some important discriminative information may be discarded during the process of obtaining the black clothing image, and this information is hidden in the original image. And it is not feasible to directly extract identity features irrelevant to the clothing from the original image. According to the knowledge distillation algorithm, we train the branch of the original image under the guidance of the pre-trained black clothing branch. Specifically, we use the mean squared error (MSE) loss to regularize the training of the original branch to train more identity-related but clothing-irrelevant features. The mean squared error loss L mse is defined as

[0078]

[0079]

[0080] L mse = L mse-opg + L mse-opl

[0081] Among which, F l1 , F g1 respectively represent the local feature and the global feature obtained from the black branch. F l2 , F g2 respectively represent the connected local feature and the global feature obtained from the original branch guided by the black branch. L mse-opg , L mse-opl respectively represent the global feature F opg = [F g1 , F g2 and the mean square error loss corresponding to the local feature F opl = [F l1 , F l2 .

[0082] 4. Extraction of Head Features

[0083] For the input original image, we use the pre-trained HRNet to obtain the head image part. The obtained head image is put into imViT3 to obtain the merged local feature F l3 and the global feature F g3 , where the former is composed of four local features F l31 , F l32 , F l33 , F l34 merged together. Similarly, the original branch guided by the black branch can obtain the connected local feature F l2 and the global feature F g2 . To obtain more uncertain clothing-related features, we perform element-wise summation on F l2 and F l3 , F g2 and F g3 , and their definitions are as follows:

[0084] F ohg = wF g2 + (1 - w)F g3

[0085] F ohl = wF l2 + (1 - w)F l3

[0086] Among which, F ohg , F ohl respectively represent the global overall feature and the local overall feature related to the pedestrian but not related to the clothing. w is the weight coefficient, and w ∈ (0, 1).

[0087] And the triplet loss and the ID loss are applied to them, which are L id-ohg , L tri-ohg, L id-ohl and L tri-ohl , where L id-ohg , L tri-ohg Represents F ohg Corresponding ID loss, triple loss, L id-ohl , L tri-ohl Represents F ohl Corresponding ID loss and triplet loss.

[0088] 5. Joint training

[0089] There are two stages in training. First, we feed the black clothes image into the black clothes branch and and We use ID loss and triplet loss on Figure 1 These two losses are omitted in the black-cloth branch of the training set. Then we get the pre-trained black-cloth branch.

[0090] Second, we fix the learning weights of the black branch and jointly train the other branches. Therefore, the total loss function is defined as:

[0091] L total (θ)=l1L id (θ)+l2L tri (θ)+l3L mse (θ)

[0092] Among them, L id (θ) represents the ID loss, L tri (θ) represents the triplet loss, L mse (θ) represents the mean squared error loss. l1, l2, and l3 are trade-off parameters that balance the contribution of each loss. In the experiments of this paper, l1, l2, and l3 are set to 0.25, 0.25, and 0.5, respectively.

[0093] In order to verify the effect of the present invention, the following experiments were performed:

[0094] We conducted experiments on the PRCC dataset (PRCC includes 33,698 images of 221 people from three different angles, and also provides silhouette sketches to facilitate the extraction of outline information) in both different clothing settings and the same clothing setting. In the different clothing setting, images from camera A were used for the candidate gallery, while images from camera C were used for the query set. In the same clothing setting, gallery images were also from camera A, but query images were from camera B. The experimental results are shown in Table 1.

[0095] Table 1 Experimental results

[0096]

[0097]

[0098] We made some comparisons between the method of the present invention and some state-of-the-art methods on the PRCC dataset, including representative traditional person re-identification methods (PCB [Y. Sun, L. Zheng, Y. Yang, Q. Tian, and S. Wang, “Beyond part models: Person retrieval with refined part pooling (and a strong convolutional baseline),” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 480–496.], Zheng et al’s method [Z. Zheng, L. Zheng, and Y. Yang, “A discriminatively learned cnn embedding for person re-identification,” ACM transactions on multimedia computing, communications, and applications (TOMM), vol. 14, no. 1, pp. 1–20, 2017.], HPM [Y. Fu, Y. Wei, Y. Zhou, H. Shi, G. Huang, X. Wang, Z. Yao, and T. Huang, “Horizontal pyramid matching for person re-identification,” in Proceedings of the AAAI conference on artificial intelligence, vol. 33, no. 01, 2019, pp. 8295–8302.]), HACNN [W. Li, X. Zhu, and S. Gong, “Harmonious attention network for person re-identification,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 2285–2294.]) and the person re-identification method for people wearing different clothes (PRCC(sketch) [Q. Yang, A. Wu, and W.-S.Zheng, "Person re-identification by contour sketch under moderate clothing change," IEEE transactions on pattern analysis and machine intelligence, vol. 43, no. 6, pp. 2029–2046, 2019.],GI-ReID(OSNet)[X. Jin, T. He, K. Zheng, Z. Yin, X. Shen, Z. Huang, R. Feng, J. Huang, Z. Chen, and X.-S. Hua, "Cloth-changing person re-identification from a single image with gait prediction and regularization," in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 14 278–14 287.],LightMBN[F. Herzog, X. Ji, T. Teepe, S. J. Gilg, and G. Rigoll, "Lightweight multi-branch network for person re-identification," in 2021 IEEE International Conference on Image Processing (ICIP). IEEE, 2021, pp. 1129–1133.])。All the results of the comparison methods are from their published papers.

[0099] As can be seen from Table 1, the method of the present invention has achieved very good performance. On the same piece of clothing, although the method of the present invention is slightly inferior to LightMBN, it almost exceeds all other methods. This indicates that the method of the present invention has better generalization ability in pedestrian re-identification on the same clothing. In the setting of different clothes, the method of the present invention shows obvious excellent performance, with the mAP (mAP is used to evaluate the overall effect of the pedestrian re-identification algorithm; where AP refers to the average precision of a query sample, representing the effect of the model on a certain sample, and mAP is the average value of the APs of all query samples, representing the overall effect of the model on all query samples) exceeding LightMBN by 4.7%, and the accuracy at rank-1 (i.e., R@1 in Table 1) exceeding by 9.6%. This may be because the method of the present invention can extract more identity features independent of clothing.

[0100] For intuitive analysis, we randomly selected 3 images from PRCC and Figure 3 shows the corresponding activation maps (CAM) captured by the baseline and the method of the present invention. As can be seen from the first row, the activation regions of the baseline are concentrated on the clothing and the background, and many background and clothing regions are activated and used to identify a person, which may confuse the recognition results. In addition, the baseline pays less attention to the head. As can be found from the second row, the activation points in the method we proposed are more concentrated, and there are fewer background and clothing regions, which reduces the influence of the background and clothing on recognition. In addition, we also noticed that the model we proposed pays more attention to the shapes of the head and the body.

[0101] Based on the above embodiments, as Figure 4 shown, the present invention also proposes a clothing-changing pedestrian re-identification device guided by black clothing and head images, including:

[0102] A black clothing image obtaining module, configured to delete clothing features from the original pedestrian image to obtain a black-clothed pedestrian image;

[0103] A head image obtaining module, configured to process the original pedestrian image using a pre-trained HRNet to obtain the pedestrian head image in the original pedestrian image;

[0104] A clothing-changing pedestrian re-identification network construction module, configured to construct a clothing-changing pedestrian re-identification network, which consists of three network branches, namely the original branch, the black clothing branch, and the head branch, respectively used to learn the features of the original pedestrian image, the black-clothed pedestrian image, and the pedestrian head image; the backbone network structures of the three network branches are the same, but the parameters are not shared;

[0105] The black-clothes branch training module is used to input the black-clothed pedestrian images obtained by the black-clothes image obtaining module into the black-clothes branch for training, and obtain a pre-trained black-clothes branch;

[0106] The original branch training module is used to train the original branch under the guidance of the pre-trained black-clothes branch, and obtain features related to the pedestrian but not related to the clothes in the original pedestrian image;

[0107] The head branch training module is used to input the pedestrian head images obtained by the head image obtaining module into the head branch for learning, and combine the learned pedestrian head image features with the features related to the pedestrian but not related to the clothes obtained by the original branch training module, so as to obtain the overall features related to the pedestrian but not related to the clothes, and complete the training of the clothes-changing pedestrian re-identification network;

[0108] The clothes-changing pedestrian re-identification module is used to perform clothes-changing pedestrian re-identification based on the trained clothes-changing pedestrian re-identification network.

[0109] Furthermore, the black-clothes image obtaining module is specifically used for:

[0110] Adopt a pre-trained human parsing model to obtain the images of each body part of the pedestrian in the original pedestrian image, and recombine the obtained images of each body part to obtain six parts: background, head, upper clothes, trousers, arms and legs. Extract the pixels of the upper clothes and trousers images to form a clothes area; set all the pixels in the clothes area to zero to obtain a black-clothed pedestrian image.

[0111] Furthermore, the backbone networks of the three network branches are all imViT networks. In each imViT network, there are two outputs respectively used to extract global features and local features. The backbone network is optimized by triplet loss and ID loss on global features and local features respectively, where the ID loss is cross-entropy loss without label smoothing.

[0112] Furthermore, the original branch training module is specifically used for:

[0113] According to the knowledge distillation algorithm, train the branch of the original image under the guidance of the pre-trained black-clothes branch, and use the mean squared error loss to standardize the training of the original branch, so as to train more features related to the identity but not related to the clothes.

[0114] In summary, compared with the method of using GAN to generate a large number of clothing-changing images to expand the dataset, the present invention uses a non-GAN method. The upper and lower body parts of the human body are obtained by using a human semantic parsing model, and the method of the present invention is used to occlude them. In this way, the model can more intensively learn features unrelated to clothing without expanding the dataset, saving both space and time. Compared with other methods that separate clothing features and identity features, the present invention occludes the clothing of all pedestrians to obtain black-clothed images, which makes the clothing colors of pedestrians unified, enabling the model to pay more attention to parts other than clothing colors, thereby improving the robustness of the model. Compared with directly learning features from images after excluding clothing, the present invention uses a black-clothed branch to guide the original branch to directly learn identity features unrelated to clothing from the original RGB images, which can effectively utilize the information in the original images, effectively reduce information loss in the process of obtaining black-clothed images, and improve the robustness of features. Compared with the method of directly extracting features unrelated to clothing from the original images, the present invention can extract more discriminative fine-grained features by adding special processing of head image patches, which has a good complementary effect on global features. The test results on the PRCC dataset show that the method of the present invention achieves excellent clothing-changing pedestrian re-identification results.

[0115] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for re-identifying people who have changed clothes guided by black clothes and head images, characterized by: include: Step 1: Delete clothing features from the original pedestrian image to obtain a black clothing pedestrian image; Step 2: Use the pre-trained HRNet to process the original pedestrian image to obtain the pedestrian head image in the original pedestrian image; Step 3: Construct a clothing-changing pedestrian re-identification network. The network consists of three network branches: the original branch, the black clothing branch, and the head branch. They are used to learn the original pedestrian image features, the black clothing pedestrian image features, and the pedestrian head image features, respectively. The three network branches have the same backbone network structure, but do not share parameters; Step 4: Input the black clothing pedestrian image obtained in step 1 into the black clothing branch for training to obtain a pre-trained black clothing branch; Step 5: Train the original branch under the guidance of the pre-trained black clothing branch to obtain features related to pedestrians but not related to clothes in the original pedestrian image; Step 6: Input the pedestrian head image obtained in step 2 into the head branch for learning, and combine the learned pedestrian head image features with the pedestrian-related but clothing-unrelated features obtained in step 5 to obtain the overall features related to the pedestrian but not to the clothing, completing the training of the pedestrian re-identification network for changing clothes; Step 7: Perform clothing-changing pedestrian re-identification based on the trained clothing-changing pedestrian re-identification network; The step 5 comprises: According to the knowledge distillation algorithm, the original image branch is trained under the guidance of the pre-trained black clothing branch, and the mean square error loss is used to standardize the training of the original branch to train more features related to identity but not related to clothing. mse is defined as Where B is the batch size, F l1 、F g1 Represent the local features and global features obtained by the black branches, F l2 、F g2 They represent the local features and global features connected by the original branches guided by the black branches, L mse-opg 、L mse-opl Represent the global features F opg =[F g1 ,F g2 ] and local features F opl =[F l1 ,F l2 ] corresponds to the mean square error loss.

2. The method for re-identifying people who have changed clothes based on black clothes and head images according to claim 1, characterized in that: The step 1 comprises: A pre-trained human body parsing model is used to obtain the images of various body parts of the pedestrian in the original pedestrian image, and the obtained body part images are recombined to obtain six parts: background, head, top, pants, arms and legs. The pixels of the top and pants images are extracted from them to form the clothing area; all pixels in the clothing area are set to zero to obtain a black clothing pedestrian image.

3. The method for re-identifying people who have changed clothes based on black clothes and head images according to claim 1, characterized in that: The backbone networks of the three network branches are all imViT networks. In each imViT network, there are two outputs for extracting global features and local features respectively. The backbone network is optimized by triple loss and ID loss on global features and local features respectively, where the ID loss is the cross entropy loss without label smoothing.

4. A device for re-identifying pedestrians who have changed clothes guided by black clothes and head images, characterized in that: include: The black clothing image extraction module is used to remove clothing features from the original pedestrian image to obtain a black clothing pedestrian image; The head image derivation module is used to process the original pedestrian image using the pre-trained HRNet to obtain the pedestrian head image in the original pedestrian image; The clothing-changing pedestrian re-identification network construction module is used to build a clothing-changing pedestrian re-identification network. The network consists of three network branches: the original branch, the black clothing branch, and the head branch, which are used to learn the original pedestrian image features, the black clothing pedestrian image features, and the pedestrian head image features respectively; The three network branches have the same backbone network structure, but do not share parameters; A black clothing branch training module is used to input the black clothing pedestrian image obtained by the black clothing image extraction module into the black clothing branch for training, thereby obtaining a pre-trained black clothing branch; The original branch training module is used to train the original branch under the guidance of the pre-trained black clothing branch to obtain features related to pedestrians but not related to clothes in the original pedestrian image; The head branch training module is used to input the pedestrian head image obtained by the head image extraction module into the head branch for learning. The learned pedestrian head image features are combined with the pedestrian-related but clothing-irrelevant features obtained by the original branch training module to obtain the overall features related to the pedestrian but not related to the clothing, completing the network training for pedestrian re-identification after changing clothes; The clothing-changing pedestrian re-identification module is used to re-identify pedestrians who have changed clothes based on the trained clothing-changing pedestrian re-identification network; The original branch training module is specifically used for: According to the knowledge distillation algorithm, the original image branch is trained under the guidance of the pre-trained black clothing branch, and the mean square error loss is used to standardize the training of the original branch to train more features related to identity but not related to clothing. mse is defined as Where B is the batch size, F l1 、F g1 Represent the local features and global features obtained by the black branches, F l2 、F g2 They represent the local features and global features connected by the original branches guided by the black branches, L mse-opg 、L mse-opl Represent the global features F opg =[F g1 ,F g2 ] and local features F opl =[F l1 ,F l2 ] corresponds to the mean square error loss.

5. The device for re-identifying pedestrians who have changed clothes guided by black clothes and head images according to claim 4, characterized in that: The black clothing image deriving module is specifically used for: A pre-trained human body parsing model is used to obtain the images of various body parts of the pedestrian in the original pedestrian image, and the obtained body part images are recombined to obtain six parts: background, head, top, pants, arms and legs. The pixels of the top and pants images are extracted from them to form the clothing area; all pixels in the clothing area are set to zero to obtain a black clothing pedestrian image.

6. The device for re-identifying pedestrians who have changed clothes guided by black clothes and head images according to claim 4, characterized in that: The backbone networks of the three network branches are all imViT networks. In each imViT network, there are two outputs for extracting global features and local features respectively. The backbone network is optimized by triple loss and ID loss on global features and local features respectively, where the ID loss is the cross entropy loss without label smoothing.