A Cross-Domain Person Re-Identification Method
By using CycleGAN in pedestrian re-identification technology for image style conversion and improving feature extraction of ResNet50 network, the problem of inter-domain differences in cross-domain applications is solved, and the recognition effect and robustness of the model are improved.
Patent Information
- Application Number
- CN202310201680.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-03
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2043-03-03
AI Technical Summary
The existing pedestrian re-identification technology has inter-domain differences in cross-domain applications, resulting in a significant reduction in the recognition effect of the model in the target domain, and the existing methods are difficult to reduce inter-domain differences and improve the generalization of the model.
The image style conversion model based on CycleGAN is adopted to learn image styles in different scenarios, reduce the differences between domains, and improve the ResNet50 network, and improve the robustness of the model through uniform blocking feature extraction ideas.
It effectively reduces the differences between domains, improves the recognition effect and robustness of the pedestrian re-identification network, and ensures the good performance of the model in the target domain.
Smart Images

Figure CN116311362B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technology of pedestrian re-identification, and particularly to a cross-domain pedestrian re-identification method based on image style conversion and improved residual network. Background Art
[0002] The task objective of pedestrian re-identification is, given a target image, such as a pedestrian image, to find multiple pictures belonging to the same pedestrian as this image in a candidate image set through an algorithm system. Now, pedestrian re-identification technology is widely applied in most fields such as intelligent security, smart city, and smart commerce.
[0003] With the rapid development of deep learning and convolutional neural networks, researchers have conducted research on datasets collected in a single fixed scenario and achieved good recognition effects. However, when the model trained on a single dataset is deployed to other scenarios, its performance will be greatly reduced and cannot meet the actual application requirements. This is because there are inter-domain difference problems in cross-domain re-identification, such as different backgrounds, illuminations, etc. Although the model performance has achieved excellent performance on the source domain dataset, if the model is deployed to the target domain dataset, its performance will drop significantly. Currently, the main means to solve the inter-domain difference is to use generative adversarial networks for style conversion and to improve the model feature extraction effect by improving convolutional neural networks.
[0004] In order to solve the problem of mutual conversion between images in different domains, the Cycle Generative Adversarial Networks (CycleGAN) has been proposed to reduce the inter-domain difference and solve the cross-domain problem.
[0005] The Person Transfer GAN (PTGAN) has been proposed. This method is an improved CycleGAN that preserves the identity of pedestrians and only converts other parts of the images.
[0006] In order to solve the image changes caused by different environments under the fields of different cameras, a Hetero-Homogeneous Learning (HHL) has been designed based on StarGAN. Its design is mainly based on two assumptions: ① Images between the source domain (the dataset used for training) and the target domain (the dataset used for testing) can form negative sample pairs; ② Camera stability, that is, the styles of images under the same camera are similar.
[0007] The domain-invariant loss function has been proposed to help the network model learn the domain-invariant features in data images more accurately.
[0008] By utilizing the domain-invariant features of learning images, the classification accuracy of the model has been successfully improved in cross-domain image classification research.
[0009] Due to the problem of inter-domain differences in different datasets, the recognition effect of the model in the target domain is much lower than that in the source domain. Most of the current methods either focus on reducing inter-domain differences or improving the generalization of the model. They cannot solve the problems of both aspects at the same time. Moreover, violently using the adversarial generation network to reduce the differences between different datasets will cause the basic information of pedestrians to change, increasing the uncertainty for the model training of the pedestrian re-identification network. Simply improving the generalization of the model and only focusing on some invariant features in the image has a strong dependence on specific features, which causes the model to ignore other feature information in the image and has poor robustness. Summary of the Invention
[0010] The main purpose of the present invention is to provide a cross-domain pedestrian re-identification method. Aiming at the problem of inter-domain differences in cross-domain re-identification, a cross-domain pedestrian re-identification method based on image style conversion and improved residual network is proposed. This method first uses CycleGAN to construct an image style conversion model, performs style conversion on data in different scenarios, and then uses the improved ResNet50 to train and test the converted data. The role of CycleGAN is to learn the image styles of image data in different scenarios and reduce the inter-domain differences before training the pedestrian re-identification network. Aiming at the disadvantages of ResNet50, such as poor adaptability and low generalization ability in cross-domain re-identification problems, the feature extraction idea of uniform block division is used to improve it, reduce the dependence on specificity, and improve the robustness of the model.
[0011] The technical solution adopted by the present invention is: a cross-domain pedestrian re-identification method, including:
[0012] S1, learning the image style of the target domain through the style conversion model, so that the image data in the source domain can simulate the target domain after conversion;
[0013] S2, completing the pedestrian re-identification task. The pedestrian re-identification network model is improved based on the ResNet50 network to make the extracted features more robust.
[0014] Further, the step S1 includes:
[0015] Constructing an image style conversion model based on CycleGAN. This model draws on the way of retaining the basic information of pedestrians in PTGAN, extracts the basic information of pedestrians using PSPNet before training the style conversion model, and protects this part during training; the loss calculation of the style conversion model is as follows:
[0016] Let the image dataset under domain M be where xi ∈ M, the image dataset under domain N is where y j ∈ N; The cycle generative adversarial network has two generators, and it is necessary to learn two mapping functions G: M → N and F: N → M, and then by the discriminator D N and D M judge whether the current image is the original image or the converted one;
[0017] Formula (1) is the loss function of one of the unidirectional GANs, which is constructed by the generator G and the discriminator D N Construct:
[0018]
[0019] In the formula: L GAN (G, D N , M, N) represents the objective function of the discriminator; y represents the real image under domain N and follows P data distribution; D N (y) represents the probability that y is a real image under domain N; x represents the real image under domain M and follows P data distribution; D N (G(x)) represents the probability that the image under domain M is generated by the generator G into an image under domain N;
[0020] Similarly, the other unidirectional GAN is constructed by another generator F and discriminator D M and its loss function is shown in Formula (2):
[0021]
[0022] The cycle generative adversarial network uses a consistency loss function to ensure that the images between two different domains will not all be converted into the style of a certain picture, forcing F(G(x)) ≈ x and G(F(y)) ≈ y. The cycle consistency loss function is defined as shown in Formula (3):
[0023]
[0024] Adding the ID loss of the pedestrian into the loss of CycleGAN, and forcing to retain the original identity of the pedestrian during the conversion. Its identity loss is shown in Formula (4):
[0025]
[0026] In the formula: FMP(x) represents the pedestrian foreground mask pixel information extracted by the PSPNet network from the image x;
[0027] In summary, the loss function of the style conversion model is shown in Formula (5):
[0028]
[0029] where λ1 and λ2 are balance factors.
[0030] Further, the step S2 includes:
[0031] The input image will be adjusted and then sent into the backbone network for feature extraction. After the network extracts the features of the image, the original average pooling layer and fully connected layer are discarded; after Conv5_x ends, the feature with a size of 2048×24×12 is evenly divided horizontally into six feature vectors with a size of 2048×4×12; then average pooling operations are respectively performed on these six feature vectors to obtain six feature vectors with a size of 2048×1×1; then dimensionality reduction is respectively performed by 1×1 convolutional kernels, and these six 256-dimensional feature vectors are then calculated by six cross-entropy loss functions through the fully connected layer;
[0032] To calculate the cross-entropy loss, first, the probabilities of the current picture belonging to each category need to be obtained, and the calculation method is shown in formula (6):
[0033]
[0034] In the formula: i represents the category label; W i and b i respectively represent the weight and bias of the feature on the i-th output in the fully connected layer; x represents the feature value; p i represents the probability that the feature of the current picture belongs to the i-th category;
[0035] After calculating the probabilities of belonging to each category, the cross-entropy loss is shown in formula (7):
[0036]
[0037] In the formula: i represents the label of the sample picture; y i represents the label of the i-th sample picture; N represents the total number of samples;
[0038] The features extracted in the network are evenly divided into six blocks horizontally and the cross-entropy losses are respectively calculated. Finally, the average of these six losses is taken as the total loss L reID of the person re-identification network, and the specific calculation is shown in formula eight:
[0039]
[0040] In the formula: L crossEntropy (h i ) represents calculating the cross-entropy loss for the i-th block; num represents the constant 6;
[0041] The number of classifications obtained by the classifier depends on the number of pedestrian IDs in the dataset. During testing, the six 256-dimensional feature strings extracted are concatenated together as the final feature vector of the image. The Euclidean distance is used as the metric distance function. The concatenated feature vector is used to measure the similarity with the features of the images in the image library, and a distance order table is obtained.
[0042] Advantages of the present invention:
[0043] The present invention uses a cyclic adversarial generation network to reduce the inter-domain difference and improve the recognition effect of the pedestrian re-identification network.
[0044] Using the method in PTGAN that preserves the basic information of pedestrians, the basic information of pedestrians is protected during style conversion to ensure the invariance of the basic features of pedestrians.
[0045] Based on the feature extraction network improved from ResNet50, the image features are trained in blocks to ensure that the network learns global information, reduces the dependence on certain specific information, and further improves the robustness of the model.
[0046] The cross-entropy loss function can better classify pedestrian IDs.
[0047] The image style conversion model and the pedestrian re-identification network improved based on ResNet50 can be respectively transplanted into other cross-domain re-identification tasks.
[0048] In addition to the purposes, features and advantages described above, the present invention has other purposes, features and advantages. The present invention will be further described in detail below with reference to the drawings. Brief Description of the Drawings
[0049] The drawings constituting a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0050] Figure 1 is the image style conversion flowchart of the present invention;
[0051] Figure 2 is the structural diagram of the pedestrian re-identification network model based on the residual network of the present invention. Detailed Embodiments
[0052] In order to make the purposes, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0053] The cross-domain person re-identification method based on image style conversion and improved residual network in the present invention mainly includes two stages:
[0054] (1) The purpose of the first stage is to learn the image style of the target domain through the style conversion model, so that the image data in the source domain can simulate the target domain after conversion. The specific implementation process is as Figure 1 shown. Its specific implementation is to construct an image style conversion model based on CycleGAN. This model draws on the method in PTGAN to retain the basic information of pedestrians. Before training the style conversion model, PSPNet is used to extract the basic information of pedestrians, and this part is protected during training. The loss calculation of the style conversion model is as follows.
[0055] Let the image data set under domain M be where x i ∈M. The image data set under domain N is where y j ∈N. The cycle generative adversarial network has two generators, and it is necessary to learn two mapping functions G: M→N and F: N→M. Then, the discriminators D N and D M judge whether the current image is the original image or the converted image. Equation (1) is the loss function of one of the one-way GANs, which is constructed by the generator G and the discriminator D N .
[0056]
[0057] In the formula: L GAN (G, D N , M, N) represents the objective function of the discriminator; y represents the real image under domain N and follows the P data distribution; D N (y) represents the probability that y is a real image under domain N; x represents the real image under domain M and follows the P data distribution; D N (G(x)) represents the probability that the image under domain M is generated by the generator G into an image under domain N.
[0058] Similarly, the other one-way GAN is constructed by another generator F and discriminator D M , and its loss function is shown in Equation (2).
[0059]
[0060] The cycle generative adversarial network uses a consistency loss function to ensure that the images between two different domains will not all be converted into the style of a certain picture, forcing F(G(x))≈x and G(F(y))≈y. The cycle consistency loss function is defined as shown in Equation (3).
[0061]
[0062] Finally, to ensure that after the GAN converts the pedestrian images in the source domain, the original pose is maintained and only the image style is changed to be close to the target domain, the ID loss of the pedestrian is added to the loss of CycleGAN, and the original identity of the pedestrian is forced to be retained during the conversion. The identity loss is shown in Equation (4).
[0063]
[0064] In the formula: FMP(x) represents the pedestrian foreground mask pixel information extracted by the PSPNet network from the image x.
[0065] In summary, the loss function of the style conversion model is shown in Equation (5):
[0066] L cycleGAN = L GAN (G, D N , M, N) + L GAN (F, D M , N, M) + λ1L cyc (G, F) + λ2L ID (G, F) (5)
[0067] Where λ1 and λ2 are balance factors.
[0068] When performing style conversion, since there are multiple cameras in each dataset, in order to better reduce the inter-domain differences, according to the number of cameras in the source domain and the target domain, a style conversion model is trained for each pair of cameras. Specifically, the number of images generated after passing through the style conversion model for each image in the source domain is equal to the number of cameras in the target domain. Taking the Market1501 dataset as the source domain and the DukeMTMC-reID as the target domain as an example, a style conversion model is trained between each of the six cameras in Market1501 and the eight cameras in DukeMTMC, and a total of 48 (6×8 = 48) style conversion models need to be obtained.
[0069] (2) The purpose of the second stage is to complete the person re-identification task. The person re-identification network model is improved based on the ResNet50 network to make the extracted features more robust. The training of the person re-identification network model is completed on the image dataset after the source domain transformation in the first stage, and the testing of the model is completed on the target domain. The reason for this design is that in the cross-domain scenario, the training image data obtained by the person re-identification network model in the training stage is often small, but the image data generated at any time in actual applications is rich and diverse. If the training dataset is not processed and the trained model is directly deployed to the application scenario, the recognition effect will be much lower than the effect achieved during model training. After the first stage, using the image style transfer model to learn the image styles in some application scenarios not only makes the style of the training samples closer to the actual scenario, but also enriches the image dataset, so that the effect of the model trained on the transformed image data and deployed to the application scenario will be significantly improved.
[0070] Among them, the network model for completing the person re-identification task is improved and optimized based on ResNet50. To adapt to the challenges of a large number of identities, complex lighting, and scene changes that occur at any time in the actual scenario, the person re-identification network model will learn the idea of uniformly dividing the image features proposed by Sun et al. to improve ResNet50. The network model is as Figure 2 shown. The input image will be adjusted and then passed into the backbone network for feature extraction. After the network extracts the features of the image, the original average pooling layer and fully connected layer are discarded. After Conv5_x ends, the feature with a size of 2048×24×12 is evenly divided horizontally to obtain six feature vectors with a size of 2048×4×12. Then, average pooling operations are performed on these six feature vectors respectively to obtain six feature vectors with a size of 2048×1×1. Then, dimensionality reduction is performed by 1×1 convolutional kernels respectively. Finally, these six 256-dimensional feature vectors are passed through the fully connected layer and calculated by six cross-entropy loss functions.
[0071] Among them, to calculate the cross-entropy loss, first, the probability that the current picture belongs to each category (person ID) needs to be obtained, and its calculation method is shown in formula (6).
[0072]
[0073] In the formula: i represents the category label; W i and b i respectively represent the weight and bias of the feature on the i-th output in the fully connected layer; x represents the feature value; p i represents the probability that the feature of the current picture belongs to the i-th category.
[0074] After calculating the probabilities belonging to each category, the cross-entropy loss is shown in formula (7):
[0075]
[0076] where: i represents the label number of the sample picture; y i represents the label of the i-th sample picture; N represents the total number of samples.
[0077] The features obtained from the features extracted in the network are evenly divided into six blocks horizontally, and the cross-entropy loss is calculated for each block respectively. Finally, the average value of these six losses is taken as the total loss L of the person re-identification network reID , and the specific calculation is shown in Equation (8):
[0078]
[0079] where: L crossEntropy (h i ) represents calculating the cross-entropy loss for the i-th block; num represents the constant 6.
[0080] The number of classifications obtained by the classifier depends on the number of pedestrian IDs in the dataset. During testing, the six 256-dimensional feature vectors extracted are concatenated together as the final feature vector of the image. The Euclidean distance is used as the distance metric function. The concatenated feature vector is used to measure the similarity with the features of the images in the image library, and a distance order table is obtained.
[0081] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A cross-domain pedestrian re-identification method, characterized in that Including: S1. Learning the image style of the target domain through a style transfer model, so that the image data in the source domain can simulate the target domain after transformation; Constructing an image style transfer model based on CycleGAN. This model draws on the method in PTGAN to retain the basic information of pedestrians. Before training the style transfer model, PSPNet is used to extract the basic information of pedestrians, and this part is protected during training; S2. Completing the person re-identification task. The person re-identification network model is improved based on the ResNet50 network to make the extracted features more robust; The input image will be adjusted and then passed into the backbone network for feature extraction. After the network extracts the features of the image, the original average pooling layer and fully connected layer are discarded; after Conv5_x ends, the feature with a size of 2048×24×12 is evenly divided horizontally to obtain six feature vectors with a size of 2048×4×12; then average pooling operations are performed on these six feature vectors respectively to obtain six feature vectors with a size of 2048×1×1; then dimensionality reduction is performed by 1×1 convolutional kernels respectively. These six 256-dimensional feature vectors are then calculated by six cross-entropy loss functions through a fully connected layer.
2. The cross-domain pedestrian re-identification method according to claim 1, wherein The step S1 includes: The loss calculation of the style transfer model is as follows: Let the image dataset under domain M be , where , and the image dataset under domain N be , where ; Since the cycle - generative adversarial network has two generators, it is necessary to learn two mapping functions G: M→N and F: N→M, and then the discriminators and judge whether the current image is the original image or the converted one; Formula (1) is the loss function of one of the one-way GANs, which is composed of the generator G and the discriminator Construction: (1) In the formula: represents the objective function of the discriminator; represents the real image in domain N and follows distribution; represents which is the probability of the real image in domain N; represents the real image in domain M and follows distribution; represents the probability that the image in domain M is generated by the generator G into the image in domain N; Similarly, another one-way GAN is constructed by another generator F and discriminator whose loss function is shown in Equation (2): (2) The cyclic generative adversarial network uses a consistency loss function to ensure that images between two different domains are not all converted into the style of a single picture, forcing and , and the cyclic consistency loss function is defined as shown in Equation (3): (3) Adding the ID loss of pedestrians into the loss of CycleGAN to forcibly retain the original identity of pedestrians during the transformation. The identity loss is shown in formula (4): (4) In the formula: represents the pedestrian foreground mask pixel information extracted by the PSPNet network from the image ; In summary, the loss function of the style transfer model is shown in formula (5): (5) Among them and are balance factors.
3. The cross-domain pedestrian re-identification method according to claim 1, characterized in that The step S2 includes: To calculate the cross-entropy loss, first obtain the probabilities of the current picture belonging to each category, and the calculation method is shown in formula (6): (6) Wherein: represents the class label; and respectively represent the weight and bias of the feature on the th output in the fully connected layer; represents the feature value; represents the probability that the feature of the current image belongs to the th class; After calculating the probabilities of belonging to each category, the cross-entropy loss is shown in formula (7): (7) In the formula: represents the label of the sample picture; represents the th label of the sample picture; represents the total number of samples; The features obtained from the features extracted in the network are evenly divided into six blocks horizontally, and the cross-entropy loss is calculated for each block respectively. Finally, the average value of these six losses is taken as the total loss of the person re-identification network , and the specific calculation is shown in Equation (8): (8) In the formula: represents the cross-entropy loss for the th block; represents the constant 6; The number of classifications obtained by the classifier depends on the number of pedestrian IDs in the dataset. During testing, the six 256-dimensional features extracted are concatenated together as the final feature vector of the image. The Euclidean distance is used as the distance metric function. The similarity between the concatenated feature vector and the features of the images in the image library is measured to obtain a distance order list.
Citation Information
Patent Citations
Pedestrian re-identification method based on improved YOLOv3 network and feature fusion
CN111783576A
Cross-domain re-identification method based on shallow texture extraction and related equipment
CN115170836A