A pedestrian re-identification method based on 3D reconstruction and image generation
Virtual pedestrian images are generated through three-dimensional reconstruction and image generation technology, and an enhanced image set is constructed, and the pedestrian re-identification model is trained on this set, which solves the problem of insufficient accuracy of pedestrian re-identification under scarcity and uneven distribution, and improves the performance and robustness of the model.
Patent Information
- Application Number
- CN202510148344.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-11
AI Technical Summary
The existing pedestrian re-identification methods have significantly reduced the model performance under the scarce data and uneven distribution, resulting in insufficient accuracy of pedestrian re-identification results.
By utilizing three-dimensional reconstruction and image generation technology, more virtual pedestrian images are generated based on the original pedestrian image set, and enhanced image sets are constructed, and pedestrian re-identification models are trained on the enhanced image set in combination with classification loss and triple loss.
It improves the accuracy and robustness of pedestrian re-identification, solves the problems of insufficient data volume and uneven distribution, and enhances the performance of the model.
Smart Images

Figure CN119649410B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of image processing and machine vision, and particularly to a pedestrian re-identification method based on three-dimensional reconstruction and image generation. Background Art
[0002] The pedestrian re-identification technology is widely applied in multiple fields, including security monitoring, retail, intelligent transportation, urban planning, and event safety management, etc.
[0003] Existing pedestrian re-identification methods mainly rely on extracting features from two-dimensional images for matching, and a large amount of labeled pedestrian image data is required for model training. However, in actual scenarios, there are often problems of less training data and uneven distribution of training data. Especially for data from different camera perspectives or different environments, there may be a large distribution shift. Therefore, in complex scenarios (such as perspective changes, occlusions, and lighting differences), the performance of the model significantly decreases, resulting in insufficient accuracy of pedestrian re-identification results.
[0004] Therefore, how to improve the accuracy and robustness of pedestrian re-identification in the case of scarce and unevenly distributed data is a technical problem that needs to be solved. Summary of the Invention
[0005] In view of this, the present invention provides a pedestrian re-identification method based on three-dimensional reconstruction and image generation. By using the methods of three-dimensional reconstruction and image generation, more virtual pedestrian images are generated on the basis of the original pedestrian image set to construct an enhanced image set, and a pedestrian re-identification model is trained based on the enhanced image set to achieve improving the accuracy of pedestrian re-identification in the case of insufficient data volume and uneven data distribution.
[0006] For this purpose, the present invention provides the following technical solutions:
[0007] A pedestrian re-identification method based on three-dimensional reconstruction and image generation, comprising:
[0008] Generating a three-dimensional human model by using the pedestrian images in the original pedestrian image set;
[0009] Rotating the three-dimensional human model to a preset angle to generate a two-dimensional pedestrian rendering;
[0010] Combining the two-dimensional pedestrian rendering and the original pedestrian images to generate virtual pedestrian images;
[0011] Adding the virtual pedestrian images to the original pedestrian image set to construct an enhanced image set;
[0012] Training a pedestrian re-identification model on the enhanced image set by combining classification loss and triplet loss;
[0013] Input the pedestrian image to be detected into the trained pedestrian re-identification model to obtain the pedestrian recognition result.
[0014] Further, generating a three-dimensional human body model by using the pedestrian images in the original pedestrian image set includes:
[0015] Extract features from the original pedestrian image I through a convolutional neural network ;
[0016] Generate a three-dimensional human body model by using the features through a multi-human linear model based on skin vertices :
[0017]
[0018] Wherein, and are the shape parameter and the pose parameter respectively.
[0019] Further, combining the two-dimensional pedestrian rendering image and the original pedestrian image to generate a virtual pedestrian image includes:
[0020] Extract the identity feature through the original pedestrian image;
[0021] Extract the pose feature through the two-dimensional pedestrian rendering image;
[0022] Recombine the identity feature and the pose feature to generate a virtual pedestrian image.
[0023] Further, recombining the identity feature and the pose feature to generate a virtual pedestrian image through feature decoupling and a dual-generator network includes:
[0024] The generator generates a generated image combining the identity feature and the pose feature;
[0025] The discriminator distinguishes between real images and generated images;
[0026] The generator and the discriminator perform adversarial training to optimize the generated image, and output the optimized generated image as the virtual image.
[0027] Further, the generator includes: an identity generator and an attribute generator;
[0028] The identity generator is responsible for maintaining the identity information of the pedestrian image;
[0029] The attribute generator combines different pose features to change the pose of the generated image.
[0030] Further, the pedestrian re-identification model includes:
[0031] Extract the global feature and the local feature of the pedestrian image in the enhanced image set through the backbone network;
[0032] Dimensionality reduction processing is performed on the global features and local features, which are respectively used for the classification task and the metric learning task.
[0033] Furthermore, the classification task is optimized through the classification loss.
[0034] Furthermore, the metric learning task is optimized through the triplet loss.
[0035] Advantages and positive effects of the present invention:
[0036] The present invention constructs an enhanced image set by generating pedestrian images from multiple angles, and trains a pedestrian re-identification model based on the enhanced image set, solving the problem of scarce data for pedestrian re-identification, thereby improving the accuracy and robustness of pedestrian re-identification; the present invention extracts image features and constructs a three-dimensional human body model through a score-guided human body mesh recovery method, improving the generation quality of virtual pedestrian images; the present invention combines the classification loss and the triplet loss to train and update the model on the enhanced data set, ensuring that the model can correctly learn the generated image information, optimizing the model in the case of uneven data distribution, and improving the accuracy of pedestrian re-identification. Brief Description of the Drawings
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 It is a flowchart of the method in the embodiment of the present invention;
[0039] Figure 2 It is a trend chart of the simulation experiment index changes in the embodiment of the present invention. Detailed Embodiments
[0040] In order to enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0041] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0042] The present invention provides a pedestrian re-identification method based on 3D reconstruction and image generation. First, virtual pedestrian images are generated from the original pedestrian images to expand the pedestrian re-identification public dataset for data augmentation: the original pedestrian images in the pedestrian re-identification public dataset are 3D modeled, and these 3D human models are rotated by multiple angles. The 2D projections of the 3D human models at specific angles are selected as new structures to simulate the postures of pedestrians in different directions. Then, the structures at the new angles are combined with the original pedestrian images to generate virtual pedestrian images for data augmentation. Then, the virtual pedestrian images are merged into the original pedestrian image set to form a new training dataset with richer data. Finally, the model is trained and updated on the training dataset: the model is trained and updated on the training dataset by combining the classification loss and the triplet loss to ensure that the model can correctly learn the features of the augmented images, thereby solving the problem of insufficient data volume for pedestrian re-identification and improving the accuracy of pedestrian re-identification.
[0043] Combined with Figure 1 Yes, the method of the present invention is further described as follows:
[0044] S1. Generate a 3D human model from a single original pedestrian image, and select a specific angle of the 3D human model to generate a 2D pedestrian rendering.
[0045] In this embodiment, a 3D human model is extracted from the postures of the characters in the original pedestrian image set by the ScoreHMR (Score-Guided Human Mesh Recovery) method.
[0046] Specifically, features are extracted from the input image I through a convolutional neural network , and the features are passed to the SMPL model (Skinned Multi-Person Linear Model) to generate a 3D mesh , that is, a 3D human model:
[0047]
[0048] Among them, and are the shape and pose parameters respectively.
[0049] Introduce a scoring mechanism to optimize the generation process of the 3D human model, and gradually adjust the parameters of the 3D mesh, including posture, shape, perspective, etc., so that the generated 3D human model is more consistent with the pedestrians in the image.
[0050] And rotate the 3D human model by a specific angle to render and generate a 2D pedestrian rendering.
[0051] S2. Use the 2D pedestrian rendering and the original pedestrian image to generate a virtual pedestrian image to achieve data augmentation.
[0052] Extract the identity features through the original pedestrian image; extract the pose features through the 2D pedestrian rendering; recombine the identity features and the pose features to generate a virtual image.
[0053] In this embodiment, virtual image generation is performed through feature decoupling and a dual generation network, including:
[0054] The generator generates an image that conforms to the target pose and identity features, where the identity generator is responsible for maintaining the identity information of the original pedestrian image; the attribute generator combines different pose features to change the pose of the generated image.
[0055] Optimize the generated image through a generative adversarial network to improve its authenticity, including:
[0056] Through the adversarial training between the generator and the discriminator, the generator generates an image that conforms to the target pose and identity features, making the generated image as realistic as possible; the discriminator distinguishes between real images and generated images and improves the generation quality by continuously providing feedback to the generator. The formula is expressed as:
[0057]
[0058] Among them, x is the real image, z is the input noise of the generator, is the image generated by the generator, is the probability that the discriminator determines that the image x is real; represents the mathematical expectation.
[0059] During the adversarial training process, the model will use a variety of loss functions, including identity consistency loss, reconstruction loss, and style consistency loss, to ensure that the generated image is consistent with the original pedestrian image in terms of details and style.
[0060] S3. Incorporate the generated virtual pedestrian images into the original pedestrian image set to construct an enhanced image set;
[0061] In this embodiment, the rule for incorporating the generated virtual pedestrian images into the original pedestrian image set is to increase the numerical part representing the number of frames of the person image in the file name, and traverse all images to increment this part of the number until a non-conflicting name is found so that it does not duplicate the existing files.
[0062] S4. Construct a deep learning network and train and update the model on the enhanced image set.
[0063] In this embodiment, the ResNet-50 is used as the backbone network during the training process, which includes three branches for extracting global and local features.
[0064] The training process includes the following steps:
[0065] Feature extraction: Extract image features through the backbone network and process them through three branches respectively. The global branch extracts the overall features, and the two local branches extract local features respectively, which are processed through different segmentation strategies.
[0066] Feature processing: The features of the global and local branches are respectively used for the classification task and the metric learning task after dimensionality reduction.
[0067] The classification task is optimized through the Softmax Loss, and the metric learning task is optimized through the triplet loss. The loss function formula is as follows:
[0068]
[0069]
[0070] Where: N represents the total number of training samples, C represents the total number of categories, i represents the index of the i-th sample, represents the true category index of the i-th sample, represents the feature vector of the i-th sample, represents the weight vector of category K, represents the bias term of category k, and represent the weight vector and bias term corresponding to the true category of the sample; P represents the number of sampled pedestrian identity categories, K represents the number of images sampled for each pedestrian identity, a represents the sample index of the anchor, p represents the sample index of the positive sample, n represents the sample index of the negative sample, represents the feature vector of the anchor sample a of identity i, represents the feature vector of the anchor sample p of identity i, The feature vector of the negative sample n with identity j (j ≠ i), denotes the minimum margin between positive and negative samples.
[0071] Joint optimization: By jointly optimizing two loss functions, the Softmax Loss is used to improve feature discrimination, and the triplet loss optimizes the feature ranking performance, finally generating a feature representation for person re-identification.
[0072] The effect of the method of the present invention is further illustrated by the following simulation experiments:
[0073] The experimental dataset is the person re-identification dataset Market1501;
[0074] 1) Input image size: The input image is adjusted to a fixed size of 384×128 to capture more pedestrian detail information.
[0075] 2) Pre-trained weights: The ResNet-50 backbone network is initialized with weights pre-trained on the LUP dataset.
[0076] 3) Batch sampling strategy: mini-batch structure: Randomly select P identities from the training set, and randomly sample K images for each identity. Set P = 16 and K = 4, that is, each mini-batch contains P×K = 64 images. This meets the sample selection requirements of the triplet loss and ensures that there are enough positive and negative samples in each batch of data.
[0077] 4) Dataset setting: Use the Market1501 dataset for image generation, select the two-dimensional pedestrian rendering images of the human model in the original image rotated horizontally by 45°, and merge the generated images into the original dataset, and use the new generated dataset for model training.
[0078] 5) Evaluation metrics:
[0079] Rank-k accuracy: Indicates whether the correct matching result appears in the top k candidate results returned;
[0080] mAP: The average of all average precisions (AP), which measures the result ranking quality in the retrieval task.
[0081] During the training process, an evaluation is performed every 50 epochs. From Figure 2 it can be seen that all indicators are continuously rising and tend to be unchanged after 350 epochs. Therefore, the total number of training epochs is selected as 400 epochs.
[0082] The method proposed in the present invention was compared with other pedestrian re-identification methods under the same conditions. As shown in Table 1, the comprehensive performance of the method of the present invention in terms of Rank-1, Rank-5, Rank-10, and mAP is better in the comparison, superior to the previous methods, which indicates that this method is effective and improves the accuracy of pedestrian re-identification. Among them, the mAP has a significant improvement, indicating that the model performs very well in the global ranking in the retrieval task, can well rank the correct samples in the front, and is reasonably distributed in the entire retrieval list, showing that the method model of the present invention has good comprehensive performance.
[0083] Table 1
[0084]
[0085] In summary, the present invention solves the problem of scarce data for pedestrian re-identification by data augmentation to form a new and more data-rich training dataset and training the model on the augmented dataset through joint training; it ensures that the model can correctly learn the generated image information and improves the accuracy of pedestrian re-identification.
[0086] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A pedestrian re-identification method based on three-dimensional reconstruction and image generation, characterized in that: include: Generate a 3D human body model using pedestrian images in an original pedestrian image set; Rotating the three-dimensional human body model to a preset angle to generate a two-dimensional pedestrian rendering; Generate a virtual pedestrian image by combining the two-dimensional pedestrian rendering and the original pedestrian image; The virtual pedestrian images are added to the original pedestrian image set to construct an enhanced image set; Combining classification loss and triplet loss to train the person re-identification model on the augmented image set; Input the pedestrian image to be detected into the trained pedestrian re-identification model to obtain the pedestrian recognition result; The method of generating a three-dimensional human body model by using pedestrian images in the original pedestrian image set includes: Extract features from the original pedestrian image I through a convolutional neural network ; Utilize the features through a multi-body linear model based on skin vertices Generate a 3D human body model: in, and They are shape parameters and posture parameters respectively; The step of combining the two-dimensional pedestrian rendering image and the original pedestrian image to generate a virtual pedestrian image includes: Extract identity features from original pedestrian images; Extract posture features through 2D pedestrian renderings; Recombining identity features and posture features to generate virtual pedestrian images.
2. According to claim 1, a pedestrian re-identification method based on three-dimensional reconstruction and image generation is characterized in that: The virtual pedestrian image is generated by recombining the identity feature and the posture feature through feature decoupling and dual generation network, including: The generator generates a generated image that combines identity features and posture features; The discriminator distinguishes between real images and generated images; The generator and the discriminator are trained adversarially to optimize the generated image, and the optimized generated image is output as a virtual image.
3. According to claim 2, a pedestrian re-identification method based on three-dimensional reconstruction and image generation is characterized in that: The generator includes: an identity generator and an attribute generator; The identity generator is responsible for maintaining the identity information of the pedestrian image; The attribute generator combines different pose features to change the pose of the generated image.
4. According to claim 1, a pedestrian re-identification method based on three-dimensional reconstruction and image generation is characterized in that: The pedestrian re-identification model includes: The global and local features of pedestrian images in the enhanced image set are extracted through the backbone network; The global features and local features are reduced in dimension and used for classification tasks and metric learning tasks respectively.
5. According to claim 4, a pedestrian re-identification method based on three-dimensional reconstruction and image generation is characterized in that: The classification task is optimized by the described classification loss.
6. The pedestrian re-identification method based on three-dimensional reconstruction and image generation according to claim 4, characterized in that: The metric learning task is optimized by the triplet loss.
Citation Information
Patent Citations
Pedestrian re-identification method based on identity migration generative adversarial network
CN115205903A
Pedestrian re-identification method
CN116052218A