A method and system for re-identifying pedestrians who change clothes based on strengthening common features
By adopting a method of strengthening common characteristics in pedestrian re-identification technology, combining the clothing mixing and matching module, human identity enhancement flow and common feature extraction module, the problem of pedestrian re-identification is solved, and higher recognition accuracy and stability are achieved.
Patent Information
- Application Number
- CN202510151987.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-02-12
AI Technical Summary
The existing pedestrian re-identification technology is difficult to effectively identify pedestrians who change clothes across cameras and scenes, and the multimodal method requires additional models and equipment, so learning to de-entangle representation takes a long time.
Using a method based on strengthening common characteristics, the robust and discriminant characteristics adapted to the clothes change scenario are extracted through joint training of clothing mixing and matching module, human identity enhancement flow, common feature extraction module and backbone network.
The pedestrian re-identification performance and stability improvement in the dressing change scenario is achieved, and the inherent characteristics of the human body can be extracted and utilized more effectively and the recognition accuracy can be improved.
Smart Images

Figure CN119625792B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a method and system for re-identifying pedestrians who have changed clothes based on enhanced common features. Background Art
[0002] In view of the increasingly urgent need for public safety and the rapid expansion of intelligent surveillance networks, pedestrian re-identification technology is increasingly closely integrated with technologies such as face recognition and pedestrian detection. Its wide application in the fields of smart security, target tracking, and smart retail has become a key link between artificial intelligence technology and actual industries, attracting a large number of researchers and scholars to explore in depth. The core of pedestrian re-identification lies in solving the problem of matching pedestrians across cameras and scenes, that is, accurately identifying the same pedestrian from images of multiple cameras. However, in practical applications, factors such as image resolution, viewing angle, lighting, and posture changes often interfere with the extraction of effective features by pedestrian re-identification algorithms, making this task extremely challenging.
[0003] So far, researchers have proposed a variety of person re-identification methods. Most of the current algorithms rely on convolutional neural networks to extract image features for classification or feature matching. Some methods, such as clothing-changing person re-identification and gait recognition, try to solve the problem from multimodal inputs (such as skeletons, contours, radio signals, contour sketches, and 3D shapes) or by decomposing representation learning. However, these multimodal methods often require additional models or devices to capture multimodal information, and the process of learning disentangled representations is usually time-consuming. In fact, the original RGB images contain rich information that is not related to clothing, but this information is largely ignored by current methods. As for some clothing-changing person re-identification methods, although they use powerful backbone networks (such as ResNet) to extract features from the original images, due to the lack of a well-designed loss function, the learned feature maps often only focus on some simple clothing-independent information (such as faces), while ignoring other key clothing-independent information.
[0004] These research works mainly focus on scenes with a short time span, usually assuming that pedestrians wear the same clothes under different cameras, so the visual appearance remains highly similar. In this scenario, the appearance of pedestrians' clothes is still an important basis for model feature extraction. However, in practical scenarios, the application of these methods is limited because pedestrians may change clothes, or different pedestrians may wear the same clothes. For example, people will change to suitable clothes in different locations and different weather; club members may wear the same uniform when participating in activities. Therefore, pedestrian re-identification methods for the clothing change problem are a more challenging and urgent problem to be solved. Summary of the invention
[0005] In order to solve the above problems, the present invention provides a method and system for re-identifying pedestrians who have changed clothes based on strengthening common features.
[0006] In order to achieve the above object, the present invention is implemented through the following technical solutions:
[0007] In a first aspect, the present invention provides a method for re-identifying a pedestrian who has changed clothes based on enhanced common features, comprising the following steps:
[0008] S1. Obtain the pedestrian image dataset to be retrieved and determine the original image in the dataset and images with the same identity label as the original image ;
[0009] S2. Image and The input is sent to the clothing mixing and matching CMM module, and the images of the same person wearing different clothes are synthesized through the human semantic parsing model, human skeleton key point detection model and perspective change method, and then passed through the ResNet50 backbone network to obtain the clothing mixing flow features. ;
[0010] S3. The original image The image is blackened by the human identity enhancement flow HIE module, and the human identity enhancement flow features are obtained by passing through the ResNet50 backbone network that shares parameters with step S2. ;
[0011] S4. The original image , clothing mixed flow characteristics And the human identity enhancement flow features Input to the common feature extraction CFE module for feature extraction to obtain the first saliency map , the second saliency map And the third saliency map ;
[0012] S5. The first saliency map , the second saliency map And the third saliency map Input into the classifier to obtain a classification result, wherein the classifier includes a normalization layer, a 1×1 convolution layer and a global pooling layer;
[0013] S6. Iteratively optimize the process of steps S2-S5 through the loss function to obtain the trained ResNet50 model;
[0014] S7. Input the image to be detected into the trained model to obtain the retrieval features of the image to be detected;
[0015] S8. Perform similarity matching between the retrieval features and the pedestrian image features in the retrieval library, sort the pedestrian images in descending order of similarity scores, and output the sorting results to obtain pedestrian re-identification results.
[0016] Furthermore, step S2 specifically includes:
[0017] S21. The original image and images with the same identity label as the original image Input them into the human body semantic parsing model respectively to get the original top mask image and the same identity label top mask image ;
[0018] S22. The original image and images with the same identity label as the original image Input into the Alpha-Pose model to extract the coordinates of the pedestrian's left shoulder, right shoulder, left hip bone and right hip bone in the two images respectively, and obtain the coordinates of the pedestrian in the original image and the coordinates of the pedestrian in the image with the same identity label;
[0019] S23. Calculate the pedestrian coordinates using the original image and the pedestrian coordinates of the same identity label image Transformation Matrix , through the transformation matrix For the original image and the original top mask image Perform perspective transformation to obtain the first perspective transformed image and the second perspective transformed image , the first perspective transformed image and images with the same identity label as the original image Multiply the features and extract the features through the ResNet50 backbone network to obtain the mixed flow features of the clothes ;
[0020] Furthermore, step S3 specifically includes: converting the original top mask image and the original image The upper garment is blackened by element-by-element multiplication, and the feature is extracted through the ResNet50 backbone network to obtain the human identity enhancement flow feature. ;
[0021] Furthermore, step S4 specifically includes:
[0022] The original image , clothing mixed flow characteristics and human identity enhancement stream features Input to the common feature extraction module, after 1×1 convolution layer and Activate to get the first attention map , the formula is as follows:
[0023] ,
[0024] in, represents the convolution operation of the 1×1 convolutional layer, Represent the height of the channel, the width of the channel and the number of channels respectively. Similarly, the second attention map is obtained and the third attention map ; The first attention map and the original image Multiply in the spatial dimension to get the first common feature , similarly, we get the second common feature And the third common feature ; First common feature , the second common feature And the third common feature After being processed by the batch normalization layer and the 1×1 convolution layer, the activation maps of the first category are obtained respectively. , the second category activation map and the third category activation map , ,in, represents the number of identity categories; the first category activation map , the second category activation map And the third category activation map Perform average pooling operations along the channel dimension to obtain the first saliency map , the second saliency map And the third saliency map , .
[0025] Furthermore, step S5 specifically includes:
[0026] The first saliency map , the second saliency map And the third saliency map Input into the classifier and obtain the first saliency map classification results respectively , the second saliency map classification results And the third saliency map classification results , through the common feature extraction loss Classification results of the first saliency map , the second saliency map classification results And the third saliency map classification results Constraints are performed to make the classification results of the three salient maps consistent, and the retrieval features are obtained; the common feature extraction loss The formula is as follows:
[0027] ,
[0028] in, represents the dot product operation, Indicates modulo.
[0029] Further, step S6 specifically includes:
[0030] In addition to the common feature extraction loss in steps S4 and S5 , the loss function also includes the recognition loss in steps S2 and S3 , CosFace loss and clothing-based adversarial loss , the CosFace function formula is as follows:
[0031] ,
[0032] in, represents the scale factor, Indicates the ratio used to adjust the loss. Indicates that it is used to introduce intervals on the target class to increase the class The margin of separability between represents the angle between the input feature vector and the class weight vector, represents the actual category label of the input feature vector, Indicates the index of the current sample, Represents the index of the category; the recognition loss function formula is as follows:
[0033] ,
[0034] in, represents the batch size, Representation sample After passing through the backbone network, the feature representation Indicates The classification weights of the identity categories, represents the identity label of the sample, represents the total number of identity categories in the training set, represents the hyperparameter used to scale log; the clothes-based adversarial loss function formula is as follows:
[0035] ,
[0036] in, represents the total number of clothing categories in the training set, Indicates The classification weights of clothing categories, represents the hyperparameter used to adjust the categorical distribution, and Respectively represent the samples Same identity and sample Set of clothing categories for different identities, represents the weight of the cross entropy loss, The weights are defined as follows:
[0037] ,
[0038] in, Representation sample The clothing category label, express The number of classes in ,when When , it means that the weight distribution returns to all positive classes sharing the same weight, and the weight is set to 0.5. , corresponding to the sample The weight of the positive class of the same clothing is set to be greater than 0.5, while the weight of other clothing categories in the positive class is set to be less than 0.5.
[0039] Furthermore, step S6 further includes:
[0040] The model training consists of 60 epochs, the first 25 epochs are the first step, and all epochs after the 25th epoch are the second step; in the first step, the loss function is defined as:
[0041] ,
[0042] Among them, min means to find the minimum value, represents the coefficient of the common feature extraction loss; in the second step, the final total loss is defined as:
[0043] .
[0044] In a second aspect, the present invention further provides a system for re-identifying pedestrians who have changed clothes based on enhanced common features, comprising:
[0045] Data acquisition module: used to obtain the pedestrian image dataset to be retrieved, and determine the original image and the image with the same identity label as the original image in the dataset;
[0046] Clothing Mixing and Matching CMM module: It is used to synthesize the original image and the image with the same identity label as the original image into an image of the same person wearing different clothes, and then pass it through the ResNet50 backbone network to obtain the clothing mixing flow features;
[0047] Human identity enhancement stream HIE module: It is used to blacken the top of the original image, and obtain the human identity enhancement stream features through the ResNet50 backbone network that shares parameters with the clothing mixing and matching CMM module;
[0048] Common feature extraction CFE module: used to extract features from the original image, clothing mixed stream features, and human identity enhancement stream features to obtain the first salient map, the second salient map, and the third salient map;
[0049] Classification unit: used for classifying the first saliency map, the second saliency map and the third saliency map to obtain a classification result;
[0050] Training module: used to iteratively optimize the process of steps S2-S5 through the loss function to obtain the trained ResNet50 model;
[0051] Retrieval feature unit: used to input the image to be detected into the trained model to obtain the retrieval features of the image to be detected;
[0052] Pedestrian re-identification unit: used to perform similarity matching between the retrieval features and the pedestrian image features in the retrieval library, sort the pedestrian images in descending order of similarity scores, and output the sorting results to obtain pedestrian re-identification results.
[0053] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in a method for re-identifying a pedestrian who has changed clothes based on enhanced common features as described in the first aspect.
[0054] In a fourth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in a method for re-identifying a pedestrian who has changed clothes based on enhanced common features as described in the first aspect are implemented.
[0055] The advantages of the present invention are:
[0056] The present invention provides a method for re-identifying people who change clothes based on strengthening common features. Through joint training of clothing mixing and matching module, human identity enhancement flow, common feature extraction module, backbone network, and clothing-based adversarial loss, the model pays more attention to the inherent features of the human body through a penalty mechanism, and can extract more robust and discriminative features suitable for clothing-changing scenarios. The method for re-identifying people who change clothes based on strengthening common features of the present invention has achieved excellent results in related clothing-changing pedestrian re-identification data sets, with higher recognition performance and stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0058] Figure 1 A flow chart of the steps of the method of the present invention;
[0059] Figure 2 This is an experimental comparison chart of the method of the present invention and the baseline model. DETAILED DESCRIPTION
[0060] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0061] Example 1
[0062] In this embodiment, Figure 1 As shown in FIG. 1 , a method for re-identifying a pedestrian after changing clothes based on strengthening common features is provided, and the specific steps include:
[0063] S1. Obtain the pedestrian image dataset to be retrieved and determine the original image in the dataset and images with the same identity label as the original image ;
[0064] S2. Image and The input is sent to the clothing mixing and matching CMM module, and the images of the same person wearing different clothes are synthesized through the human semantic parsing model, human skeleton key point detection model and perspective change method, and then passed through the ResNet50 backbone network to obtain the clothing mixing flow features. ;
[0065] Specifically, S21. The original image and images with the same identity label as the original image Input them into the human body semantic parsing model respectively to get the original top mask image and the same identity label top mask image ;
[0066] S22. The original image and images with the same identity label as the original image Input into the Alpha-Pose model to extract the coordinates of the pedestrian's left shoulder, right shoulder, left hip bone and right hip bone in the two images respectively, and obtain the coordinates of the pedestrian in the original image and the coordinates of the pedestrian in the image with the same identity label;
[0067] S23. Calculate the pedestrian coordinates using the original image and the pedestrian coordinates of the same identity label image Transformation Matrix , through the transformation matrix For the original image and the original top mask image Perform perspective transformation to obtain the first perspective transformed image and the second perspective transformed image , the first perspective transformed image and images with the same identity label as the original image Multiply the features and extract the features through the ResNet50 backbone network to obtain the mixed flow features of the clothes ;
[0068] S3. The original image The image is blackened by the human identity enhancement flow HIE module, and the human identity enhancement flow features are obtained by passing through the ResNet50 backbone network that shares parameters with step S2. ;
[0069] Specifically, the original shirt mask image and the original image The upper garment is blackened by element-by-element multiplication, and the feature is extracted through the ResNet50 backbone network to obtain the human identity enhancement flow feature. ;
[0070] S4. The original image , clothing mixed flow characteristics And the human identity enhancement flow features Input to the common feature extraction CFE module for feature extraction to obtain the first saliency map , the second saliency map And the third saliency map ;
[0071] Specifically, the original image , clothing mixed flow characteristics and human identity enhancement stream features Input to the common feature extraction module, after 1×1 convolution layer and Activate to get the first attention map , the formula is as follows:
[0072] ,
[0073] in, represents the convolution operation of the 1×1 convolutional layer, Represent the height of the channel, the width of the channel and the number of channels respectively. Similarly, the second attention map is obtained and the third attention map ; The first attention map and the original image Multiply in the spatial dimension to get the first common feature , similarly, we get the second common feature And the third common feature ; First common feature , the second common feature And the third common feature After being processed by the batch normalization layer and the 1×1 convolution layer, the activation maps of the first category are obtained respectively. , the second category activation map and the third category activation map , ,in, represents the number of identity categories; the first category activation map , the second category activation map And the third category activation map Perform average pooling operations along the channel dimension to obtain the first saliency map , the second saliency map And the third saliency map , .
[0074] S5. The first saliency map , the second saliency map And the third saliency map Input into the classifier to obtain a classification result, wherein the classifier includes a normalization layer, a 1×1 convolution layer and a global pooling layer;
[0075] Specifically, the first saliency map , the second saliency map And the third saliency map Input into the classifier and obtain the first saliency map classification results respectively , the second saliency map classification results And the third saliency map classification results , through the common feature extraction loss Classification results of the first saliency map , the second saliency map classification results And the third saliency map classification results Constraints are performed to make the classification results of the three salient maps consistent, and the retrieval features are obtained; the common feature extraction loss The formula is as follows:
[0076] ,
[0077] in, represents the dot product operation, Indicates modulus.
[0078] S6. Iteratively optimize the process of steps S2-S5 through the loss function to obtain the trained ResNet50 model;
[0079] Specifically, in addition to the common feature extraction loss in steps S4 and S5 , the loss function also includes the recognition loss in steps S2 and S3 , CosFace loss and clothing-based adversarial loss , the CosFace function formula is as follows:
[0080] ,
[0081] in, represents the scale factor, Indicates the ratio used to adjust the loss. Indicates that it is used to introduce intervals on the target class to increase the class The margin of separability between represents the angle between the input feature vector and the class weight vector, represents the actual category label of the input feature vector, Represents the index of the current sample, Represents the index of the category. In classification tasks, there are usually multiple categories, and It is used to distinguish these categories; the recognition loss function formula is as follows:
[0082] ,
[0083] in, represents the batch size, Representation sample After passing through the backbone network, the feature representation Indicates The classification weights of the identity categories, represents the identity label of the sample, represents the total number of identity categories in the training set, represents the hyperparameter used to scale log; the clothes-based adversarial loss function formula is as follows:
[0084] ,
[0085] in, represents the total number of clothing categories in the training set, Indicates The classification weights of clothing categories, represents the hyperparameter used to adjust the categorical distribution, and Respectively represent the samples Same identity and sample Set of clothing categories for different identities, represents the weight of the cross entropy loss, The weights are defined as follows:
[0086] ,
[0087] in, Representation sample The clothing category label, express The number of classes in ,when When , it means that the weight distribution returns to all positive classes sharing the same weight, and the weight is set to 0.5. , corresponding to the sample The weight of the positive class of the same clothing is set to be greater than 0.5, while the weight of other clothing categories in the positive class is set to be less than 0.5.
[0088] The model training consists of 60 epochs, the first 25 epochs are the first step, and all epochs after the 25th epoch are the second step; in the first step, the loss function is defined as:
[0089] ,
[0090] Among them, min means to find the minimum value, represents the coefficient of the common feature extraction loss; in the second step, the final total loss is defined as:
[0091] .
[0092] S7. Input the image to be detected into the trained model to obtain the retrieval features of the image to be detected;
[0093] S8. Perform similarity matching between the retrieval features and the pedestrian image features in the retrieval library, sort the pedestrian images in descending order of similarity scores, and output the sorting results to obtain pedestrian re-identification results.
[0094] Example 2
[0095] In this embodiment, in order to verify the effectiveness of the present invention, Figure 2 As shown in the figure, the retrieval results of the model invented in this paper and the baseline model (simple ResNet50 network) on the PRCC dataset are shown; the first picture in each row of (a), (b), and (c) is a retrieval example, and the retrieval examples (first pictures) of the baseline model and the method of the present invention are the same. The subsequent retrieval process is given 10 query graphs (because it is necessary to query ten times, Rank-1-Rank-10), calculate the similarity distance between all images in the retrieval library and the query graph, and return the top ten most similar images. All returned results may be correct or incorrect, and they are confirmed according to the labels. The correct retrieval results are marked in green in the visualization results, and the incorrect retrieval results are marked in red. It can be seen that the correct results retrieved by the method of the present invention can be ranked at the top of the retrieval list, and there are more correct results in the retrieval results, which proves that the clothing-changing pedestrian re-identification method based on enhanced common features is very useful for clothing-changing pedestrian re-identification, and the extracted features are effective and robust.
[0096] Example 3
[0097] This embodiment provides a system for re-identifying pedestrians who have changed clothes based on enhanced common features, including:
[0098] Data acquisition module: used to obtain the pedestrian image dataset to be retrieved, and determine the original image and the image with the same identity label as the original image in the dataset;
[0099] Clothing Mixing and Matching CMM module: It is used to synthesize the original image and the image with the same identity label as the original image into an image of the same person wearing different clothes, and then pass it through the ResNet50 backbone network to obtain the clothing mixing flow features;
[0100] Human identity enhancement stream HIE module: It is used to blacken the top of the original image, and obtain the human identity enhancement stream features through the ResNet50 backbone network that shares parameters with the clothing mixing and matching CMM module;
[0101] Common feature extraction CFE module: used to extract features from the original image, clothing mixed stream features, and human identity enhancement stream features to obtain the first salient map, the second salient map, and the third salient map;
[0102] Classification unit: used for classifying the first saliency map, the second saliency map and the third saliency map to obtain a classification result;
[0103] Training module: used to iteratively optimize the process of steps S2-S5 through the loss function to obtain the trained ResNet50 model;
[0104] Retrieval feature unit: used to input the image to be detected into the trained model to obtain the retrieval features of the image to be detected;
[0105] Pedestrian re-identification unit: used to perform similarity matching between the retrieval features and the pedestrian image features in the retrieval library, sort the pedestrian images in descending order of similarity scores, and output the sorting results to obtain pedestrian re-identification results.
[0106] Example 4
[0107] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in the method for re-identifying pedestrians who change clothes based on enhanced common features as described in Embodiment 1 are implemented;
[0108] Storage media include: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks or optical disks, and other media that can store programs.
[0109] Example 5
[0110] This embodiment provides a computer device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps in the method for re-identifying a person who changes clothes based on enhanced common features as described in Embodiment 1 are implemented;
[0111] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0112] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or replace some of the technical features therein by equivalents. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for re-identifying pedestrians who have changed clothes based on strengthening common features, characterized in that: The following steps are involved: S1. Obtain the pedestrian image dataset to be retrieved and determine the original image in the dataset and images with the same identity label as the original image ; S2. Image and The input is sent to the clothing mixing and matching CMM module, and the images of the same person wearing different clothes are synthesized through the human semantic parsing model, human skeleton key point detection model and perspective change method, and then passed through the ResNet50 backbone network to obtain the clothing mixing flow features. ; S3. The original image The image is blackened by the human identity enhancement flow HIE module, and the human identity enhancement flow features are obtained by passing through the ResNet50 backbone network that shares parameters with step S2. ; S4. The original image , clothing mixed flow characteristics And the human identity enhancement flow features Input to the common feature extraction CFE module for feature extraction to obtain the first saliency map , the second saliency map And the third saliency map ; S5. The first saliency map , the second saliency map And the third saliency map Input into the classifier to obtain a classification result, wherein the classifier includes a normalization layer, a 1×1 convolution layer and a global pooling layer; S6. Iteratively optimize the process of steps S2-S5 through the loss function to obtain the trained ResNet50 model; S7. Input the image to be detected into the trained model to obtain the retrieval features of the image to be detected; S8. Perform similarity matching between the retrieval features and the pedestrian image features in the retrieval library, sort the pedestrian images in descending order of similarity scores, and output the sorting results to obtain pedestrian re-identification results.
2. According to claim 1, a method for re-identifying pedestrians who change clothes based on enhanced common features is characterized in that: Step S2 specifically includes: S21. The original image and images with the same identity label as the original image Input them into the human body semantic parsing model respectively to get the original top mask image and the same identity label top mask image ; S22. The original image and images with the same identity label as the original image Input into the Alpha-Pose model to extract the coordinates of the pedestrian's left shoulder, right shoulder, left hip bone and right hip bone in the two images respectively, and obtain the coordinates of the pedestrian in the original image and the coordinates of the pedestrian in the image with the same identity label; S23. Calculate the pedestrian coordinates using the original image and the pedestrian coordinates of the same identity label image Transformation Matrix , through the transformation matrix For the original image and the original top mask image Perform perspective transformation to obtain the first perspective transformed image and the second perspective transformed image , the first perspective transformed image and images with the same identity label as the original image Multiply the features and extract the features through the ResNet50 backbone network to obtain the mixed flow features of the clothes .
3. According to claim 2, a method for re-identifying pedestrians who change clothes based on enhanced common features is characterized in that: Step S3 specifically includes: transforming the original top mask image and the original image The upper garment is blackened by element-by-element multiplication, and the feature is extracted through the ResNet50 backbone network to obtain the human identity enhancement flow feature. .
4. According to claim 3, a method for re-identifying pedestrians who change clothes based on strengthening common features is characterized in that: Step S4 specifically includes: The original image , clothing mixed flow characteristics and human identity enhancement stream features Input to the common feature extraction module, after 1×1 convolution layer and Activate to get the first attention map , the formula is as follows: , in, represents the convolution operation of the 1×1 convolutional layer, Represent the height of the channel, the width of the channel and the number of channels respectively. Similarly, the second attention map is obtained And the third attention map ; The first attention map and the original image Multiply in the spatial dimension to get the first common feature , similarly, we get the second common feature And the third common feature ; First common feature , the second common feature And the third common feature After being processed by the batch normalization layer and the 1×1 convolution layer, the activation maps of the first category are obtained respectively. , the second category activation map and the third category activation map , ,in, represents the number of identity categories; the first category activation map , the second category activation map And the third category activation map Perform average pooling operations along the channel dimension to obtain the first saliency map , the second saliency map And the third saliency map , .
5. The method for re-identifying pedestrians who change clothes based on enhanced common features according to claim 4 is characterized in that: Step S5 specifically includes: The first saliency map , the second saliency map And the third saliency map Input into the classifier and obtain the first saliency map classification results respectively , the second saliency map classification results And the third saliency map classification results , through the common feature extraction loss Classification results of the first saliency map , the second saliency map classification results And the third saliency map classification results Constraints are performed to make the classification results of the three salient maps consistent, and the retrieval features are obtained; the common feature extraction loss The formula is as follows: , in, represents the dot product operation, Indicates modulus.
6. The method for re-identifying pedestrians who change clothes based on enhanced common features according to claim 5, characterized in that: Step S6 specifically includes: In addition to the common feature extraction loss in steps S4 and S5 , the loss function also includes the recognition loss in steps S2 and S3 , CosFace loss and clothing-based adversarial loss , the CosFace function formula is as follows: , in, represents the scale factor, Indicates the ratio used to adjust the loss. Indicates that it is used to introduce intervals on the target class to increase the class The margin of separability between represents the angle between the input feature vector and the class weight vector, represents the actual category label of the input feature vector, Represents the index of the current sample, Represents the index of the category; the recognition loss function formula is as follows: , in, represents the batch size, Representation sample After passing through the backbone network, the feature representation Indicates The classification weights of the identity categories, represents the identity label of the sample, represents the total number of identity categories in the training set, represents the hyperparameter used to scale log; the clothes-based adversarial loss function formula is as follows: , in, represents the total number of clothing categories in the training set, Indicates The classification weights of clothing categories, represents the hyperparameter used to adjust the categorical distribution, and Respectively represent the samples Same identity and sample Set of clothing categories for different identities, represents the weight of the cross entropy loss, The weights are defined as follows: , in, Representation sample The clothing category label, express The number of classes in ,when When , it means that the weight distribution returns to all positive classes sharing the same weight, and the weight is set to 0.
5. , corresponding to the sample The weight of the positive class of the same clothing is set to be greater than 0.5, while the weight of other clothing categories in the positive class is set to be less than 0.
5.
7. The method for re-identifying pedestrians who change clothes based on enhanced common features according to claim 6, characterized in that: Step S6 also includes: The model training consists of 60 epochs, the first 25 epochs are the first step, and all epochs after the 25th epoch are the second step; in the first step, the loss function is defined as: , Among them, min means to find the minimum value, represents the coefficient of the common feature extraction loss; in the second step, the final total loss is defined as: 。 8. A system for re-identifying pedestrians changing clothes based on strengthening common features, which implements the method for re-identifying pedestrians changing clothes based on strengthening common features as claimed in claim 1, characterized in that: include: Data acquisition module: used to obtain the pedestrian image dataset to be retrieved, and determine the original image and the image with the same identity label as the original image in the dataset; Clothing Mixing and Matching CMM module: It is used to synthesize the original image and the image with the same identity label as the original image into an image of the same person wearing different clothes, and then pass it through the ResNet50 backbone network to obtain the clothing mixing flow features; Human identity enhancement stream HIE module: It is used to blacken the top of the original image, and obtain the human identity enhancement stream features through the ResNet50 backbone network that shares parameters with the clothing mixing and matching CMM module; Common feature extraction CFE module: used to extract features from the original image, clothing mixed stream features, and human identity enhancement stream features to obtain the first salient map, the second salient map, and the third salient map; Classification unit: used for classifying the first saliency map, the second saliency map and the third saliency map to obtain a classification result; Training module: used to iteratively optimize the process of steps S2-S5 through the loss function to obtain the trained ResNet50 model; Retrieval feature unit: used to input the image to be detected into the trained model to obtain the retrieval features of the image to be detected; Pedestrian re-identification unit: used to perform similarity matching between the retrieval features and the pedestrian image features in the retrieval library, sort the pedestrian images in descending order of similarity scores, and output the sorting results to obtain pedestrian re-identification results.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in a method for re-identifying a pedestrian who has changed clothes based on enhanced common features as described in any one of claims 1 to 7 are implemented.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method for re-identifying pedestrians who change clothes based on enhanced common features as described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Combined learning clothes changing pedestrian re-identification method and system based on identity guidance
CN116129473A
Clothes changing pedestrian re-identification method based on clothes unification and variety fusion
CN117333901A