A Person Re-identification Method Based on an Improved Multi-Scale Hard Triplet Loss Function
Through the ResNeSt50 network and the improved multi-scale difficult triple loss function, the problem of insufficient mining of difficult samples in pedestrian re-identification is solved, the feature expression and recognition accuracy of the model is improved, and more efficient pedestrian retrieval is achieved.
Patent Information
- Application Number
- CN202210597107.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-30
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-05-30
AI Technical Summary
The existing pedestrian recognition method cannot effectively mine the pair of positive and negative samples online during network training, resulting in the intra-class distance being greater than the inter-class distance and the model performance is degraded.
The ResNeSt50 network was used to extract pedestrian features of different scales, and the improved multi-scale difficult triple loss function and the cross-entropy loss function of label smooth regularization was combined with the loss equilibrium weighting strategy to improve the model's feature expression ability and sample data dependence mitigation.
The performance of pedestrian re-identification is improved, and the recognition accuracy and efficiency of the model are improved through online mining of multi-scale difficult triple loss functions and weight adjustment of loss function.
Smart Images

Figure CN117237979B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to the field of pedestrian re-identification. Background Art
[0002] Pedestrian re-identification is a technology for determining whether a specific pedestrian exists in an image or video sequence under two non-overlapping surveillance cameras in space. Pedestrian re-identification is a relatively popular field in current computer vision and has a wide range of applications in many directions, such as autonomous driving, intelligent transportation, and intelligent security. With the in-depth development of artificial intelligence, pedestrian re-identification methods have gradually transitioned from mainly traditional machine learning methods to current deep learning methods. Currently, the main methods for pedestrian re-identification tasks include pedestrian re-identification methods based on representation learning, pedestrian re-identification methods based on metric learning, and pedestrian re-identification methods based on local segmentation, etc., and good recognition effects have been achieved. However, during the network training process, the problem of ineffective online mining of difficult positive and negative sample pairs often occurs, that is, when performing similarity distance measurement, the distance between the same ID may be greater than the distance between different IDs (intra-class distance is greater than inter-class distance), and the model will erroneously regard negative samples as positive samples of the same ID. Especially when this situation is serious, the performance of the network model drops significantly. Therefore, by designing a pedestrian re-identification method, this problem can be better improved, and then a specific pedestrian target can be retrieved more efficiently, which is of great significance for improving the overall performance of the model. Summary of the Invention
[0003] (1) Technical problems to be solved:
[0004] Currently, the main methods for pedestrian re-identification tasks include pedestrian re-identification methods based on representation learning, pedestrian re-identification methods based on metric learning, and pedestrian re-identification methods based on local segmentation, etc., and good recognition effects have been achieved. However, during the network training process, the problem of ineffective online mining of difficult positive and negative sample pairs often occurs, that is, when performing similarity distance measurement, the distance between the same ID may be greater than the distance between different IDs (intra-class distance is greater than inter-class distance), and the model will erroneously regard negative samples as positive samples of the same ID. Especially when this situation is serious, the performance of the network model drops significantly.
[0005] (2) Technical solutions:
[0006] To achieve the above object, the present invention proposes a pedestrian re-identification method based on an improved multi-scale hard triplet loss function. The schematic diagram of this method is as shown in the appendix Figure 1 as follows, and the specific implementation of this method is as follows:
[0007] First, in the basic feature acquisition stage, the model uses modules such as the conv1, conv2_x, conv3_x, and conv4_x (where x represents the collective term for the subsequences of this layer) of ResNeSt50 as the feature acquisition network. Among them, conv2_x, conv3_x, and conv4_x contain 3, 4, and 6 Bottlenecks respectively. The second convolutional network module of each Bottleneck replaces the Conv2d module with the SplAtConv2d module on the basis of ResNet50. The SplAtConv2d module includes three groups of conv2d with convolutional kernels of 3×3, 1×1, and 1×1 respectively. The pedestrian sample features participating in distance calculation in the low-scale feature space come from the features extracted by the backbone network conv3_x. The pedestrian sample features participating in distance calculation in the medium-scale feature space come from the features extracted by the backbone network conv4_x. The pedestrian sample features participating in distance calculation in the high-scale feature space come from the features extracted by the backbone network conv5_x. When classifying the pedestrian ID, a cross-entropy loss function with label smoothing regularization is adopted, and the loss function used for similarity metric optimization is the multi-scale hard triplet loss function proposed in the present invention. Finally, a strategy of loss balance weighting is adopted to jointly optimize the two loss functions to achieve the purpose of pedestrian retrieval for a specific ID. The schematic diagram of this method is shown in the appendix Figure 1 。
[0008] The principle of the multi-scale hard triplet loss function proposed in the present invention is as follows:
[0009] The present invention requires that for the set low, medium, and high three different-scale feature spaces, the anchor samples respectively correspond to positive and negative samples at different scales The pedestrian sample features participating in distance calculation in the low-scale feature space come from the features extracted by the backbone network conv3_x. The pedestrian sample features participating in distance calculation in the medium-scale feature space come from the features extracted by the backbone network conv4_x. The pedestrian sample features participating in distance calculation in the high-scale feature space come from the features extracted by the backbone network conv5_x.
[0010] Then P×K samples in a batch jointly form 3×P×K sample pairs Define the distance functions of the three-scale feature spaces as D L (·), D M (·), D H (·), and calculate the average distance for the three-scale feature spaces respectively.
[0011] Let d L represent the average distance of the low-scale space, and the calculation is as follows:
[0012]
[0013] Let d M represent the average distance in the mesoscale space, and it is calculated as follows:
[0014]
[0015]
[0016] Let d H represent the average distance in the high-scale space, and it is calculated as follows:
[0017]
[0018] Next, calculate the normalized distances corresponding to each scale feature space, that is, divide the distances of the samples corresponding to each scale space by their respective average distances, so as to facilitate standardizing and unifying the scales when making numerical comparisons.
[0019] Let d′ L represent the normalized distance of the corresponding sample in the low-scale space, and it is calculated as follows:
[0020]
[0021] Let d′ M represent the normalized distance of the corresponding sample in the mesoscale space, and it is calculated as follows:
[0022]
[0023] Let d′ H represent the normalized distance of the corresponding sample in the high-scale space, and it is calculated as follows:
[0024]
[0025] By comparing the numerical values of d′ L 、d′ M and d′ H select the group with the smallest numerical value as the difficult positive and negative samples, denoted as and respectively, and calculate the loss values corresponding to the scales. After traversing P×K samples, repeat steps (4)-(6), and finally obtain the loss function, and its formula is as follows:
[0026]
[0027] (III) Beneficial effects:
[0028] The present invention provides a pedestrian re-identification method based on an improved multi-scale hard triplet loss function, having the following beneficial effects:
[0029] 1. The present invention uses the ResNeSt50 network as the feature extraction network to further extract higher fine-grained pedestrian features.
[0030] 2. The present invention performs online mining of positive and negative hard sample pairs for different-scale feature spaces through the improved multi-scale hard triplet loss function, further enhancing the feature expression ability of the model.
[0031] 3. The present invention further adjusts the weight ratio of different loss functions in a targeted manner through a loss balance weighting strategy, and uses a label smoothing strategy to alleviate the dependence of the model on sample data.
[0032] 4. The present invention uses the Hamming distance function for similarity calculation, improving the calculation efficiency. Description of the Drawings
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art.
[0034] Figure 1 The drawings are the schematic diagram of the pedestrian re-identification model based on the improved multi-scale hard triplet loss function;
[0035] Figure 2 The drawings are the flowchart of the core method for pedestrian re-identification based on the improved multi-scale hard triplet loss function. Detailed Embodiments
[0036] The present invention is a pedestrian re-identification method based on an improved multi-scale hard triplet loss function. The core flowchart of this method is shown in the appendix Figure 2 , and the specific implementation steps are as follows:
[0037] Step S1: Perform data preprocessing operations on the pedestrian dataset, including data augmentation and data normalization operations. Among them, data augmentation includes random size cropping, transposition, horizontal flipping, vertical flipping, random affine transformation, and normalization.
[0038] Step S2: Input the preprocessed picture data into the ResNeSt50 pre-trained model on ImageNet for training to extract pedestrian features with certain semantic information.
[0039] Step S3: For the set low, medium, and high three different-scale feature spaces, the anchor samples respectively correspond to positive and negative samples at different scales Among them, the pedestrian sample features participating in distance calculation in the low-scale feature space are from the features extracted by the backbone network conv3_x, the pedestrian sample features participating in distance calculation in the middle-scale feature space are from the features extracted by the backbone network conv4_x, and the pedestrian sample features participating in distance calculation in the high-scale feature space are from the features extracted by the backbone network conv5_x.
[0040] Step S4: Then a batch of P×K samples together form 3×P×K sample pairs Define the distance functions for the three-scale feature spaces as D L (·), D M (·), D H (·), and calculate the average distance for each of the three-scale feature spaces respectively.
[0041] Step S5: Let d L represent the average distance of the low-scale space, and the calculation is as follows:
[0042]
[0043] Step S6: Let d M represent the average distance of the middle-scale space, and the calculation is as follows:
[0044]
[0045] Step S7: Let d H represent the average distance of the high-scale space, and the calculation is as follows:
[0046]
[0047] Step S8: Next, calculate the normalized distances corresponding to each scale feature space, that is, divide the distances of the samples corresponding to each scale space by their respective average distances, which is convenient to standardize and unify the scale when making numerical comparisons.
[0048] Step S9: Let d′ L represent the normalized distance of the low-scale space of the corresponding sample, and the calculation is as follows:
[0049]
[0050] Step S10: Let d′ M represent the normalized distance of the middle-scale space of the corresponding sample, and the calculation is as follows:
[0051]
[0052] Step S11: Let d′ Hrepresents the normalized distance of the high-scale space corresponding to the representative sample, and is calculated as follows:
[0053]
[0054] Step S12: As shown in the accompanying drawings Figure 1 shown, by comparing the values of d′ L 、d′ M and d′ H select the group with the smallest value as the hard positive and negative samples, denoted as and respectively, and calculate the loss value of the corresponding scale. After traversing P×K samples, repeat steps S9 - S11, and finally obtain the loss function, the formula of which is as follows:
[0055]
[0056] Step S13: Introduce a loss balance weighting strategy to adjust the weight ratio of the improved multi-scale hard triplet loss function and the cross-entropy loss function with label smoothing regularization, and then perform joint optimization. Use the two loss functions to supervise the training of the model according to the set number of iteration rounds.
[0057] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. The above are only the preferred embodiments of the present invention, and any changes and modifications made within the scope defined by the claims of the invention fall within the protection scope of the present invention.
Claims
1. A pedestrian re-identification method based on an improved multi-scale hard triplet loss function, characterized in that The specific implementation steps are as follows: Step S1: Perform data preprocessing operations on the pedestrian dataset, including data augmentation and data normalization operations. Among them, data augmentation includes random size cropping, transposition, horizontal flipping, vertical flipping, random affine transformation, and normalization; Step S2: Input the preprocessed image data into the ResNeSt50 pre-trained model on ImageNet for training to extract pedestrian features with certain semantic information; Step S3: For the feature spaces of three different scales of low, medium, and high set, the anchor samples respectively correspond to positive and negative samples at different scales. Among them, the pedestrian sample features participating in distance calculation in the low-scale feature space are the features extracted from conv3_x of the backbone network, the pedestrian sample features participating in distance calculation in the medium-scale feature space are the features extracted from conv4_x of the backbone network, and the pedestrian sample features participating in distance calculation in the high-scale feature space are the features extracted from conv5_x of the backbone network. Step S4: Then, a batch of P×K samples together form 3×P×K sample pairs Define the distance functions of the three scale feature spaces as D L (·), D M (·), D H (·), and calculate the average distance for each of the three scale feature spaces respectively; Step S5: Using d L to represent the average distance in the low-scale space, the calculation is as follows: Step S6: Using d M to represent the average distance in the mesoscale space, which is calculated as follows: Step S7: Using d H to represent the average distance in the high-scale space, the calculation is as follows: Step S8: Next, calculate the normalized distances corresponding to each scale feature space, that is, divide the distances of the samples corresponding to each scale space by their respective average distances, which facilitates standardizing and unifying the scale when performing numerical comparisons; Step S9: Using d' L to represent the normalized distance of the low-scale space corresponding to the sample, the calculation is as follows: Step S10: Using d' M to represent the normalized distance in the mesoscale space of the corresponding sample, the calculation is as follows: Step S11: Using d' H to represent the normalized distance of the high-scale space corresponding to the corresponding sample, which is calculated as follows: Step S12: As shown in Figure 1 of the accompanying drawings, by comparing the values of d' L , d' M and d' H , select the group with the smallest value as the difficult positive and negative samples, denoted as and respectively, and calculate the loss value corresponding to the scale. After traversing P×K samples, repeat Steps S9 - S11. Finally, obtain the loss function, and its formula is as follows: Step S13: Introduce a loss balance weighting strategy to jointly optimize the improved multi-scale hard triplet loss function and the cross-entropy loss function with label smoothing regularization.
2. The person re-identification method based on an improved multi-scale hard triplet loss function according to claim 1, characterized in that Through the improved multi-scale hard triplet loss function, online mining of positive and negative hard sample pairs is performed for different scale feature spaces. That is, in each set feature scale space, the distance metric is calculated for it. Finally, the distance metric values corresponding to the three feature spaces are normalized and compared, and the group with the smallest value is selected as the feature vector corresponding to the hard positive and negative sample pairs to participate in the calculation of the metric loss value. The model weight value is adjusted through gradient backpropagation to achieve the purpose of supervising the optimization training of the model.
3. A person re-identification method based on an improved multi-scale hard triplet loss function according to claim 1, characterized in that, Through the loss balance weighting strategy, the weight ratio of the cross-entropy loss function with label smoothing and the improved multi-scale hard triplet loss function is adjusted and summed.
Citation Information
Patent Citations
Pedestrian re-identification method based on component power set and multi-scale features
CN110633631A
Multi-scale pedestrian re-identification method based on multi-granularity depth feature fusion
CN112818931A