Pedestrian Re-identification Network Model Data Augmentation and Training Method, Training Device

By using the horizontal stripe segmentation-shuffling method in the pedestrian re-identification network model for data augmentation and using the extended triple loss function to process the decimal similarity label, the problem of poor data augmentation in the existing technology is solved, and higher quality training data and better model generalization capabilities are achieved.

CN115546583BActive Publication Date: 2025-06-24GUANGDONG GAOHANG INTELLECTUAL PROPERTY OPERATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211236485.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-10
Publication Date
2025-06-24
Estimated Expiration
2042-10-10

AI Technical Summary

Technical Problem

The existing pedestrian re-identification network model has problems with data augmentation. The mixed images generated by cutmix are of poor quality and cannot effectively handle the decimal similarity labels in the pedestrian re-identification task.

Method used

The horizontal strip segmentation-shuffling method (Strip-Cutmix) is used for data augmentation, and an extended triple loss function (extend-triplet-loss) is designed to handle the decimal similarity label.

Benefits of technology

The quantity and quality of training data are improved, overfitting is avoided, and the generalization ability of the network model is enhanced, thereby obtaining better performance pedestrian re-identification network model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115546583B_ABST
    Figure CN115546583B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and discloses a method for data augmentation and training of a person re-identification network model, including the following steps: S101: Obtain M training images and the annotation data of the M training images. The M training images include pedestrians, and the annotation data of each training image includes the bounding box where the pedestrian in each training image is located and the pedestrian identity identification information; S102: Apply a set sampling strategy to select a batch of training images from the M training images as a batch of training samples, and apply the horizontal strip segmentation-shuffle method to perform data augmentation on the batch of training samples to obtain data-augmented batch training samples. The extended triple loss function designed for the person re-identification network model in the present invention can handle decimal similarity labels, so that it can be jointly applied to the training of the person re-identification network model with the data augmentation Strip-Cutmix method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and specifically to a method and device for data augmentation and training of a person re-identification network model. Background Art

[0002] Person re-identification is a very challenging image retrieval task, the goal of which is to match images of the same person in different cameras, and it has great application value in intelligent video surveillance, autonomous driving, unmanned aerial vehicle autonomous driving, sports event broadcasting, etc.

[0003] A person re-identification network model usually uses a deep neural network to implement, and it requires a large amount of training data to ensure the generalization ability of the model. However, the existing person re-identification datasets do not have a large enough amount of data, and collecting and annotating data requires a large amount of manpower, so data augmentation methods are usually used to increase the amount of training data.

[0004] The data augmentation method cutmix generates a new mixed image by combining two images. Its basic operation is to intercept an area of the same size as another picture and fill the same position area of this picture, and it is often used for data augmentation of deep neural networks.

[0005] However, cutmix has not been used in the person re-identification task, and the reasons are as follows:

[0006] 1) The quality of the mixed image generated by cutmix is poor, seriously affecting the accuracy and performance of subsequent deep neural network training.

[0007] 2) Cutmix combines two images to generate a new mixed image that contains multiple non-integer labels, and the triplet-loss used in the person re-identification task cannot handle fractional similarity labels.

[0008] Therefore, we propose a method and device for data augmentation and training of a person re-identification network model. Summary of the Invention

[0009] The purpose of the present invention is to provide a method and device for data augmentation and training of a person re-identification network model, and solve the problems raised in the above background art.

[0010] To achieve the above purpose, the present invention provides the following technical solution: A method for data augmentation and training of a person re-identification network model, including the following steps:

[0011] S101: Obtain M training images and the annotation data of the M training images. The M training images include pedestrians, and the annotation data of each training image includes the bounding box where the pedestrian is located in each training image and the pedestrian identity identification information;

[0012] S102: Apply the set sampling strategy to select a batch of training images from the M training images as a batch of training samples, and apply the horizontal strip segmentation-shuffle method to augment the data of this batch of training samples to obtain the data-augmented batch of training samples;

[0013] S103: Input the data-augmented batch of training samples into the person re-identification network model according to the set sampling strategy for feature extraction to obtain the feature vectors of this batch of sample pairs;

[0014] S104: Apply the extended triplet loss function to calculate the function value of the loss function corresponding to the feature vectors of this batch of samples;

[0015] S105: Update the network parameters of the person re-identification network model according to the function value of the loss function;

[0016] S106: Repeat the above steps S102, S103, S104, and S105 to train each batch of training samples until the function value of the extended triplet loss function meets the preset conditions, thereby completing the training of the person re-identification network model and obtaining the person re-identification network model that meets the preset conditions. The preset conditions are at least one of the following conditions: the number of training times of the person re-identification network is greater than or equal to the preset number of times, and the function value of the extended triplet loss function is less than or equal to the preset threshold.

[0017] Preferably, the specific steps of the horizontal strip segmentation-shuffle method in step S102 are as follows:

[0018] 1) From the original training samples within a batch, apply the set sampling strategy to select some images, perform horizontal strip cutting on the selected images at the same position, shuffle the cut image blocks, and then paste them back into the original images to generate synthetic images, where: each synthetic image is composed of some original image blocks and a pasted image block;

[0019] 2) The original training samples within a batch and the above synthetic images together constitute the data-augmented batch of training samples;

[0020] 3) Calculate two labels for each image in the data-augmented batch of training samples: one label is the class similarity label S dass , and the label value is equal to the area ratio of the image blocks of each identity; the other label is the sample pair similarity label S pair , which is used to compare the identity labels of the image blocks at the same position in two images, and the label value is equal to the ratio of the image blocks with the same identity label between the two.

[0021] Preferably, the specific steps of the sampling strategy set in step S102 are as follows:

[0022] a) In the data-augmented batch training samples described in claim 2, there are two types of images: original sample images and synthetic images, and any two images form a sample pair;

[0023] b) The sample pairs are divided into 4 groups: The first group is the Clean-Clean type, including two types of sample pairs: Clean-Clean: positive sample - positive sample, Clean-Clean: negative sample - negative sample; The second group is the Clean-Mixed: negative sample pair type, including two types of sample pairs: Clean-Mixed: negative sample - negative sample, Mixed-Clean: negative sample - negative sample; The third group is the Clean-Mixed: positive sample pair type, including two types: Clean-Mixed: positive sample - positive sample, Mixed-Clean: positive sample - positive sample; The fourth group is the Mixed-Mixed type, including two types: Mixed-Mixed: positive sample - positive sample, Mixed-Mixed: negative sample - negative sample;

[0024] c) In the original training samples within a batch, the PK probability sampling strategy is used, where: P represents the number of pedestrian identities in each batch, and K represents how many images each pedestrian has;

[0025] d) Select all P and half of K from the data-augmented batch training samples to participate in iterative training. The selection strategy is: The sample pair types for each iteration include 25% of the Clean-Clean type, 25% of the Mixed-Mixed type, and a 50% combination of the Clean-Mixed: negative sample pair + Clean-Mixed: positive sample pair type.

[0026] Preferably, the extended triplet loss function in step S104 is specifically: Given an anchor image a, the set of positive samples P(a) with the same identity as the anchor, and the set of negative samples N(a) with a different identity from the anchor, and m is the relaxation term:

[0027] s p = s(a, p) p ∈ P(a);

[0028] s n = s(a, n) n ∈ N(a);

[0029] Where s(a, p) represents the similarity of the positive sample pair, and s(a, n) represents the similarity of the negative sample pair. The smaller Sn is and the larger Sp is, the smaller the loss;

[0030] The standard Triplet Loss is expressed as: L triplet = s n - s p+m, which includes a relaxation condition that as long as s n -s p >m, that is, the positive sample similarity is greater than the negative sample similarity by a relaxation term m, then this set of s will no longer be optimized n and s p ;

[0031] From the perspective of making the similarity output by network training close to the similarity label, the extended triplet loss function (extend-triplet-loss) is extended as follows:

[0032] a1) Dynamically determine the optimization direction of the sample pair similarity. If the similarity output by network training is less than the similarity label, the similarity needs to be increased and optimized in the positive direction. If it is greater than the label similarity, the similarity needs to be decreased and optimized in the negative direction. Since the optimization direction is determined according to the positive and negative sample pairs, the dynamic determination of the optimization direction is achieved by dynamically dividing the positive and negative sample pairs. The dynamic division of the positive and negative sample pairs is as follows:

[0033] P(a) = {i = 1, 2,..., M | s(a, i) ≤ y(a, i) and y(a, i) ≠ 0};

[0034] N(a) = {i = 1, 2,..., M | s(a, i) > y(a, i) or y(a, i) = 0};

[0035] If the sample pair similarity s(a, i) obtained by network output is greater than the sample pair similarity label y(a, i), then this sample pair is a negative sample pair, and the similarity s(a, i) needs to be optimized in the negative direction to reduce the similarity. If s(a, i) is less than or equal to the sample pair similarity label y(a, i), then this sample pair is a positive sample pair, and the similarity s(a, i) needs to be optimized in the positive direction to increase the similarity;

[0036] a2) Extend the relaxation condition and extend the standard Triplet Loss so that its function value is distributed around the sample pair similarity label:

[0037]

[0038]

[0039]

[0040]

[0041] When and , this triplet will no longer be optimized;

[0042] where: yn is the similarity label for negative sample pairs, y p is the similarity label for positive sample pairs, is the relaxation term, let so that when y n = 0 and y p = 1, is compatible with the standard Triplet Loss.

[0043] The pedestrian re-identification network model data augmentation training device includes:

[0044] An acquisition module: used to acquire M training images and the annotation data of the M training images; the M training images include pedestrians, and the annotation data of each training image includes the bounding box where the pedestrian in each training image is located and the pedestrian identity identification information;

[0045] A data augmentation module: used to select a batch of training images from the M training images as a batch of training samples according to the set sampling strategy, and apply the horizontal strip segmentation-shuffle method to perform data augmentation on the batch of training sample pairs to obtain data-augmented batch training samples;

[0046] A sampling module: used to sample the data-augmented batch training samples according to the set sampling strategy to obtain batch training data;

[0047] A training module: used to execute the following steps one to four, including:

[0048] Step one: Input the batch of training sample pairs into the pedestrian re-identification network model for feature extraction to obtain the feature vectors of the batch of sample pairs;

[0049] Step two: Apply the extended triplet loss function to calculate the function value of the loss function corresponding to the feature vectors of the batch of sample pairs;

[0050] Step three: Update the network parameters of the pedestrian re-identification network model according to the function value of the loss function;

[0051] Step four: Repeat step one, step two, and step three to train each batch of training sample pairs until the function value of the extended triplet loss function meets the preset conditions, so as to complete the training of the pedestrian re-identification network model and obtain a pedestrian re-identification network model that meets the preset conditions.

[0052] The present invention provides a pedestrian re-identification network model data augmentation and training method and a training device. The pedestrian re-identification network model data augmentation and training method and training device have the following beneficial effects:

[0053] (1) The data augmentation Strip-Cutmix method designed according to the characteristics of person re-identification in the present invention improves the quantity and quality of training data. Under the condition that the amount of data in the training dataset of the person re-identification network model is limited, it allows the network model to be trained for a longer time while avoiding overfitting, improves the generalization ability of the network model, and thus obtains a person re-identification network model with better performance.

[0054] (2) The extended triplet loss function designed for the person re-identification network model in the present invention can handle fractional similarity labels, and thus can be jointly applied to the training of the person re-identification network model with the data augmentation Strip-Cutmix method. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a schematic structural diagram of the overall process of the data augmentation and training method of the person re-identification network model disclosed in the embodiment of the present invention;

[0056] Figure 2 It is an architecture diagram of the person re-identification network model disclosed in the embodiment of the present invention;

[0057] Figure 3 It is a schematic diagram of the calculation of two labels for each image in the data augmentation batch training samples disclosed in this embodiment;

[0058] Figure 4 It is a schematic block diagram of the person re-identification training device disclosed in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further elaborates the present invention with reference to the accompanying drawings and by way of examples; it should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention;

[0060] Figure 1 It is a flowchart of an embodiment of the data augmentation and training method of a person re-identification network model shown in an exemplary embodiment of the present invention, Figure 2 It is an architecture diagram of the person re-identification network model shown in an exemplary embodiment of the present invention; The following combines Figure 1 and Figure 2 to elaborate on this embodiment in detail;

[0061] As Figure 1 shown, this embodiment provides a data augmentation and training method for a person re-identification network model, including the following steps:

[0062] S101: Obtain M training images and the annotation data of the M training images. The M training images include pedestrians, and the annotation data of each training image includes the bounding box where the pedestrian in each training image is located and the pedestrian identity identification information;

[0063] S102: Apply a set sampling strategy to select a batch of training images from the M training images as a batch of training samples. Apply the horizontal strip cut - mix method (Strip - Cutmix) to perform data augmentation on this batch of training sample pairs to obtain data - augmented batch training samples;

[0064] S103. Input the data - augmented batch training samples into the person re - identification network model according to the set sampling strategy for feature extraction to obtain the feature vectors of this batch of sample pairs;

[0065] S104: Apply the extended triplet - loss function (extend - triplet - loss) to calculate the function value of the loss function corresponding to the feature vectors of this batch of samples;

[0066] S105: Update the network parameters of the person re - identification network model according to the function value of the loss function;

[0067] S106: Repeat the above steps S102, S103, S104, and S105 to train each batch of training samples until the function value of the extended triplet - loss function (extend - triplet - loss) meets the preset conditions, thereby completing the training of the person re - identification network model and obtaining a person re - identification network model that meets the preset conditions; the preset conditions are at least one of the following: the number of training times of the person re - identification network is greater than or equal to the preset number of times, and the function value of the extended triplet - loss function (extend - triplet - loss) is less than or equal to the preset threshold;

[0068] Combined Figure 2 Person re - identification network model architecture diagram: A10 executes the Strip - Cutmix data augmentation method in step S102 to perform data augmentation to increase the amount of training data; A20 executes the feature extraction in step S103 to extract the features of the image; A30 is used to execute the calculation of the extended triplet - loss function in step S104 to calculate the function value of the loss function;

[0069] Specifically, in this embodiment, ResNet50 or ResNeSt50 is used as the Backbone for feature extraction. After the extracted features are processed by POOL (pooling layer), BN (normalization layer), and FC (fully - connected layer), they are used as inputs for the calculation of the extended triplet - loss function;

[0070] ResNet50: The classic Residual Network (Deep Residual Learning for Image Recognition), which is widely used in various feature extraction applications. This method was published at CVPR in 2016;

[0071] ResNeSt50: A general improvement network based on ResNet (Resnest: Split-attention networks), which can achieve optimal performance on multiple tasks by introducing mechanisms such as attention and multi-branch. This method was published in 2020;

[0072] Combined with Figure 2 , the horizontal strip segmentation-shuffle method (Strip-Cutmix) is as follows:

[0073] 1) From the original training samples within a batch, apply the set sampling strategy to select some images. For the selected images, perform horizontal strip cutting at the same position, shuffle the cut image patches, and then paste them back into the original images to generate synthetic images; where: each synthetic image is composed of some original image patches and a pasted image patch;

[0074] 2) The original training samples within a batch and the above synthetic images together constitute the data augmentation batch training samples;

[0075] 3) Calculate two labels for each image in the data augmentation batch training samples: one label is the class similarity label s class , and the label value is equal to the area ratio of the image patches of each identity; the other label is the sample pair similarity label s pair , which is used to compare the identity labels of the image patches at the same position in two images, and the label value is equal to the ratio of the image patches with the same identity label;

[0076] Figure 3 Shown is an example of calculating the class similarity label s class and the sample pair similarity label s pair :

[0077] The class similarity labels s class of the original images A and B are one-hot encoded as s class-A = [1, 0] and s class-B = [0, 1] respectively; after images A and B are respectively subjected to 55% horizontal strip segmentation (strip-cut) and shuffle (mix), the synthetic images C and D are obtained. Then, the class similarity labels s class of the synthetic images C and D are one-hot encoded as s class-C= [0.55, 0.45] and s class-D = [0.45, 0.55]; Then calculate the sample pair similarity label s pair , then the sample pair similarity label s of the original image A and the original image B pair-A-B = [0], the sample pair similarity label s of the original image A and the synthesized image C pair-A-C = [0.55], the sample pair similarity label s of the original image A and the synthesized image D pair-A-D = [0.45], the sample pair similarity label s of the synthesized image C and the synthesized image D pair-C-D = [0];

[0078] Combined with Figure 2 , the set sampling strategy is as follows:

[0079] a) In the data augmentation batch training samples, there are two types of images: original sample images (Clean) and synthesized images (Mixed), and any two images form a sample pair;

[0080] b) The sample pairs are divided into 4 groups: The first group is the Clean-Clean type, including two types of sample pairs: Clean-Clean (positive sample - positive sample), Clean-Clean (negative sample - negative sample); The second group is the Clean-Mixed (negative sample pair) type, including two types of sample pairs: Clean-Mixed (negative sample - negative sample), Mixed-Clean (negative sample - negative sample); The third group is the Clean-Mixed (positive sample pair) type, including two types: Clean-Mixed (positive sample - positive sample), Mixed-Clean (positive sample - positive sample); The fourth group is the Mixed-Mixed type, including two types: Mixed-Mixed (positive sample - positive sample), Mixed-Mixed (negative sample - negative sample);

[0081] c) In the original training samples within a batch, use the PK probability sampling strategy (where: P represents the number of pedestrian identities in each batch, K represents how many pictures each pedestrian has): Execute the horizontal strip segmentation - shuffling method described in claim 2 with a probability ρ (for example, ρ = 0.8) to obtain the data augmentation batch training samples; The data augmentation batch training samples contain 20% original sample images (Clean) and 80% mixed images (Mixed);

[0082] d) Select all P and half of K from the data-augmented batch training samples for iterative training. The selection strategy is as follows: The sample pair types for each iteration include 25% Clean-Clean type, 25% Mixed-Mixed type, and 50% Clean-Mixed (negative sample pair) + Clean-Mixed (positive sample pair) type combinations;

[0083] Combined with Figure 2 , the extended triplet loss function (extend-triplet-loss) is:

[0084] Given an anchor image a, a set of positive samples P(a) with the same identity as the anchor, and a set of negative samples N(a) with different identities from the anchor, where m is the relaxation term:

[0085] s p = s(a, p) p ∈ P(a);

[0086] s n = s(a, n) n ∈ N(a);

[0087] Where s(a, p) represents the similarity of the positive sample pair, and s(a, n) represents the similarity of the negative sample pair; The smaller Sn is and the larger Sp is, the smaller the loss;

[0088] The standard Triplet Loss is expressed as: L triplet = s n - s p + m, which includes a relaxation condition. As long as s n - s p > m, that is, the positive sample similarity is greater than the negative sample similarity by a relaxation term m, this group of s n and s p will no longer be optimized;

[0089] From the perspective of making the similarity output by the network training close to the similarity label, the extended triplet loss function (extend-triplet-loss) is as follows:

[0090] a1) Dynamically determine the optimization direction of the sample pair similarity. If the similarity output by the network training is less than the similarity label, the similarity needs to be increased and optimized in the positive direction. If it is greater than the label similarity, the similarity needs to be decreased and optimized in the negative direction. Since the optimization direction is determined based on the positive and negative sample pairs, the dynamic determination of the optimization direction is achieved by dynamically dividing the positive and negative sample pairs. The dynamic division of the positive and negative sample pairs is as follows:

[0091] P(a) = {i = 1, 2,..., M | s(a, i) ≤ y(a, i) and y(a, i) ≠ 0};

[0092] N(a) = {i = 1, 2, ..., M | s(a, i) > y(a, i) or y(a, i) = 0};

[0093] If the similarity s(a, i) of the sample pair obtained from the network output is greater than the similarity label y(a, i) of the sample pair, then this sample pair is a negative sample pair, and the similarity s(a, i) should be optimized in the negative direction to reduce the similarity. If s(a, i) is less than or equal to the similarity label y(a, i) of the sample pair, then this sample pair is a positive sample pair, and the similarity s(a, i) should be optimized in the positive direction to increase the similarity;

[0094] Expand the relaxation condition; expand the standard Triplet Loss so that its function value is distributed around the similarity label of the sample pair:

[0095]

[0096]

[0097]

[0098]

[0099] When and then this triplet is no longer optimized;

[0100] where: y n is the similarity label of the negative sample pair, y p is the similarity label of the positive sample pair, is the relaxation term, let so that when y n = 0 and y p = 1, is compatible with the standard Triplet Loss;

[0101] In this application, the person re-identification network model obtained by using the data augmentation and training method of the person re-identification network model in the embodiments of this application can be better trained and obtain better performance. Therefore, using this person re-identification network model to process the image to be recognized can achieve better person recognition results;

[0102] The following describes the person recognition effect of the person re-identification network model in the embodiments of this application in combination with specific test results;

[0103] Table 1:

[0104]

[0105] Table 1 shows the test results of different solutions on different datasets. The test results include the mean average precision (mAP) and Rank-1. Rank-1 represents the probability that the image with the closest feature vector to the feature vector of the image to be recognized among the existing images belongs to the same pedestrian as the image to be recognized.

[0106] In Table 1 above, the datasets are DukeMTMC, Market1501, and MSMT17 respectively.

[0107] Among them, DukeMTMC contains 36,411 images of 1,812 pedestrian identities. The training set contains 16,522 images of 702 pedestrian identities, and the test set consists of 2,228 query images of 702 identities and 17,661 gallery images of 1,110 identities. Market-1501 is captured by six non-overlapping cameras. It contains 32,668 images of 1,501 people. These images are divided into a training set and a test set. The training set contains 12,936 images of 751 people, and the test set contains 19,732 images of 750 people. MSMT17 consists of 32,621 training set images of 1,041 identities, 11,659 query images, and 82,161 gallery images.

[0108] Existing solution 1: Optimization techniques and backbone network for person re-identification (Bag of Tricks and A Strong Baseline for Deep Person Re-identification), which was published at CVPR in 2019.

[0109] Existing solution 2: Non-local attention mechanism, weighted normalized triplet loss, and backbone network for person re-identification (Deep Learning for Person Re-identification: A Survey and Outlook), which was published in IEEE Transactions on Pattern Analysis and Machine Intelligence in 2021.

[0110] Existing solution 3: Combining many other optimization techniques on the basis of BoT to improve the accuracy of person re-identification (FastReID: A Pytorch Toolbox for Real-world Person Re-identification), which was published by JD AI Research Institute in 2020.

[0111] As can be seen from Table 1, the person re-identification network model obtained after being trained by the solution of the present application has better mAP and Rank-1 performance than the existing solutions on the datasets DukeMTMC, Market1501, and MSMT17, and has a good recognition effect;

[0112] Figure 4 is a schematic block diagram of the person re-identification training device according to an embodiment of the present application; Figure 4 The person re-identification training device 10 shown includes an acquisition module 110, a data augmentation module 120, a sampling module 130, and a training module 140;

[0113] The person re-identification training device 10 can be used to execute the data augmentation and training method of the person re-identification network model according to an embodiment of the present application;

[0114] Specifically, the acquisition module 110 is used to acquire M training images and the annotation data of the M training images; the M training images include pedestrians, and the annotation data of each training image includes the bounding box where the pedestrian in each training image is located and the pedestrian identity identification information;

[0115] Specifically, the data augmentation module 120 is used to select a batch of training images from the M training images as a batch of training samples according to the set sampling strategy, and apply the horizontal strip segmentation-shuffle method (Strip-Cutmix) to perform data augmentation on this batch of training sample pairs to obtain data-augmented batch training samples;

[0116] Optionally, the sampling strategy set in this embodiment is: in the original training samples within a batch, use the PK probability sampling strategy (where: P represents the number of pedestrian identities in each batch, and K represents how many images each pedestrian has): execute the horizontal strip segmentation-shuffle method described in claim 2 with a probability ρ (for example, ρ = 0.8) to obtain the data-augmented batch training samples; 20% of the original sample images (Clean) and 80% of the mixed images (Mixed) are included in the data-augmented batch training samples;

[0117] Specifically, the sampling module 130 is used to sample the data-augmented batch training samples according to the set sampling strategy to obtain batch training data;

[0118] Optionally, the sampling strategy that can be set in this embodiment is: select all P and half of K from the data-augmented batch training samples to participate in iterative training, and the selection strategy is: the sample pair types for each iteration include 25% of the Clean-Clean type, 25% of the Mixed-Mixed type, and 50% of the Clean-Mixed (negative sample pair) + Clean-Mixed (positive sample pair) type combination;

[0119] Specifically, the training module 140 is used to execute the following steps 1 to 4, including:

[0120] Step 1: Input the batch of training sample pairs into the person re-identification network model for feature extraction to obtain the feature vectors of this batch of sample pairs;

[0121] Step 2: Apply the extended triplet loss function (extend-triplet-loss) to calculate the function value of the loss function corresponding to the feature vectors of this batch of sample pairs;

[0122] Step 3: Update the network parameters of the person re-identification network model according to the function value of the loss function;

[0123] Step 4: Repeat Step 1, Step 2, and Step 3 to train each batch of training sample pairs until the function value of the extended triplet loss function (extend-triplet-loss) meets the preset conditions, thereby completing the training of the person re-identification network model and obtaining a person re-identification network model that meets the preset conditions;

[0124] Optionally, the preset conditions are at least one of the following: the number of training times of the person re-identification network is greater than or equal to the preset number of times, and the function value of the extended triplet loss function (extend-triplet-loss) is less than or equal to the preset threshold;

[0125] It should be understood that those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A method for data augmentation and training of a pedestrian re-identification network model, characterized in that, Including the following steps: S101: Obtain M training images and the annotation data of the M training images. The M training images include pedestrians. The annotation data of each training image includes the bounding box where the pedestrian is located in each training image and the pedestrian identity identification information; S102: Apply the set sampling strategy to select a batch of training images from the M training images as a batch of training samples. Apply the horizontal strip segmentation-shuffle method to perform data augmentation on this batch of training sample pairs to obtain data-augmented batch training samples; S103: Input the data-augmented batch training samples into the person re-identification network model according to the set sampling strategy for feature extraction to obtain the feature vectors of this batch of training sample pairs; S104: Apply the extended triplet loss function to calculate the function value of the loss function corresponding to the feature vectors of this batch of training samples; The extended triplet loss function is specifically: Given an anchor image a, a set of positive samples P(a) with the same identity as the anchor, a set of negative samples N(a) with different identities from the anchor, and m is the relaxation term: s p = s(a, p) for p ∈ P(a); s n = s(a, n) where n ∈ N(a); where s(a, p) represents the similarity of positive sample pairs, s(a, n) represents the similarity of negative sample pairs, and the smaller S n is, and the larger S p is, the smaller the loss is; The standard Triplet Loss is expressed as: L triplet = s n - s p + m, which includes a relaxation condition. As long as S p - S n > m, that is, the similarity of the positive sample is greater than the similarity of the negative sample by a relaxation term m, then this set of S n and S p ; Extend the relaxation condition, extend the standard Triplet Loss, so that its function value is distributed around the similarity label of the sample pair: When and this triple will no longer be optimized; Where: y n is the similarity label for negative sample pairs, y p is the similarity label for positive sample pairs, is the relaxation term, let so that when y n = 0 and y p = 1, is compatible with the standard Triplet Loss; S105: Update the network parameters of the person re-identification network model according to the function value of the loss function; S106: Repeat the above steps S102, S103, S104, and S105 to train each batch of training samples until the function value of the extended triplet loss function meets the preset conditions, thereby completing the training of the person re-identification network model to obtain a person re-identification network model that meets the preset conditions. The preset conditions are at least one of the following conditions: The number of training times of the person re-identification network is greater than or equal to the preset number of times, and the function value of the extended triplet loss function is less than or equal to the preset threshold.

2. The method for data augmentation and training of the pedestrian re-identification network model according to claim 1, characterized in that: The specific steps of the horizontal strip segmentation-shuffle method in step S102 are: 1) From the original training samples in a batch, apply the set sampling strategy to select some images. For the selected part of the images, perform horizontal strip cutting at the same position, shuffle the cut image blocks, and then paste them back into the original images to generate synthetic images, where: Each synthetic image is composed of some original image blocks and a pasted image block; 2) The original training samples in a batch and the above synthetic images together constitute the data-augmented batch training samples; 3) Calculate two types of labels for each image in the data augmentation batch training samples: one label is the class similarity label S dass , and the label value is equal to the area ratio of the image patches of each identity; the other label is the sample pair similarity label S pair , which is used to compare the identity labels of the image patches at the same position in two images, and the label value is equal to the ratio of the image patches with the same identity label between the two.

3. The pedestrian re-identification network model data augmentation and training method according to claim 2, wherein: The specific steps of the sampling strategy set in step S102 are: a) Among the data-augmented batch training samples, there are two types of images: original sample images and synthetic images. Any two images form a sample pair; b) The sample pairs are divided into 4 types: The first group is the Clean-Clean type, including two types of sample pairs: Clean-Clean: positive sample - positive sample, Clean-Clean: negative sample - negative sample; The second group is the Clean-Mixed: negative sample pair type, including two types of sample pairs: Clean-Mixed: negative sample - negative sample, Mixed-Clean: negative sample - negative sample; The third group is the Clean-Mixed: positive sample pair type, including two types: Clean-Mixed: positive sample - positive sample, Mixed-Clean: positive sample - positive sample; The fourth group is the Mixed-Mixed type, including two types: Mixed-Mixed: positive sample - positive sample, Mixed-Mixed: negative sample - negative sample; c) In the original training samples within one batch, the PK probability sampling strategy is used, where: P represents the number of pedestrian identities in each batch, and K represents the number of images for each pedestrian; d) Select all P and half of K from the data-augmented batch training samples to participate in iterative training. The selection strategy is: The sample pair types for each iteration include 25% of the Clean-Clean type, 25% of the Mixed-Mixed type, and a 50% combination of the Clean-Mixed: negative sample pair + Clean-Mixed: positive sample pair type.

4. The method for data augmentation and training of the pedestrian re-identification network model according to claim 1, wherein: From the perspective of making the similarity output by network training close to the similarity label, the extended triplet loss function (extend-triplet-loss) is extended as follows: Dynamically determine the optimization direction of the sample pair similarity. If the similarity output by network training is less than the similarity label, the similarity needs to be increased and optimized in the positive direction. If it is greater than the label similarity, the similarity needs to be decreased and optimized in the negative direction. Since the optimization direction is determined based on positive and negative sample pairs, the dynamic determination of the optimization direction is achieved by dynamically dividing positive and negative sample pairs. The dynamic division of positive and negative sample pairs is as follows: P(a) = {i = 1, 2,..., M | s(a,i) ≤ y(a,i) and y(a,i) ≠ 0}; N(a) = {i = 1, 2,..., M | s(a,i) > y(a,i) or y(a,i) = 0}; If the sample pair similarity s(a,i) obtained from network output is greater than the sample pair similarity label y(a,i), then this sample pair is a negative sample pair, and the similarity s(a,i) needs to be optimized in the negative direction to reduce the similarity. If s(a,i) is less than or equal to the sample pair similarity label y(a,i), then this sample pair is a positive sample pair, and the similarity s(a,i) needs to be optimized in the positive direction to increase the similarity.

5. A pedestrian re-identification network model data augmentation and training device for implementing the pedestrian re-identification network model data augmentation and training method according to claim 1, characterized in that, Including: Acquisition module: used to acquire M training images and the annotation data of these M training images; The M training images include pedestrians, and the annotation data of each training image includes the bounding box where the pedestrian is located in each training image and the pedestrian identity identification information; Data augmentation module: Used to select a batch of training images from the M training images as a batch of training samples according to the set sampling strategy, and apply the horizontal strip segmentation-shuffle method to perform data augmentation on this batch of training sample pairs to obtain data-augmented batch training samples; Sampling module: Used to sample the data-augmented batch training samples according to the set sampling strategy to obtain batch training data; Training module: Used to execute the following steps 1 to 4, including: Step 1: Input the batch of training sample pairs into the person re-identification network model for feature extraction to obtain the feature vectors of this batch of training sample pairs; Step 2: Apply the extended triplet loss function to calculate the function value of the loss function corresponding to the feature vectors of this batch of training sample pairs; Step 3: Update the network parameters of the person re-identification network model according to the function value of the loss function; Step 4: Repeat Step 1, Step 2, and Step 3 to train each batch of training sample pairs until the function value of the extended triplet loss function meets the preset conditions, thereby completing the training of the person re-identification network model and obtaining a person re-identification network model that meets the preset conditions.

Citation Information

Patent Citations

  • A method for generating pedestrian re-identification data under different illumination conditions based on an adversarial network

    CN109726669A

  • Training method of pedestrian re-identification network, pedestrian re-identification method and pedestrian re-identification device

    CN112446270A