Pedestrian detection algorithm training method based on pedestrian re-identification and text feature matching supervision
By adding pedestrian re-identification and text feature matching branches in the pedestrian detection algorithm model, and combining calibration information to optimize feature extraction and loss calculation, the error detection and missed detection problems in pedestrian detection are solved, and the detection accuracy and accuracy are improved.
Patent Information
- Application Number
- CN202510538651.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-04-27
AI Technical Summary
The existing pedestrian detection algorithm has high false detection rates in complex environments and has high missed detection rates for pedestrian posture diversity, making it difficult to balance detection accuracy.
Pedestrian re-identification and text feature matching branches are added at each stage of the backbone network of the pedestrian detection algorithm model. Through freeze training, it operates independently, and combines calibration information and feature matching loss calculation to optimize feature extraction and prediction accuracy.
It improves the sensitivity of the pedestrian detection model to features, reduces the false detection rate and missed detection rate, and improves the detection accuracy and accuracy.
Smart Images

Figure CN120495978A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target recognition, and in particular to a pedestrian detection algorithm training method based on pedestrian re-identification and text feature matching supervision. Background Art
[0002] There are two issues with deep learning for pedestrian detection. Problem 1: The complex environment of pedestrian detection leads to a high false positive rate. Furthermore, current target detection algorithms continuously enhance the target detection model's ability to extract target features, making the target detection model highly sensitive to targets and misidentifying some human-like objects as targets. Problem 2: Pedestrians have a wide variety of postures, including standing, standing, and other postures, resulting in a large number of missed detections. This is because the target detection model lacks sensitivity to targets during feature extraction, leading to feature loss. Summary of the Invention
[0003] The present invention mainly solves the technical problems of easy false detection or missed detection and difficult balance in the existing technology, and provides a pedestrian detection algorithm training method based on pedestrian re-identification and text feature matching supervision that can reduce false detection and missed detection and has higher accuracy.
[0004] The present invention solves the above technical problems mainly through the following technical solutions: a pedestrian detection algorithm training method based on pedestrian re-identification and text feature matching supervision, comprising the following steps:
[0005] S1. Constructing the training structure of the pedestrian detection algorithm model: Add a pedestrian re-identification (reid) branch and a text feature matching branch to the output end of each stage of the backbone network of the pedestrian detection algorithm model. After training the added pedestrian re-identification branch and text feature matching branch, freeze the two branches. The branches added at each stage of the backbone network do not modify the original input-output relationship of the backbone network. For example, the output of the first stage originally entered the second stage, and it remains the same after the addition, except that the output of each stage also enters the two newly added branches. The parameters of the branches added at each stage are independent after training. The subsequent training process is based on the pedestrian detection algorithm model training structure. After the training is completed, the original pedestrian detection algorithm model is still used for formal reasoning, and the two newly added branches are not needed. The backbone network adopts ResNet50.
[0006] S2. Input the training images of the training set into the backbone network to obtain the stage feature map output by each stage of the backbone network; obtain the target feature information based on the calibration information of the training images, and assign the background features in the stage feature map to 0; the training data set includes the training images and the calibration information of the training images, and the calibration information includes the bounding box and category of the target; the target in this solution is the pedestrian; obtaining the target feature information based on the calibration information means selecting the correct coordinates and type information in the feature map based on the calibration information, and erroneous coordinates are classified as background features and ignored; the calibration information is generally manually calibrated, and the existing pre-trained high-accuracy model can also be used for machine calibration;
[0007] S3. Copy each stage feature map n-1 times, where n is the number of target categories in the training image. In each stage feature map, only one category of the target remains with a value, and the categories with values in each stage feature map are different. The targets of other categories are assigned 0; that is, the stage feature map and the target category are one-to-one corresponding.
[0008] S4, scaling the stage feature map obtained in step S3 to a specified width and height;
[0009] S5. Input the scaled stage feature maps one by one into the person re-identification branch and text feature matching branch of the corresponding stage, and each stage of the backbone network obtains n person re-identification feature vectors; input the target categories of the stage feature maps one by one into the text feature matching branch, and each stage of the backbone network obtains n text matching feature vectors; for example, the stage feature map outputted by the backbone network at stage i is inputted into the person re-identification branch added at stage i after the previous processing, and the text feature matching branch corresponds similarly; the target category of the stage feature map is the category of the target that is not assigned 0 in the stage feature map; each feature map corresponds to a text matching feature vector;
[0010] S6. Calculate the person re-identification loss at each stage based on the person re-identification feature vector and the target category in the calibration information of the training image;
[0011] S7. Calculate the text feature matching loss at each stage based on the text matching feature vector and the person re-identification feature vector;
[0012] S8. Input the feature map finally output by the backbone network into the detection head of the pedestrian detection algorithm model to predict the bounding box and category of the target;
[0013] S9, calculating a detection head loss based on the predicted bounding box and category and the bounding box and category of the target in the calibration information of the training image;
[0014] S10, correct the detection head loss based on pedestrian re-identification loss and text feature matching loss;
[0015] S11, add the pedestrian re-identification loss, text feature matching loss and the corrected detection head loss to obtain the total loss;
[0016] S12. Based on the total loss, update the parameters of the pedestrian detection algorithm model through the back propagation algorithm;
[0017] S13. Repeat steps S2 to S12 until the total loss is stable and the pedestrian detection algorithm model training is completed.
[0018] As a preference, in step S6, the pedestrian re-identification loss Obtained by the following formula:
[0019]
[0020] Where K is the total number of feature maps in this iteration. For example, if the backbone network has L stages, K = n*L. Indicates the reid loss of the kth stage feature map of this iteration, seq_list same Represents the set of feature sequences in the reid feature database that have the same target category as the k-th stage feature map, seq k Represents the pedestrian re-identification feature vector of the k-th feature map, N dif Indicates the number of categories in the reid feature database that are different from the target category of the k-th stage feature map. For example, there are T categories in the reid feature database, N dif It is T-1; It represents the set of feature sequences of different categories, that is, the set of feature sequences of the i-th category in the reid feature database that is different from the category of the k-th stage feature map. The max function is to compare each sequence of the same category in the database with seq k Calculate the cosine similarity (cos function in the formula) and take the maximum value.
[0021] As a preference, in step S7, the text feature matching loss Obtained by the following formula:
[0022]
[0023] Where, represents the text feature matching loss of the k-th stage feature map, Represents the text matching feature vector of the k-th stage feature map, Represents the text matching feature vector of the stage feature map with different target categories from the i-th stage feature map and the k-th stage feature map in this iteration, N dif_tIndicates the number of categories in this iteration that are different from the target category of the k-th stage feature map. For example, there are M target categories in this iteration, N dif-t This is M-1. After obtaining all text matching feature vectors, duplicate removal is performed. Because the same category corresponds to the same text features, the total number of target categories and the final total number of text matching features are the same.
[0024] Preferably, step S10 is specifically as follows:
[0025] S101. Map the predicted target coordinates to the feature map output by each stage of the backbone network, extract the corresponding feature map, and copy the feature map to the same number of copies as the target category. Each feature map retains a target mapping area of the category that is not repeated in other feature maps without modification, and assigns 0 to other areas. Then, scale the feature map to the specified size and input it to the pedestrian re-identification branch added in the corresponding stage to obtain the predicted pedestrian re-identification feature vector. The predicted pedestrian re-identification feature vector corresponding to the mth feature map is The target category of the predicted feature map is input into the text feature matching branch added in the corresponding stage to obtain the predicted text matching feature vector. The predicted text matching feature vector corresponding to the target category of the mth feature map is The whole process is similar to steps S2-S5, the main difference is that the mapping here is based on the prediction result output by the detection head, while the mapping in step S2 (i.e., obtaining target feature information according to the calibration information) is based on the calibration information;
[0026] S102: Perform IOU matching between the predicted target and the calibrated target. If the IOU is greater than or equal to the intersection-over-union (IOU) threshold, proceed to step S103; if the IOU is less than the IOU threshold, proceed to step S106.
[0027] S103, determine whether the predicted target category is the same as the calibrated target category, if they are the same, proceed to step S104, if they are different, proceed to step S105;
[0028] S104. Correct the detection head loss according to the following formula:
[0029]
[0030] Where, represents the corrected detection head loss, represents the most dissimilar person re-identification similarity value at each stage, Indicates the most dissimilar text matching similarity value at each stage, Represents the set of feature sequences in the re-id feature database that have the same target category as the feature map of the mth stage; loss clsIndicates the target detection conventional category loss, such as: cross entropy loss, loss bbox Indicates the conventional regression loss of target detection, such as Smooth L1 Loss; list indicates a list;
[0031] S105. Correct the detection head loss according to the following formula:
[0032]
[0033] Where, represents the corrected detection head loss, represents the most dissimilar person re-identification similarity value at each stage, Indicates the most dissimilar text matching similarity value at each stage, Represents the set of feature sequences in the re-id feature database that have the same target category as the feature map of the mth stage; loss cls Indicates the target detection conventional category loss, such as: cross entropy loss, loss bbox Represents the conventional regression loss for target detection, such as Smooth L1 Loss. Conventional category loss and conventional regression loss for target detection are both well-known technologies and will not be expanded here.
[0034] S106. Correct the detection head loss according to the following formula:
[0035]
[0036] Where, represents the corrected detection head loss, represents the most dissimilar person re-identification similarity value at each stage, Indicates the most dissimilar text matching similarity value at each stage, Represents the set of feature sequences in the re-id feature database that have the same target category as the feature map of the mth stage; loss cls Represents the conventional category loss of target detection, such as cross entropy loss.
[0037] Preferably, the re-identification feature database is obtained by inputting calibrated training and test set images into a pedestrian detection algorithm model, matching the predicted results with the calibration, inputting the correct matched results into a trained person re-identification branch to obtain a feature sequence, constructing the feature sequence into a two-dimensional matrix, and performing PCA dimensionality reduction to obtain the re-identification feature database. The re-identification feature database is established after the training of the person re-identification branch is completed.
[0038] Preferably, the person re-identification branch includes the following sequentially connected branches:
[0039] Self-attention layer: performs self-attention on the feature map to obtain self-attention feature values;
[0040] Feature fusion layer: Add the feature map and the self-attention feature value to obtain a fused feature map;
[0041] Convolution sliding frame layer: Perform convolution sliding frame operation on the fused feature map to extract convolution features;
[0042] Average pooling layer: performs average pooling operation on the convolution features to obtain pooling features;
[0043] Fully connected layer: Perform a fully connected operation on the pooled features to obtain pedestrian re-identification features.
[0044] As a preferred method, the training process of the person re-identification branch is as follows:
[0045] A1. Freeze the backbone network parameters of the pedestrian detection algorithm model and exclude them from the training process. Before training, assign the existing backbone network parameters of the pedestrian detection target to the backbone network variables.
[0046] A2. Input the training set into the pedestrian detection algorithm model to obtain the prediction results. Pair the prediction results with the calibration results and perform IOU matching first. If IOU ≥ 0.5, then perform category matching. If the categories are consistent, assign 0 to the part of the stage feature map output by each stage of the backbone network except the target. If IOU < 0.5 or the categories are inconsistent, skip this training image. If the target of the prediction result has n different categories, n ≥ 2, copy the feature map into n copies. Each stage feature map corresponds to one type of target and there is no duplication. The target corresponding to the stage feature map is retained, and the other parts are assigned 0.
[0047] A3. Scale the stage feature map to the specified width and height, and then input it into the person re-identification branch. Each stage of the backbone network obtains n training person re-identification feature vectors. The training person re-identification feature vector corresponding to the jth stage feature map is seq j ;
[0048] A4. Calculate the person re-identification training loss using the following formula
[0049]
[0050] Where N all Indicates the total number of feature maps after replication in this iteration. It is the same as the previous K concept, but is applied to different training stages. Represents the similarity loss value of the same category of the j-th stage feature map, N sameIndicates the number of stage feature maps with the same category as the jth stage feature map in this iteration, seq i Represents the training person re-identification feature vector of the i-th feature map in the stage feature map set with the same target category as the j-th stage feature map, seq p N represents the training person re-identification feature vector of the p-th stage feature map in the stage feature map set with a different target category from the j-th stage feature map; dif_j Indicates the number of stage feature maps in this iteration that are different from the j-th stage feature map category;
[0051] A5. Based on the person re-identification training loss, the parameters of the person re-identification branch are updated through the back-propagation algorithm;
[0052] A6. Repeat steps A2 to A5 until the person re-ID training loss is stable and the person re-ID branch training is completed.
[0053] Preferably, the text feature matching branch is a CLIP encoding head, and the training process of the text feature matching branch is:
[0054] B1. Freeze the backbone network parameters and person re-identification branch parameters of the pedestrian detection algorithm model and exclude them from the training process.
[0055] B2. Input the training set into the pedestrian detection algorithm model to obtain the prediction results. Pair the prediction results with the calibration results and perform IOU matching first. If IOU ≥ 0.5, then perform category matching. If the categories are consistent, assign the value of 0 to the part of the stage feature map output by each stage of the backbone network except the target. If IOU < 0.5 or the categories are inconsistent, skip this training image. If there are n different category results predicted, n ≥ 2, copy the stage feature map into n copies, each of which corresponds to one target and is not repeated. The target corresponding to the stage feature map is retained, and the other parts are assigned 0.
[0056] B3. Scale the stage feature map to the specified width and height, and then input it into the person re-identification branch. Each stage of the backbone network corresponds to n training person re-identification feature vectors. The training person re-identification feature vector corresponding to the jth stage feature map is seq j ; Input the target category of the stage feature map into the text feature matching branch one by one, and obtain n training text matching feature vectors in each stage. Then, the training text matching feature vectors are deduplicated. The training text matching feature vector corresponding to the jth stage feature map is
[0057] B4. Calculate the text feature matching training loss using the following formula
[0058]
[0059] Where, Indicates the total number of stage feature maps after this iteration, Represents the text similarity loss value of the j-th stage feature map, Indicates the number of feature maps with different target categories from the j-th stage feature map in this iteration, seq r The text matching feature vector representing the feature map with different target categories between the rth and jth stage feature maps in this iteration; Represents the similarity loss value of different categories of the j-th feature map; That is, the loss value of the image after all correct targets are copied;
[0060] B5. Based on the text feature matching training loss, update the parameters of the text feature matching branch through the backpropagation algorithm;
[0061] B6. Repeat steps B2 to B5 until the text feature matching training is stable and the text feature matching branch training is completed.
[0062] The substantial effects brought by the present invention are: ① The backbone network (this solution adopts ResNet50) outputs a feature map in each stage to perform re-id and text feature supervision in the existing calibration frame, find the stage where the model is weak in feature extraction, and strengthen the feature extraction capability, thereby improving the model's sensitivity to features; ② The prediction results of the pedestrian detection algorithm model are compared with the existing calibration frame in terms of category and coordinates, and the region of interest is reversely calculated to obtain feature value alignment for re-id supervision and text feature supervision, thereby improving the model's ability to general features and reducing the interference of individual features, thereby improving the model's overall ability to extract effective features and reducing false detection; ③ Increasing the loss calculation of each stage to improve the feature capability of the stage; ④ The prediction results reversely calculate the accuracy of each prediction at each stage, strengthen the learning of erroneous targets, and thus improve the prediction capability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a structural schematic diagram of the present invention. DETAILED DESCRIPTION
[0064] The technical solution of the present invention will be further specifically described below through embodiments and in conjunction with the accompanying drawings.
[0065] Example: This embodiment is a pedestrian detection algorithm training method based on pedestrian re-identification and text feature matching supervision, such as Figure 1 As shown, the following steps are included:
[0066] S1. Constructing the training structure of the pedestrian detection algorithm model: Add a pedestrian re-identification branch and a text feature matching branch to the output end of each stage of the backbone network of the pedestrian detection algorithm model. After training the added pedestrian re-identification branch and text feature matching branch, freeze the two branches. The branches added at each stage of the backbone network do not modify the original input-output relationship of the backbone network. For example, the output of the first stage originally entered the second stage, and it remains the same after the addition, except that the output of each stage also enters the two newly added branches. The parameters of the branches added at each stage are independent after training. The subsequent training process is based on the pedestrian detection algorithm model training structure. After the training is completed, the original pedestrian detection algorithm model is still used for formal reasoning, and the two newly added branches are not needed. The backbone network adopts ResNet50.
[0067] S2. Input the training images of the training set into the backbone network to obtain the stage feature map output by each stage of the backbone network; obtain the target feature information based on the calibration information of the training images, and assign the background features in the stage feature map to 0; the training data set includes the training images and the calibration information of the training images, and the calibration information includes the bounding box and category of the target; the target in this solution is the pedestrian; obtaining the target feature information based on the calibration information means selecting the correct coordinates and type information in the feature map based on the calibration information, and erroneous coordinates are classified as background features and ignored; the calibration information is generally manually calibrated, and the existing pre-trained high-accuracy model can also be used for machine calibration;
[0068] S3. Copy each stage feature map n-1 times, where n is the number of target categories in the training image. In each stage feature map, only one category of the target remains with a value, and the categories with values in each stage feature map are different. The targets of other categories are assigned 0; that is, the stage feature map and the target category are one-to-one corresponding.
[0069] S4. Scale the stage feature map obtained in step S3 to the specified width and height. This solution uses an upsampling method to keep the feature map of each stage consistent with the first stage.
[0070] S5. Input the scaled stage feature maps one by one into the person re-identification branch and text feature matching branch of the corresponding stage. Each stage of the backbone network obtains n person re-identification feature vectors. Input the target categories of the stage feature maps one by one into the text feature matching branch. Each stage of the backbone network obtains n text matching feature vectors. For example, the stage feature map outputted by the backbone network at stage i is inputted into the person re-identification branch added at stage i after the previous processing. The text feature matching branch is similar. The target category of the stage feature map is the category of the target that is not assigned a value of 0 in the stage feature map. Each feature map corresponds to a text matching feature vector.
[0071] S6. Calculate the person re-identification loss at each stage based on the person re-identification feature vector and the target category in the calibration information of the training image;
[0072] S7. Calculate the text feature matching loss at each stage based on the text matching feature vector and the person re-identification feature vector;
[0073] S8. Input the feature map finally output by the backbone network into the detection head of the pedestrian detection algorithm model to predict the bounding box and category of the target;
[0074] S9, calculating a detection head loss based on the predicted bounding box and category and the bounding box and category of the target in the calibration information of the training image;
[0075] S10, correct the detection head loss based on pedestrian re-identification loss and text feature matching loss;
[0076] S11, add the pedestrian re-identification loss, text feature matching loss and the corrected detection head loss to obtain the total loss;
[0077] S12. Based on the total loss, update the parameters of the pedestrian detection algorithm model through the back propagation algorithm;
[0078] S13. Repeat steps S2 to S12 until the total loss is stable and the pedestrian detection algorithm model training is completed.
[0079] In step S6, the pedestrian re-identification loss Obtained by the following formula:
[0080]
[0081] Where K is the total number of feature maps in this iteration. For example, if the backbone network has L stages, K = n*L. Indicates the reid loss of the kth stage feature map of this iteration, seq_list same Represents the set of feature sequences in the reid feature database that have the same target category as the k-th stage feature map, seq k Represents the pedestrian re-identification feature vector of the kth feature map, N dif Indicates the number of categories in the reid feature database that are different from the target category of the k-th stage feature map. For example, there are T categories in the reid feature database, N dif It is T-1; It represents the set of feature sequences of different categories, that is, the set of feature sequences of the i-th category in the reid feature database that is different from the category of the k-th stage feature map. The max function is to compare each sequence of the same category in the database with seqk Calculate the cosine similarity (cos function in the formula) and take the maximum value.
[0082] In step S7, the text feature matching loss Obtained by the following formula:
[0083]
[0084] Where, represents the text feature matching loss of the k-th stage feature map, Represents the text matching feature vector of the k-th stage feature map, Represents the text matching feature vector of the stage feature map with different target categories from the i-th stage feature map and the k-th stage feature map in this iteration, N dif_t Indicates the number of categories in this iteration that are different from the target category of the k-th stage feature map. For example, there are M target categories in this iteration, N dif-t This is M-1. After obtaining all text matching feature vectors, duplicate removal is performed. Because the same category corresponds to the same text features, the total number of target categories and the final total number of text matching features are the same.
[0085] Step S10 is specifically as follows:
[0086] S101. Map the predicted target coordinates to the feature map output by each stage of the backbone network, extract the corresponding feature map, and copy the feature map to the same number of copies as the target category. Each feature map retains a target mapping area of the category that is not repeated in other feature maps without modification, and assigns the value of other areas to 0. Then, scale the feature map to the specified size and input it to the pedestrian re-identification branch added in the corresponding stage to obtain the predicted pedestrian re-identification feature vector. The predicted pedestrian re-identification feature vector corresponding to the mth feature map is The target category of the predicted feature map is input into the text feature matching branch added in the corresponding stage to obtain the predicted text matching feature vector. The predicted text matching feature vector corresponding to the target category of the mth feature map is The whole process is similar to steps S2-S5, the main difference is that the mapping here is based on the prediction result output by the detection head, while the mapping in step S2 (i.e., obtaining target feature information according to the calibration information) is based on the calibration information;
[0087] S102: Perform IOU matching between the predicted target and the calibrated target. If the IOU is greater than or equal to the intersection-over-union (IOU) threshold, proceed to step S103; if the IOU is less than the IOU threshold, proceed to step S106.
[0088] S103, determine whether the predicted target category is the same as the calibrated target category, if they are the same, proceed to step S104, if they are different, proceed to step S105;
[0089] S104. Correct the detection head loss according to the following formula:
[0090]
[0091] Where, represents the corrected detection head loss, represents the most dissimilar person re-identification similarity value at each stage, Indicates the most dissimilar text matching similarity value at each stage, Represents the set of feature sequences in the re-id feature database that have the same target category as the feature map of the mth stage; loss cls Indicates the target detection conventional category loss, such as: cross entropy loss, loss bbox Indicates the conventional regression loss of target detection, such as Smooth L1 Loss; list indicates a list;
[0092] S105. Correct the detection head loss according to the following formula:
[0093]
[0094] Where, represents the corrected detection head loss, represents the most dissimilar person re-identification similarity value at each stage, Indicates the most dissimilar text matching similarity value at each stage, Represents the set of feature sequences in the re-id feature database that have the same target category as the feature map of the mth stage; loss cls Indicates the target detection conventional category loss, such as: cross entropy loss, loss bbox Represents the conventional regression loss for target detection, such as Smooth L1 Loss. Conventional category loss and conventional regression loss for target detection are both well-known technologies and will not be expanded here.
[0095] S106. Correct the detection head loss according to the following formula:
[0096]
[0097] Where, represents the corrected detection head loss, represents the most dissimilar person re-identification similarity value at each stage, Indicates the most dissimilar text matching similarity value at each stage, Represents the set of feature sequences in the re-id feature database that have the same target category as the feature map of the mth stage; loss cls Represents the conventional category loss of target detection, such as cross entropy loss.
[0098] The ReID feature database is generated by inputting calibrated training and test set images into a person detection algorithm model, matching the predicted results with the calibration, and then inputting the correct matched results into the trained person re-identification branch to obtain a feature sequence. This feature sequence is then constructed into a two-dimensional matrix and subjected to PCA dimensionality reduction to obtain the ReID feature database. The ReID feature database is established after the person re-identification branch training is completed.
[0099] The person re-identification branch includes the following connected in sequence:
[0100] Self-attention layer: The self-attention mechanism is performed on the feature map to obtain the self-attention feature value; the self-attention mechanism is performed on the feature map at each stage, so that each feature value of the feature map has correlation and global correlation
[0101] Feature fusion layer: Add the feature map and the self-attention feature value to obtain a fused feature map;
[0102] Convolution sliding frame layer: Perform convolution sliding frame operation on the fused feature map to extract convolution features;
[0103] Average pooling layer: performs average pooling operation on the convolution features to obtain pooling features;
[0104] Fully connected layer: Perform a fully connected operation on the pooled features to obtain pedestrian re-identification features.
[0105] The training process of the person re-identification branch is:
[0106] A1. Freeze the backbone network parameters of the pedestrian detection algorithm model and exclude them from the training process. Before training, assign the existing backbone network parameters of the pedestrian detection target to the backbone network variables.
[0107] A2. Input the training set into the pedestrian detection algorithm model to obtain the prediction results. Pair the prediction results with the calibration results and perform IOU matching first. If IOU ≥ 0.5, then perform category matching. If the categories are consistent, assign 0 to the part of the stage feature map output by each stage of the backbone network except the target. If IOU < 0.5 or the categories are inconsistent, skip this training image. If the target of the prediction result has n different categories, n ≥ 2, copy the feature map into n copies. Each stage feature map corresponds to one type of target and there is no duplication. The target corresponding to the stage feature map is retained, and the other parts are assigned 0.
[0108] A3. Scale the stage feature map to the specified width and height, and then input it into the person re-identification branch. Each stage of the backbone network obtains n training person re-identification feature vectors. The training person re-identification feature vector corresponding to the jth stage feature map is seq j ;
[0109] A4. Calculate the person re-identification training loss using the following formula
[0110]
[0111] Where N all Indicates the total number of feature maps after replication in this iteration. It is the same as the previous K concept, but is applied to different training stages. Represents the similarity loss value of the same category of the j-th stage feature map, N same Indicates the number of stage feature maps with the same category as the jth stage feature map in this iteration, seq i Represents the training person re-identification feature vector of the i-th feature map in the stage feature map set with the same target category as the j-th stage feature map, seq p N represents the training person re-identification feature vector of the p-th stage feature map in the stage feature map set with a different target category from the j-th stage feature map; dif_j Indicates the number of stage feature maps in this iteration that are different from the j-th stage feature map category;
[0112] A5. Based on the person re-identification training loss, the parameters of the person re-identification branch are updated through the back-propagation algorithm;
[0113] A6. Repeat steps A2 to A5 until the person re-ID training loss is stable and the person re-ID branch training is completed.
[0114] The text feature matching branch is the CLIP encoding head. The training process of the text feature matching branch is as follows:
[0115] B1. Freeze the backbone network parameters and person re-identification branch parameters of the pedestrian detection algorithm model and exclude them from the training process.
[0116] B2. Input the training set into the pedestrian detection algorithm model to obtain the prediction results. Pair the prediction results with the calibration results and perform IOU matching first. If IOU ≥ 0.5, then perform category matching. If the categories are consistent, assign the value of 0 to the part of the stage feature map output by each stage of the backbone network except the target. If IOU < 0.5 or the categories are inconsistent, skip this training image. If there are n different category results predicted, n ≥ 2, copy the stage feature map into n copies, each of which corresponds to one target and is not repeated. The target corresponding to the stage feature map is retained, and the other parts are assigned 0.
[0117] B3. Scale the stage feature map to the specified width and height, and then input it into the person re-identification branch. Each stage of the backbone network corresponds to n training person re-identification feature vectors. The training person re-identification feature vector corresponding to the jth stage feature map is seq j ; Input the target category of the stage feature map into the text feature matching branch one by one, and obtain n training text matching feature vectors in each stage. Then, the training text matching feature vectors are deduplicated. The training text matching feature vector corresponding to the jth stage feature map is
[0118] B4. Calculate the text feature matching training loss using the following formula
[0119]
[0120] Where, Indicates the total number of stage feature maps after this iteration, Represents the text similarity loss value of the j-th stage feature map, Indicates the number of categories in this iteration that are different from the target category of the j-th stage feature map, seq r Represents the training text matching feature vector of the rth target category in the target category set that is different from the target category of the jth stage feature map in this iteration; Represents the similarity loss value of different categories of the j-th feature map; That is, the loss value of the image after all correct targets are copied;
[0121] B5. Based on the text feature matching training loss, update the parameters of the text feature matching branch through the backpropagation algorithm;
[0122] B6. Repeat steps B2 to B5 until the text feature matching training is stable and the text feature matching branch training is completed.
[0123] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Persons skilled in the art may make various modifications, additions, or substitutions to the described specific embodiments without departing from the spirit of the present invention or exceeding the scope of the appended claims.
[0124] Although this article frequently uses terms such as person re-identification and text feature matching, the use of other terms is not excluded. These terms are used solely to more conveniently describe and explain the essence of the present invention; interpreting them as any additional limitations is contrary to the spirit of the present invention.
Claims
1. A pedestrian detection algorithm training method based on pedestrian re-identification and text feature matching supervision, characterized in that: The following steps are involved: S1. Constructing a pedestrian detection algorithm model training structure: Add a pedestrian re-identification branch and a text feature matching branch to the output of each stage of the pedestrian detection algorithm model's backbone network. After training the added pedestrian re-identification branch and text feature matching branch, freeze the two branches. S2. Input the training images of the training set into the backbone network to obtain the stage feature map output by each stage of the backbone network; obtain the target feature information based on the calibration information of the training images, and assign the background features in the stage feature map to 0; S3. Copy each stage feature map n-1 times, where n is the number of target categories in the training image. In each stage feature map, only one category of the target remains with a value, and the categories of the target in each stage feature map are different. The targets of other categories are assigned a value of 0. S4, scaling the stage feature map obtained in step S3 to a specified width and height; S5. Input the scaled stage feature maps one by one into the person re-identification branch and text feature matching branch of the corresponding stage, and obtain n person re-identification feature vectors at each stage of the backbone network; input the target categories of the stage feature maps one by one into the text feature matching branch, and obtain n text matching feature vectors at each stage of the backbone network; S6. Calculate the person re-identification loss at each stage based on the person re-identification feature vector and the target category in the calibration information of the training image; S7. Calculate the text feature matching loss at each stage based on the text matching feature vector and the person re-identification feature vector; S8. Input the feature map finally output by the backbone network into the detection head of the pedestrian detection algorithm model to predict the bounding box and category of the target; S9, calculating a detection head loss based on the predicted bounding box and category and the bounding box and category of the target in the calibration information of the training image; S10, correct the detection head loss based on pedestrian re-identification loss and text feature matching loss; S11, add the pedestrian re-identification loss, text feature matching loss and the corrected detection head loss to obtain the total loss; S12. Based on the total loss, update the parameters of the pedestrian detection algorithm model through the back propagation algorithm; S13. Repeat steps S2 to S12 until the total loss is stable and the pedestrian detection algorithm model training is completed.
2. The method for training a pedestrian detection algorithm based on person re-identification and text feature matching supervision according to claim 1, characterized in that: In step S6, the pedestrian re-identification loss Obtained by the following formula: Where K is the total number of feature maps in this iteration, Indicates the reid loss of the kth stage feature map of this iteration, seq_list same Represents the set of feature sequences in the reid feature database that have the same target category as the k-th stage feature map, seq k Represents the pedestrian re-identification feature vector of the kth feature map, N dif Indicates the number of categories in the reid feature database that are different from the target category of the k-th stage feature map; Represents the i-th set of feature sequences of different categories.
3. The method for training a pedestrian detection algorithm based on person re-identification and text feature matching supervision according to claim 2, characterized in that: In step S7, the text feature matching loss Obtained by the following formula: Where, represents the text feature matching loss of the k-th stage feature map, Represents the text matching feature vector of the k-th stage feature map, Represents the text matching feature vector of the stage feature map with different target categories from the i-th stage feature map and the k-th stage feature map in this iteration, N dif_t Indicates the number of feature maps in this iteration that have different target categories from the k-th stage feature map.
4. The method for training a pedestrian detection algorithm based on person re-identification and text feature matching supervision according to claim 3, characterized in that: Step S10 is specifically as follows: S101. Map the predicted target coordinates to the feature map output by each stage of the backbone network, extract the corresponding feature map, and copy the feature map to the same number of copies as the target category. Each feature map retains a target mapping area of the category that is not repeated in other feature maps without modification, and assigns the value of other areas to 0. Then, scale the feature map to the specified size and input it to the pedestrian re-identification branch added in the corresponding stage to obtain the predicted pedestrian re-identification feature vector. The predicted pedestrian re-identification feature vector corresponding to the mth feature map is The target category of the predicted feature map is input into the text feature matching branch added in the corresponding stage to obtain the predicted text matching feature vector. The predicted text matching feature vector corresponding to the target category of the mth feature map is S102: Perform IOU matching between the predicted target and the calibrated target. If the IOU is greater than or equal to the intersection-over-union (IOU) threshold, proceed to step S103; if the IOU is less than the IOU threshold, proceed to step S106. S103, determine whether the predicted target category is the same as the calibrated target category, if they are the same, proceed to step S104, if they are different, proceed to step S105; S104. Correct the detection head loss according to the following formula: Where, represents the corrected detection head loss, represents the most dissimilar person re-identification similarity value at each stage, Indicates the most dissimilar text matching similarity value at each stage, Represents the set of feature sequences in the re-id feature database that have the same target category as the feature map of the mth stage; loss cls Represents the target detection general category loss, loss bbox represents the conventional regression loss for target detection; S105. Correct the detection head loss according to the following formula: Where, represents the corrected detection head loss, represents the most dissimilar person re-identification similarity value at each stage, Indicates the most dissimilar text matching similarity value at each stage, Represents the set of feature sequences in the re-id feature database that have the same target category as the feature map of the mth stage; loss cls Represents the target detection general category loss, loss bbox represents the conventional regression loss for target detection; S106. Correct the detection head loss according to the following formula: Where, represents the corrected detection head loss, represents the most dissimilar person re-identification similarity value at each stage, Indicates the most dissimilar text matching similarity value at each stage, Represents the set of feature sequences in the re-id feature database that have the same target category as the feature map of the mth stage; loss cls Represents the general category loss for object detection.
5. The method for training a pedestrian detection algorithm based on person re-identification and text feature matching supervision according to claim 2, 3 or 4, characterized in that: The re-identification feature database is obtained by inputting the calibrated training set and test set images into the pedestrian detection algorithm model, matching the obtained prediction results with the calibration, matching the correct input to the trained pedestrian re-identification branch to obtain a feature sequence, constructing the feature sequence into a two-dimensional matrix and performing PCA dimensionality reduction to obtain the re-identification feature database.
6. The method for training a pedestrian detection algorithm based on person re-identification and text feature matching supervision according to claim 1, characterized in that: The person re-identification branch includes the following connected in sequence: Self-attention layer: performs self-attention on the feature map to obtain self-attention feature values; Feature fusion layer: Add the feature map and the self-attention feature value to obtain a fused feature map; Convolution sliding frame layer: Perform convolution sliding frame operation on the fused feature map to extract convolution features; Average pooling layer: performs average pooling operation on the convolution features to obtain pooling features; Fully connected layer: Perform a fully connected operation on the pooled features to obtain pedestrian re-identification features.
7. The method for training a pedestrian detection algorithm based on person re-identification and text feature matching supervision according to claim 6, characterized in that: The training process of the person re-identification branch is: A1. Freeze the backbone network parameters of the pedestrian detection algorithm model and exclude them from the training process. A2. Input the training set into the pedestrian detection algorithm model to obtain the prediction results. Pair the prediction results with the calibration results and perform IOU matching first. If IOU ≥ 0.5, then perform category matching. If the categories are consistent, assign 0 to the part of the stage feature map output by each stage of the backbone network except the target. If IOU < 0.5 or the categories are inconsistent, skip this training image. If the target of the prediction result has n different categories, n ≥ 2, copy the feature map into n copies. Each stage feature map corresponds to one type of target and there is no duplication. The target corresponding to the stage feature map is retained, and the other parts are assigned 0. A3. Scale the stage feature map to the specified width and height, and then input it into the person re-identification branch. Each stage of the backbone network obtains n training person re-identification feature vectors. The training person re-identification feature vector corresponding to the jth stage feature map is seq j ; A4. Calculate the person re-identification training loss using the following formula Where N all Indicates the total number of stage feature maps after replication in this iteration, Represents the similarity loss value of the same category of the j-th stage feature map, N same Indicates the number of stage feature maps with the same category as the jth stage feature map in this iteration, seq i Represents the training person re-identification feature vector of the i-th feature map in the stage feature map set with the same target category as the j-th stage feature map, seq p N represents the training person re-identification feature vector of the p-th stage feature map in the stage feature map set with a different target category from the j-th stage feature map in this iteration; dif_j Indicates the number of stage feature maps in this iteration that are different from the j-th stage feature map category; A5. Based on the person re-identification training loss, the parameters of the person re-identification branch are updated through the back-propagation algorithm; A6. Repeat steps A2 to A5 until the person re-ID training loss is stable and the person re-ID branch training is completed.
8. The method for training a pedestrian detection algorithm based on person re-identification and text feature matching supervision according to claim 7, characterized in that: The text feature matching branch is the CLIP encoding head. The training process of the text feature matching branch is as follows: B1. Freeze the backbone network parameters and person re-identification branch parameters of the pedestrian detection algorithm model and exclude them from the training process. B2. Input the training set into the pedestrian detection algorithm model to obtain the prediction results. Pair the prediction results with the calibration results and perform IOU matching first. If IOU ≥ 0.5, then perform category matching. If the categories are consistent, assign the value of 0 to the part of the stage feature map output by each stage of the backbone network except the target. If IOU < 0.5 or the categories are inconsistent, skip this training image. If there are n different category results predicted, n ≥ 2, copy the stage feature map into n copies, each of which corresponds to one target and is not repeated. The target corresponding to the stage feature map is retained, and the other parts are assigned 0. B3. Scale the stage feature map to the specified width and height, and then input it into the person re-identification branch. Each stage of the backbone network corresponds to n training person re-identification feature vectors. The training person re-identification feature vector corresponding to the jth stage feature map is seq j ; Input the target categories of the stage feature graphs into the text feature matching branch one by one, and obtain n training text matching feature vectors in each stage. The training text matching feature vector corresponding to the jth stage feature graph is B4. Calculate the text feature matching training loss using the following formula Where, Indicates the total number of stage feature maps after this iteration, Represents the text similarity loss value of the j-th stage feature map, Indicates the number of feature maps with different target categories from the j-th stage feature map in this iteration, seq r The text matching feature vector representing the feature map with different target categories between the rth and jth stage feature maps in this iteration; Represents the similarity loss value of different categories of the j-th feature map; B5. Based on the text feature matching training loss, update the parameters of the text feature matching branch through the backpropagation algorithm; B6. Repeat steps B2 to B5 until the text feature matching training is stable and the text feature matching branch training is completed.
Citation Information
Patent Citations
Pedestrian re-identification method based on image and text dual-channel combination
CN114612927A
Implicit category prompt learning method for pedestrian image lacking text information
CN117711070A
Text-image pedestrian re-identification method based on multi-scale information interaction network
CN117727069A
Person re-identification method and apparatus for fusing global features with ladder-shaped local features
WO2024021394A1