Training method, pedestrian re-identification method, medium and electronic device

By conducting two trainings on the neural network model, combining source domain and target domain data, generating pseudo-labels and performing feature fusion, the generalization and accuracy of the neural network model in cross-domain data recognition is solved, and the generalization of the model and the recognition ability of cross-domain data are improved.

CN114445775BActive Publication Date: 2025-07-11WINNERYUN (SHANGHAI DATA SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210054428.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-18
Publication Date
2025-07-11
Estimated Expiration
2042-01-18

AI Technical Summary

Technical Problem

Existing neural network models are not generalized when there is insufficient data volume, resulting in low retrieval accuracy when cross-domain data.

Method used

By obtaining the source domain data and the target domain data, first training is performed using the initial neural network model, joint features are obtained, and classification is performed based on the feature to generate the target domain data of the pseudo-label, and then a second training is performed on the initial neural network model to obtain the second neural network model.

Benefits of technology

The generalization of neural network model and the identification accuracy of cross-domain data are improved. By combining domain information and pedestrian information, the target domain data with pseudo-labels is expanded as training data, and the recognition effect of the model is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114445775B_ABST
    Figure CN114445775B_ABST
Patent Text Reader

Abstract

The present invention provides a training method, a pedestrian re-identification method, a medium and an electronic device. The training method includes: obtaining source domain data and target domain data; performing first training on an initial neural network model based on the source domain data to obtain a first neural network model; processing the target domain data through the first neural network model to obtain joint features of the target domain data, where the joint features are obtained by weighted fusion of pedestrian features and domain features of the target domain data; classifying the target domain data based on the joint features, and obtaining the target domain data with pseudo-labels according to the classification results; performing second training on the first neural network model based on the source domain data and the target domain data with pseudo-labels to obtain a second neural network model. This method can improve the generalization of the neural network model and the accuracy of the neural network model when processing cross-domain data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a training method for a neural network, and in particular to a training method, a person re-identification method, a medium, and an electronic device. Background Art

[0002] With the development of many technologies related to video surveillance systems, the person re-identification technology has been widely applied. The person re-identification technology is a technology that uses an algorithm to search for a target person in an image library. Among many current algorithms, extracting features from a person image through a neural network model can realize the digitalization of the retrieval object. This method is one of the most popular and common methods, mainly due to the high generalization and excellent efficiency of the neural network model, and it can quickly realize the retrieval function in a large data group without the requirement of face recognition.

[0003] In a person re-identification system with a neural network as the core, data is a very important link. In the currently popular neural network re-identification system, due to the lack of a sufficient amount of data, the generalization of the neural network model is not high, and when facing cross-domain data, the probability that the neural network model hits the correct retrieval is small. Summary of the Invention

[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a training method, a person re-identification method, a medium, and an electronic device, which are used to solve the problems of low generalization of the neural network model and low accuracy when processing cross-domain data in the prior art.

[0005] To achieve the above object and other related objects, a first aspect of the present invention provides a training method, which includes obtaining source domain data and target domain data, where the source domain data is labeled person data, and the target domain data is unlabeled person data; performing first training on an initial neural network model based on the source domain data to obtain a first neural network model; processing the target domain data through the first neural network model to obtain joint features of the target domain data, where the joint features are obtained by weighted fusion of the person features and domain features of the target domain data; classifying the target domain data based on the joint features, and obtaining the target domain data with pseudo-labels according to the classification result; performing second training on the first neural network model based on the source domain data and the target domain data with pseudo-labels to obtain a second neural network model.

[0006] In an embodiment of the first aspect, a method for obtaining a first neural network model by performing a first training on an initial neural network model based on the source domain data includes: processing the source domain data through the initial neural network model to obtain pedestrian features, domain features, the number of pedestrians, and the number of domains of the source domain data; obtaining a pedestrian category prediction result based on the pedestrian features; obtaining a pedestrian category prediction loss through a first classification loss function based on the pedestrian category prediction result and the number of pedestrians; obtaining a domain category prediction result based on the domain features; obtaining a domain category prediction loss through a second classification loss function based on the domain category prediction result and the number of domains; obtaining a combined feature of the pedestrian features and the domain features by weighted fusion of the pedestrian features and the domain features; obtaining a combined feature loss through a metric learning loss function based on the combined feature; and updating parameters of the initial network model based on the pedestrian category prediction loss, the domain category prediction loss, and the combined feature loss.

[0007] In an embodiment of the first aspect, a method for obtaining a second neural network model by performing a second training on the first neural network model based on the source domain data and the target domain data with pseudo-labels includes: obtaining pedestrian features, domain features, the number of pedestrians, and the number of domains of the source domain data and the target domain data with pseudo-labels; obtaining a pedestrian category prediction result based on the pedestrian features; obtaining a pedestrian category prediction loss through a third classification loss function based on the pedestrian category prediction result and the number of pedestrians; obtaining a domain category prediction result based on the domain features; obtaining a domain category prediction loss through a fourth classification loss function based on the domain category prediction result and the number of domains; obtaining a combined feature of the pedestrian features and the domain features by weighted fusion of the pedestrian features and the domain features; obtaining a combined feature loss through a metric learning loss function based on the combined feature; and updating parameters of the first neural network model based on the pedestrian category prediction loss, the domain category prediction loss, and the combined feature loss.

[0008] In an embodiment of the first aspect, the third classification loss function is expressed by the following formula:

[0009]

[0010] where K is the number of pedestrians, a is a hyperparameter, and the value range of a is 0 - 1, and the Loss CrossEntropy can be expressed by the following formula:

[0011]

[0012] where the P i represents the true probability, q i represents the predicted probability, and the P iIt can be expressed by the following formula:

[0013]

[0014] In an embodiment of the first aspect, the classification loss function is a cross - entropy loss function, and the metric learning loss function is a triplet loss function.

[0015] In an embodiment of the first aspect, the pedestrian data includes pedestrian images, and the training method further includes: deleting pedestrian images with less than 4 pedestrians in the target domain data with pseudo - labels.

[0016] In an embodiment of the first aspect, the initial neural network model includes a common feature extraction layer, a pedestrian feature extraction layer, a domain feature extraction layer, and a feature fusion layer. Among them, the common feature extraction layer is used to obtain common features, the pedestrian feature extraction layer is used to obtain pedestrian features of the input data according to the common features, the domain feature extraction layer is used to obtain domain features of the input data according to the common features, and the feature fusion layer is used to perform weighted fusion according to the pedestrian features and the domain features of the input data to obtain joint features of the input data, where the input data includes the source domain data and / or the target domain data.

[0017] The second aspect of the present invention provides a pedestrian re - identification method. The pedestrian re - identification method includes: obtaining target domain data, where the target domain data is unlabeled pedestrian data; processing the target domain data through a neural network model to obtain joint features of the target domain data, and the neural network model executes the training method according to any one of the first aspects of the present invention during training. Classifying the target domain data based on the joint features, and obtaining labels of the target domain data according to the classification results.

[0018] The third aspect of the present invention provides a computer - readable storage medium, and when the computer program is executed by a processor, it implements the training method according to any one of the first aspects of the present invention and / or the pedestrian re - identification method according to the third aspect.

[0019] The fourth aspect of the present invention provides an electronic device. The electronic device includes: a memory on which a computer program is stored; a processor communicatively connected to the memory, and when calling the computer program, it executes the training method according to any one of the first aspects of the present invention and / or the pedestrian re - identification method according to the third aspect.

[0020] As described above, the training method, pedestrian re - identification method, medium, and electronic device of the present invention have the following beneficial effects:

[0021] The training method can obtain a first neural network model by performing a first training on an initial neural network model based on source domain data. Processing the target domain data through the first neural network model can obtain the joint features of the target domain data, where the joint features are obtained by weighted fusion of the pedestrian features and domain features of the target domain data. Classifying the target domain features based on the joint features can obtain corresponding classification results, and thus obtain the target domain data with pseudo-labels. Performing a second training on the first neural network model based on the source domain data and the target domain data with pseudo-labels can obtain a second neural network model. The source domain data is labeled pedestrian data, and the target domain data is unlabeled pedestrian data. It can be seen that in the training method of the present invention, by combining domain information and pedestrian information, the recognition effect of the second neural network model on cross-domain data can be improved, and in the second training of the first neural network model, by expanding the target domain data with pseudo-labels as training data, the generalization of the second neural network model is improved. In summary, the training method can improve the generalization of the neural network model and the accuracy of the neural network model when processing cross-domain data. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It shows a flowchart of the training method of the present invention in a specific embodiment.

[0023] Figure 2 It shows a flowchart of performing a first training on an initial neural network model in a specific embodiment of the present invention.

[0024] Figure 3 It shows a flowchart of performing a second training on the first neural network model in a specific embodiment of the present invention.

[0025] Figure 4 It shows a flowchart of the training method of the present invention in a specific embodiment.

[0026] Figure 5 It shows a flowchart of the person re-identification method of the present invention in a specific embodiment.

[0027] Figure 6 It shows a schematic structural diagram of the electronic device of the present invention in a specific embodiment.

[0028] ELEMENT LABEL DESCRIPTION

[0029] 600 Electronic device

[0030] 610 Memory

[0031] 620 Processor

[0032] Steps S11 - S15

[0033] Steps S21 - S28

[0034] Steps S31 - S38

[0035] Steps S41 - S45

[0036] Steps S51 - S53 Specific embodiments

[0037] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0038] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Therefore, only the components related to the present invention are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0039] In the existing neural network re - identification system, due to the lack of a sufficient amount of data during the training of the neural network model, the generalization of the neural network model is not high, and when facing cross - domain data, the probability that the neural network model hits the correct retrieval is relatively small.

[0040] In view of the above problems, the present invention provides a training method. The first training of the initial neural network model based on the source domain data can obtain a first neural network model. Processing the target domain data through the first neural network model can obtain the joint features of the target domain data, and the joint features are obtained by weighted fusion of the pedestrian features and domain features of the target domain data. Classifying the target domain features based on the joint features can obtain corresponding classification results, and then obtain the target domain data with pseudo-labels. The second training of the first neural network model based on the source domain data and the target domain data with pseudo-labels can obtain a second neural network model. The source domain data is labeled pedestrian data, and the target domain data is unlabeled pedestrian data. It can be seen that in the training method of the present invention, by combining domain information and pedestrian information, the recognition effect of the second neural network model on cross-domain data can be improved, and in the second training of the first neural network model, by expanding the target domain data with pseudo-labels as training data, the generalization of the second neural network model is improved. In summary, the training method can improve the generalization of the neural network model and the accuracy of the neural network model when processing cross-domain data.

[0041] Please refer to Figure 1 , in an embodiment of the present invention, the training method includes:

[0042] S11, obtaining source domain data and target domain data. The source domain data is labeled pedestrian data, and the target domain data is unlabeled pedestrian data. The labels of the pedestrian data can be obtained by means such as manual annotation, but the present invention is not limited thereto.

[0043] S12, performing first training on the initial neural network model based on the source domain data to obtain a first neural network model. Among them, the source domain data can be labeled pedestrian images.

[0044] Optionally, before step S12, the training method may further include: preprocessing the pedestrian images. Among them, the preprocessing includes randomly flipping the pedestrian images left and right, and performing processing such as scaling, Gaussian blur, motion blur, illumination enhancement, and contrast enhancement within a certain range.

[0045] Optionally, during the process of performing first training on the initial neural network model based on the source domain data, the loss of the initial neural network model can be determined through a loss function, and when the loss no longer decreases significantly, the first training is completed and the first neural network model is obtained.

[0046] S13. Process the target domain data through the first neural network model to obtain the joint features of the target domain data, where the joint features are obtained by weighted fusion of the pedestrian features and domain features of the target domain data. Optionally, the dimension ratio of the pedestrian features and the domain features can be used as the weight fusion ratio of the pedestrian features and the domain features. For example, if the pedestrian features are a feature vector of 1×1×2048 and the domain features are a feature vector of 1×1×256, the dimension ratio 2048:256 of the pedestrian features and the domain features can be used as the weight fusion ratio of the pedestrian features and the domain features. Among them, the joint features are the overall features of the pedestrian expression output by the first neural network model.

[0047] S14. Classify the target domain data based on the joint features, and obtain the target domain data with pseudo-labels according to the classification results.

[0048] Optionally, in this embodiment, the target domain data can be classified through a clustering algorithm based on the joint features to obtain the classification results.

[0049] Optionally, the target domain data can be unlabeled pedestrian images. At this time, the pedestrian images with less than 4 pedestrians can be deleted from the target domain data with pseudo-labels to reduce the noise impact of the classification results and obtain the completely processed target domain data with pseudo-labels.

[0050] S15. Perform second training on the first neural network model based on the source domain data and the target domain data with pseudo-labels to obtain a second neural network model.

[0051] Optionally, during the process of performing second training on the first neural network model based on the source domain data and the target domain data with pseudo-labels, the loss of the first neural network model can be determined through a loss function. When the loss no longer decreases significantly, the second training is completed and the second neural network model is obtained.

[0052] According to the above description, the training method described in this embodiment can obtain a first neural network model by performing a first training on an initial neural network model based on source domain data. Processing the target domain data through the first neural network model can obtain the joint features of the target domain data, where the joint features are obtained by weighted fusion of the pedestrian features and domain features of the target domain data. Classifying the target domain features based on the joint features can obtain corresponding classification results, and then obtain the target domain data with pseudo-labels. Performing a second training on the first neural network model based on the source domain data and the target domain data with pseudo-labels can obtain a second neural network model. The source domain data is labeled pedestrian data, and the target domain data is unlabeled pedestrian data. Among them, by combining domain information and pedestrian information during the training process, the recognition effect of the second neural network model on cross-domain data can be improved. And in the second training of the first neural network model, by expanding the target domain data with pseudo-labels as training data, the generalization of the second neural network model is improved. In summary, the training method can improve the generalization of the neural network model and the accuracy of the neural network model when processing cross-domain data.

[0053] Please refer to Figure 2 , in an embodiment of the present invention, the implementation method of performing a first training on an initial neural network model based on the source domain data to obtain a first neural network model includes:

[0054] S21. Process the source domain data through the initial neural network model to obtain the pedestrian features, domain features, number of pedestrians, and number of domains of the source domain data.

[0055] Optionally, the initial neural network model includes a common feature extraction layer, a pedestrian feature extraction layer, a domain feature extraction layer, and a feature fusion layer. Among them, the common feature extraction layer is used to obtain common features, the pedestrian feature extraction layer is used to obtain the pedestrian features of the input data according to the common features, the domain feature extraction layer is used to obtain the domain features of the input data according to the common features, and the feature fusion layer is used to perform weighted fusion according to the pedestrian features and domain features of the input data to obtain the joint features of the input data, where the input data is source domain data.

[0056] S22. Obtain the pedestrian category prediction result based on the pedestrian features. The pedestrian category prediction result can be that the source domain data is predicted correctly and the source domain data is predicted incorrectly.

[0057] S23. Obtain the pedestrian category prediction loss through a first classification loss function based on the pedestrian category prediction result and the number of pedestrians. The first classification loss function can be a cross-entropy loss function.

[0058] S24. Obtain a domain class prediction result based on the domain feature. Among them, the domain class prediction result can be that the source domain data is predicted correctly and the source domain data is predicted wrongly.

[0059] S25. Obtain a domain class prediction loss based on the domain class prediction result and the number of domains through a second classification loss function. Among them, the second classification loss function can be a cross-entropy loss function.

[0060] S26. Obtain a joint feature of the pedestrian feature and the domain feature by weighted fusion of the pedestrian feature and the domain feature. Optionally, the dimension ratio of the pedestrian feature and the domain feature can be used as the weight fusion ratio of the pedestrian feature and the domain feature.

[0061] S27. Obtain a joint feature loss based on the joint feature through a metric learning loss function. Among them, the metric learning loss function can complete the metric learning of the joint feature by using a triplet loss function.

[0062] S28. Update the parameters of the initial network model based on the pedestrian class prediction loss, the domain class prediction loss, and the joint feature loss.

[0063] Optionally, in this embodiment, the pedestrian class prediction loss, the domain class prediction loss, and the joint feature loss can be weighted and fused to obtain a total loss, and the parameters of the initial network model are updated based on the total loss. When the total loss no longer decreases significantly, the first training is completed and the first neural network model is obtained.

[0064] According to the above description, in this embodiment, the pedestrian feature and the domain feature are weighted and fused to obtain a joint feature as the overall expression feature of the pedestrian data. This method can improve the recognition effect of the neural network model on cross-domain data. In addition, in this embodiment, the parameters of the neural network model can also be updated by the pedestrian class prediction loss, the domain class prediction loss, and the joint feature loss, so that the neural network model has stable and good recognition ability.

[0065] Please refer to Figure 3 , in an embodiment of the present invention, the implementation method for performing a second training on the first neural network model based on the source domain data and the target domain data with pseudo-labels to obtain a second neural network model includes:

[0066] S31. Obtain the pedestrian features, domain features, number of pedestrians, and number of domains of the source domain data and the target domain data with pseudo-labels. Optionally, the first neural network model includes a common feature extraction layer, a pedestrian feature extraction layer, a domain feature extraction layer, and a feature fusion layer. Among them, the common feature extraction layer is used to obtain common features, the pedestrian feature extraction layer is used to obtain the pedestrian features of the input data according to the common features, the domain feature extraction layer is used to obtain the domain features of the input data according to the common features, and the feature fusion layer is used to perform weighted fusion according to the pedestrian features and the domain features of the input data to obtain the joint features of the input data, where the input data is the source domain data and the target domain data with pseudo-labels.

[0067] S32. Based on the pedestrian features, obtain the pedestrian category prediction result. Among them, the pedestrian category prediction result can be that the source domain data is predicted correctly, the source domain data is predicted incorrectly, the target domain data is predicted correctly, and the target domain data is predicted incorrectly.

[0068] S33. Based on the pedestrian category prediction result and the number of pedestrians, obtain the pedestrian category prediction loss through the third classification loss function. Among them, the third classification loss function can be an improved cross-entropy loss function, and the improved cross-entropy loss function can be expressed by the following formula:

[0069]

[0070] Among them, K is the number of pedestrians, a is a hyperparameter, and the value range of a is 0 - 1. The Loss CrossEntropy can be expressed by the following formula:

[0071]

[0072] Among them, the P i represents the true probability, q i represents the predicted probability, and the P i can be expressed by the following formula:

[0073]

[0074] In this embodiment, the improved cross-entropy loss function assigns higher confidence to the source domain data and lower confidence to the target domain data with pseudo-labels, realizes re-focusing on the source domain data with less noise and light-focusing on the target domain data with more noise, reduces the impact of incorrect pseudo-labels generated by the target domain data, and improves the recall rate and accuracy of the neural network model for the target domain data.

[0075] S34. Obtain the domain class prediction result based on the domain feature. Wherein, the domain class prediction result can be that the source domain data is predicted correctly, the source domain data is predicted wrongly, the target domain data is predicted correctly, and the target domain data is predicted wrongly.

[0076] S35. Obtain the domain class prediction loss based on the domain class prediction result and the domain number through a fourth classification loss function. Wherein, the fourth classification loss function can be a cross-entropy loss function.

[0077] S36. Obtain the joint feature of the pedestrian feature and the domain feature by weighted fusion of the pedestrian feature and the domain feature.

[0078] Optionally, the dimension ratio of the pedestrian feature and the feature can be used as the weight fusion ratio of the pedestrian feature and the domain feature.

[0079] S37. Obtain the joint feature loss based on the joint feature through a metric learning loss function. Wherein, the metric learning loss function can complete the metric learning of the joint feature by using a triplet loss function.

[0080] S38. Update the parameters of the first neural network model based on the pedestrian class prediction loss, the domain class prediction loss, and the joint feature loss.

[0081] Optionally, the pedestrian class prediction loss, the domain class prediction loss, and the joint feature loss can be weighted and fused to obtain a total loss, and the parameters of the first neural network model are updated based on the total loss. When the total loss no longer decreases significantly, the second training is completed and the second neural network model is obtained.

[0082] According to the above description, in this embodiment, by using the target domain data with pseudo-labels as the effectively augmented training data, the generalization and recognition ability of the neural network model can be effectively improved. By weighting the domain feature and the pedestrian feature, the influence brought by the wrong label can be weakened, and the influence of noise generation can be reduced.

[0083] Please refer to Figure 4 , in an embodiment of the present invention, the training method includes:

[0084] S41. Obtain source domain data and target domain data. The source domain data is labeled pedestrian data, and the target domain data is unlabeled pedestrian data. Wherein, the source domain data and the target domain data can be pedestrian images.

[0085] Optionally, the training method may further include: preprocessing the pedestrian image. The preprocessing includes randomly flipping the pedestrian image left and right, and performing processing such as scaling within a certain range, Gaussian blur, motion blur, illumination enhancement, and contrast enhancement on the pedestrian image.

[0086] S42. Based on the source domain data, perform a first training on the initial neural network model to obtain a first neural network model. The method for obtaining the initial neural network may be: building a network and randomly initializing it; loading the ResNet-50 network, where the initial weights are the pre-trained weights of VGG-16 on ImageNet; initializing the parameters of other parts of the network structure by setting the mean to 0, the mean square deviation to 0.01, and the bias to 0.

[0087] During the first training, the domain classification task and the pedestrian features can be trained together as a joint task, and according to the different feature dimensions of the pedestrian features and the domain features, adjust the weights of the pedestrian features and the domain features in the joint features. For example, if the pedestrian features are a feature vector of 1×1×2048 and the domain features are a feature vector of 1×1×256, the dimension ratio 2048:256 of the pedestrian features and the domain features can be used as the weight fusion ratio of the pedestrian features and the domain features.

[0088] In this embodiment, the common feature extraction layer of the initial neural network is used for extracting common features, the pedestrian feature extraction layer is used for obtaining pedestrian features, and the domain feature extraction layer is used for obtaining domain features. The pooling layer of the initial neural network can reduce the dimension on the channel, making the above features become a feature vector of 1×1×N. By normalizing the features through different output layers, features with the dimensions required for different tasks can be obtained. The feature fusion layer of the initial neural network is used to output a joint feature vector as the overall feature of the pedestrian expression, and the pedestrian features and the domain features are respectively used to predict the pedestrian category and the domain category through the corresponding fully connected layers.

[0089] Optionally, during the first training, for the loss of pedestrian category prediction, the cross-entropy loss function can be used for measurement; for the loss of domain category prediction, the cross-entropy loss function can be used for measurement; for the metric learning of pedestrian features, the triplet loss function can be used for measurement. By using the above loss functions, the neural network model can have better recognition ability for data in different domains and generate pseudo-labels with better quality.

[0090] S43. Process the target domain data through the first neural network model to obtain the joint features of the target domain data, where the joint features are obtained by weighted fusion of the pedestrian features and the domain features of the target domain data.

[0091] S44. Classify the target domain data based on the combined features, and obtain the target domain data with pseudo-labels according to the classification results. Optionally, the target domain data with pseudo-labels can be generated according to the combined features in combination with DBSCAN (Density-Based Spatial Clustering of Applications with Noise). In addition, in this embodiment, pedestrian images with less than 4 people can also be deleted to reduce the noise impact of outliers, and the completely processed target domain data with pseudo-labels can be obtained.

[0092] S45. Perform second training on the first neural network model based on the source domain data and the target domain data with pseudo-labels to obtain a second neural network model. Among them, for the loss of the pedestrian category, due to the existence of pseudo-label data with noise, it is not appropriate to directly use the loss function in the first training. It is necessary to assign higher confidence to the pedestrian data with labels and lower confidence to the pedestrian data with pseudo-labels. Therefore, the loss function in the first training can be improved, and label smoothing is performed on the predicted category probability of the pseudo-labels. The corresponding predicted probability changes are as follows:

[0093]

[0094] where K is the number of pedestrians, a is a hyperparameter, and the value range of a is 0-1. The Loss CrossEntropy can be expressed by the following formula:

[0095]

[0096] where the P i represents the true probability, q i represents the predicted probability, and the P i can be expressed by the following formula:

[0097]

[0098] For the metric learning of pedestrian features, the same triplet loss function as in the first training can be used. For the loss of the predicted category of the domain category, since the domain information is known, the same cross-entropy loss function as in the first training can be used.

[0099] According to the above description, the training method described in this embodiment can obtain a first neural network model by performing first training on an initial neural network model based on source domain data. Processing the target domain data through the first neural network model can obtain the joint features of the target domain data, and the joint features are obtained by weighted fusion of the pedestrian features and domain features of the target domain data. Classifying the target domain features based on the joint features can obtain corresponding classification results, and then obtain the target domain data with pseudo-labels. Performing second training on the first neural network model based on the source domain data and the target domain data with pseudo-labels can obtain a second neural network model. The source domain data is labeled pedestrian data, and the target domain data is unlabeled pedestrian data. Among them, the initial neural network model has undergone two trainings. By expanding the target domain data with pseudo-labels as training data, the generalization ability of the neural network model can be improved. Since there are often many situations where multiple people are in one category and one person has multiple categories in the pseudo-label data obtained after the first training, if it is directly used for the second training, it will cause a significant decline in the recognition ability of the neural network model. Therefore, by performing weighted processing on the domain features and pedestrian features during the second round of training, the influence brought by the wrong labels can be weakened, and the loss function is used to weaken the noise ambiguity, improving the recognition ability of the neural network model for cross-domain data.

[0100] Please refer to Figure 5 , in an embodiment of the present invention, the person re-identification method includes:

[0101] S51, obtaining target domain data, where the target domain data is unlabeled pedestrian data.

[0102] S52, processing the target domain data through a neural network model to obtain the joint features of the target domain data. When training the neural network model, perform Figure 1 or Figure 4 the training method shown.

[0103] S53, classifying the target domain data based on the joint features and obtaining the labels of the target domain data according to the classification results.

[0104] According to the above description, the person re-identification method described in this embodiment effectively expands the pedestrian data from different domains by using Figure 1 or Figure 4 the training method shown, and has a gain on the generalization and accuracy of the person re-identification method.

[0105] Based on the above description of the training method and the person re-identification method, the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements Figure 1 orFigure 4 The training method shown and / or Figure 5 the person re-identification method shown.

[0106] Based on the above descriptions of the training method and the person re-identification method, the present invention further provides an electronic device. Please refer to Figure 6 , in an embodiment of the present invention, the electronic device 600 includes: a memory 610, on which a computer program is stored; a processor 620, communicatively connected to the memory 610, for executing the computer program and implementing Figure 1 or Figure 4 the training method shown and / or Figure 5 the person re-identification method shown.

[0107] The protection scope of the training method of the present invention is not limited to the execution order of the steps listed in this embodiment. Any solution achieved by adding or subtracting steps of the prior art and replacing steps according to the principle of the present invention is included in the protection scope of the present invention.

[0108] In summary, the training method, the person re-identification method, the medium and the electronic device of the present invention are used to improve the generalization of the neural network model and the accuracy of the neural network model when processing cross-domain data. Therefore, the present invention effectively overcomes various disadvantages in the prior art and has high industrial utilization value.

[0109] The above embodiments are only illustrative of the principles and effects of the present invention, and are not used to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A training method, characterized in that, The training method includes: Obtain source domain data and target domain data, where the source domain data is labeled pedestrian data and the target domain data is unlabeled pedestrian data; Perform first training on an initial neural network model based on the source domain data to obtain a first neural network model; Process the target domain data through the first neural network model to obtain the joint features of the target domain data, where the joint features are obtained by weighted fusion of the pedestrian features and domain features of the target domain data; Classify the target domain data based on the joint features and obtain the target domain data with pseudo-labels according to the classification results; Perform second training on the first neural network model based on the source domain data and the target domain data with pseudo-labels to obtain a second neural network model.

2. The training method according to claim 1, wherein The implementation method of performing first training on an initial neural network model based on the source domain data to obtain a first neural network model includes: Process the source domain data through the initial neural network model to obtain the pedestrian features, domain features, number of pedestrians, and number of domains of the source domain data; Obtain the pedestrian category prediction result based on the pedestrian features; Obtain the pedestrian category prediction loss through a first classification loss function based on the pedestrian category prediction result and the number of pedestrians; Obtain the domain category prediction result based on the domain features; Obtain the domain category prediction loss through a second classification loss function based on the domain category prediction result and the number of domains; Weightedly fuse the pedestrian features and the domain features to obtain the joint features of the pedestrian features and the domain features; Obtain the joint feature loss through a metric learning loss function based on the joint features; Update the parameters of the initial neural network model based on the pedestrian category prediction loss, the domain category prediction loss, and the joint feature loss.

3. The training method according to claim 1, characterized in that, The implementation method of performing second training on the first neural network model based on the source domain data and the target domain data with pseudo-labels to obtain a second neural network model includes: Obtain the pedestrian features, domain features, number of pedestrians, and number of domains of the source domain data and the target domain data with pseudo-labels; Obtain the pedestrian category prediction result based on the pedestrian features; Obtain the pedestrian category prediction loss through a third classification loss function based on the pedestrian category prediction result and the number of pedestrians; Obtain the domain category prediction result based on the domain features; Obtain the domain category prediction loss through a fourth classification loss function based on the domain category prediction result and the number of domains; Weightedly fuse the pedestrian features and the domain features to obtain the joint features of the pedestrian features and the domain features; Obtain the joint feature loss through a metric learning loss function based on the joint features; Update the parameters of the first neural network model based on the pedestrian category prediction loss, the domain category prediction loss, and the joint feature loss.

4. The training method according to claim 3, wherein The classification loss function is a cross-entropy loss function, and the metric learning loss function is a triplet loss function.

5. The training method according to claim 1, wherein The pedestrian data includes pedestrian images, and the training method further includes: Delete pedestrian images with less than 4 pedestrians in the target domain data with pseudo-labels.

6. The training method according to claim 1, wherein The initial neural network model includes a common feature extraction layer, a pedestrian feature extraction layer, a domain feature extraction layer, and a feature fusion layer. Among them, the common feature extraction layer is used to obtain common features, the pedestrian feature extraction layer is used to obtain pedestrian features of input data according to the common features, the domain feature extraction layer is used to obtain domain features of the input data according to the common features, and the feature fusion layer is used to perform weighted fusion according to the pedestrian features and the domain features of the input data to obtain joint features of the input data, where the input data includes the source domain data and / or the target domain data.

7. A pedestrian re-identification method, characterized in that, The pedestrian re-identification method includes: Obtaining target domain data, where the target domain data is unlabeled pedestrian data; Processing the target domain data through a neural network model to obtain joint features of the target domain data, and the neural network model executes the training method according to any one of claims 1-6 during training; Classifying the target domain data based on the joint features, and obtaining labels of the target domain data according to the classification results.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the training method according to any one of claims 1-6 and / or the pedestrian re-identification method according to claim 7.

9. An electronic device, characterized in that, The electronic device includes: A memory storing a computer program; A processor communicatively connected to the memory, and when calling the computer program, executes the training method according to any one of claims 1-6 and / or the pedestrian re-identification method according to claim 7.

Citation Information

Patent Citations

  • Cross-domain pedestrian re-recognition method based on complementary pseudo tags

    CN112016687A

  • Pedestrian re-recognition model training method and a pedestrian re-recognition method and system

    CN113869193A