An Enhanced and Optimized Cross-Domain Person Re-identification Method

The proposed method optimizes cross-domain person re-identification by leveraging auxiliary and target data sets with dynamic neighbor exploration and pseudo-labeling, addressing the challenge of deploying supervised methods and neglecting challenging samples, thereby enhancing model performance in real-world scenarios.

CN114332916BActive Publication Date: 2025-07-15JIANGXI COMWORTH INTERNET OF THINGS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111465470.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-07-15
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

In the existing unsupervised domain adaptation pedestrian recognition method, the network model is complex and difficult to ignore, resulting in insufficient matching performance of the model in the target domain.

Method used

By building an enhanced optimization cross-domain pedestrian re-identification method, using auxiliary data and target data sets, it is divided into reliable training data sets and enhanced training data sets, dynamically explores nearest neighbor samples, and uses deep neural networks to perform feature mining and loss function optimization to improve the matching performance of the model in the target domain.

Benefits of technology

Effective utilization of difficult sample information improves the model's re-identification performance in the target domain and improves the matching accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114332916B_ABST
    Figure CN114332916B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for enhancing and optimizing cross-domain person re-identification, belonging to the field of computer vision technology. The method includes: construction of auxiliary data and target data sets, division of reliable training data sets and enhanced training data sets, classification loss of auxiliary data and reliable training data, exploration of dynamic nearest neighbor samples of reliable training data, and mining of dynamic nearest neighbor samples of enhanced training data; The present invention makes full use of the information of out-of-cluster hard samples and the features of in-cluster samples, mines the potential information of the target domain, and performs alternating training with the auxiliary data, thereby improving the matching performance of the model in the target domain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and relates to a method for enhancing and optimizing cross-domain person re-identification. Background Art

[0002] The person re-identification technology belongs to the category of image retrieval, aiming to retrieve the person image that matches the given query image from the displayed images. With the development of intelligent technology, the person re-identification technology is widely applied in fields such as suspect image retrieval and behavior analysis. The person re-identification method based on supervised learning is difficult to be deployed in actual scenarios because it requires a large amount of training data with identity information. As one of the solutions, the unsupervised domain adaptation person re-identification method involves auxiliary data with labeled information and target domain data without labels, which greatly alleviates the pressure of data annotation. Most of the best-performing unsupervised domain adaptation person re-identification methods use teaching models and clustering methods to assign pseudo-labels. However, the involved network models are complex, and the difficult samples outside the clusters are directly discarded, ignoring the guiding role of difficult samples for the model. Therefore, by constructing a simple model for enhancing and optimizing cross-domain person re-identification method, making full use of the difficult sample information and reliable sample features, and mining the potential information in the target domain, the matching performance of the model in the target domain can be improved. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a method for enhancing and optimizing cross-domain person re-identification. The method includes: construction of auxiliary data and target data sets, division of reliable training data set and enhanced training data set, classification loss of auxiliary data and reliable training data, exploration of dynamic nearest neighbor samples of reliable training data, and mining of dynamic nearest neighbor samples of enhanced training data.

[0004] To achieve the above purpose, the present invention provides the following technical solutions:

[0005] A method for enhancing and optimizing cross-domain person re-identification, the method includes the following steps:

[0006] S1: Obtain auxiliary data and target data. The identities of pedestrians in the two types of data are completely different and come from multiple cameras; the auxiliary data contains pedestrian identity information, and the target data does not contain identity information;

[0007] S2: Send the target data into a deep neural network, specify reliable training data and enhanced training data, and assign pseudo-labels to the reliable training data; the enhanced training data consists of difficult samples and randomly sampled reliable samples;

[0008] S3: Input the auxiliary data and reliable training data into the deep neural network in sequence, calculate the classification loss of the auxiliary data and reliable training data, and represent them by L src and L r_id respectively;

[0009] S4: Store the reliable training data features output by the deep neural network into the feature dictionary E r ; Explore the dynamic neighboring samples of the query image according to the feature dictionary, make similar features close to each other, and calculate the reliable similarity loss L r_near ;

[0010] S5: Input the augmented training data in step S2 into the deep neural network, store the output features into the augmented feature dictionary E e , mine the dynamic neighboring samples, optimize the model's recognition of difficult samples, and calculate the augmented training loss L enh_near ;

[0011] S6: According to S3 - S5, weight the classification loss, similarity loss of reliable training data, and augmented similarity loss of augmented training data to obtain the total loss of the target data as follows:

[0012] L tgt = ηL r_id + L r_near + L enh_near

[0013] S7: The total loss function of the model includes: the weighted sum of the auxiliary data classification loss and the total loss of the target data;

[0014] L = (1 - λ)L src + λL tgt

[0015] S8: Repeat S2 - S7 until the iteration number reaches the set maximum iteration number, then the model training is completed;

[0016] S9: Input the query image and the display image into the pedestrian re - identification model trained in step S8, and output the matching sorted list and performance results.

[0017] Optionally, the selection of the reliable training data and the augmented training data includes: clustering the target data, using the data within the cluster as the reliable training data, and using the data outside the cluster as the difficult samples of the augmented training data; dynamically assigning pseudo - labels to the reliable training data, and randomly selecting 1 / n samples from the clusters with the number of samples within the cluster greater than n to augment the augmented training data.

[0018] Optionally, the selection of the dynamic neighboring samples includes: dynamically determining the number of neighboring samples; performing weighted averaging on the samples in this batch according to the number of samples whose similarity score between the input feature and the feature dictionary is greater than a proportion σ of the highest score to obtain the number of neighboring samples k that changes with this batch of data; dynamically selecting neighboring samples according to the number of neighboring samples and the feature similarity score matrix in the feature dictionary.

[0019] Optionally, the method for processing the dynamic neighboring samples is as follows: the input image and the neighboring samples are assigned the same pseudo-identity, and a weight of 1 / k is multiplied when calculating the reliable similarity loss.

[0020] Optionally, the augmented training data is used to enhance the construction of the feature dictionary and to enhance the model's exploration of difficult-to-match samples in the target domain;

[0021] In the target domain, the expression of the loss function is:

[0022]

[0023] The beneficial effects of the present invention are as follows:

[0024] (1) The present invention uses an unsupervised clustering method to divide the target data into a reliable training data set and an augmented training data set, and randomly selects 1 / n samples from each cluster to augment the augmented training set. According to the clustering results, pseudo-labels are assigned to the reliable training data, providing pseudo-identity constraints, and constructing the feature dictionary E r , and dynamically explores the neighboring samples in the reliable training data set. An augmented feature dictionary E is constructed for the augmented training data set e , and dynamically mines the neighboring samples in the difficult samples. Through these three constraints, the ability of the enhanced model to mine dynamic neighboring samples is enhanced, effectively improving the re-identification performance of the model in the target domain;

[0025] (2) The present invention designs a dynamic neighboring selection method as a constraint for exploring neighboring samples in reliable training and augmented training. By dynamically selecting the number of neighboring samples in different batches, the distance between neighboring samples is effectively reduced, and the matching accuracy of the model is improved.

[0026] Other advantages, objectives, and features of the present invention will, to some extent, be described in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on an examination of the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings

[0027] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0028] Figure 1 is a flowchart of the method of the present invention;

[0029] Figure 2 is a schematic block diagram of the principle of the method of the present invention. Detailed Embodiments

[0030] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0031] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.

[0032] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0033] The present invention proposes an enhanced and optimized cross-domain person re-identification method, as Figure 1 、 2 shown. The main contents of this method include: construction of auxiliary data and target data sets, division of reliable training data sets and enhanced training data sets, classification losses of auxiliary data and reliable training data, exploration of dynamic nearest neighbor samples of reliable training data, and mining of dynamic nearest neighbor samples of enhanced training data.

[0034] The process of training a person re-identification model includes:

[0035] S1: Obtain auxiliary data and target data. The identities of the pedestrians in the two data sets are completely different and come from multiple cameras; the auxiliary data contains pedestrian identity information, and the target data does not contain identity information;

[0036] S2: Feed the target data into the deep neural network to find reliable training data and augmented training data, and assign pseudo-labels to the reliable training data. The reliable training data and augmented training data are divided using the DBSCAN clustering algorithm. The data within the cluster is used as reliable training data, and pseudo-labels are assigned according to the number of clusters. The outlier samples are used as hard samples for the augmented training data. To avoid difficult-to-mine information in the augmented training data, randomly select 1 / n samples from the clusters with a sample size greater than n and supplement them to the augmented training data set.

[0037] S3: Input the auxiliary data and reliable training data into the deep neural network in sequence, calculate the classification losses of the auxiliary data and reliable training data, denoted by L src and L r_id respectively.

[0038] S4: Store the features of the reliable training data output by the deep neural network into the feature dictionary E r ; Dynamically obtain the nearest neighbor samples according to the feature dictionary E r and calculate the reliable similarity loss L r_near . The calculation formula of the loss function is as follows:

[0039]

[0040] where i is the sample index, y r is the index of the reliable training data, w(y r ) is the label weight of the reliable sample, n r is the batch size of the samples, N r represents the size of the reliable training data, E r is the feature dictionary, E r [y r represents the sample feature corresponding to the index y r , is the feature representation of the sample , and α is the balance factor.

[0041] The nearest neighbor samples are selected according to the similarity scores between the input sample features and the feature dictionary E r . The number is determined by the weighted average of the number of samples whose scores in the batch of samples are greater than the proportion of the highest score. The formula for determining the number of dynamic nearest neighbor samples is as follows:

[0042]

[0043]

[0044]

[0045] where σ is the ratio, is The similarity score between and

[0046] The label weight formula is as follows:

[0047]

[0048] where is the index of the dynamic neighbor sample, and y r represents the index of the input query image.

[0049] S5: Input the enhanced training data from step S2 into the deep neural network, store the output features in the enhanced feature dictionary E e , obtain the dynamic neighbor samples and calculate the enhanced training loss L enh_near , and the calculation method is similar to the reliable similarity loss.

[0050] S6: According to steps S3 - S5, weight the classification loss, similarity loss of the reliable training data, and the enhanced training loss of the enhanced training data to obtain the total loss of the target domain as:

[0051]

[0052] S7: The total loss function of the model includes: weighting the auxiliary data classification loss and the total loss of the target data, which is

[0053] L=(1 - λ)L src +λL tgt

[0054] S8: Repeat steps S2 - S7 until the iteration number reaches the set maximum iteration number, then the model training is completed;

[0055] S9: Input the query image and the display image into the person re - identification model trained in step S8, and output the matching sorted list and performance results.

[0056] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer - readable storage medium. The storage medium can include: ROM, RAM, disk, optical disc, etc.

[0057] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An enhanced and optimized cross-domain pedestrian re-identification method, characterized in that: The method includes the following steps: S1: Obtain auxiliary data and target data. The pedestrian identities of the two types of data are completely different and come from multiple cameras; the auxiliary data contains pedestrian identity information, and the target data does not contain identity information; S2: Feed the target data into a deep neural network, specify reliable training data and augmented training data, and assign pseudo-labels to the reliable training data; the augmented training data consists of hard samples and randomly sampled reliable samples; S3: Input the auxiliary data and the reliable training data into the deep neural network in sequence, calculate the classification losses of the auxiliary data and the reliable training data, and denote them as L src and L r_id respectively; S4: Store the reliable training data features output by the deep neural network into the feature dictionary E r ; Explore the dynamic nearest neighbor samples of the query image according to the feature dictionary to make similar features close to each other, and calculate the reliable similarity loss L r_near ; S5: Input the enhanced training data from step S2 into the deep neural network, store the output features in the enhanced feature dictionary E e , mine dynamic neighbor samples, optimize the model's recognition of difficult samples, and calculate the enhanced training loss L enh_near ; S6: According to S3 - S5, weight the classification loss, similarity loss of the reliable training data, and augmented similarity loss of the augmented training data to obtain the total loss of the target data as: L tgt = ηL r_id + L r_near + L enh_near S7: The total loss function of the model includes: the weighted sum of the classification loss of the auxiliary data and the total loss of the target data; L = (1 - λ)L src + λL tgt S8: Repeat S2 - S7 until the iteration number reaches the set maximum iteration number, and the model training is completed; S9: Input the query image and the display image into the pedestrian re-identification model trained in step S8, and output the matching sorted list and performance results.

2. The enhanced and optimized cross-domain person re-identification method according to claim 1, characterized in that: The selection of the reliable training data and the augmented training data includes: clustering the target data before each iteration, using the data within the cluster as the reliable training data, and using the data outside the cluster as the hard samples of the augmented training data; dynamically assign pseudo-labels to the reliable training data, and randomly select 1 / n samples from the clusters with the number of samples within the cluster greater than n to augment the augmented training data.

3. An enhanced and optimized cross-domain pedestrian re-identification method according to claim 1, characterized in that: The feature dictionary E r and the enhanced feature dictionary E e are updated in different ways; E r uses the momentum update method to store the features of reliable training data and retains the historical information of the samples; E e directly stores the features learned by the model this time.

4. An enhanced and optimized cross-domain pedestrian re-identification method according to claim 1, characterized in that: The selection of the dynamic nearest neighbor samples includes: dynamically determining the number of nearest neighbor samples; performing weighted averaging on the number of samples in this batch whose similarity score between the input feature and the feature dictionary is greater than the highest score multiplied by a ratio σ to obtain the number of nearest neighbor samples k that changes with this batch of data; dynamically select the nearest neighbor samples according to the number of nearest neighbor samples and the feature similarity score matrix in the feature dictionary.

5. The enhanced and optimized cross-domain pedestrian re-identification method according to claim 3, characterized in that: The processing method for the dynamic nearest neighbor samples is: assign the same pseudo-identity to the input image and the nearest neighbor samples, and multiply by the weight 1 / k when calculating the reliable similarity loss.

6. An enhanced and optimized cross-domain pedestrian re-identification method according to claim 3, characterized in that: The augmented training data is used to enhance the construction of the feature dictionary and to enhance the exploration of the re-identification model for hard-to-match samples in the target domain; In the target domain, the expression of the loss function is: