A Label Association Method for Unsupervised Visible-Infrared Person Re-identification
By constructing an affinity matrix and obtaining the optimal soft pseudo-label, the identification complexity and time-consuming problems caused by modal differences in visible-infrared pedestrian re-identification technology under unsupervised conditions are solved, and efficient pedestrian re-identification is achieved.
Patent Information
- Application Number
- CN202410102504.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-24
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-01-24
AI Technical Summary
The existing visible-infrared pedestrian re-identification technology has a large modal difference, resulting in a time-consuming and complex identification process, which cannot effectively solve the problem of pedestrian re-identification under unsupervised conditions.
By acquiring pedestrian image features of visible and infrared light modes, affinity matrix is constructed to characterize homogeneous and heterogeneous affinities, thereby obtaining the optimal soft pseudo-label of visible and infrared light modes and applying it to unsupervised visible-infrared pedestrian re-identification.
The alignment of soft and pseudo-labels under different modes is achieved, providing rich and reliable supervision signals for unsupervised visible-infrared pedestrian re-identification, alleviating the complexity and time-consuming problems of the identification process.
Smart Images

Figure CN118135648B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image detection, and specifically, to a label association method for unsupervised visible-infrared pedestrian re-identification. Background Art
[0002] Due to its practical application in intelligent monitoring systems, visible-infrared pedestrian re-identification has attracted increasing attention. Visible-infrared pedestrian re-identification technology refers to the technology of using dual-modal data of visible light and infrared light for pedestrian identification. In visible light images, information such as the appearance features and clothing colors of pedestrians is more clearly visible, while in infrared images, due to its strong penetration power, it can penetrate some obstacles, such as fog, smoke and dust, etc., so relatively clear images can be obtained even under harsh lighting conditions. Therefore, visible-infrared pedestrian re-identification technology is to retrieve the same person from a set of visible / infrared gallery images when given an image from the infrared / visible modality. However, these methods all require a labeled dataset with modality sharing as the basis for network training, and due to the huge modality differences between visible light images and infrared images, the visible-infrared pedestrian re-identification technology is time-consuming and the recognition process is complex in practical applications. Summary of the Invention
[0003] In order to solve the above problems existing in the prior art, the present invention provides a label association method for unsupervised visible-infrared pedestrian re-identification.
[0004] According to the first aspect of the embodiments of the present invention, a label association method for unsupervised visible-infrared pedestrian re-identification is provided, and the method includes:
[0005] Obtain an image dataset; wherein, the image dataset includes pedestrian images in the visible light modality and pedestrian images in the infrared light modality;
[0006] According to the image dataset, obtain a first feature corresponding to the pedestrian image in the visible light modality and a second feature corresponding to the pedestrian image in the infrared light modality;
[0007] Construct an affinity matrix according to the first feature and the second feature to represent homogeneous affinity and heterogeneous affinity; wherein, the homogeneous affinity represents the affinity relationship within images of the same modality, and the heterogeneous affinity represents the affinity relationship between images of different modalities;
[0008] Obtain a first optimal soft pseudo-label corresponding to the visible light modality and a second optimal soft pseudo-label corresponding to the infrared light modality according to the affinity matrix;
[0009] Use the first optimal soft pseudo-label and the second optimal soft pseudo-label as supervision signals to be applied to unsupervised visible-infrared pedestrian re-identification.
[0010] Optionally, before obtaining the first optimal soft pseudo-label corresponding to the visible light modality and the second optimal soft pseudo-label corresponding to the infrared light modality according to the affinity matrix, the method further includes:
[0011] Obtaining a clustering center according to the first feature and the image data set;
[0012] Obtaining a first repository according to the clustering center; wherein, the first repository contains a plurality of the first features and pseudo-labels corresponding to the plurality of the first features;
[0013] Obtaining a first initial rough label according to the first repository and the first feature;
[0014] Assigning a label to the second feature by using the first repository as a second initial rough label.
[0015] Optionally, the obtaining manner of the first initial rough label is as follows:
[0016]
[0017] wherein, represents the first initial rough label, represents the feature f v the predicted probability distribution output from the first repository, K represents the number of features, and τ is the temperature coefficient;
[0018] The obtaining manner of the second initial rough label is as follows:
[0019]
[0020] wherein, OTLA is the optimal transport label assignment algorithm, represents assigning the pseudo-labels in the first repository to the second feature f r as the second initial rough label
[0021] Optionally, constructing the affinity matrix according to the first feature and the second feature includes:
[0022] Constructing the homogeneous affinity according to the first feature and the second feature by using the Jaccard similarity;
[0023] Constructing the transport cost between the first feature and the second feature according to the first feature and the second feature by using the optimal transport model;
[0024] Obtaining the heterogeneous affinity according to the transport cost;
[0025] The affinity matrix is obtained based on the homogeneous affinity and the heterogeneous affinity.
[0026] Optionally, the homogeneous affinity is obtained according to the following formula:
[0027]
[0028] where S ho(e) is S ho(v) or S ho(r) , representing the homogeneous affinity within the visible light modality or the homogeneous affinity within the infrared light modality, that is, the affinity relationship between and within the same modality, represents 's κ - mutual nearest neighbor, represents 's κ - mutual nearest neighbor, S ho(v) is the affinity relationship within the image of the visible light modality, S ho(r) is the affinity relationship within the image of the infrared light modality, and are two different features in the same modality.
[0029] Optionally, the heterogeneous affinity is obtained according to the following formula:
[0030]
[0031] where S he represents the transportation plan between the images of different modalities, C he is the cost matrix constructed based on the Euclidean distance between the features of different modalities, <a, b> represents the Frobenius dot product of matrix a and matrix b, λ is a hyperparameter, and the optimal transportation plan S he* serves as the heterogeneous affinity, represents the i - th feature of the visible modality, represents the j - th feature of the infrared light modality.
[0032] Optionally, obtaining the first optimal soft pseudo - label corresponding to the visible light modality and the second optimal soft pseudo - label corresponding to the infrared light modality based on the affinity matrix includes:
[0033] Obtaining the homogeneous inconsistency and the heterogeneous inconsistency based on the Euclidean distance between the pseudo - labels of the image features in the image dataset and the affinity matrix;
[0034] Obtaining the self - inconsistency based on the labels of the image features, the first initial coarse label, and the second initial coarse label;
[0035] The first optimal soft pseudo-label and the second optimal soft pseudo-label are obtained according to the homogeneous inconsistency, the heterogeneous inconsistency, and the self-inconsistency.
[0036] Optionally, the homogeneous inconsistency can be obtained according to the following formula:
[0037]
[0038] where represents the homogeneous inconsistency of the visible modality or the infrared light modality, represents the pseudo-label of the corresponding feature in the visible modality or the infrared light modality, represents the pseudo-label of the feature in the visible modality, represents the pseudo-label of the feature in the infrared light modality, N e represents the number of images corresponding to modality e in the image dataset, represents the homogeneous affinity between the i'-th feature and the j'-th feature of the visible modality or the infrared light modality, represents the pseudo-label corresponding to the i'-th feature of the visible modality or the infrared light modality, represents the pseudo-label corresponding to the j'-th feature of the visible modality or the infrared light modality;
[0039] The heterogeneous inconsistency can be represented similarly:
[0040]
[0041] where represents the heterogeneous inconsistency of the visible modality, N v represents the number of images corresponding to the visible modality v in the image dataset, N r represents the number of images corresponding to the infrared light modality r in the image dataset, represents the affinity matrix from the visible modality to the infrared light modality, represents the pseudo-label corresponding to the i''-th feature of the visible modality, represents the pseudo-label corresponding to the j''-th feature of the infrared light modality; represents the heterogeneous inconsistency of the infrared light modality, represents the affinity matrix from the infrared light modality to the visible modality, represents the pseudo-label corresponding to the i'''-th feature of the infrared light modality; represents the pseudo-label corresponding to the j'''-th feature of the visible modality;
[0042] The self-inconsistency can be represented similarly:
[0043]
[0044] wherein, represents the pseudo-label corresponding to the i1-th feature of the visible modality or the infrared light modality, represents the initial rough label of the i1-th feature of the visible modality or the infrared light modality, N e represents the number of images corresponding to modality e in the image dataset.
[0045] Optionally, the method further includes:
[0046] Constructing a second repository and a third repository; wherein, the second repository contains a plurality of the first features and the pseudo-labels corresponding to the plurality of the first features, a plurality of the second features and the pseudo-labels corresponding to the plurality of the second features in the visible light space, and the third repository stores a plurality of the second features and the pseudo-labels corresponding to the plurality of the second features in the visible light space and the clustering centers corresponding to the plurality of the second features;
[0047] Obtaining a first contrast loss corresponding to the visible light modality according to the first repository;
[0048] Obtaining a second contrast loss corresponding to the infrared light modality according to the third repository;
[0049] Obtaining a cross-modal contrast loss between the visible light modality and the infrared light modality according to the second repository, denoted as a third contrast loss;
[0050] Obtaining a first label correction loss corresponding to the visible light modality and a second label correction loss corresponding to the infrared light modality according to the first repository, the second repository and the third repository;
[0051] Obtaining a loss function according to the first contrast loss, the second contrast loss, the third contrast loss, the first label correction loss and the second label correction loss;
[0052] Optimizing the first optimal soft pseudo-label and the second optimal soft pseudo-label by using the loss function to optimize the unsupervised visible-infrared pedestrian re-identification.
[0053] Optionally, the first contrast loss is obtained according to the following formula:
[0054]
[0055] wherein, is the first contrast loss, dividing the image dataset into multiple batches, the number of training data in each batch is B, and is is the cross - entropy function, represents the feature predicted probability distribution output from the first repository , τ is the temperature coefficient, represents the corresponding pseudo - label;
[0056] The second contrast loss is obtained according to the following formula:
[0057]
[0058] where, is the second contrast loss, is the cross - entropy function, represents the feature predicted probability distribution output from the original repository of the infrared light modality , represents the corresponding pseudo - label; represents the feature predicted probability distribution output from the third repository of the infrared light modality , represents the corresponding pseudo - label;
[0059] The third contrast loss is obtained according to the following formula:
[0060]
[0061] where, represents the third contrast loss, represents the feature predicted probability distribution output from the second repository M a , represents the corresponding pseudo - label; represents the feature predicted probability distribution output from the second repository M a , represents the corresponding pseudo - label;
[0062] The first label correction loss is obtained according to the following formula:
[0063]
[0064] where, is the first label correction loss, represents the feature predicted probability distribution output from the second repository M aThe predicted probability distribution output indicating features from the first repository The predicted probability distribution output indicating features from the third repository The predicted probability distribution output;
[0065] The second label correction loss is obtained according to the following formula:
[0066]
[0067] wherein is the second label correction loss indicating features from the second repository M a The predicted probability distribution output indicating features from the first repository The predicted probability distribution output indicating features from the third repository The predicted probability distribution output;
[0068] The loss function is obtained according to the following formula:
[0069]
[0070] wherein is the loss function
[0071] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0072] In the above technical solution, an image dataset is obtained, where the image dataset includes pedestrian images in the visible light modality and pedestrian images in the infrared light modality. According to the image dataset, a first feature corresponding to the pedestrian image in the visible light modality and a second feature corresponding to the pedestrian image in the infrared light modality are obtained. An affinity matrix is constructed based on the first feature and the second feature to represent homogeneous affinity and heterogeneous affinity, where the homogeneous affinity represents the affinity relationship within images of the same modality, and the heterogeneous affinity represents the affinity relationship between images of different modalities. The first optimal soft pseudo-label corresponding to the visible light modality and the second optimal soft pseudo-label corresponding to the infrared light modality are obtained according to the affinity matrix. The first optimal soft pseudo-label and the second optimal soft pseudo-label are used as supervision signals and applied to unsupervised visible-infrared pedestrian re-identification. Through the above technical solution, an affinity matrix representing the homogeneous affinity and heterogeneous affinity of pedestrian images in the visible light modality and the infrared light modality is constructed by using the features corresponding to the pedestrian images in the visible light modality and the infrared light modality, so as to obtain the first optimal soft pseudo-label and the second optimal soft pseudo-label with consistent feature space structures, realizing the alignment of soft pseudo-labels in different modalities, providing rich and reliable supervision signals for the unsupervised visible-infrared pedestrian re-identification method, and thus alleviating the problems of time-consuming and complex recognition process in the practical application of the visible-infrared pedestrian re-identification technology.
[0073] Other features and advantages of the present invention will be described in detail in the following specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the following specific implementation, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:
[0075] Figure 1 is a flowchart of a label association method for unsupervised visible-infrared pedestrian re-identification shown according to an exemplary embodiment.
[0076] Figure 2 is a flowchart of another label association method for unsupervised visible-infrared pedestrian re-identification shown according to an exemplary embodiment.
[0077] Figure 3 is a flowchart of yet another label association method for unsupervised visible-infrared pedestrian re-identification shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0078] To facilitate understanding of the solution of the present invention, the relevant situation of the prior art and the inventive concept of the present invention are briefly described first.
[0079] In recent years, with the rapid development of computer vision and deep learning technologies, great progress has been made in pedestrian re-identification technology. By extracting image features and using deep learning models for feature learning and matching, visible-infrared pedestrian re-identification technology has been widely applied in fields such as video surveillance, intelligent security, and intelligent transportation. Visible-infrared pedestrian re-identification aims to retrieve the same person from a set of visible / infrared images when given an image from the infrared / visible modality. Existing supervised visible-infrared pedestrian re-identification methods have achieved remarkable performance through deep neural networks. However, these methods all require a labeled dataset with modality sharing as the basis for network training, which is time-consuming and labor-intensive in practical scenarios.
[0080] Although many methods have achieved excellent performance in the field of unsupervised single-modal pedestrian re-identification, such as Cluster-Contrast, Image Search Engine (ISE), Probabilistic Principal Component Analysis (PPLR), due to the huge modality differences between visible light images and infrared images, these methods cannot be directly applied to visible-infrared pedestrian re-identification. If the above methods are adopted, the same pedestrian in different modalities cannot be assigned the same pseudo-label. Therefore, associating pseudo-labels in different modalities is an important challenge in solving the problem of unsupervised visible-infrared pedestrian re-identification. Some existing methods have made attempts at cross-modal pseudo-label association. Some of these methods use a global perspective to associate clusters in different modalities. However, these methods ignore the complex fine-grained structure information at the sample level, resulting in inaccurate association. To utilize the sample-level information, the Optimal Transport for Label Assignment (OTLA) algorithm formulates the association between cross-modal samples and clusters as an optimal transport problem. However, the Optimal Transport for Label Assignment algorithm ignores the homogeneous structural consistency of visible / infrared modality images, resulting in samples within the same cluster in one modality being scattered into multiple clusters in the other modality.
[0081] Based on the above analysis, due to the huge modality differences between the visible / infrared modalities, it is difficult to assign images of the same pedestrian in different modalities to the same pseudo-label. Therefore, the present invention proposes a label association method for unsupervised visible-infrared pedestrian re-identification to solve the above technical problems.
[0082] Figure 1 is a flowchart of a label association method for unsupervised visible-infrared pedestrian re-identification shown according to an exemplary embodiment, as Figure 1 shown, the method includes the following steps.
[0083] S101, obtain an image dataset; wherein, the image dataset contains pedestrian images in the visible light modality and pedestrian images in the infrared light modality.
[0084] Among them, the visible light modality refers to image capture and sensor response by using light within the visible light range. This modality generally refers to the spectral range that can be seen by the human naked eye. The infrared light modality refers to image capture and sensor response by using light within the infrared spectral range.
[0085] S102, according to the image dataset, obtain a first feature corresponding to the pedestrian image in the visible light modality and a second feature corresponding to the pedestrian image in the infrared light modality.
[0086] In one implementation, when obtaining the first feature corresponding to the pedestrian image in the visible light modality and the second feature corresponding to the pedestrian image in the infrared light modality according to the image dataset, an image encoder can be used to extract the features in the image dataset. Correspondingly, the features in the image dataset can also be obtained by other means.
[0087] S103, construct an affinity matrix according to the first feature and the second feature to characterize homogeneous affinity and heterogeneous affinity; wherein, the homogeneous affinity characterizes the affinity relationship within images of the same modality, and the heterogeneous affinity characterizes the affinity relationship between images of different modalities.
[0088] It can be understood that in the field of image processing, in multi-modal image analysis and fusion, an affinity matrix is a matrix representing the similarity between different regions of an image. Among them, the homogeneous affinity represents the similarity or matching degree between different regions within an image of the same modality. The heterogeneous affinity represents the similarity and matching ability between corresponding regions between images of different modalities. The affinity matrix can achieve effective registration of images within a modality or between different modalities.
[0089] S104, obtain a first optimal soft pseudo-label corresponding to the visible light modality and a second optimal soft pseudo-label corresponding to the infrared light modality according to the affinity matrix.
[0090] It is understandable that soft pseudo-labels are labels that are opposite to hard pseudo-labels. Soft pseudo-labels contain the probability distribution generated by the algorithm for each unlabeled sample data, that is, the confidence scores for each category. When using soft pseudo-labels, the model provides a probability vector for the sample data, where each element represents the probability that the sample data belongs to each category. These soft labels can then be used as supervision information to further train the model. Utilizing soft pseudo-labels can utilize the potential structural information in the unlabeled data, thereby helping to improve the model performance. And soft pseudo-labels can contain the uncertainty of the sample data, and using soft labels can perform iterative training and self-correction more robustly.
[0091] S105, use the first optimal soft pseudo-label and the second optimal soft pseudo-label as supervision signals and apply them to unsupervised visible-infrared pedestrian re-identification.
[0092] It is understandable that during the training process of unsupervised visible-infrared pedestrian re-identification, the features extracted from the training set are clustered, and the samples with similar features are grouped into the same cluster. For each clustering cluster, the model calculates the probability distribution of the sample features within the cluster belonging to the cluster center, and uses the first optimal soft pseudo-label and the second optimal soft pseudo-label as the supervision signals for the model, which are used in the subsequent training iteration of the model.
[0093] In the above technical solution, an image data set is obtained; wherein, the image data set contains pedestrian images in the visible light modality and pedestrian images in the infrared light modality; according to the image data set, the first feature corresponding to the pedestrian image in the visible light modality and the second feature corresponding to the pedestrian image in the infrared light modality are obtained; an affinity matrix is constructed according to the first feature and the second feature to represent the homogeneous affinity and the heterogeneous affinity; wherein, the homogeneous affinity represents the affinity relationship within the images of the same modality, and the heterogeneous affinity represents the affinity relationship between the images of different modalities; the first optimal soft pseudo-label corresponding to the visible light modality and the second optimal soft pseudo-label corresponding to the infrared light modality are obtained according to the affinity matrix; the first optimal soft pseudo-label and the second optimal soft pseudo-label are used as supervision signals and applied to unsupervised visible-infrared pedestrian re-identification. Through the above technical solution, an affinity matrix representing the homogeneous affinity and the heterogeneous affinity of the pedestrian images in the visible light modality and the infrared light modality is constructed by using the features corresponding to the pedestrian images in the visible light modality and the infrared light modality, so as to obtain the first optimal soft pseudo-label and the second optimal soft pseudo-label with consistent feature space structures, realizing the alignment of soft pseudo-labels in different modalities, providing rich and reliable supervision signals for the unsupervised visible-infrared pedestrian re-identification method, and thus alleviating the problems of time-consuming and complex recognition process in the practical application of the visible-infrared pedestrian re-identification technology.
[0094] Optionally,Figure 2 is a flowchart of another label association method for unsupervised visible-infrared pedestrian re-identification shown according to an exemplary embodiment. As Figure 2 shown, before step S104, the method further includes the following steps.
[0095] S106, obtaining a clustering center according to the first feature and the image data set.
[0096] S107, obtaining a first repository according to the clustering center, wherein the first repository contains a plurality of the first features and pseudo-labels corresponding to the plurality of the first features.
[0097] S108, obtaining a first initial rough label according to the first repository and the first feature.
[0098] S109, using the first repository to assign a label to the second feature as a second initial rough label.
[0099] It can be understood that after the first feature and the second feature are extracted, a clustering method can be used to cluster the first feature corresponding to the visible light modality and the second feature corresponding to the infrared light modality, such as a clustering method like the DBSCAN method. After clustering, similar features will be grouped into the same cluster, obtaining a clustering center corresponding to the visible light modality and a clustering center corresponding to the infrared light modality, and initializing a first repository according to the clustering center and initializing an infrared light initial repository according to the second clustering center. The first repository contains each first feature and pseudo-labels corresponding to each first feature obtained according to the clustering center of the visible light modality, and the infrared light initial repository contains each second feature and pseudo-labels corresponding to each second feature obtained according to the clustering center of the infrared light modality.
[0100] In addition, the initial rough label is a rough and preliminary class label assigned to the samples in the image data set in a fast, low-cost or automated manner.
[0101] Optionally, in one implementation, the obtaining manner of the first initial rough label is as follows:
[0102]
[0103] Wherein, represents the first initial rough label, represents the feature f v the predicted probability distribution output from the first repository K represents the number of features, and τ is the temperature coefficient;
[0104] It is worth mentioning that the specific calculation method of can refer to:
[0105]
[0106] Among them, where c l represents the cluster center corresponding to the l-th cluster in the first repository, and f in the first repository, f T represents the transpose of the feature, and τ′ is the temperature coefficient;
[0107] The method for obtaining the second initial rough label is as follows:
[0108]
[0109] Among them, OTLA is the optimal transport label assignment algorithm, which means assigning the pseudo-labels in the first repository to the second feature f r to serve as the second initial rough label
[0110] Optionally, S103 may include:
[0111] Constructing the homogeneous affinity according to the first feature and the second feature using the Jaccard similarity;
[0112] Constructing the transport cost between the first feature and the second feature according to the first feature and the second feature using the optimal transport model;
[0113] Obtaining the heterogeneous affinity according to the transport cost;
[0114] Obtaining the affinity matrix according to the homogeneous affinity and the heterogeneous affinity.
[0115] Optionally, to solve the modality bias problem and achieve the purpose that all samples within one modality have equal total affinity for another modality. Therefore, the heterogeneous affinity can be modeled by constructing an optimal transport (OT) model.
[0116] Optionally, the homogeneous affinity is obtained according to the following formula:
[0117]
[0118] Among them, S ho(e) is S ho(v) or S ho(r) , representing the homogeneous affinity within the visible light modality or the homogeneous affinity within the infrared light modality, that is, the affinity relationship between and within the same modality, represents the κ-mutual nearest neighbor of express κ-mutual nearest neighbors, S ho(v) is the affinity relationship within the image of the visible light modality, S ho(r) is the affinity relationship within the image of the infrared light modality, and are two different features in the same mode, that is, different features in two visible light modes or different features in two infrared light modes.
[0119] Optionally, the heterogeneous affinity is obtained according to the following formula:
[0120]
[0121] Among them, S he represents the transmission plan between the images of different modalities, C he is a cost matrix constructed based on the Euclidean distances between the features of the different modalities,<a,b> represents the Frobenius dot product of matrix a and matrix b, λ is a hyperparameter, and the optimal transmission plan S obtained according to the above formula he* As the heterogeneous affinity, represents the i-th feature of the visible mode, represents the j-th feature of the infrared light modality.
[0122] It is worth mentioning that for heterogeneous affinity, the following constraints can be added:
[0123]
[0124] Among them, N v Represents the number of samples of the visible light mode, N r represents the number of samples of the infrared light modality; represents a column vector of all 1s. This constraint constrains the transmission plan S he The marginal distribution of follows a uniform distribution, which ensures that the total affinity of all samples in the same modality is the same as that of all samples in the other modality.
[0125] The optimal transmission plan S obtained according to the formula of heterogeneous affinity he* can be regarded as the affinity between two heterogeneous features. Then, the heterogeneous affinity relationship S he* The affinity matrix from the visible light mode to the infrared light mode is obtained and the affinity matrix from infrared light modality to visible light modality The above affinity relationships are row normalized and the final affinity matrix S = S ho(v) ,S ho(r) ,S he(vr) ,S he(rv)。
[0126] Optionally, S104 may include:
[0127] Obtaining homogeneous inconsistency and heterogeneous inconsistency according to the Euclidean distance between the pseudo-labels of the image features in the image dataset and the affinity matrix;
[0128] Obtaining self-inconsistency according to the labels of the image features, the first initial coarse label, and the second initial coarse label;
[0129] Obtaining the first optimal soft pseudo-label and the second optimal soft pseudo-label according to the homogeneous inconsistency, the heterogeneous inconsistency, and the self-inconsistency.
[0130] It can be understood that the Euclidean distance between the pseudo-labels of the image features in the image dataset and the affinity matrix are used as their structural inconsistencies, and the structural inconsistencies include homogeneous inconsistency and heterogeneous inconsistency.
[0131] The homogeneous inconsistency can be obtained according to the following formula:
[0132]
[0133] Where is or represents the homogeneous inconsistency of the visible modality or the infrared light modality, represents the pseudo-label of the corresponding feature in the visible modality or the infrared light modality, represents the pseudo-label of the feature in the visible modality, represents the pseudo-label of the feature in the infrared light modality, N e represents the number of images corresponding to modality e in the image dataset, represents the homogeneous affinity between the i'-th feature and the j'-th feature of the visible modality or the infrared light modality, represents the pseudo-label corresponding to the i'-th feature of the visible modality or the infrared light modality, represents the pseudo-label corresponding to the j'-th feature of the visible modality or the infrared light modality;
[0134] The heterogeneous inconsistency can be expressed similarly:
[0135]
[0136] Where represents the heterogeneous inconsistency of the visible modality, N vDenote the number of images corresponding to the visible modality v in the image dataset, N r Denote the number of images corresponding to the infrared light modality r in the image dataset Denote the affinity matrix from the visible modality to the infrared light modality Denote the pseudo-label corresponding to the i''-th feature of the visible modality Denote the pseudo-label corresponding to the j''-th feature of the infrared light modality Denote the heterogeneous inconsistency of the infrared light modality Denote the affinity matrix from the infrared light modality to the visible modality Denote the pseudo-label corresponding to the i''-th feature of the infrared light modality Denote the pseudo-label corresponding to the j'''-th feature of the visible modality
[0137] The self-inconsistency can be represented similarly
[0138]
[0139] Among them Denote the pseudo-label corresponding to the i1-th feature of the visible modality or the infrared light modality Denote the initial coarse label of the i1-th feature of the visible modality or the infrared light modality, N e Denote the number of images corresponding to the modality e in the image dataset. Since simply minimizing the above homogeneous inconsistency and heterogeneous inconsistency terms may cause all samples to eventually have almost the same label, self-inconsistency can be introduced to avoid this problem
[0140] To obtain the first optimal soft pseudo-label and the second optimal soft pseudo-label according to the homogeneous inconsistency, the heterogeneous inconsistency and the self-inconsistency, the above homogeneous inconsistency, the heterogeneous inconsistency and the self-inconsistency can be minimized, and specifically, the following formula can be used to obtain them
[0141]
[0142] Among them, α is a preset trade-off parameter Denote the pseudo-label of the feature corresponding to the visible light modality or the infrared light modality is or is or is or When the sum of the three terms of and in the above formula reaches the minimum, according to the y at this timee* Obtain the first optimal soft pseudo-label and the second optimal soft pseudo-label. Specifically, when the above-mentioned homogeneous inconsistency, heterogeneous inconsistency, and self-inconsistency are within the visible light modality, the obtained y e* is the first optimal soft pseudo-label within the visible light modality; when the above-mentioned homogeneous inconsistency, heterogeneous inconsistency, and self-inconsistency are within the infrared light modality, the obtained y e* is the first optimal soft pseudo-label within the infrared light modality.
[0143] It is worth mentioning that when minimizing the sum of and the following iterative optimization formula for pseudo-labels can be obtained:
[0144]
[0145]
[0146] wherein, represents the pseudo-label corresponding to the visible light modality at the (t + 1)-th iteration in the iterative optimization process, represents the pseudo-label corresponding to the infrared light modality at the (t + 1)-th iteration in the iterative optimization process, α is a preset trade-off parameter, represents the first initial coarse label, is the second initial coarse label, S ho(r) is the homogeneous affinity within the infrared light modality, S ho(v) is the homogeneous affinity within the visible light modality, S he(vr) is the affinity matrix from the visible light modality to the infrared light modality, S he(rv) is the affinity matrix from the infrared light modality to the visible light modality. Finally, the optimized soft pseudo-label and the linear combination of the one-hot hard pseudo-label based on this soft pseudo-label can be used as the final optimal soft pseudo-label to provide a supervision signal for unsupervised visible-infrared pedestrian re-identification training:
[0147]
[0148] wherein, y e* is the first optimal soft pseudo-label or the second optimal soft pseudo-label, β is a parameter used to control the smoothness of the pseudo-label, and the hard pseudo-label y is obtained from the transmitted label hard and the soft pseudo-label y soft .
[0149] Optionally, Figure 3 is a flowchart of another label association method for unsupervised visible-infrared pedestrian re-identification shown according to an exemplary embodiment. As Figure 3 shown, this method may further include:
[0150] S110, construct a second repository and a third repository; wherein, the second repository contains a plurality of the first features and the pseudo-labels corresponding to the plurality of the first features, a plurality of the second features, and the pseudo-labels corresponding to the plurality of the second features in the visible light space, and the third repository stores a plurality of the second features, the pseudo-labels corresponding to the plurality of the second features in the visible light space, and the cluster centers corresponding to the plurality of the second features.
[0151] S111, obtain a first contrastive loss corresponding to the visible light modality according to the first repository.
[0152] S112, obtain a second contrastive loss corresponding to the infrared light modality according to the third repository.
[0153] S113, obtain a cross-modal contrastive loss between the visible light modality and the infrared light modality according to the second repository, denoted as a third contrastive loss.
[0154] S114, obtain a first label correction loss corresponding to the visible light modality and a second label correction loss corresponding to the infrared light modality according to the first repository, the second repository, and the third repository.
[0155] S115, obtain a loss function according to the first contrastive loss, the second contrastive loss, the third contrastive loss, the first label correction loss, and the second label correction loss.
[0156] S116, optimize the first optimal soft pseudo-label and the second optimal soft pseudo-label by using the loss function to optimize the unsupervised visible-infrared person re-identification.
[0157] It can be understood that, in order to improve the robustness to noise, a cross-repository label correction loss can be added during the training process. In the present invention, contrastive learning is performed during the forward propagation (FP), and each repository is updated during the backward propagation (BP). Since the third repository stores a plurality of the second features, the pseudo-labels corresponding to the plurality of the second features in the visible light space, and the cluster centers corresponding to the plurality of the second features, the third repository can be updated by the pseudo-labels of the infrared images transferred into the visible light label space.
[0158] Optionally, the first contrastive loss is obtained according to the following formula:
[0159]
[0160] Wherein, For the first contrast loss, the image dataset is divided into multiple batches, and the number B of training data in each batch is is the cross-entropy function, represents the feature from the first repository The predicted probability distribution output, τ is the temperature coefficient, represents the corresponding pseudo-label; The first contrast loss is the contrast loss within the visible light modality.
[0161] The second contrast loss is obtained according to the following formula:
[0162]
[0163] where is the second contrast loss, is the cross-entropy function, represents the feature from the original repository of the infrared light modality The predicted probability distribution output, represents the corresponding pseudo-label; represents the feature from the third repository of the infrared light modality The predicted probability distribution output, represents the corresponding pseudo-label; The second contrast loss is the contrast loss within the infrared light modality;
[0164] The third contrast loss is obtained according to the following formula:
[0165]
[0166] where represents the third contrast loss, represents the feature from the second repository M a The predicted probability distribution output, represents the corresponding pseudo-label; represents the feature from the second repository M a The predicted probability distribution output, represents the corresponding pseudo-label; The third contrast loss is the contrast loss between the visible light modality and the infrared light modality;
[0167] In the backpropagation stage, the features that perform contrast learning in the forward propagation can perform momentum updates on their corresponding repositories:
[0168]
[0169] Among them, μ is the momentum update factor. is the label in the repository of the prototype, while is the current input feature with the label It should be noted that consists of soft pseudo-labels and hard pseudo-labels, that is, a linear combination of soft pseudo-labels and one-hot hard pseudo-labels according to the soft pseudo-labels.
[0170] Furthermore, in order to improve the robustness to noise, an online label correction loss across repositories can be proposed during training, and the prediction of the intra-modal repository is used to correct the prediction of the inter-modal repository. The first label correction loss is obtained according to the following formula:
[0171]
[0172] Among them, is the first label correction loss, represents the feature the predicted probability distribution output from the second repository M a ; represents the feature the predicted probability distribution output from the first repository ; represents the feature the predicted probability distribution output from the third repository ;
[0173] Correspondingly, the second label correction loss is obtained according to the following formula:
[0174]
[0175] Among them, is the second label correction loss, represents the feature the predicted probability distribution output from the second repository M a ; represents the feature the predicted probability distribution output from the first repository ; represents the feature the predicted probability distribution output from the third repository ;
[0176] The loss function is obtained according to the following formula:
[0177]
[0178] Among them, is the loss function.
[0179] It can be understood that by continuously training and updating the generation of the first optimal soft pseudo-label and the second optimal soft pseudo-label according to the above loss function, better first and second optimal soft pseudo-labels can be obtained, thereby optimizing unsupervised visible-infrared pedestrian re-identification.
[0180] In one implementation, the test set is put into an unsupervised visible-infrared pedestrian re-identification model that applies the optimal soft pseudo-label in the present invention, and the pedestrian image features in the visible light and infrared modalities in the test set are respectively extracted using this model, and the pedestrian re-identification function is realized based on the Euclidean distance.
[0181] It is implemented using PyTorch, and the two-stream ResNet50 pre-trained on ImageNet is used as the backbone network. In a mini-batch training data, the number of classes P and the number of samples K for each class are both 12. All input images are resized to 288*144. The image enhancement method follows the ADCA algorithm. This method is trained for 90 epochs in total. The Adam optimizer is used for model training, and the weight decay is 5e-4. The initial learning rate is set to 3.5e-4 and reduced by 10 times at the 20th, 60th, and 80th epochs. The momentum factor μ is 0.1, the temperature coefficient τ is 0.05, and the parameter κ for κ-mutual nearest neighbors is set to 30. The hyperparameter λ for constructing the optimal transport problem is set to 25. The trade-off parameter α is set to 0.2, and the hyperparameter β for controlling the label smoothing degree is set to 0.7.
[0182] To evaluate this method, the dataset and evaluation metrics are introduced.
[0183] Evaluations are carried out on two public visible-infrared datasets, namely the SYSU-MM01 dataset and the RegDB dataset. The SYSU-MM01 dataset contains 395 identities, among which there are 22258 visible images and 11909 infrared images, and these images are captured by indoor and outdoor cameras. The RegDB dataset is a smaller dataset obtained by a pair of aligned visible and infrared cameras. It contains 412 identities, and each identity has 10 visible images and 10 infrared images.
[0184] All experiments follow the common evaluation protocol for visible-infrared. The evaluation metrics include Cumulative Matching Characteristic (CMC), mean Average Precision (mAP), and mean Inverse Negative Penalty (mINP). For the SYSU-MM01 dataset, the proposed method is evaluated in two search modes: full search mode and indoor search mode. As for the RegDB dataset, the method is evaluated in two test modes: visible light modality to infrared light modality, and infrared light modality to visible light modality. 206 identities are randomly selected for training, and the remaining 206 identities are used for testing. Table 1 shows the evaluation results of the present invention and the supervised visible-infrared pedestrian re-identification and unsupervised visible-infrared pedestrian re-identification methods in the existing methods on the SYSU-MM01 dataset, and Table 2 shows the evaluation results of the present invention and the supervised visible-infrared pedestrian re-identification and unsupervised visible-infrared pedestrian re-identification methods in the existing methods on the RegDB dataset. The specific evaluation results are shown in Table 1 and Table 2.
[0185] Table 1
[0186]
[0187] Table 2
[0188]
[0189]
[0190] Through the above technical solutions, the relationships at the sample level for homogeneous and heterogeneous in the visible light modality and the infrared light modality are considered, and the structural inconsistency is defined. Finally, the optimal soft pseudo-labels consistent with the feature space structure are obtained. The optimal soft pseudo-labels achieve the alignment of pseudo-labels under different modalities while protecting the clustering structure under the same modality, providing rich and reliable supervision signals for network training. Moreover, the present invention constructs intra-modal and inter-modal repositories for joint contrast learning, utilizes the optimal pseudo-labels, proposes label correction for the pseudo-label noise in unsupervised learning, and improves the robustness of the model to noise.
[0191] The following details the specific embodiments of the present invention in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.
[0192] It should be noted that all actions of obtaining signals, information, or data in the present invention are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where the location is located and obtaining the authorization given by the owner of the corresponding device.
[0193] The preferred embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solution of the present invention, and these simple modifications all fall within the protection scope of the present invention.
[0194] In addition, it should be noted that, in the various specific technical features described in the above specific embodiments, they can be combined in any suitable manner without conflict. To avoid unnecessary repetition, the present invention will not separately describe various possible combination methods.
[0195] Furthermore, any combination can be made between various different embodiments of the present invention as long as it does not violate the idea of the present invention, and it should also be regarded as the content disclosed by the present invention.
Claims
1. A label association method for unsupervised visible-infrared person re-identification, characterized in that: The method comprises: Acquire an image data set; wherein the image data set includes pedestrian images in a visible light modality and pedestrian images in an infrared light modality; According to the image data set, a first feature corresponding to the pedestrian image in the visible light modality and a second feature corresponding to the pedestrian image in the infrared light modality are obtained; An affinity matrix is constructed according to the first feature and the second feature to characterize homogeneous affinity and heterogeneous affinity; wherein the homogeneous affinity characterizes the affinity relationship within the image of the same modality, and the heterogeneous affinity characterizes the affinity relationship between the images of different modalities; Obtaining a first optimal soft false label corresponding to the visible light modality and a second optimal soft false label corresponding to the infrared light modality according to the affinity matrix specifically includes: obtaining a cluster center according to the first feature and the image data set; Acquire a first repository according to the cluster center; wherein the first repository contains a plurality of first features and pseudo labels corresponding to the plurality of first features; Acquire a first initial rough label according to the first repository and the first feature; Assigning a label to the second feature using the first repository and the optimal transmission label assignment algorithm as a second initial rough label; Minimize homogeneous inconsistency, heterogeneous inconsistency and self-inconsistency, iteratively optimize the pseudo-labels, and obtain the first optimal soft pseudo-label and the second optimal soft pseudo-label; The formula for homogeneous inconsistency is as follows: Among them, y e represents the pseudo label of the feature in the visible light modality or infrared light modality, N e represents the number of images corresponding to modality e in the image dataset, represents the homogeneity affinity between the i′th feature and the j′th feature of the visible light modality or the infrared light modality, represents the pseudo label corresponding to the i′th feature of the visible light modality or infrared light modality, represents the pseudo label corresponding to the j′th feature of the visible light modality or infrared light modality; The heterogeneous inconsistency formula is as follows: Among them, N v Indicates the number of images corresponding to the visible light modality v in the image dataset, N r represents the number of images corresponding to the infrared light modality r in the image dataset, represents the affinity matrix from the visible light modality to the infrared light modality, represents the pseudo label corresponding to the i″th feature of the visible light modality, represents the jth infrared light mode ′′ Pseudo labels corresponding to features; represents the heterogeneous inconsistency of the infrared light modality, represents the affinity matrix from infrared light modality to visible light modality, Indicates the i-th infrared light mode ′′′ Pseudo labels corresponding to features; Indicates the pseudo label corresponding to the j″′th feature of the visible light modality; The self-inconsistency formula is as follows: in, represents the pseudo label corresponding to the i1th feature of the visible light modality or infrared light modality, represents the initial coarse label of the i1th feature of the visible light modality or infrared light modality, N e Indicates the number of images corresponding to modality e in the image dataset; The first best soft pseudo label and the second best soft pseudo label are used as supervisory signals for unsupervised visible-infrared person re-identification.
2. The label association method for unsupervised visible-infrared person re-identification according to claim 1, characterized in that: The first initial coarse label acquisition method is as follows: in, represents the first initial coarse label, Represents feature f v From the first repository The predicted probability distribution of the output, K represents the number of features, and τ is the temperature coefficient; The second initial coarse label acquisition method is as follows: Among them, OTLA is the optimal transmission label allocation algorithm. Indicates that the first repository The pseudo-label in is assigned to the second feature f r , as the second initial coarse label 3. The label association method for unsupervised visible-infrared person re-identification according to claim 1, characterized in that: An affinity matrix is constructed based on the first feature and the second feature, including: Based on the first and second features, homogeneity affinity is constructed using Jaccard similarity; According to the first feature and the second feature, constructing a transmission cost between the first feature and the second feature using an optimal transmission model; Heterogeneous affinity is obtained based on transmission cost; The affinity matrix is obtained based on homogeneous affinity and heterogeneous affinity.
4. The label association method for unsupervised visible-infrared pedestrian re-identification according to claim 3, characterized in that: Homogeneity affinity is obtained according to the following formula: Among them, S ho(e) For S ho(v) or S ho(r) , which indicates homogeneous affinity within the visible light mode or homogeneous affinity within the infrared light mode, that is, within the same mode and The affinity between express The κ-nearest neighbors of express κ-mutual nearest neighbors, S ho(v) is the affinity relationship within the image of the visible light modality, S ho(r) is the affinity relationship within the image of the infrared light modality, and are two different features in the same mode; The heterogeneous affinity is obtained according to the following formula: Among them, S he represents the transfer plan between images of different modalities, C he is the cost matrix constructed based on the Euclidean distance between the features of different modes, <> represents the Frobenius dot product, λ is a hyperparameter, and the optimal transmission plan S obtained according to the above formula he* As heterogeneous affinity, represents the i-th feature of the visible light mode, Represents the jth feature of the infrared light modality.
5. The label association method for unsupervised visible-infrared person re-identification according to claim 1, characterized in that: The method further comprises: Constructing a second repository and a third repository; wherein the second repository contains a plurality of first features and pseudo labels corresponding to the plurality of first features, a plurality of second features and pseudo labels corresponding to the plurality of second features in the visible light space, and the third repository stores a plurality of second features and pseudo labels corresponding to the plurality of second features in the visible light space and cluster centers corresponding to the plurality of second features; Obtaining a first contrast loss corresponding to a visible light modality according to a first storage memory; Obtaining a second contrast loss corresponding to the infrared light modality according to the third storage memory; The cross-modal contrast loss of the visible light modality and the infrared light modality is obtained according to the second storage library, which is recorded as the third contrast loss; Obtaining a first label correction loss corresponding to the visible light modality and a second label correction loss corresponding to the infrared light modality according to the first storage library, the second storage library, and the third storage library; A loss function is obtained according to the first contrast loss, the second contrast loss, the third contrast loss, the first label correction loss, and the second label correction loss; The loss function is used to optimize the first best soft false label and the second best soft false label to optimize the unsupervised visible-infrared pedestrian re-identification.
6. The label association method for unsupervised visible-infrared person re-identification according to claim 5, characterized in that: The first contrast loss is obtained according to the following formula: in, For the first contrast loss, the image dataset is divided into multiple batches, and the number of training data in each batch is B. is the cross entropy function, Representation characteristics From the first repository The predicted probability distribution of the output, τ is the temperature coefficient, express The corresponding pseudo labels; The second contrast loss is obtained according to the following formula: in, is the second contrast loss, Representation characteristics From the original repository of infrared light modalities The predicted probability distribution of the output, express The corresponding pseudo labels; Representation characteristics From the third repository of infrared light modalities The predicted probability distribution of the output, express The corresponding pseudo labels; The third contrast loss is obtained according to the following formula: in, represents the third contrast loss, Representation characteristics From the second repository M a The predicted probability distribution of the output, express The corresponding pseudo labels; Representation characteristics From the second repository M a The predicted probability distribution of the output, express The corresponding pseudo labels; The first label correction loss is obtained according to the following formula: in, Corrected loss for the first label, Representation characteristics From the second repository M a The predicted probability distribution of the output, Representation characteristics From the first repository The predicted probability distribution of the output, Representation characteristics From the third repository The predicted probability distribution of the output; The second label correction loss is obtained according to the following formula: in, Correct the loss for the second label, Representation characteristics From the second repository M a The predicted probability distribution of the output, Representation characteristics From the first repository The predicted probability distribution of the output, Representation characteristics From the third repository The predicted probability distribution of the output; The loss function is obtained according to the following formula: in, is the loss function.
Citation Information
Patent Citations
Unsupervised cross-modal pedestrian re-identification method and system based on hierarchical difference
CN117351518A