Video-based unsupervised visible light infrared pedestrian re-identification method

Through the video-based unsupervised visible light infrared pedestrian re-identification method, the feature extraction module, clustering module and progressive pseudo-label correction module are used to solve the problem of pseudo-label noise samples, improve the robustness and accuracy of the model, and are suitable for pedestrian re-identification in complex scenarios.

CN120452021APending Publication Date: 2025-08-08CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510563355.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the existing unsupervised pedestrian re-identification method, pseudo label noise samples cannot effectively participate in model training, resulting in difficulty in learning features of the model and lack of difficult sample representation ability, forming a vicious cycle.

Method used

The unsupervised visible light infrared pedestrian re-identification method based on video is adopted, and the feature extraction module, clustering module and progressive pseudo-label correction module are used to gradually restore the pseudo-label to an effective label through in-modular and inter-modular correction techniques, thereby enhancing the robustness of the model.

Benefits of technology

By utilizing sequence information and a progressive pseudo-label correction module, the robustness and performance of the model in complex scenarios are improved, the modal differences are reduced, and the accuracy of pedestrian re-identification is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452021A_ABST
    Figure CN120452021A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of pedestrian re-identification, and relates to a video-based unsupervised visible light infrared pedestrian re-identification method, which comprises the following steps: acquiring query data and a data set, inputting the query data and the data set into a trained re-identification model to obtain query features and a feature set, and matching the query features with the feature set to obtain an identification result; the training process of the re-identification model comprises the following steps: acquiring visible light data SV and infrared data ST; inputting the SV and the ST into a feature extraction module to obtain visible light and infrared features FV and FT; inputting the FV and the FT into a clustering module to obtain a clustering result; inputting the clustering result into a progressive false label correction module to obtain a corrected clustering result; inputting the SV and the ST into a feature extraction module to obtain visible light and infrared features qV and qT; updating model parameters according to the qV, the qT and the corrected clustering result until a trained re-identification model is obtained; according to the method, noise samples are recovered into effective labels through intra-modal correction and inter-modal correction, and robustness is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of visible light-infrared pedestrian re-identification, and relates to a video-based unsupervised visible light-infrared pedestrian re-identification method. Background Art

[0002] Research on person re-identification has undergone a paradigm shift from traditional hand-crafted features to deep learning-driven approaches. With technological advancements, it has gradually overcome the challenges of complex real-world scenarios. Early research relied on manually designed shallow features such as color and texture, but was limited in their expressiveness and environmental adaptability was also extremely challenging. With the rise of deep learning, convolutional neural networks have significantly improved the discriminability of pedestrian representations through end-to-end feature learning, and their cross-camera generalization capabilities have been greatly enhanced. To achieve successful practical applications, the need for all-weather intelligent monitoring must be addressed. Consequently, visible-light-infrared cross-modal re-identification has become a research focus, dedicated to reducing the difference between day and night modalities and alleviating challenges such as sudden changes in illumination.

[0003] On this basis, in order to reduce the dependence on massive labeled data, unsupervised learning solutions came into being. Through clustering to generate pseudo labels, cross-modal alignment and other technologies, it achieved performance close to supervised learning on unlabeled data, further promoting the evolution of pedestrian re-identification technology towards practical application.

[0004] Existing unsupervised methods typically use clustered pseudo-labels to simulate true labels to constrain model learning. However, the pseudo-labels contain a large amount of unavoidable noise, which prevents these samples from participating in model training. This makes it difficult for the model to learn difficult features, which in turn leads to a loss of representation for difficult samples, creating a vicious cycle. Summary of the Invention

[0005] To address the above-mentioned problems in the prior art, the present invention adopts a video-based unsupervised visible-infrared person re-identification method, comprising: obtaining visible-infrared query data and a visible-infrared dataset, inputting the visible-infrared query data and the visible-infrared dataset into a trained person re-identification model, obtaining query features and a feature set, and matching the query features with all features in the feature set to obtain a person recognition result; the person re-identification model includes: a feature extraction module, a clustering module, and a progressive pseudo-label correction module;

[0006] The training process of the person re-identification model includes:

[0007] S1. Obtain a visible light-infrared dataset, which includes a visible light data sequence S V and infrared data sequence S TThe visible light data sequence and the infrared data sequence include multiple visible light samples and multiple infrared samples, respectively. Each visible light sample includes multiple frames of visible light images, and each infrared sample includes multiple frames of infrared images.

[0008] S2, the visible light data sequence S V and infrared data sequence S T Input the feature extraction module respectively to obtain the visible light feature sequence F V and infrared characteristic sequence F T ;

[0009] S3, the visible light feature sequence F V and infrared characteristic sequence F T Input the clustering module respectively to obtain the visible light feature sequence F V and infrared characteristic sequence F t The clustering results of

[0010] S4, the visible light feature sequence F V and infrared characteristic sequence F t The clustering result is input into the progressive pseudo-label correction module for correction to obtain the visible light feature sequence F V and infrared characteristic sequence F T Corrected clustering results;

[0011] S5, the visible light data sequence S V and infrared data sequence S T Input the feature extraction module respectively to obtain the visible light query feature sequence q V and infrared query feature sequence q T ;

[0012] S6. Query feature sequence q based on visible light V , infrared query feature sequence q T The overall loss function value is calculated based on the corrected clustering results, and the model parameters are updated according to the overall loss function value. When the overall loss function value is minimized, the trained pedestrian re-identification model is obtained.

[0013] Beneficial effects:

[0014] 1. The present invention makes full use of sequence information. Compared with single-frame images, sequence information captures the temporal and spatial information of pedestrians, has richer features, and makes up for the shortcomings of single-frame features; 2. The progressive pseudo-label correction module of the present invention gradually restores the pseudo-labels of noise samples to valid labels through simple frame-guided intra-modal correction and positive sample sequence-guided inter-modal correction, thereby providing an opportunity for model training to learn noise samples and enhancing the robustness of the model; 3. The present invention additionally provides simple prototypes for correctable noise samples and difficult prototypes for non-noise samples. The intra-modal difficulty and ease mutual learning contrast loss is calculated based on the simple prototypes and difficult prototypes. According to the intra-modal difficulty and ease mutual learning contrast loss model, the non-noise samples and noise samples within the modality can be gradually aligned to improve model performance; the present invention uses cross-modal sequence contrast loss and cross-modal frame contrast loss to align cross-modal positive samples, reduce modal differences, and improve model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A training flow chart of an unsupervised visible light infrared person re-identification method based on video provided by an embodiment of the present invention;

[0016] Figure 2 An overall flow chart of the training and testing of an unsupervised visible light infrared person re-identification method based on video provided by an embodiment of the present invention;

[0017] Figure 3 This is a structural diagram of a video-based unsupervised visible light infrared pedestrian re-identification method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0019] like Figure 1 、 Figure 2 、 Figure 3As shown, the present invention adopts a video-based unsupervised visible light infrared pedestrian re-identification method, including: obtaining visible light-infrared query data and a visible light-infrared dataset, inputting the visible light-infrared query data and the visible light-infrared dataset into a trained pedestrian re-identification model, obtaining query features of the visible light-infrared query data and a feature set of the visible light-infrared dataset, matching the query features with features in the feature set to obtain a matching result, i.e., a pedestrian recognition result; wherein the visible light-infrared query data is a visible light-infrared image of the pedestrian to be queried, and the visible light-infrared dataset is a visible light-infrared image dataset of the pedestrian recorded by a security system, for example; the pedestrian re-identification model includes: a feature extraction module, a clustering module, and a progressive pseudo-label correction module;

[0020] Matching the query feature with the features in the feature set includes: calculating the similarity between the query feature and each feature in the feature set, and taking the labels corresponding to the top-k features with the highest similarity in the feature set as the pedestrian recognition results.

[0021] The training process of the person re-identification model includes:

[0022] S1. Obtain a visible light-infrared dataset, which includes a visible light data sequence. and infrared data series The visible light data sequence and the infrared data sequence include multiple visible light samples and multiple infrared samples respectively, each visible light sample includes multiple frames of visible light images, and each infrared sample includes multiple frames of infrared images; wherein, is the jth visible light image of the i-th visible light sample, is the jth infrared image frame of the i-th infrared sample, N is the number of visible light samples, M is the number of infrared samples, and J is the number of frames.

[0023] S2, the visible light data sequence S V and infrared data sequence S T Input the feature extraction module respectively to obtain the visible light feature sequence F V and infrared characteristic sequence F T ;

[0024] The feature extraction module includes: a visible light feature extraction module and an infrared feature extraction module; the feature extraction module extracts the visible light data sequence S V and infrared data sequence S T The processing includes:

[0025] S21, the visible light data sequence S V Input the visible light feature extraction module to obtain the visible light feature sequence F V and pseudo-infrared characteristic sequence FVC ;

[0026] The visible light feature extraction module extracts the visible light data sequence S V The processing includes: processing the visible light data sequence S V Perform channel enhancement to obtain pixel-aligned pseudo-infrared data sequence S VC ; The visible light data sequence S V and the pseudo-infrared data sequence S VC Input the visible light feature extraction module respectively to obtain the visible light feature sequence and pseudo-infrared characteristic sequences F VC and F ′V Combine to obtain the final visible light feature sequence

[0027] S22, the infrared data sequence S T Input infrared feature extraction module to obtain infrared feature sequence

[0028] S3, the visible light feature sequence F V and infrared characteristic sequence F T Input the clustering module respectively to obtain the visible light feature sequence F V and infrared characteristic sequence F T The clustering results of

[0029] Specifically, the clustering module includes: a frame-level clustering module and a sequence-level clustering module; the visible light feature sequence F V and infrared characteristic sequence F T Input the frame-level clustering module respectively to obtain the visible light feature sequence F V and infrared characteristic sequence F T The frame-level clustering results of the visible light feature sequence F V and infrared characteristic sequence F T Input the sequence level clustering module respectively to obtain the visible light feature sequence F V and infrared characteristic sequence F T Sequence-level clustering results.

[0030] The sequence-level clustering module is used to classify the visible light feature sequence F V and infrared characteristic sequence F T Processing includes:

[0031] S31, visible light characteristic sequence F V and infrared characteristic sequence F T Perform average time series pooling respectively to obtain the pooled visible light feature sequence And the pooled infrared feature sequence

[0032] Using the simplest temporal pooling operation, the pooled visible light feature sequence is Infrared feature sequence after pooling in, is the visible light sample characteristic, Infrared sample characteristics.

[0033] S32, use DBSCAN clusterer to cluster the visible light sequence features after pooling And the infrared sequence features after pooling The features of each sample in the sequence are clustered to obtain the pooled visible light sequence features And the infrared sequence features after pooling The sequence-level clustering results include: multiple visible light clusters and their cluster centers and labels, multiple infrared clusters and their cluster centers and labels, and pseudo labels for each visible light sample and each infrared sample. The pseudo label for each visible light sample and each infrared sample is the label of the cluster to which its feature belongs, and samples with a pseudo label of -1 are noise samples.

[0034] Pseudo-labeling in, is the pseudo label of the i-th visible light sample, is the pseudo label of the j-th infrared sample.

[0035] Visible light cluster center Infrared Cluster Center in, is the kth visible light cluster, is the feature of the nth sample of the kth cluster; is the lth infrared cluster, is the feature of the nth sample in the lth cluster.

[0036] The frame-level clustering module is used to classify the visible light feature sequence F V and infrared characteristic sequence F T Processing includes:

[0037] The DBSCAN clusterer is used to cluster the visible light feature sequence F V and infrared characteristic sequence F T Each frame of each sample in the cluster is clustered to obtain the visible light feature sequence F V and infrared characteristic sequence F TThe frame-level clustering results include multiple visible light clusters and their cluster centers and labels, multiple infrared clusters and their cluster centers and labels, and pseudo labels for each visible light sample and each infrared sample frame; the pseudo labels for each visible light sample frame and the pseudo labels for each infrared sample frame are the labels of the clusters to which they belong; the pseudo labels for each visible light sample frame are the pseudo labels for each visible light sample frame. Pseudo labels for each frame of infrared samples

[0038] During the clustering process, some samples may be mistakenly labeled as noise samples (i.e., the pseudo label is -1) due to feature degradation or noise interference, and the remaining labels represent the categories of pedestrians.

[0039] Because the clusterer independently clusters frame-level and sequence-level features, the frame label and sequence label for the same sample may not necessarily be the same. However, since they belong to the same sample, there is a direct mapping between them. This mapping is the key foundation for noise correction. In other words, at the pseudo-label level, the acceptable pseudo-label for a noisy sequence is the valid sequence pseudo-label corresponding to the pseudo-label of the most representative valid frame in the sequence.

[0040] S4, the visible light feature sequence F V and infrared characteristic sequence F T The clustering result is input into the progressive pseudo-label correction module for correction to obtain the visible light feature sequence F V and infrared characteristic sequence F T Corrected clustering results;

[0041] The progressive pseudo-label correction module includes: an intra-modality correction module and an inter-modality correction module; the progressive pseudo-label correction module corrects the visible light feature sequence F V and infrared characteristic sequence F T Correction of clustering results includes:

[0042] S41, the visible light feature sequence F V and infrared characteristic sequence F t The clustering results are input into the intra-modal correction module to obtain the intra-modal corrected visible light feature sequence F V and infrared characteristic sequence F T The clustering results of

[0043] The intra-modality correction module analyzes the frame-level pseudo-labels of the noisy sequence and identifies simple frames with strong discriminability (i.e., frames that are close to the cluster center of the positive samples). Based on the pseudo-labels of these simple frames, the module can infer the true pseudo-labels of the noisy sequence and restore them to valid sequence-level pseudo-labels.

[0044] Specifically, the intra-modal correction module corrects the visible light feature sequence FV and infrared characteristic sequence F T The clustering results are processed as follows:

[0045] S411, based on pseudo labels Construct a visible light noise sample sequence for the visible light sample of -1 According to the pseudo-label Construct infrared noise sample sequence for infrared sample of -1

[0046] The label pair of each frame in the visible light noise sample sequence is Frame-level labels There are many cases, some frames are noise frames (may appear as severely occluded frames), the pseudo label is -1, while some frames (simple frames) have a normal label; the samples with simple frames in the noise sample sequence are simple samples, and the samples without simple frames are difficult samples;

[0047] S412, recovering a pseudo label corresponding to each simple sample according to a simple frame of each simple sample in the visible light noise sample sequence; recovering a pseudo label corresponding to each simple sample according to a simple frame of each simple sample in the infrared noise sample sequence;

[0048] The pseudo labels corresponding to the simple samples recovered from the simple frames of the simple samples in the visible light noise sample sequence include:

[0049] Select a simple sample All simple frames (i.e. ), calculate the distance between the features of each simple frame and its cluster center Set the distance threshold τ, filter out simple frames whose distance from the cluster center is less than the distance threshold τ, count the most pseudo labels in all the filtered simple frames, and take the most pseudo labels as simple samples Pseudo-label of ; where d is the distance metric function (such as Euclidean distance), For simple frames The cluster center of

[0050] Specifically, the pseudo-label information of the simple frame is used to restore the pseudo-label of the noise sequence to a valid label consistent with its simple frame. The restored pseudo-label can be expressed as Where mode is the mode function, which is used to select the pseudo label with the highest frequency.

[0051] Similarly, the pseudo labels of the simple samples can be recovered based on the simple frames of the simple samples in the infrared noise sample sequence.

[0052] Through the above steps, the module can restore the noisy labels of most simple samples into valid labels, thereby significantly improving the discriminability of sequence-level features.

[0053] S42, the visible light feature sequence F after the modal correction V and infrared characteristic sequence F T The clustering results are input into the inter-modal correction module to obtain the final corrected visible light feature sequence F V and infrared characteristic sequence F T The clustering results.

[0054] In the task of person re-identification based on video sequences, some difficult samples may have significantly degraded feature discriminability across all frames due to severe occlusion, sudden illumination changes, or other complex scene factors. Both their frame-level and sequence-level pseudo-labels are marked as noise (i.e., the pseudo-labels are invalid values). For such samples, correction methods guided by simple intra-modality frames cannot be directly applied because all their frames are difficult frames and lack direct association with positive samples. To address this issue, the positive sequence-guided inter-modality correction module (PSC) indirectly recovers the pseudo-labels of the entire difficult sequence through the mapping relationship between cross-modal positive samples.

[0055] Specifically, the inter-modality correction module corrects the visible light feature sequence F after intra-modality correction. V and infrared characteristic sequence F T The clustering results are processed as follows:

[0056] S421, based on pseudo labels The characteristic of infrared samples with a value of -1 Constructing infrared noise sample feature sequence According to the pseudo-label Characteristics of visible light samples that are not -1 Constructing visible light sample feature sequence The infrared noise sample feature sequence Visible light sample characteristic sequence Perform similarity matching to obtain the cross-modal (visible light modality) positive sample of each infrared sample in the infrared noise sample feature sequence;

[0057] Similarity matching includes: Where, Calculate the similarity between infrared features and visible light features. The range is [-1, 1]. The larger the value, the higher the similarity. Represents the feature vector of the i-th sample in the infrared noise sample feature sequence, Represents the feature vector of the jth non-noise sample in the visible light sample feature sequence, Represents the feature vector The L2 norm of (Euclidean distance normalization).

[0058] S422, based on pseudo labels The characteristic of the visible light sample is -1 Constructing a feature sequence of visible light noise samples According to the pseudo-label Characteristics of infrared samples that are not -1 Construct infrared sample feature sequence The visible light noise sample feature sequence and infrared sample characteristic sequences Perform similarity matching to obtain the cross-modal (infrared mode) positive sample of each visible light sample in the visible light noise sample feature sequence;

[0059] S423: Map the pseudo labels of the cross-modal (visible light modality) positive samples of each infrared sample to obtain the cross-modal mapped pseudo labels of each infrared sample. The pseudo label of each infrared sample is restored to the corresponding cross-modal mapping pseudo label, that is, Similarly, the pseudo labels of the cross-modal (infrared modality) positive samples of each visible light sample are mapped to obtain the cross-modal mapping pseudo labels of each visible light sample. The pseudo label of each visible light sample is restored to the corresponding cross-modal mapping pseudo label, that is, Among them, PGM (Progressive Graph Matching) is a progressive graph matching algorithm and a general cross-modal label mapping algorithm;

[0060] S424: Combine the features of visible light samples with the same pseudo-label to obtain corrected visible light clustering. Combine the features of infrared samples with the same pseudo-label to obtain the corrected infrared clustering

[0061] Through the above steps, the module can indirectly recover pseudo-labels for all difficult sequences by leveraging the mapping relationship between cross-modal positive samples, enabling them to participate in model training. Adding a positive sample sequence-guided inter-modal correction module further leverages cross-modal information. By mining the positive sample mapping relationship between visible and infrared modalities, the module can effectively utilize cross-modal information to recover pseudo-labels for all difficult sequences. This addresses difficult sequences that cannot be directly corrected within the modality, significantly improving the model's robustness in complex scenarios.

[0062] S5, the visible light data sequence S V and infrared data sequence S T Input the feature extraction module respectively to obtain the visible light query feature sequence and infrared query feature sequence in, is the feature of the jth frame of the i-th visible light sample, is the feature of the jth frame of the i-th infrared sample;

[0063] The specific feature extraction process is the same as step 2.

[0064] S6. Query feature sequence q based on visible light V , infrared query feature sequence q T The overall loss function value is calculated based on the corrected clustering results, and the model parameters are updated according to the overall loss function value. When the overall loss function value is minimized, the trained pedestrian re-identification model is obtained.

[0065] Calculating the overall loss function value includes:

[0066] S51. Clustering based on corrected visible light and infrared clustering Calculate the corrected visible light cluster centers separately and infrared cluster centers

[0067] S52, according to the cluster center Calculate each visible light cluster and infrared clustering Prototype

[0068] Among them, for visible light clustering whose pseudo label is not -1 Select visible light clustering Distance visible light clustering among all visible light samples The cluster center The features of the nearest frame are used as visible light clusters Prototype (ie easy prototype); for visible light clustering with pseudo label -1 Select visible light clustering All visible light samples in the distance visible light clustering The cluster center The features of the farthest frame are clustered as visible light Prototype (i.e. difficult prototype); similarly, calculate the prototype of each infrared cluster.

[0069] S53, query the visible light feature sequence q V and infrared query feature sequence q T Perform average time series pooling to obtain the pooled visible light query feature sequence and infrared query feature sequence in, is the feature of the i-th visible light sample, is the feature of the i-th infrared sample;

[0070] S54. Query feature sequence based on pooled visible light and infrared query feature sequence And the prototype calculation of visible light clustering and infrared clustering modality difficulty mutual learning contrast loss L intra ;

[0071]

[0072] in, The contrast loss calculated for the visible light query feature and the visible light sample pair, is the contrast loss calculated for the infrared query feature and infrared sample pair, is the prototype of the visible light cluster corresponding to the pseudo label of the i-th visible light sample, τ is the temperature parameter, K is the number of visible light sample clusters, is the prototype of the visible light cluster corresponding to the pseudo label of the i-th infrared sample, and L is the number of infrared sample clusters.

[0073] S53, query feature sequence based on pooled visible light and infrared query feature sequence And the prototypes of visible light clustering and infrared clustering calculate the sequence-level cross-modal contrast loss L inter ;

[0074]

[0075] Among them, Lq V2T Lq is the contrast loss calculated for the visible light query feature and infrared sample pair. T2V Contrastive loss computed for infrared query feature and visible light sample pairs.

[0076] S54. Query feature sequence q based on visible light V , infrared query feature sequence q T And the prototypes of visible light clustering and infrared clustering calculate the frame-level cross-modal contrast loss L frm ;

[0077]

[0078] Among them, Lq V2T,frm is the frame-level contrast loss calculated for the visible light query feature and infrared sample pair, Lq T2V ,frmThe frame-level contrast loss calculated for infrared query features and visible light samples, is the prototype of the visible light cluster corresponding to the pseudo label of the jth frame of the i-th visible light sample, τ is the temperature parameter, K is the number of visible light sample clusters, is the prototype of the visible light cluster corresponding to the pseudo label of the jth frame of the i-th infrared sample.

[0079] S55, combined contrast loss L intra , L inter and L frm , get the overall loss function value L total .

[0080] L total =L intra +L inter +L frm

[0081] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation plans of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A video-based unsupervised visible light infrared pedestrian re-identification method, characterized by: include: Obtain visible light-infrared query data and a visible light-infrared dataset, input the visible light-infrared query data and the visible light-infrared dataset into the trained person re-identification model respectively, obtain query features and feature sets, match the query features with all features in the feature set, and obtain the person recognition result; The pedestrian re-identification model includes: feature extraction module, clustering module and progressive pseudo-label correction module; The training process of the person re-identification model includes: S1. Obtain a visible light-infrared dataset, which includes a visible light data sequence S V and infrared data sequence S T The visible light data sequence and the infrared data sequence include multiple visible light samples and multiple infrared samples, respectively. Each visible light sample includes multiple frames of visible light images, and each infrared sample includes multiple frames of infrared images. S2, the visible light data sequence S V and infrared data sequence S T Input the feature extraction module respectively to obtain the visible light feature sequence F V and infrared characteristic sequence F T ; S3, the visible light feature sequence F V and infrared characteristic sequence F T Input the clustering module respectively to obtain the visible light feature sequence F V and infrared characteristic sequence F t The clustering results of S4, the visible light feature sequence F V and infrared characteristic sequence F t The clustering result is input into the progressive pseudo-label correction module for correction to obtain the visible light feature sequence F V and infrared characteristic sequence F T Corrected clustering results; S5, the visible light data sequence S V and infrared data sequence S T Input the feature extraction module respectively to obtain the visible light query feature sequence q V and infrared query feature sequence q T ; S6. Query feature sequence q based on visible light V , infrared query feature sequence q T The overall loss function value is calculated based on the corrected clustering results, and the model parameters are updated according to the overall loss function value. When the overall loss function value is minimized, the trained pedestrian re-identification model is obtained.

2. The unsupervised visible light infrared person re-identification method based on video according to claim 1 is characterized in that: The feature extraction module includes: a visible light feature extraction module and an infrared feature extraction module; the feature extraction module extracts the visible light data sequence S V and infrared data sequence S T The processing includes: converting the visible light data sequence S V Input the visible light feature extraction module to obtain the visible light feature sequence F V ; The infrared data sequence S T Input the infrared feature extraction module to obtain the infrared feature sequence F T .

3. The unsupervised visible light infrared pedestrian re-identification method based on video according to claim 1 is characterized in that: The clustering module includes: frame level clustering module and sequence level clustering module; the clustering module is used to cluster the visible light feature sequence F V and infrared characteristic sequence F T Clustering includes: The visible light feature sequence F V and infrared characteristic sequence F T Input the frame-level clustering module respectively to obtain the visible light feature sequence F V and infrared characteristic sequence F T The frame-level clustering results of the visible light feature sequence F V and infrared characteristic sequence F T Input the sequence level clustering module respectively to obtain the visible light feature sequence F V and infrared characteristic sequence F T Sequence-level clustering results.

4. The video-based unsupervised visible light infrared pedestrian re-identification method according to claim 3 is characterized in that: The sequence-level clustering module is used to classify the visible light feature sequence F V and infrared characteristic sequence F T Processing includes: For the visible light characteristic sequence F V and infrared characteristic sequence F T Perform average time series pooling respectively to obtain the pooled visible light feature sequence And the pooled infrared feature sequence Visible light feature sequence after pooling Includes the features of each visible light sample and the pooled infrared feature sequence Includes characteristics of each infrared sample; The visible light sequence features after pooling are And the infrared sequence features after pooling Cluster the features of each sample in to obtain the pooled visible light sequence features And the infrared sequence features after pooling The sequence-level clustering results include: multiple visible light clusters and their cluster center labels, multiple infrared clusters and their cluster centers and labels, and pseudo labels for each visible light sample and each infrared sample. The pseudo label for each visible light sample and each infrared sample is the label of the cluster to which its feature belongs, and samples with a pseudo label of -1 are noise samples.

5. The video-based unsupervised visible light infrared pedestrian re-identification method according to claim 4 is characterized in that: Visible light characteristic sequence F V and infrared characteristic sequence F T Including the features of each frame of the corresponding sample; the frame level clustering module is used to cluster the visible light feature sequence F V and infrared characteristic sequence F T The processing includes: respectively processing the visible light feature sequence F V and infrared characteristic sequence F T Each frame of each sample in the cluster is clustered to obtain the visible light feature sequence F V and infrared characteristic sequence F T The frame-level clustering results include multiple visible light clusters and their cluster centers and labels, multiple infrared clusters and their cluster centers and labels, and pseudo labels for each visible light sample and each infrared sample frame; wherein, the pseudo labels for each visible light sample frame and the pseudo labels for each infrared sample frame are the labels of the clusters to which their features belong.

6. The video-based unsupervised visible light infrared pedestrian re-identification method according to claim 5, characterized in that: The progressive pseudo-label correction module includes: an intra-modality correction module and an inter-modality correction module; the progressive pseudo-label correction module corrects the visible light feature sequence F V and infrared characteristic sequence F T Correction of clustering results includes: S41, the visible light feature sequence F V and infrared characteristic sequence F t The clustering results are input into the intra-modal correction module to obtain the intra-modal corrected visible light feature sequence F V and infrared characteristic sequence F T The clustering results of S42, the visible light feature sequence F after the modal correction V and infrared characteristic sequence F T The clustering results are input into the inter-modal correction module to obtain the final corrected visible light feature sequence F V and infrared characteristic sequence F T The clustering results.

7. The video-based unsupervised visible light infrared pedestrian re-identification method according to claim 6, characterized in that: The intra-modal correction module is used to correct the visible light feature sequence F V and infrared characteristic sequence F T The clustering results are processed as follows: S411, based on pseudo labels Construct a visible light noise sample sequence for the visible light sample of -1, according to the pseudo label Construct an infrared noise sample sequence for the infrared sample with a pseudo label of -1; among them, the frames whose pseudo labels are not -1 in the sample are simple frames, the samples with simple frames are simple samples, and the samples without simple frames are difficult samples; S412 , recovering a pseudo label corresponding to each simple sample according to the simple frame of each simple sample in the visible light noise sample sequence; and recovering a pseudo label corresponding to each simple sample according to the simple frame of each simple sample in the infrared noise sample sequence.

8. The video-based unsupervised visible light infrared pedestrian re-identification method according to claim 7, characterized in that: Restoring the pseudo labels of simple samples according to the simple frames of simple samples in the visible light noise sample sequence includes: selecting all simple frames of the simple samples, calculating the distance between the features of each simple frame and its cluster center; setting a distance threshold τ, screening simple frames whose distance to the cluster center is less than the distance threshold τ, counting the most pseudo labels in all the screened simple frames, and taking the most pseudo labels as the pseudo labels of the simple samples.

9. The video-based unsupervised visible light infrared pedestrian re-identification method according to claim 6, characterized in that: The inter-modal correction module corrects the visible light feature sequence F after intra-modal correction V and infrared characteristic sequence F T The clustering results are processed as follows: S421, based on pseudo labels Construct an infrared noise sample feature sequence for the infrared sample with a value of -1; The features of the visible light samples that are not -1 are used to construct a visible light sample feature sequence; the infrared noise sample feature sequence is matched with the visible light sample feature sequence for similarity to obtain a cross-modal positive sample of each infrared sample in the infrared noise sample feature sequence; S422, based on pseudo labels Construct a visible light noise sample feature sequence for the feature of the visible light sample of -1; according to the pseudo label The features of infrared samples that are not -1 are used to construct an infrared sample feature sequence; the visible light noise sample feature sequence and the infrared sample feature sequence are matched for similarity to obtain a cross-modal positive sample of each visible light sample in the visible light noise sample feature sequence; S423. Map the pseudo labels of the cross-modal positive samples of each infrared sample in the infrared noise sample feature sequence to obtain a cross-modal mapping pseudo label of each infrared sample; restore the pseudo label of each infrared sample to the corresponding cross-modal mapping pseudo label; map the pseudo labels of the cross-modal positive samples of each visible light sample in the visible light noise sample feature sequence to obtain a cross-modal mapping pseudo label of each visible light sample, and restore the pseudo label of each visible light sample to the corresponding cross-modal mapping pseudo label; S424 , combining features of visible light samples with the same pseudo-label to obtain corrected visible light clusters, and combining features of infrared samples with the same pseudo-label to obtain corrected infrared clusters.

10. The video-based unsupervised visible light infrared pedestrian re-identification method according to claim 1, characterized in that: Corrected clustering results include corrected visible light clustering and infrared clustering Calculating the loss function value includes: S51. Clustering based on corrected visible light and infrared clustering Calculate the corrected visible light cluster centers separately and infrared cluster centers Where k and l represent the indexes of visible light clustering and infrared clustering respectively; S52, according to the cluster center Calculate each visible light cluster and infrared clustering Prototype S53, query the visible light feature sequence q V and infrared query feature sequence q T Perform average time series pooling to obtain the pooled visible light query feature sequence and infrared query feature sequence S54. Query feature sequence based on pooled visible light and infrared query feature sequence and prototypes of visible light clustering and infrared clustering Calculate the intra-modality difficulty-easy mutual learning contrast loss L intra ; S55. Query feature sequence based on pooled visible light and infrared query feature sequence and prototypes of visible light clustering and infrared clustering Calculate the sequence-level cross-modal contrast loss L inter ; S56. Query feature sequence q based on visible light V , infrared query feature sequence q T and prototypes of visible light clustering and infrared clustering Calculate the cross-modal contrast loss L at the frame level frm ; S57, combined contrast loss L intra 、L inter and L frm , and get the overall loss function value.

Citation Information

Cited By

  • Audio and video analysis method based on noise label learning

    CN121682445A

  • An audio and video parsing method based on noise label learning

    CN121682445B

  • Unsupervised cross-modal pedestrian re-identification method based on hierarchical clustering and edge noise learning

    CN122157315A