Re-recognition model training method and system based on multi-modal collaborative learning framework

Through the multimodal collaborative learning framework, pseudo-labels are generated using visible and infrared light image features and dynamic comparison learning is solved, and the problem of insufficient semantic relationship capture and noise influence in visible-infrared pedestrian re-identification is achieved, achieving a more stable recognition effect.

CN120260075APending Publication Date: 2025-07-04HUNAN NORMAL UNIVERSITY
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510279339.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing visible-infrared pedestrian re-identification training method cannot fully capture the semantic relationship between samples and is easily affected by noise, resulting in unstable model performance.

Method used

Using a method based on a multimodal collaborative learning framework, by obtaining visible and infrared light image features, generating pseudo-labels and using pre-trained CLIP models to generate text semantic information, dynamic comparison learning and weighting calculations are carried out, and the contrast loss function from image to text is constructed for model training.

Benefits of technology

Effectively capture higher-order semantic information as a supervision signal, enhance the cluster reliability and distinction of homomodal learning, improve the robustness and stability of the model, and improve the recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260075A_ABST
    Figure CN120260075A_ABST
Patent Text Reader

Abstract

The invention relates to the field of re-recognition model processing, and discloses a re-recognition model training method and system based on a multi-mode collaborative learning framework, and the method comprises the steps: generating visible light text semantic information and infrared light text semantic information based on a pre-trained CLIP model; performing dynamic contrast learning based on a sample input into the visible-infrared pedestrian re-recognition model to obtain a dynamic contrast loss function; calculating the text contrast loss of the visible light image, calculating the text contrast loss of the infrared light image, and constructing a contrast loss function from the image to the text according to the text contrast loss of the visible light image and the text contrast loss of the infrared light image; training a visible-infrared pedestrian re-recognition model by using the dynamic comparison loss function and the image-to-text comparison loss function; according to the method, the problem that the existing visible-infrared pedestrian re-recognition training cannot fully capture the semantic relationship between samples and is easily influenced by noise is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of re-identification model processing, and in particular to a method and system for training a re-identification model based on a multi-modal collaborative learning framework. Background Art

[0002] Visible-infrared person re-identification is a cross-modal person re-identification task aimed at matching images of the same person taken under visible light and infrared conditions. Its supervised methods require a large amount of labeled data for training, which brings high costs and practical challenges. The advantage of unsupervised visible-infrared person re-identification is that it can learn modality-specific information and modality-invariant information without cross-modal annotation. Therefore, many scholars have conducted extensive research and exploration in this field. However, these methods only use one-hot encoded vectors as supervision signals and do not consider the potential of text information as a supervision signal, thus limiting their effectiveness. Large-scale pre-trained vision-language models, such as CLIP, have shown excellent transfer capabilities and can perform downstream tasks through text-image descriptions rather than relying solely on discrete labels. However, current methods only generate text descriptions for a single modality and do not consider the differences between different modalities in cross-modal learning tasks. It can be seen that the existing visible-infrared person re-identification training has problems of being unable to fully capture the semantic relationships between samples and being easily affected by noise. Summary of the Invention

[0003] The present invention provides a method and system for training a re-identification model based on a multi-modal collaborative learning framework to solve the problems that the existing visible-infrared person re-identification training cannot fully capture the semantic relationships between samples and is easily affected by noise.

[0004] To achieve the above object, the present invention is realized by the following technical solutions: In a first aspect, the present invention provides a method for training a re-identification model based on a multi-modal collaborative learning framework, including: Obtaining visible light image features of a target visible light image and infrared light image features of a target infrared light image, and respectively clustering the visible light image features and the infrared light image features to generate visible light image pseudo-labels and infrared light image pseudo-labels; Constructing a pre-trained CLIP model, generating visible light prompts and infrared light prompts for the target visible light image and the target infrared light image respectively, and generating visible light text semantic information and infrared light text semantic information respectively based on the pre-trained CLIP model; Fuse the visible light text semantic information and the infrared light text semantic information to obtain fused text semantic information, and perform weighted calculation on the fused text semantic information to obtain visible light weighted text semantic information and infrared light weighted text semantic information; Use the loss function of the visible-infrared person re-identification model to train the visible light prompt word and the infrared light prompt word respectively to obtain the trained visible light prompt word and the trained infrared light prompt word, and perform dynamic contrast learning based on the samples input to the visible-infrared person re-identification model to obtain a dynamic contrast loss function; Use the trained visible light prompt word and the trained infrared light prompt word to generate accurate visible light text semantic information and accurate infrared light text semantic information respectively, and generate optimal visible light text semantic information and optimal infrared light text semantic information based on the accurate visible light text semantic information and the accurate infrared light text semantic information; Calculate the visible light image-text contrast loss based on the optimal visible light text semantic information, calculate the infrared light image-text contrast loss based on the optimal infrared light text semantic information, construct an image-to-text contrast loss function according to the visible light image-text contrast loss and the infrared light image-text contrast loss, and use the dynamic contrast loss function and the image-to-text contrast loss function to train the visible-infrared person re-identification model.

[0005] Optionally, obtaining the visible light image feature of the target visible light image and the infrared light image feature of the target infrared light image includes: Input the target visible light image and the target infrared light image into a visual encoder to extract the visible light image feature and the infrared light image feature respectively; wherein, the target visible light image and the target infrared light image are both located in the training sets of the SYSU-MM01 and RegDB datasets.

[0006] Optionally, clustering the visible light image feature and the infrared light image feature respectively to generate a visible light image pseudo-label and an infrared light image pseudo-label, including: Use the DBSCAN algorithm to cluster the visible light image feature and the infrared light image feature respectively, and generate a visible light image pseudo-label for the target visible light image based on the clustering result, and generate an infrared light image pseudo-label for the target infrared light image based on the clustering result, wherein the visible light image pseudo-label and the infrared light image pseudo-label satisfy the following relational expression: ; ; In the formula, represents the pseudo-label of the visible light image, represents the pseudo-label of the infrared light image, represents the th clustering label of the visible light image, represents the th clustering label of the infrared image.

[0007] Optionally, generate a visible light prompt and an infrared light prompt for the target visible light image and the target infrared light image respectively, including: Define the visible light image text prompt according to the clustering label in the pseudo-label of the visible light image, and use the learnable prompt in the visible light image text prompt as the visible light prompt; Define the infrared light image text prompt according to the clustering label in the pseudo-label of the infrared light image, and use the learnable prompt in the infrared light image text prompt as the infrared light prompt; Among them, the visible light image text prompt satisfies the following: A visible photo of a person; The infrared light image text prompt satisfies the following: An infrared photo of a person; In the text, represents the randomly initialized visible light prompt of the th clustering, represents the randomly initialized infrared light prompt of the th clustering, represents the number of learnable prompts; Generate visible light text semantic information and infrared light text semantic information respectively based on the pre-trained CLIP model, including: Input the visible light image text prompt into the text encoder of the pre-trained CLIP model to obtain visible light text semantic information ; Input the infrared light image text prompt into the text encoder of the pre-trained CLIP model to obtain infrared light text semantic information .

[0008] Optionally, fuse the visible light text semantic information and the infrared light text semantic information to obtain fused text semantic information, including: Using the modal perception semantic fusion module, the optical text semantic information can be used as a query, the infrared optical text semantic information can be used as keys and values, and the visible light text semantic information and the infrared optical text semantic information are fused to obtain fused text semantic information. The fusion process satisfies the following relational expression: , ; ; ; In the formula, represents the query, represents the key, represents the value, , , , respectively represent four fully connected layers, represents the optical text semantic information, represents the infrared optical text semantic information, represents the attention, represents the activation function, represents the dimension of the text features, represents the fused text semantic information after fusion; Performing weighted calculation on the fused text semantic information to obtain visible light weighted text semantic information and infrared light weighted text semantic information, including: Based on the weighted of the modality, the visible light features and the infrared light features are further refined, and the weighted calculation is performed on the fused text semantic information to obtain the optical weighted text semantic information and the infrared light weighted text semantic information. The weighted calculation satisfies the following relational expression: ; ; In the formula, represents the visible light weighted text semantic information, represents the infrared light weighted text semantic information, represents the hyperparameter of the weighted calculation.

[0009] Optionally, using the loss function of the visible-infrared pedestrian re-identification model to train the visible light prompt word and the infrared light prompt word respectively to obtain the trained visible light prompt word and the trained infrared light prompt word, including: Putting the visible light weighted text semantic information and the into the loss function of the visible-infrared pedestrian re-identification model to train the visible light prompt word and the infrared light prompt word respectively, and obtaining the trained visible light prompt word and the trained infrared light prompt word. The training satisfies the following relational expression:

[0010] ;

[0011] ; In the formula, represents the trained visible light prompt, represents the trained infrared light prompt, represents the number of a batch, represents the visible light feature, represents the visible light weighted text semantic information, represents the infrared light feature, represents the infrared light weighted text semantic information, represents the hyperparameter, represents the set composed of samples with identity and represents the cosine similarity between two vectors.

[0012] Optionally, a dynamic contrast loss function is obtained through dynamic contrast learning based on the samples of the input visible-infrared pedestrian re-identification model, including: For each sample input to the model, use the visual encoder to extract the corresponding visible light image feature of the sample and the infrared light image feature , and respectively determine the top nearest neighbor features of the visible light image feature and the top nearest neighbor features of the infrared light image feature ; Calculate the mean values of the top nearest neighbor features of the visible light image feature and the top nearest neighbor features of the infrared light image feature respectively, and obtain the visible light target embedding and the infrared light target embedding respectively. The mean value calculation satisfies the following relational expression: ; ; In the formula, represents the visible light target embedding, represents the infrared light target embedding, represents the top nearest neighbor features of the visible light image feature, represents the top The nearest neighbor feature, represents the visible light image feature, and represents the infrared light image feature; Introduce the Gaussian kernel function to embed the visible light target and the distance between the visible light image feature and adjust the distance between the infrared light target embedding and the infrared light image feature . The Gaussian kernel function satisfies the following relational expression: ; ; In the formula, , respectively represent the Gaussian kernel function corresponding to visible light and the Gaussian kernel function corresponding to infrared light, and respectively represent the Euclidean distance between the visible light target embedding and the visible light image feature and the Euclidean distance between the infrared light target embedding and the infrared light image feature , represents the bandwidth controlling the Gaussian kernel; Introduce the attraction term function to reduce the Euclidean distance in similar samples. The attraction term function satisfies the following relational expression: ; ; In the formula, , respectively represent the attraction term loss function of the visible light image and the attraction term loss function of the infrared light image, represents the number of a batch; Introduce the repulsion term function to increase the Euclidean distance in different samples. The repulsion term function satisfies the following relational expression: ; ; In the formula, , respectively represent the repulsion term loss function of the visible light image and the repulsion term loss function of the infrared light image, represents the number of a batch, represents the predefined threshold; Obtain the dynamic contrast loss function according to the repulsion term function and the attraction term function. The dynamic contrast loss function satisfies the following relational expression: + + ; In the formula, represents the dynamic contrast loss function.

[0013] Optionally, generating accurate visible light text semantic information and accurate infrared light text semantic information by using the trained visible light prompt and the trained infrared light prompt respectively, including: Inputting the trained visible light prompt into the text encoder of the pre-trained CLIP model to obtain accurate visible light text semantic information ; Inputting the trained infrared light prompt into the text encoder of the pre-trained CLIP model to obtain accurate infrared light text semantic information ; Generating optimal visible light text semantic information and optimal infrared light text semantic information based on the accurate visible light text semantic information and the accurate infrared light text semantic information, including: Using the modality-aware semantic fusion module to take the accurate visible light text semantic information as the query, and the accurate infrared light text semantic information as the key and value, and fusing the accurate visible light text semantic information and the accurate infrared light text semantic information to obtain the fused text semantic information; Performing weighted calculation on the fused text semantic information to obtain the optimal visible light text semantic information and the optimal visible light text semantic information; Calculating the visible light image-text contrast loss based on the optimal visible light text semantic information, calculating the infrared light image-text contrast loss based on the optimal infrared light text semantic information, and constructing an image-to-text contrast loss function according to the visible light image-text contrast loss and the infrared light image-text contrast loss, including: Calculating the visible light image-text contrast loss based on the optimal visible light text semantic information, and its calculation satisfies the following relational expression: ; In the formula, represents the visible light image-to-text contrast loss function, represents the number of a batch, represents the one-hot encoded vector of the identity label, represents the visible light feature, represents the optimal visible light text semantic information, represents the identity label; Calculating the infrared light image-text contrast loss based on the optimal infrared light text semantic information, and its calculation satisfies the following relational expression: ; In the formula, represents the infrared light image-to-text contrast loss function, Represents the quantity of a batch, Indicates the one-hot encoded vector of the identity label, Indicates the visible light feature, Indicates the optimal visible light text semantic information, Indicates the identity label; Construct an image-to-text contrast loss function based on the visible light image text contrast loss and the infrared light image text contrast loss, and its construction satisfies the following relational expression: ; In the formula, Indicates the image-to-text contrast loss function.

[0014] Optionally, the method further includes: Extract the features of the query image and the gallery image in the test set, and input them into the trained visible-infrared person re-identification model to evaluate the performance of the trained visible-infrared person re-identification model. Among them, the evaluation method includes: According to the output result of the trained visible-infrared person re-identification model, use the cumulative matching curve, average query precision, and average inverse negative sample penalty rate to evaluate the performance of the model.

[0015] In a second aspect, an embodiment of the present application provides a re-identification model training system based on a multi-modal collaborative learning framework, including a processor and a memory; The memory is used to store computer programs; The processor, when executing the program stored on the memory, implements any of the method steps in the first aspect.

[0016] Beneficial effects: The re-identification model training method based on the multi-modal collaborative learning framework provided by the present invention captures high-order semantic information as the supervision signal for the unsupervised VI-ReID task by integrating text features from visible light and infrared language descriptions, effectively solving the limitations of unsupervised visible-infrared person re-identification that only relies on a single pseudo-label as the supervision model and a single-modal generated text description.

[0017] At the same time, by dynamically measuring the distance between the sample and its neighboring centroids, the same-modal learning is enhanced, effectively improving the reliability of clustering and the distinguishability of the same-modal features, and achieving the current optimal effect on various data sets, effectively solving the influence of noise on the model performance, and improving the robustness and stability of the model. Description of the drawings

[0018] Figure 1 Is the flowchart of the re-identification model training based on the multi-modal collaborative learning framework of the embodiment of the present invention; Figure 2 Schematic diagram of the training of the re-identification model based on the multi-modal collaborative learning framework according to the embodiment of the present invention; Figure 3 Supplementary diagram of the schematic diagram of the training of the re-identification model based on the multi-modal collaborative learning framework according to the embodiment of the present invention. In the figure, (a) is part a of the schematic diagram of the training of the re-identification model based on the multi-modal collaborative learning framework, and (b) is part b of the schematic diagram of the training of the re-identification model based on the multi-modal collaborative learning framework; Figure 4 ARI index measurement result diagram of the dynamic contrast learning module according to the embodiment of the present invention; Figure 5 t-SNE visualization schematic diagram according to the embodiment of the present invention. Detailed implementation manners

[0019] The technical solutions of the present invention will be described clearly and completely below. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without any creative efforts shall fall within the protection scope of the present invention.

[0020] Unless otherwise defined, the technical terms or scientific terms used in the present invention shall have the ordinary meanings understood by those of ordinary skill in the art to which the present invention belongs. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, the terms such as "a" or "one" do not denote a quantity limitation, but mean that there is at least one. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left" and "right" are only used to represent relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship also changes accordingly.

[0021] As Figure 1 shown, the embodiment of the present application provides a method for training a re-identification model based on a multi-modal collaborative learning framework, including: Obtaining the visible light image features of the target visible light image and the infrared light image features of the target infrared light image, and respectively clustering the visible light image features and the infrared light image features to generate visible light image pseudo-labels and infrared light image pseudo-labels; Build a pre-trained CLIP model to generate visible light prompts and infrared light prompts for the target visible light image and the target infrared light image respectively, and generate visible light text semantic information and infrared light text semantic information based on the pre-trained CLIP model respectively; Fuse the visible light text semantic information and the infrared light text semantic information to obtain fused text semantic information, and perform weighted calculation on the fused text semantic information to obtain visible light weighted text semantic information and infrared light weighted text semantic information; Use the loss function of the visible-infrared person re-identification model to train the visible light prompt and the infrared light prompt respectively to obtain the trained visible light prompt and the trained infrared light prompt, and perform dynamic contrast learning based on the samples input to the visible-infrared person re-identification model to obtain a dynamic contrast loss function; Use the trained visible light prompt and the trained infrared light prompt to generate accurate visible light text semantic information and accurate infrared light text semantic information respectively, and generate optimal visible light text semantic information and optimal infrared light text semantic information based on the accurate visible light text semantic information and the accurate infrared light text semantic information; Calculate the visible light image-text contrast loss based on the optimal visible light text semantic information, calculate the infrared light image-text contrast loss based on the optimal infrared light text semantic information, construct an image-to-text contrast loss function according to the visible light image-text contrast loss and the infrared light image-text contrast loss, and use the dynamic contrast loss function and the image-to-text contrast loss function to train the visible-infrared person re-identification model.

[0022] In the above embodiment, the re-identification model training method based on the multi-modal collaborative learning framework can be mainly divided into a training stage and an application stage, as Figures 2 - 3 shown, and the specific steps of each stage are as follows: Training stage: S1: Input the visible light and infrared light images into the model, and extract image features through the visual encoder; All visible light images in the training set and infrared images are input into the visual encoder to extract visible light features and infrared features , the light and infrared light can be in the training sets of SYSU-MM01 and RegDB datasets; the SYSU-MM01 dataset contains images from six cameras, four of which are visible light cameras and two are infrared cameras. The dataset is divided into a training set and a test set. The training set contains 395 identities, the test set contains 96 identities, and there are a total of 30,071 visible light images and 22,258 infrared images; the RegDB dataset is a small-scale visible light-infrared pedestrian re-identification benchmark, containing 8,240 images of 412 identities, with each identity consisting of 10 visible light images and 10 infrared images. All images are taken by a single camera in each modality. According to the standard evaluation protocol, 2,060 visible light images and 2,060 infrared images are randomly selected from 206 identities as the training set, and the remaining images are used for testing.

[0023] S2: Cluster the extracted image features using the DBSCAN algorithm to generate pseudo-labels for visible light and infrared light images respectively; denoted as and . Among them, represents the th clustering label of the visible light image, and represents the th clustering label of the infrared image; In S2, the DBSCAN algorithm is used to assign pseudo-labels to all images, denoted as and . Among them, represents the th clustering label of the visible light image, and represents the th clustering label of the infrared image.

[0024] S3: Input these pseudo-labels into the pre-trained CLIP model, introduce the prompt learning strategy, first generate learnable prompt words for visible light and infrared light images, and then generate corresponding text semantic information according to the clustering pseudo-labels of visible light and infrared modality data respectively; S3.1: According to each clustering or obtained by DBSCAN in S2, define the text prompt words as "A visible photo of a person." and "An infrared photo of a person.".

[0025] Here, and both belong to the Randomly initialized learnable prompt words for each cluster. For ease of distinction, represents the randomly initialized visible light prompt words for the th cluster, and represents the randomly initialized infrared light prompt words for the th cluster. And under other specific training conditions, and can be shared, while represents the number of learnable prompt words; S3.2: According to the text prompt word template obtained in S3.1, input it into the text encoder of the pre-trained CLIP model to obtain unique text class semantic representations for each cluster, respectively represented as and , where represents the visible light modality, and represents the infrared modality.

[0026] S4: Fuse the text semantic information generated by CLIP. Through the modality-aware semantic fusion module, use the text semantic information of the visible light image as the query, and the text semantic information of the infrared light image as the key and value to calculate the fusion matching score. Output the fused higher-order text semantic information through the fully connected layer, and then weight the visible light and infrared light text semantic information respectively to obtain the fused visible light and infrared light text semantic information; S4.1: According to and obtained in S4, use the modality-aware semantic fusion module to fuse them into a comprehensive class semantic prompt. Use as the query, and at the same time use as the key and value. The text semantic fusion process is as follows: ,

[0027]

[0028]

[0029] where is the dimension of the text feature, and corresponds to four fully connected layers. represents the fused text semantic information; S4.2: Based on the fused text semantic information obtained in S4.1, further refine the visible light and infrared features based on modality weighting to form modality-aware semantic representations, respectively represented as and .

[0030]

[0031]

[0032] Among them, is a weighted hyperparameter.

[0033] S5: Train the learnable prompt words through the loss function so that the prompt words can accurately and efficiently represent the pseudo-labels of visible light images and infrared light images; Train the learnable prompt words through the loss function, and the formula is and :

[0034]

[0035]

[0036]

[0037] Among them, is the number of a batch, is a hyperparameter, represents the set composed of samples with identity and represents the cosine similarity between two vectors.

[0038] S6: After the prompt words are trained, train the overall model. For each sample of the input image, extract the features of each image through the visual encoder, perform dynamic contrast learning, find the k-nearest neighbors of each sample through the Jaccard distance, and then calculate through the contrast loss function to enhance the consistency of the features within each modality, making similar features closer together and pushing dissimilar features further apart; S6.1: For each sample of the input image, extract the features of each image through the visual encoder, and calculate the features based on the Jaccard similarity and of the first nearest neighbors and ; S6.2: According to the and obtained in S6.1, by calculating the mean of the selected first nearest neighbors, obtain the target embeddings and , and the calculation formula is:

[0039]

[0040] S6.3: Introduce a Gaussian kernel function to adjust the target embedding and the corresponding source features and The distance between them is calculated as follows:

[0041]

[0042] where and are and and also and The Euclidean distance between. The parameter controls the bandwidth of the Gaussian kernel; S6.4: According to S6.3, if the source features and target features share a large amount of collaborative information, the distance between them should be minimized to ensure that the source samples are more closely aligned with their neighboring samples, thereby enhancing collaborative consistency. To this end, an attractive term is introduced to reduce the distance between similar sample pairs:

[0043]

[0044] S6.5: According to S6.4, when the collaborative information shared by the source features and target features is minimal, it is beneficial for them to maintain a large distance in the feature space. A repulsive term is added to increase the distance between different sample pairs, especially applying the repulsive force to negative sample pairs:

[0045]

[0046] where represents the number of batches, represents a predefined threshold, indicating the minimum distance that needs to be maintained between different embeddings; S6.6: According to S6.4 and S6.5, the final dynamic contrast loss function is obtained for the overall model training, and the formula is: + +

[0047] In the formula, Represents the dynamic contrast loss function.

[0048] S7: Use the trained prompt words to generate the most accurate visible light text semantic information and infrared light image text semantic information according to the pseudo-labels, and perform modal-aware semantic fusion on these text semantic information again to obtain the fused visible light and infrared light text semantic information for training the overall model; S7.1: According to S3, the trained prompt words are regenerated to obtain the text semantic information features of visible light and infrared light as and ; S7.2: According to S4, the obtained and are fused to obtain the optimal text semantic information as and ; S7.3: According to S7.2, use the fused text information to calculate the image-to-text contrast loss to further improve the alignment between modalities and train the overall model:

[0049]

[0050]

[0051] Here, represents the one-hot encoded vector of the identity label of.

[0052] The present invention uses a modified ResNet-50 pre-trained by CLIP as the feature extraction network, sets the size of the input image to 288 × 144, and applies data augmentation techniques such as random flipping, random grayscaling, and random erasing to enhance the diversity of training data. The training process is divided into two different optimization stages. Stage 1: Modal-aware semantic fusion: This stage lasts for 120 epochs and focuses on optimizing the prompt word modality. Stage 2: ReID model optimization: In this stage, the ReID model is trained for 120 epochs, using the high-order semantic supervision signals generated in Stage 1 and combining with the dynamic contrast learning module for training. In all stages, the Adam optimizer is used, and the cosine learning rate scheduler is adopted, with an initial learning rate of 3.5e-4. The number of samples in a batch is set to 128. According to empirical observations, the k value of the RegDB dataset is set to 5, and the k value of the SYSU-MM01 dataset is set to 20. All methods are implemented using PyTorch on a single NVIDIA A800 80GB GPU and are carried out using the pytorch framework; Application stage S8: Extract the features of the query images and gallery images in the test set and input them into the trained model. Evaluate the performance of the unsupervised visible-infrared person re-identification model according to the output results for the unsupervised visible-infrared person re-identification test set, and use the cumulative match curve (CMC), mean average precision (mAP), and mean inverse negative penalty rate (mINP) to evaluate the performance of the model.

[0053] In the application stage on the SYSU-MM01 dataset, after adding the dynamic contrast learning module and the modality-aware semantic fusion module, the R1, mAP, and mINP metrics have all been greatly improved. The modality-aware semantic fusion module integrates complementary text features from the visible light and infrared modalities, strengthens the cross-modal supervision signal, provides a strong co-training signal, and significantly improves the recognition performance. The dynamic contrast learning reduces label noise and improves the alignment of intra-modal features, focusing on improving robustness during the feature learning process, while enhancing semantic alignment, making them jointly effective in unsupervised visible-infrared person re-identification; In the RegDB dataset, after adding the dynamic contrast learning module and the modality-aware semantic fusion module, the R1, mAP, and mINP metrics have all been greatly improved. The modality-aware semantic fusion module integrates complementary text features from the visible light and infrared modalities, strengthens the cross-modal supervision signal, provides a strong co-training signal, and significantly improves the recognition performance. The dynamic contrast learning reduces label noise and improves the alignment of intra-modal features, focusing on improving robustness during the feature learning process, while enhancing semantic alignment, making them jointly effective in unsupervised visible-infrared person re-identification.

[0054] As Figure 4 shown, to evaluate the effectiveness of the proposed dynamic contrast learning module, this paper conducts a detailed pseudo-label evaluation experiment on the SYSU-MM01 dataset, using the adjusted Rand index (ARI) as the evaluation metric to quantify the pseudo-label accuracy of the visible light and infrared modalities. ARI measures the similarity between the clustering results and the true labels, and a higher score indicates more accurate clustering and higher-quality pseudo-labels. In this experiment, the baseline model only optimizes the loss function, while the dynamic contrast learning module is added as an enhancement method. After adding the dynamic contrast module, the ARI score is significantly improved: it increases by 6% for the visible light modality and 16% for the infrared modality. This proves that the dynamic contrast learning greatly optimizes the clustering process by reducing label noise and improving the alignment of intra-modal features. The enhanced clustering accuracy by dynamic contrast learning promotes more stable pseudo-label generation, thus improving the overall stability and recognition performance of the model.

[0055] As Figure 5 shown, this paper uses t-SNE to visualize the 2D feature distributions of visible and infrared modalities of 10 randomly selected identities. Compared with the baseline method, the method in this paper produces significantly more compact feature clusters within a single modality. In addition, the feature distributions of the same identity in different modalities are also closer.

[0056] The embodiment of the present disclosure also provides a re-identification model training system based on a multi-modal collaborative learning framework, including a processor and a memory; The memory is used to store computer programs; The processor is used to implement any of the method steps in the re-identification model training method based on the multi-modal collaborative learning framework when executing the program stored on the memory.

[0057] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art shall fall within the protection scope determined by the claims.

Claims

1. A method for training a re-identification model based on a multi-modal collaborative learning framework, characterized in that, Including: Obtain the visible light image features of the target visible light image and the infrared light image features of the target infrared light image, and respectively cluster the visible light image features and the infrared light image features to generate visible light image pseudo-labels and infrared light image pseudo-labels; Construct a pre-trained CLIP model, generate visible light prompt words and infrared light prompt words for the target visible light image and the target infrared light image respectively, and generate visible light text semantic information and infrared light text semantic information based on the pre-trained CLIP model respectively; Fuse the visible light text semantic information and the infrared light text semantic information to obtain fused text semantic information, and perform weighted calculation on the fused text semantic information to obtain visible light weighted text semantic information and infrared light weighted text semantic information; Use the loss function of the visible-infrared person re-identification model to train the visible light prompt words and the infrared light prompt words respectively to obtain trained visible light prompt words and trained infrared light prompt words, and perform dynamic contrast learning based on the samples input into the visible-infrared person re-identification model to obtain a dynamic contrast loss function; Use the trained visible light prompt words and the trained infrared light prompt words to generate accurate visible light text semantic information and accurate infrared light text semantic information respectively, and generate optimal visible light text semantic information and optimal infrared light text semantic information based on the accurate visible light text semantic information and the accurate infrared light text semantic information; Calculate the visible light image text contrast loss based on the optimal visible light text semantic information, calculate the infrared light image text contrast loss based on the optimal infrared light text semantic information, construct an image-to-text contrast loss function according to the visible light image text contrast loss and the infrared light image text contrast loss, and use the dynamic contrast loss function and the image-to-text contrast loss function to train the visible-infrared person re-identification model.

2. The method for training a re-identification model based on a multi-modal collaborative learning framework according to claim 1, wherein Obtain the visible light image features of the target visible light image and the infrared light image features of the target infrared light image, including: Input the target visible light image and the target infrared light image into the visual encoder to respectively extract the visible light image features and the infrared light image features ; Among them, the target visible light image and the target infrared light image are both located in the training sets of the SYSU-MM01 and RegDB datasets.

3. The re-identification model training method based on the multi-modal collaborative learning framework according to claim 1, characterized in that Respectively cluster the visible light image features and the infrared light image features to generate visible light image pseudo-labels and infrared light image pseudo-labels, including: Use the DBSCAN algorithm to cluster the visible light image features and the infrared light image features respectively, and generate visible light image pseudo-labels for the target visible light image based on the clustering results, and generate infrared light image pseudo-labels for the target infrared light image based on the clustering results, where the visible light image pseudo-labels and the infrared light image pseudo-labels satisfy the following relational formula: ; ; In the formula, represents the pseudo-label of the visible light image, represents the pseudo-label of the infrared light image, represents the th clustering label of the visible light image, represents the th clustering label of the infrared image.

4. The re-identification model training method based on the multi-modal collaborative learning framework according to claim 1, wherein Generate visible light prompt words and infrared light prompt words for the target visible light image and the target infrared light image respectively, including: Define visible light image text prompt words according to the clustering labels in the visible light image pseudo-labels, and use the learnable prompt words in the visible light image text prompt words as visible light prompt words; Define infrared light image text prompt words according to the clustering labels in the infrared light image pseudo-labels, and use the learnable prompt words in the infrared light image text prompt words as infrared light prompt words; Wherein, the visible light image text prompt words satisfy the following: A visible photo of a person; The infrared light image text prompt satisfies the following: An infrared photo of a person; In the text, represents the randomly initialized visible light prompt words for the th cluster, represents the randomly initialized infrared light prompt words for the th cluster, represents the number of learnable prompt words; Based on the pre-trained CLIP model, visible light text semantic information and infrared light text semantic information are generated respectively, including: Input the visible light image text prompt into the text encoder of the pre-trained CLIP model to obtain visible light text semantic information ; Input the infrared light image text prompt into the text encoder of the pre-trained CLIP model to obtain infrared light text semantic information 。 5. The method for training a re-identification model based on a multi-modal collaborative learning framework according to claim 1, wherein Fusing the visible light text semantic information and the infrared light text semantic information to obtain fused text semantic information, including: Using the modality-aware semantic fusion module, taking the visible light text semantic information as the query, and the infrared light text semantic information as the key and value, to fuse the visible light text semantic information and the infrared light text semantic information to obtain fused text semantic information, and its fusion process satisfies the following relational expression: , ; ; ; In the formula, represents a query, represents a key, represents a value, , , , respectively represent four fully connected layers, represents the optical text semantic information that can be represents the infrared optical text semantic information, represents attention, represents an activation function, represents the dimension of the text feature, represents the fused text semantic information after fusion; Performing weighted calculation on the fused text semantic information to obtain visible light weighted text semantic information and infrared light weighted text semantic information, including: Based on modality-based weighting, further refining the visible light features and infrared light features, performing weighted calculation on the fused text semantic information to obtain visible light weighted text semantic information and infrared light weighted text semantic information, and its weighted calculation satisfies the following relational expression: ; ; In the formula, represents the visible light weighted text semantic information, represents the infrared light weighted text semantic information, represents the hyperparameter of the weighted calculation.

6. The method for training a re-identification model based on a multi-modal collaborative learning framework according to claim 1, wherein Using the loss function of the visible-infrared person re-identification model to train the visible light prompt and the infrared light prompt respectively to obtain the trained visible light prompt and the trained infrared light prompt, including: Putting the visible light weighted text semantic information and the infrared light weighted text semantic information into the loss function of the visible-infrared person re-identification model to train the visible light prompt and the infrared light prompt respectively, and obtaining the trained visible light prompt and the trained infrared light prompt, and its training satisfies the following relational expression: ; ; Wherein, represents the trained visible light prompt word, represents the trained infrared light prompt word, represents the number of a batch, represents the visible light feature, represents the visible light weighted text semantic information, represents the infrared light feature, represents the infrared light weighted text semantic information, represents the hyperparameter, represents the set composed of samples with identity and represents the cosine similarity between two vectors.

7. The re-identification model training method based on the multi-modal collaborative learning framework according to claim 1, characterized in that Performing dynamic contrast learning based on the samples input into the visible-infrared person re-identification model to obtain a dynamic contrast loss function, including: For each sample of the input model, use a visual encoder to extract the visible light image features corresponding to the sample and infrared light image features , and respectively determine the top nearest neighbor features of the visible light image features and the top nearest neighbor features of the infrared light image features ; The first few nearest neighbor features of the visible light image features and the first few nearest neighbor features of the infrared light image features are respectively subjected to mean calculation, and a visible light target embedding and an infrared light target embedding are respectively obtained. The mean calculation satisfies the following relational expression: ; ; Wherein, represents the visible light target embedding, represents the infrared light target embedding, represents the first nearest neighbor features of the visible light image features, represents the first nearest neighbor features of the infrared light image features, represents the visible light image features, represents the infrared light image features; Introduce the Gaussian kernel function to embed the visible light target and the visible light image features the distance between them, and the embedding of the infrared light target and the infrared light image features Adjust the distance between them. The Gaussian kernel function satisfies the following relationship: ; ; In the formula, and represent the Gaussian kernel function corresponding to visible light and the Gaussian kernel function corresponding to infrared light respectively, and represent the Euclidean distance between the visible light target embedding and the visible light image feature and the Euclidean distance between the infrared light target embedding and the infrared light image feature respectively, represents the bandwidth controlling the Gaussian kernel; Introducing an attraction term function to reduce the Euclidean distance in similar samples, and its attraction term function satisfies the following relational expression: ; ; In the formula, and respectively represent the attraction term loss function of the visible light image and the attraction term loss function of the infrared light image, represents the number of a batch; Introducing a repulsion term function to increase the Euclidean distance in different samples, and its repulsion term function satisfies the following relational expression: ; ; In the formula, and respectively represent the rejection term loss function of the visible light image and the rejection term loss function of the infrared light image, represents the number of a batch, represents a predefined threshold; Obtaining a dynamic contrast loss function according to the repulsion term function and the attraction term function, and its dynamic contrast loss function satisfies the following relational expression: + + ; In the formula, represents the dynamic contrast loss function.

8. The method for training a re-identification model based on a multi-modal collaborative learning framework according to claim 1, wherein Using the trained visible light prompt and the trained infrared light prompt to generate accurate visible light text semantic information and accurate infrared light text semantic information respectively, including: Input the trained visible light prompt into the text encoder of the pre-trained CLIP model to obtain accurate visible light text semantic information ; Input the trained infrared light prompt into the text encoder of the pre-trained CLIP model to obtain accurate infrared light text semantic information ; Generating optimal visible light text semantic information and optimal infrared light text semantic information based on the accurate visible light text semantic information and the accurate infrared light text semantic information, including: Using the modality-aware semantic fusion module, taking the accurate visible light text semantic information as the query, and the accurate infrared light text semantic information as the key and value, to fuse the accurate visible light text semantic information and the accurate infrared light text semantic information to obtain the fused text semantic information; Performing weighted calculation on the fused text semantic information to obtain optimal visible light text semantic information and optimal visible light text semantic information; Calculating the visible light image text contrast loss based on the optimal visible light text semantic information, calculating the infrared light image text contrast loss based on the optimal infrared light text semantic information, and constructing an image-to-text contrast loss function according to the visible light image text contrast loss and the infrared light image text contrast loss, including: Calculate the visible light image-text contrast loss based on the optimal visible light text semantic information, and its calculation satisfies the following relational expression: ; In the formula, represents the contrastive loss function from visible light images to text, represents the number of a batch, represents the one-hot encoded vector of the identity label, represents the visible light feature, represents the optimal visible light text semantic information, represents the identity label; Calculate the infrared light image-text contrast loss based on the optimal infrared light text semantic information, and its calculation satisfies the following relational expression: ; Wherein, represents the contrast loss function from the infrared light image to the text, represents the number of a batch, represents the one-hot encoded vector of the identity label, represents the visible light feature, represents the optimal visible light text semantic information, represents the identity label; Construct an image-to-text contrast loss function according to the visible light image-text contrast loss and the infrared light image-text contrast loss, and its construction satisfies the following relational expression: ; In the formula, represents the contrastive loss function from image to text.

9. The method for training a re-identification model based on a multi-modal collaborative learning framework according to claim 1, wherein The method further includes: Extract the features of the query images and gallery images in the test set, and input them into the trained visible-infrared pedestrian re-identification model to evaluate the performance of the trained visible-infrared pedestrian re-identification model. Among them, the evaluation method includes: According to the output results of the trained visible-infrared pedestrian re-identification model, use the cumulative matching curve, mean average precision, and mean reciprocal rank to evaluate the performance of the model.

10. A re-identification model training system based on a multi-modal collaborative learning framework, characterized in that, It includes a processor and a memory; The memory is used to store computer programs; When the processor is used to execute the programs stored on the memory, it implements the method steps described in any one of claims 1-9.

Citation Information

Cited By

  • Infrared image unknown target detection method and device based on feature migration and high-order semantics

    CN121236367A

  • Infrared image unknown target detection method and device based on feature migration and high-order semantics

    CN121236367B