Re-recognition model training method and system based on noise robust prompt learning framework
By adopting a noise-based robust prompt learning framework method in the pedestrian re-identification model, the CLIP model and a learnable mapping network are used to generate pseudo-language prompt words, combined with text embedding and cross-attention module, the problem of model performance degradation under noise labels is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510541991.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing pedestrian re-identification model has degraded performance under noise labels and failed to fully utilize the semantic information in the training data, resulting in insufficient noise robustness in complex scenarios.
Using a method based on the noise-robust prompt learning framework, the image features are extracted using pre-trained CLIP model, pseudo-language prompt words are generated through a learnable mapping network, and global and context text prompts are constructed in combination with predefined basic prompt templates, text embedding is integrated and language-guided cross-attention modules are introduced, and model parameters are jointly optimized to improve robustness to noise labels.
It effectively improves the recognition accuracy and robustness of the pedestrian re-identification model in a noise environment, can better capture semantic nuances in the data, and improves the accuracy and noise immunity of the model.
Smart Images

Figure CN120148072A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of re-identification model processing, and particularly to a method and system for training a re-identification model based on a noise-robust prompt learning framework. Background Art
[0002] In video surveillance systems, pedestrian re-identification is a key task aimed at matching the same object across different non-overlapping camera views. In recent years, deep learning-based pedestrian re-identification methods have made significant progress, but these methods usually rely on a large amount of accurately labeled data. However, obtaining high-quality manually labeled data is both expensive and time-consuming, and is prone to introducing noisy labels, which can seriously affect the recognition performance of the model. Most existing noise-robust learning methods follow the paradigm of reliable sample selection and noisy label correction. Although this method can alleviate the noise problem to a certain extent, it fails to fully utilize the semantic information in the training data, resulting in insufficient noise robustness in complex scenarios. In addition, many existing methods ignore the inherent semantic nuances in the data and cannot effectively handle the challenges in fine-grained image re-identification, especially in the absence of specific descriptive text labels. Summary of the Invention
[0003] The present invention provides a method and system for training a re-identification model based on a noise-robust prompt learning framework, which solves the problem of the performance degradation of existing pedestrian re-identification models under noisy labels and improves the recognition accuracy and robustness in a noisy environment.
[0004] To achieve the above object, the present invention is realized through the following technical solutions: In a first aspect, the present invention provides a method for training a re-identification model based on a noise-robust prompt learning framework, including: Extracting the global visual features of the input pedestrian image by using the visual encoder of the pre-trained CLIP model, retrieving the k nearest neighbor global visual features of the global visual features of the target sample in the feature space, and generating context visual features based on cosine similarity weighted aggregation; Mapping the global visual features and context visual features to global pseudo-language prompt words and context pseudo-language prompt words through a learnable mapping network; Combining the predefined basic prompt template with the global pseudo-language prompt words and context pseudo-language prompt words to construct global text prompts and context text prompts, and generating global text embeddings and context text embeddings based on the text encoder of the pre-trained CLIP model; Fusing the global text embeddings and context text embeddings based on the max-pooling strategy to form comprehensive text embeddings; Introduce a language-guided cross-attention module, use the comprehensive text embedding as the query vector, and the patch embeddings output by the visual encoder as the key vector and value vector to optimize the multi-modal feature representation; use the generalized cross-entropy loss function and the symmetric contrast loss function to jointly optimize the model parameters to improve the robustness to noisy labels; Use the comprehensive text embedding as the category semantic vector, and align the visual feature distribution of the student re-identification model with the text feature distribution of the teacher model through the prompt imitation strategy; jointly train the student model based on the identity classification loss, triplet margin loss, and semantic knowledge distillation loss; Use the visual encoder of the trained re-identification model to extract the visual features of the query image and the gallery image in the test set, and calculate the cross-modal similarity in combination with the text features generated by the teacher model; sort the gallery samples according to the similarity score and output the target identity matching the query image to complete the pedestrian re-identification task.
[0005] Optionally, use the visual encoder of the pre-trained CLIP model to extract the global visual features of the input pedestrian image, including: Input all pedestrian images in the training set into the visual encoder to generate a series of global visual features ; where, represents the identity label, and the pedestrian images are in the training sets of the Market1501 and DukeMTMC-ReID datasets; Retrieve the k nearest neighbor global visual features of the global visual features of the target sample in the feature space, and generate the context visual features based on cosine similarity weighted aggregation, including: For the global visual features of the target sample , find its k nearest neighbor global visual features in the feature space, denoted as , and obtain the context visual features from the formula. The process satisfies the following relationship: ; In the formula, , is the cosine similarity metric between two samples, represents the context visual features.
[0006] Optionally, map the global visual features and context visual features to global pseudo-language prompt words and context pseudo-language prompt words through a learnable mapping network, including: Use a learnable mapping network Map the global visual feature and the context visual feature to pseudo - language prompt words and context pseudo - language prompt words. This process satisfies the following relational expressions: ; ; In the formula, is the global visual feature, is the generated global pseudo - language prompt word, is the context visual feature, is the context pseudo - language prompt word.
[0007] Optionally, combine the predefined basic prompt template with the global pseudo - language prompt word and the context pseudo - language prompt word to construct the global text prompt and the context text prompt, including: Construct the global text prompt and the context text prompt according to the predefined basic prompt template, the global pseudo - language prompt word and the context pseudo - language prompt word; The basic prompt template satisfies the following: A photo of a person; A photo of a person; Among them, represents the global pseudo - language prompt word, represents the context pseudo - language prompt word; According to the obtained global text prompt and context text prompt, input them into the text encoder of the pre - trained CLIP model to generate the global text embedding and the context text embedding. The generation process satisfies the following relational expressions: ; ; In the formula, is the global text prompt, is the global text embedding, is the context text prompt, is the context text embedding, is the pre - trained CLIP text encoder.
[0008] Optionally, based on the max - pooling strategy, fuse the global text embedding and the context text embedding to form a comprehensive text embedding, including: According to the obtained global text embedding and context text embedding , use the max - pooling strategy to fuse the two into a comprehensive text embedding , where the process of fusing the comprehensive text embedding satisfies the following relational expression: ; In the formula, is the global text embedding, is the context text embedding, is the comprehensive text embedding, is the max pooling operation.
[0009] Optionally, a language-guided cross-attention module is introduced. The comprehensive text embedding is used as the query vector, and the tile embeddings output by the visual encoder are used as the key vector and value vector to optimize the multi-modal feature representation. The generalized cross-entropy loss function and the symmetric contrast loss function are used to jointly optimize the model parameters to improve the robustness to noisy labels, including: According to the tile embedding and the comprehensive text embedding , using the language-guided cross-attention module, the comprehensive text embedding acts as the query vector, and the full tile embedding is used as the key vector and value vector. The text semantic fusion process satisfies the following relational expression: ; ; In the formula, is the dimension of the text feature, corresponds to three fully connected layers, represents the fused text semantic information; In the prompt optimization, the generalized cross-entropy loss is used to optimize the prompt, and the optimization process satisfies the following formula: ; In the formula, is the probability that the sample is classified into the category based on its final feature representation, = 0.7 to balance robustness and performance; In this training stage, we freeze both the image encoder and the text encoder. To ensure that the learned text semantic information can effectively capture the image context, we use a symmetric contrast loss to optimize the prompt. The formula is as follows: ; ; In the formula, the parameter plays a role in adjusting the softening degree of the similarity score in the contrast loss function; According to the generalized cross-entropy loss function and the symmetric contrast loss function, the total loss formula for prompt optimization is constructed to satisfy the following relational expression: ; In the formula, represents the generalized cross-entropy loss, and represents the symmetric contrast loss function.
[0010] Optionally, the comprehensive text embeddings are aggregated into category semantic vectors, and the visual feature distribution of the student re-identification model is aligned with the text feature distribution of the teacher model through a prompt imitation strategy; the student model is jointly trained based on the identity classification loss, triplet margin loss, and semantic knowledge distillation loss, including: The text embeddings are used as category vectors to provide semantic supervision signals. For each identity, the prompts of the same identity are aggregated to form a unified text embedding, and the aggregation process satisfies the following relationship: ; In the formula, represents the unified text embedding of a specific identity after aggregation, represents the comprehensive text embedding, represents the identity category, represents the identity label; The student model calculates the probability that a given image belongs to an identity category, and the calculation process satisfies the following relationship: ; In the formula, represents the probability that the image predicted by the student model belongs to the identity , is the output visual feature of the re-identification model; According to the obtained probability distribution , the knowledge distillation loss is: ; Among them, represents the probability that the teacher model assigns the input image to the category , = 2; The re-identification model is trained to obtain a refined visual representation, and the total training loss is: ; Among them, the identity loss , the triplet loss .
[0011] Optionally, the method further includes: Extract the visual features of the query images and gallery images in the test set using the visual encoder of the trained re-identification model, and calculate the cross-modal similarity by combining the text features generated by the teacher model; sort the gallery samples according to the similarity scores, output the target identity matching the query image, and complete the pedestrian re-identification task.
[0012] In a second aspect, an embodiment of the present application provides a re-identification model training system based on a noise-robust prompt learning framework, including a processor and a memory; The memory is used to store computer programs; The processor is configured to implement the method steps described in any one of the first aspects when executing the programs stored on the memory.
[0013] Beneficial effects: The present invention designs a vision-guided global text prompt and a neighborhood-aware context text prompt mechanism, which can effectively capture semantic nuances in the data. The vision-guided global text prompt generates specific text prompts based on image features, solving the problem of lack of specific text labels; the neighborhood-aware context prompt reduces the risk of confirmation bias by capturing proximity relationship information, improving the accuracy and robustness of the model. At the same time, the present invention proposes a prompt-driven knowledge distillation strategy, using the learned comprehensive text prompts as class vectors to guide model optimization, enhancing the performance of the model under noise conditions; converting image-specific prompts into class-specific semantic prompts through a max-pooling strategy further improves the anti-noise ability of the model. Description of the drawings
[0014] Figure 1 It is a flowchart of the training of the re-identification model based on the noise-robust prompt learning framework according to an embodiment of the present invention; Figure 2 It is a logical schematic diagram of the training of the re-identification model based on the noise-robust prompt learning framework according to an embodiment of the present invention; Figure 3 It is a supplementary diagram of the logical schematic diagram of the training of the re-identification model based on the noise-robust prompt learning framework according to an embodiment of the present invention. In the figure, Figure 3 (a) is part a of the logical schematic diagram of the training of the re-identification model based on the noise-robust prompt learning framework, Figure 3 (b) is part b of the logical schematic diagram of the training of the re-identification model based on the noise-robust prompt learning framework; Detailed implementation manners
[0015] The technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0016] Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those of ordinary skill in the art to which the present invention pertains. The "first", "second" and similar terms used in the present invention do not denote any order, quantity or importance, but are only used to distinguish different components. Similarly, terms such as "a" or "one" do not denote a quantity limitation, but mean that there is at least one. The terms "connected" or "coupled" and the like are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right" and the like are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship also changes accordingly.
[0017] As Figure 1 shown, an embodiment of the present application provides a method for training a re-identification model based on a noise-robust prompt learning framework, including: Extracting the global visual features of the input pedestrian image by using the visual encoder of the pre-trained CLIP model, retrieving the k nearest neighbor global visual features of the global visual features of the target sample in the feature space, and generating context visual features based on cosine similarity weighted aggregation; Mapping the global visual features and context visual features into global pseudo-language prompt words and context pseudo-language prompt words through a learnable mapping network; Combining the predefined basic prompt template with the global pseudo-language prompt words and context pseudo-language prompt words to construct global text prompts and context text prompts, and generating global text embeddings and context text embeddings based on the text encoder of the pre-trained CLIP model; Fusing the global text embedding and the context text embedding based on the max-pooling strategy to form a comprehensive text embedding; Introducing a language-guided cross-attention module, using the comprehensive text embedding as the query vector, and the patch embeddings output by the visual encoder as the key vector and value vector to optimize the multi-modal feature representation; jointly optimizing the model parameters by using the generalized cross-entropy loss function and the symmetric contrast loss function to improve the robustness to noise labels; Embed the comprehensive text as a category semantic vector, and align the visual feature distribution of the student re-identification model with the text feature distribution of the teacher model through a prompt imitation strategy; jointly train the student model based on the identity classification loss, triplet margin loss, and semantic knowledge distillation loss; Use the visual encoder of the trained re-identification model to extract the visual features of the query images and gallery images in the test set, and calculate the cross-modal similarity in combination with the text features generated by the teacher model; sort the gallery samples according to the similarity scores, and output the target identity matching the query image to complete the pedestrian re-identification task.
[0018] In the above embodiment, the training method of the re-identification model based on the noise-robust prompt learning framework can be mainly divided into a training stage and an application stage, as Figures 2 - 3 shown, and the specific steps of each stage are as follows: S1: Use the visual encoder of the pre-trained CLIP model to extract the global visual features of the input pedestrian images, retrieve the k nearest neighbor global visual features of the global visual features of the target samples in the feature space, and generate context visual features based on cosine similarity weighted aggregation; The pedestrian images are in the training sets of the Market1501 and DukeMTMC-ReID datasets; the Market1501 dataset contains 32,668 images of 1,501 identities captured by six cameras. This dataset is divided into a training set and a test set, where the training set contains 12,936 images of 751 identities; the test set contains a gallery set with 19,732 images and a query set with 3,368 images, both of which contain 750 identities. The DukeMTMC-ReID dataset is adapted from DukeMTMC and contains 36,411 images of 1,812 identities recorded from eight cameras. The training set contains 16,522 images of 702 identities, while the test set includes a query set with 2,228 images of 702 identities and a gallery set with 17,661 images of 1,110 identities; S1.1: Use the visual encoder of the pre-trained CLIP model to extract the global visual features of the input pedestrian images; Input all the training set pedestrian images into the visual encoder to generate a series of global visual features ; S1.2: Retrieve the k nearest neighbor global visual features of the global visual features of the target samples in the feature space, and generate context visual features based on cosine similarity weighted aggregation; For the global visual features of the target sample , find its k nearest neighbor global visual features in the feature space, denoted as , obtain the context visual features from the formula, and the process satisfies the following relational expression: ; In the formula, , is the cosine similarity measure between two samples, represents the context visual features.
[0019] S2: Use the learnable mapping network to map the global visual features and context visual features into pseudo-language prompt words and context pseudo-language prompt words. This process satisfies the following relational expression: ; ; In the formula, is the global visual feature, is the generated global pseudo-language prompt word, is the context visual feature, is the context pseudo-language prompt word.
[0020] S3: Combine the predefined basic prompt template with the global pseudo-language prompt word and context pseudo-language prompt word to construct the global text prompt and context text prompt, and generate the global text embedding and context text embedding based on the text encoder of the pre-trained CLIP model; S3.1 Construct the global text prompt and context text prompt according to the predefined basic prompt template and the global pseudo-language prompt word and context pseudo-language prompt word obtained in S2; The basic prompt template is as follows: A photo of a person; A photo of a person; Among them, represents the global pseudo-language prompt word, represents the context pseudo-language prompt word; S3.2: According to the global text prompt and context text prompt obtained in S3.1, input them into the text encoder of the pre-trained CLIP model to generate the global text embedding and context text embedding. The generation process satisfies the following relational expression: ; ; In the formula, is the global text prompt, is the global text embedding, is the context text prompt, is the context text embedding, is the pre-trained CLIP text encoder.
[0021] S4: Based on the max-pooling strategy, fuse the global text embedding and the context text embedding to form a comprehensive text embedding; According to the obtained global text embedding and the context text embedding , use the max-pooling strategy to fuse the two into a comprehensive text embedding , where the process of fusing the comprehensive text embedding satisfies the following relational expression: ; In the formula, is the global text embedding, is the context text embedding, is the comprehensive text embedding, is the max-pooling operation.
[0022] S5: Introduce a language-guided cross-attention module, use the comprehensive text embedding as the query vector, and the patch embeddings output by the visual encoder as the key vector and value vector to optimize the multi-modal feature representation; use the generalized cross-entropy loss function and the symmetric contrastive loss function to jointly optimize the model parameters to improve the robustness to noisy labels; S5.1: According to the patch embeddings and the comprehensive text embedding , use the language-guided cross-attention module, and the comprehensive text embedding acts as the query vector, and the full patch embeddings are used as the key vector and value vector. The text semantic fusion process satisfies the following relational expression: ; ; In the formula, is the dimension of the text feature, corresponds to three fully connected layers, represents the fused text semantic information; S5.2: In the prompt optimization, use the generalized cross-entropy loss to optimize the prompt, and the optimization process satisfies the following formula: ; In the formula, is the probability that the sample is classified into the category based on its final feature representation, = 0.7 to balance robustness and performance; S5.3: In this training stage, both the image encoder and the text encoder are frozen. To ensure that the learned text semantic information can effectively capture the image context, we adopt a symmetric contrastive loss, and the formula is as follows: ; ; where the parameter plays a role in adjusting the softening degree of the similarity score in the contrastive loss function; S5.4: According to the generalized cross-entropy loss function and the symmetric contrastive loss function described in S5.2 and S5.3, the total loss formula for optimization is constructed to satisfy the following relationship: ; In the formula, represents the generalized cross-entropy loss, and represent the symmetric contrastive loss function.
[0023] S6: Aggregate the comprehensive text embeddings into class semantic vectors, and align the visual feature distribution of the student re-identification model with the text feature distribution of the teacher model through the prompt imitation strategy; jointly train the student model based on the identity classification loss, triplet margin loss, and semantic knowledge distillation loss; S6.1: The text embeddings provide semantic supervision signals as class vectors. For each identity, aggregate the prompts of the same identity to form a unified text embedding, and its aggregation process satisfies the following relationship: ; In the formula, represents the unified text embedding of a specific identity after aggregation, represents the comprehensive text embedding, represents the identity class, represents the identity label; S6.2: The student model calculates the probability that a given image belongs to an identity class, and the calculation process satisfies the following relationship: ; In the formula, represents the probability that the image predicted by the student model belongs to the identity , is the output visual feature of the re-identification model; S6.3: According to the probability distribution obtained in S6.2, the knowledge distillation loss is: ; where Indicates the probability that the teacher model assigns to the input image to the category is = 2; S6.4: Train the re-identification model to obtain a refined visual representation. In addition to using the identity loss and the triplet loss , introduce the knowledge distillation loss as an additional supervision signal. The total training loss is: ; where the identity loss , the triplet loss .
[0024] The present invention uses the ViT-B / 16 model pre-trained on CLIP as the visual encoder and the CLIP text Transformer as the text encoder; integrates a randomly initialized mapping network and a multi-modal interaction module, which is defined by the parameter . The mapping network is a lightweight three-layer MLP with a 512-dimensional hidden state, and a batch normalization layer is added after its last layer to improve training stability; in the text prompt generation stage, the adam optimizer is used with an initial learning rate of 3.5e-4, which is adjusted by the cosine decay schedule. The training is carried out for 30 epochs with a batch size of 64, and data augmentation is not applied. In this stage, only the mapping network and the multi-modal interaction module are updated; in the semantic-driven ReID model optimization stage, the adam optimizer is used to train the image encoder. The learning rate starts from 5e-7 and linearly increases to 5e-6 in the first 10 epochs. Subsequently, the learning rate is reduced by 0.1 times at the 30th and 50th epochs. The entire model is trained for 60 epochs with a batch size of 64, including 16 identities with 4 images per identity; to enhance robustness and diversity, image augmentation techniques such as random horizontal flipping, padding, cropping, and erasing are applied; Application stage S7: Use the visual encoder of the trained re-identification model to extract the visual features of the query images and gallery images in the test set, and calculate the cross-modal similarity by combining the text features generated by the teacher model; sort the gallery samples according to the similarity scores and output the target identity that matches the query image to complete the pedestrian re-identification task.
[0025] The embodiments of the present disclosure also provide a re-identification model training system based on a noise-robust prompt learning framework, including a processor and a memory; The memory is used to store computer programs; A processor, when executing a program stored in a memory, implements any of the method steps in the re-identification model training method based on a noise-robust prompt learning framework.
[0026] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.
Claims
1. A method for re-identification model training based on a noise robust cue learning framework, characterized in that: include: The visual encoder of the pre-trained CLIP model is used to extract the global visual features of the input pedestrian image, and the k nearest neighbor global visual features of the global visual features of the target sample are retrieved in the feature space, and the contextual visual features are generated based on the weighted aggregation of cosine similarity; Mapping global visual features and contextual visual features into global pseudo-linguistic cues and contextual pseudo-linguistic cues through a learnable mapping network; Combine the predefined basic prompt template with the global pseudo-language prompt words and the contextual pseudo-language prompt words to construct global text prompts and contextual text prompts, and generate global text embedding and contextual text embedding based on the text encoder of the pre-trained CLIP model; The global text embedding and the context text embedding are fused based on the maximum pooling strategy to form a comprehensive text embedding; A language-guided cross-attention module is introduced, and the comprehensive text embedding is used as the query vector, and the tile embedding output by the visual encoder is used as the key vector and value vector to optimize the multimodal feature representation; the generalized cross entropy loss function and the symmetric contrast loss function are used to jointly optimize the model parameters to improve the robustness to noisy labels; The comprehensive text embedding is used as a category semantic vector, and the visual feature distribution of the student re-ID model is aligned with the text feature distribution of the teacher model through a prompt imitation strategy; The student model is jointly trained based on identity classification loss, triple boundary loss and semantic knowledge distillation loss; The visual encoder of the trained re-ID model is used to extract the visual features of the query image and gallery image in the test set, and the cross-modal similarity is calculated by combining the text features generated by the teacher model. The gallery samples are sorted according to the similarity scores, and the target identity that matches the query image is output to complete the pedestrian re-ID task.
2. A method for re-identification model training based on a noise robust cue learning framework according to claim 1, characterized in that: The visual encoder of the pre-trained CLIP model is used to extract global visual features of the input pedestrian image, including: All training sets Pedestrian images Input to the visual encoder to generate a series of global visual features ; in, Represents identity labels, pedestrian images in the training sets of Market1501 and DukeMTMC-ReID datasets; The k nearest neighbor global visual features of the global visual features of the target sample are retrieved in the feature space, and the contextual visual features are generated based on the cosine similarity weighted aggregation, including: For the target sample Global visual features , find its k nearest neighbor global visual features in the feature space, denoted as , the contextual visual features are obtained by the formula, and the process satisfies the following relationship: ; In the formula, , is the cosine similarity measure between two samples, Represents contextual visual features.
3. The method for re-identification model training based on a noise robust cue learning framework according to claim 1, characterized in that: The global visual features and contextual visual features are mapped into global pseudo-language cues and contextual pseudo-language cues through a learnable mapping network, including: Using a learnable mapping network The global visual features and contextual visual features are mapped into global pseudo-language prompts and contextual pseudo-language prompts. This process satisfies the following relationship: ; ; In the formula, is the global visual feature, is the generated global pseudo-language prompt word, is the contextual visual feature, It is a contextual pseudo-linguistic cue word.
4. The method for re-identification model training based on a noise robust cue learning framework according to claim 1, characterized in that: Combine the predefined basic prompt template with the global pseudo-language prompt words and the contextual pseudo-language prompt words to construct global text prompts and contextual text prompts, including: Constructing global text prompts and contextual text prompts according to predefined basic prompt templates, global pseudo-language prompt words, and contextual pseudo-language prompt words; The basic prompt template meets the following requirements: A photo of a person; A photo of a person; in, Represents the global pseudo-language prompt word, Represents contextual pseudo-linguistic cue words; The text encoder based on the pre-trained CLIP model generates global text embeddings and contextual text embeddings, including: According to the obtained global text prompts and contextual text prompts, they are input into the text encoder of the pre-trained CLIP model to generate global text embedding and contextual text embedding. The generation process satisfies the following relationship: ; ; In the formula, For global text prompt, is the global text embedding, is a contextual text prompt, is contextual text embedding, To pre-train CLIP text encoder.
5. The method for re-identification model training based on a noise robust cue learning framework according to claim 1, characterized in that: The global text embedding and the contextual text embedding are fused based on the maximum pooling strategy to form a comprehensive text embedding, including: According to the global text embedding and contextual text embedding , using the maximum pooling strategy to merge the two into a comprehensive text embedding , where the integrated text embedding process satisfies the following relationship: ; In the formula, is the global text embedding, is contextual text embedding, For comprehensive text embedding, It is the maximum pooling operation.
6. The method for re-identification model training based on a noise robust cue learning framework according to claim 1, characterized in that: A language-guided cross-attention module is introduced, which uses the comprehensive text embedding as the query vector and the tile embedding output by the visual encoder as the key vector and value vector to optimize the multimodal feature representation. The generalized cross entropy loss function and the symmetric contrast loss function are used to jointly optimize the model parameters to improve the robustness to noisy labels, including: Embed by tile and comprehensive text embeddings , using language-guided cross-attention modules, comprehensive text embedding Acts as a query vector, full tile embedding As key vector and value vector, the text semantic fusion process satisfies the following relationship: ; ; In the formula, is the dimension of text features, Corresponding to three fully connected layers, Represents the fused text semantic information; In prompt optimization, generalized cross entropy loss is used to optimize prompts, and the optimization process satisfies the following formula: ; In the formula, It is a sample Classified into categories based on their final feature representation The probability of =0.7 balances robustness and performance; In this training phase, we freeze both the image encoder and the text encoder. To ensure that the learned text semantic information can effectively capture the image context, we use a symmetric contrast loss to optimize the hint, as follows: ; ; In the formula, the parameters It plays a role in adjusting the softening degree of the similarity score in the contrast loss function; According to the generalized cross entropy loss function and the symmetric contrast loss function, the optimized total loss formula is constructed to satisfy the following relationship: ; In the formula, represents the generalized cross entropy loss, and represents the symmetric contrast loss function.
7. The method for re-identification model training based on a noise robust cue learning framework according to claim 1, characterized in that: aggregating the comprehensive text embedding into a category semantic vector, and aligning the visual feature distribution of the student re-ID model with the text feature distribution of the teacher model through a prompt imitation strategy; The student model is trained jointly based on identity classification loss, triple boundary loss and semantic knowledge distillation loss, including: Text embedding provides semantic supervision signals as category vectors. For each identity, the prompts of the same identity are aggregated to form a unified text embedding. The aggregation process satisfies the following relationship: ; In the formula, represents the unified text embedding of a specific identity after aggregation, represents comprehensive text embedding, Indicates the identity category, Indicates identity tag; The student model calculates the probability that a given image belongs to an identity category, and the calculation process satisfies the following relationship: ; In the formula, Image representing the student model predictions Belong to identity The probability of is the output visual feature of the re-ID model; According to the probability distribution , the knowledge distillation loss is: ; in, Represents the teacher model's response to the input image Assign to category The probability of =2; The re-ID model is trained to obtain refined visual representations, and the total training loss is: ; Among them, identity loss , triplet loss .
8. The method for re-identification model training based on a noise robust cue learning framework according to claim 1, characterized in that: The visual encoder of the trained re-ID model is used to extract the visual features of the query image and gallery image in the test set, and the cross-modal similarity is calculated by combining the text features generated by the teacher model. The gallery samples are sorted according to the similarity scores, and the target identity that matches the query image is output to complete the pedestrian re-ID task.
9. A re-identification model training system based on a noise robust cue learning framework, characterized in that: Including processor and memory; Memory, used to store computer programs; The processor is used to implement a re-identification model training method based on a noise robust prompt learning framework as described in any one of claims 1-8 when executing a program stored in the memory.
Citation Information
Patent Citations
Multi-granularity clothes changing pedestrian re-identification method based on clothes desensitization network
CN112784728A
Multi-modal pre-training model migration method based on self-supervised learning
CN118097685A
Semantic knowledge guided vehicle re-identification method
CN118230321A
Generalized pedestrian re-identification method based on visual language multi-granularity distillation
CN118506267A
Partial prompt learning-based lifelong target re-identification method
CN118864825A
Cited By
Image cross-modal pedestrian re-identification method and system based on text clue alignment
CN120429469A
Modal distortion speech recognition method and system
CN120544547A
Video pedestrian re-identification method and system based on hyperbolic uncertainty recovery
CN120580646A
Video pedestrian re-identification method and system based on hyperbolic uncertainty recovery
CN120580646B
AI-based children story picture book generation method and system
CN120782921A