A Lifelong Goal Re-identification Method Based on Descriptive Text Prompt Prototype Compensation

Through the prototype compensation method based on descriptive text prompts, visual and text prototypes are generated and compensators are constructed, which solves the problem that the lifelong goal re-identification model is prone to forget old knowledge when adapting to new tasks, and achieves higher adaptability and robustness.

CN119625285BActive Publication Date: 2025-05-27NORTHWESTERN POLYTECHNICAL UNIV

Patent Information

Application Number
CN202510148167.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-27
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

Lifetime goal re-identification models tend to forget the old knowledge they learned before when adapting to new tasks, resulting in catastrophic forgetting problems.

Method used

Using a prototype compensation method based on descriptive text cues, visual and text prototypes are generated through CLIP image encoder and text encoder, and compensators are constructed using the multi-head cross attention layer and feedforward network layer to optimize visual and text features to reduce feature drift.

Benefits of technology

It effectively alleviates the problem of catastrophic forgetting, improves the adaptability and robustness of the model in multi-task lifelong learning scenarios, reduces dependence on historical samples, and reduces storage requirements and computing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119625285B_ABST
    Figure CN119625285B_ABST
Patent Text Reader

Abstract

The present invention discloses a lifelong object re-identification method based on descriptive text prompt prototype compensation, which relates to the field of computer vision. By introducing a prompt prototype compensation module, it can effectively solve the feature drift problem caused by incremental model updates. This module generates multi-modal visual and text prototypes for each object category through a vision-language pre-training model, and maps the old task features to the new task feature space through a cross-attention mechanism, thereby achieving fine-grained semantic drift compensation. In addition, a boundary-aware prototype correction module is designed to solve the boundary confusion problem caused by the update of old class prototypes. This module dynamically samples pseudo-features in the compensation distribution to correct the old class prototypes with blurred boundaries, focuses on easily confused categories to maintain clear decision boundaries, and optimizes the cross-task and intra-task correction losses to reduce category confusion, effectively alleviating the catastrophic forgetting problem and improving the adaptability and generalization ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a lifelong target re-identification method based on descriptive text prompt prototype compensation. Background Art

[0002] Target re-identification refers to the recognition and matching of the same target by comparing and analyzing the visual features of the target under different viewing angles or time points. This technology can achieve continuous tracking of specific targets under cross-camera conditions and has important applications in monitoring security, intelligent transportation, retail and other fields. Traditional target re-identification methods are mostly trained on a single data set and evaluate the model performance in a static scene. However, in real environments, target image data is rarely provided at one time, but is usually acquired step by step in the form of a sequential input stream. To this end, the target re-identification model needs to have the ability of lifelong learning, which can continuously learn new knowledge from a continuous data stream while maintaining the memory of old knowledge. This is particularly important in dynamically changing application scenarios, such as intelligent monitoring and unmanned driving systems, where the system needs to adapt to the identity information of new targets in real time without losing existing knowledge.

[0003] The main challenge facing lifelong object re-identification is the catastrophic forgetting problem, that is, the model tends to forget the old knowledge learned before when adapting to new tasks. In the process of lifelong learning, the continuous updating of the object re-identification model will cause the feature space to change continuously, which directly leads to the occurrence of catastrophic forgetting. Due to the fine-grained characteristics of the target feature distribution, the inter-class differences between different targets are small, and the distribution boundary is prone to overlapping interference. When feature drift occurs, samples at the distribution boundary are easily forgotten, and the identities of targets with similar feature distributions are also more easily confused. Summary of the invention

[0004] In view of the above-mentioned deficiencies in the prior art, the present invention provides a lifelong target re-identification method based on prototype compensation with descriptive text prompts, which solves the catastrophic forgetting problem existing in the prior art.

[0005] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: a lifelong target re-identification method based on descriptive text prompt prototype compensation, comprising the following steps:

[0006] S1: Get the pre-trained CLIP image encoder, CLIP text encoder, and A dataset, introducing visual cues and initialization;

[0007] S2: Add the initialized visual cues to each layer of the Transformer encoder of the CLIP image encoder;

[0008] S3: Set up training steps , optimize the visual prompts in the CLIP image encoder through the new task loss on the first dataset;

[0009] S4: Use the CLIP image encoder to calculate the visual prototypes and variances of each target identity category in the dataset;

[0010] S5: Optimize the descriptive text of each target identity category in the dataset through the contrastive loss to obtain the category-specific text description of each target identity;

[0011] S6: Input the optimized category-specific text description of each target identity into the CLIP text encoder to obtain the encoded text features of each target identity category as text prototypes;

[0012] S7: Training step , determine whether the current training step is , if so, obtain the text prototypes of the current dataset, complete the life-long target re-identification based on the descriptive text prompt prototype compensation, otherwise enter step S8;

[0013] S8: Optimize the visual prompts in the CLIP image encoder through the new task loss on the current dataset;

[0014] S9: Optimize the compensator through the prompt prototype compensation loss, and use the optimized compensator to update the old class visual prototypes in the previous dataset;

[0015] S10: Based on the updated results, optimize the visual prompts in the CLIP image encoder through the cross-task correction loss, the intra-task correction loss, and the prototype-based distillation loss;

[0016] S11: Based on the optimization results, use the CLIP image encoder to calculate the visual prototypes and variances of each target identity category in the current dataset;

[0017] S12: Optimize the descriptive text of each target identity category in the current dataset through the contrastive loss to obtain the category-specific text description of each target identity;

[0018] S13: Input the optimized category-specific text description of each target identity into the CLIP text encoder to obtain the encoded text features of each target identity category as text prototypes, and return to step S7.

[0019] Furthermore, the CLIP image encoder in S1 is formed based on the ViT model architecture and is used as the backbone network for feature extraction and representation learning;

[0020] The CLIP text encoder is formed based on the Transformer model architecture and is used to convert the natural language text description related to the target identity into text feature representations;

[0021] The input form of the dataset is a set of data streams divided by time steps , the th dataset is expressed as:

[0022]

[0023] where, is the target image of the th sample in the training step , is the identity label corresponding to the target image, is the training step the total number of samples in;

[0024] The dimension size of the visual prompt is ;

[0025] where, is the number of layers of the ViT model, is the sequence length of the prompt, is the feature encoding dimension.

[0026] Furthermore, in S2, for the th layer Transformer encoder, the th layer visual prompt and the output features of the previous layer of the Transformer encoder are combined to form new input features, and the formula is:

[0027]

[0028] where, is the output feature of the th layer Transformer encoder, is the model parameter of the th layer Transformer encoder, represents the concatenation operation;

[0029] After being encoded by the layer Transformer encoder, the obtained visual feature is extracted from the class token vector in the output feature of the th layer, and the formula is:

[0030]

[0031] where, is the input image in the image encoder The encoding of is the class token vector of the final output feature.

[0032] Furthermore, the new task loss is:

[0033]

[0034] where is the new task loss, is the cross-entropy loss, is the triplet loss, is a hyperparameter;

[0035]

[0036] where is the batch size, is the number of classes, marks whether the th sample belongs to the th class, is the probability that the classifier predicts the th sample as the th class, are the classifier parameters;

[0037]

[0038] where is the maximum value, is the Euclidean distance, is the feature of the anchor sample, is the feature of the positive sample, is the feature of the negative sample, is a preset boundary value.

[0039] Furthermore, the calculation formulas for the visual prototype and variance of each target identity class in the dataset are:

[0040]

[0041] where is the visual prototype of identity ; is the number of samples of identity ; is the rd visual feature of the sample, is the th identity label of the sample;

[0042]

[0043] Among them, is the variance of the identity , represents extracting the main diagonal elements of the matrix, represents the transpose operation of the matrix.

[0044] Furthermore, the contrast loss includes the image-to-text contrast loss and the text-to-image contrast loss , and the formula is:

[0045]

[0046] Among them, is to calculate the cosine similarity between the visual feature and the text feature, is the text feature of the th sample extracted by using the text encoder, is the text feature of the th sample extracted by using the text encoder;

[0047]

[0048] Among them, is the index set of all samples in the batch that have the same category as, represents the number of elements in the set, is the th sample in the batch, is the visual feature of the th sample extracted by using the image encoder, is the visual feature of the th sample extracted by using the image encoder, is the text feature corresponding to the identity category to which the th sample belongs;

[0049] By optimizing the descriptive text prompt minimize the image-to-text contrast loss and the text-to-image contrast loss , and the formula is:

[0050]

[0051] Among them, is the sum of the image-to-text contrast loss and the text-to-image contrast loss , is the number of learnable text tokens.

[0052] Furthermore, the compensator in S9 includes a multi-head cross-attention layer and a feed-forward network layer. S9 includes the following sub-steps:

[0053] S91: Calculate the drift value predicted by the old-class visual prototype and the drift value predicted by the old-class text prototype through the multi-head cross-attention layer. The formula is: and the drift value predicted by the old-class text prototype , the formula is:

[0054]

[0055]

[0056] Among them, represents multi-head cross-attention, is the visual feature extracted from the old model obtained by the training samples sampled from the current dataset in the previous training step, is the old-class visual prototype calculated in the previous training step, is the old-class text prototype calculated in the previous training step;

[0057] S92: Combine the drift value predicted by the old-class visual prototype and the drift value predicted by the old-class text prototype to obtain the drift prediction for the current training sample to adapt from the old model to the new model feature space. The formula is:

[0058]

[0059] Among them, is the predicted visual feature at the current time step, is the drift value predicted by the old-class visual or text prototype;

[0060] S93: Optimize the compensator by minimizing the mean square error between the drift prediction value and the true value. The formula is:

[0061]

[0062] Among them, is the prompt prototype compensation loss, is the training step in the th sample's predicted visual feature, is the training step in the th sample's true visual feature;

[0063] S94: Update the old-class visual prototype in the previous dataset with the optimized compensator , the formula is:

[0064]

[0065] Among them, is the old-class visual prototype after compensation from training step to training step .

[0066] Furthermore, the cross-task calibration loss in the S10 is:

[0067] The boundary fuzziness is defined by the pairwise cosine similarity between and , and the formula is:

[0068]

[0069] Among them, is the fuzziness between the current task sample feature and the th old-class sample, is the visual prototype feature after compensation for the th target category in the previous dataset, is the visual prototype feature after compensation for the th target category in the previous dataset, is the predefined temperature parameter;

[0070] Sample old classes from the fuzziness distribution , and sample pseudo-features from each old-class prototype Gaussian distribution to calculate the cross-task calibration loss :

[0071]

[0072]

[0073]

[0074] Among them, is the index set of all samples in the batch with the same class as the th sample, is the normalization term for positive samples, is the normalization term for pseudo-samples, is the sampled old-class pseudo-feature.

[0075] Furthermore, the intra-task calibration loss in the S10 is:

[0076]

[0077] Among them, is the in-task calibration loss.

[0078] Furthermore, the prototype-based distillation loss in S10 is:

[0079]

[0080] Among them, is the prototype-based distillation loss, is the Kullback-Leibler divergence, is the Softmax normalization function, is the pairwise similarity matrix between the current task samples and the compensated visual prototypes of the old classes, is from the current dataset sampled visual feature matrix composed of

[0081] The beneficial effects of the present invention are as follows:

[0082] 1. A no-replay prompt fine-tuning lifelong learning mechanism. The present invention does not need to store and replay the original image data of the old task target classes, effectively avoiding the privacy leakage problem caused by data replay. By constructing visual class prototypes, the present invention can reduce the dependence on historical samples, thereby reducing the storage requirements and computational costs. In addition, the present invention only introduces a small number of learnable prompt parameters to achieve efficient parameter fine-tuning of the pre-trained CLIP model, accelerating the model training speed while improving the model adaptability.

[0083] 2. A multimodal prototype compensation and text prompt guidance mechanism. The multimodal prototype compensation method proposed by the present invention constructs robust class prototypes with rich discriminative features by fusing visual and text information. Using the visual and text prototypes generated by the CLIP model and combining descriptive text prompts, the present invention realizes the precise compensation of fine-grained semantic feature drift, effectively alleviates the problem of feature space inconsistency, and improves the adaptability and robustness of the model in the multi-task lifelong learning scenario.

[0084] 3. A cross-task and in-task calibration mechanism. The cross-task and in-task calibration mechanism proposed by the present invention can effectively alleviate the catastrophic forgetting problem in the lifelong target re-identification task. By introducing the cross-task calibration loss and the in-task calibration loss, the present invention can maintain the model's memory of the previously learned knowledge when learning new tasks, maintain a clear decision boundary between different target identities, and effectively improve the anti-forgetting and generalization abilities of the model. Description of the Drawings

[0085] Figure 1 is a flowchart of a lifelong target re-identification method based on descriptive text prompt prototype compensation.

[0086] Figure 2 It is a structural diagram of a lifelong object re-identification network model based on descriptive text prompt prototype compensation. Specific implementation manner

[0087] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0088] The present invention proposes a lifelong object re-identification method based on descriptive text prompt prototype compensation. By introducing a prompt prototype compensation module to explicitly solve the feature drift problem caused by incremental model updates, this module generates multi-modal visual and text prototypes for each object category through the vision-language pre-training model CLIP, and maps the old task features to the new task feature space through the cross-attention mechanism, thereby achieving fine-grained semantic drift compensation. In addition, a boundary-aware prototype correction module is designed to solve the boundary confusion problem caused by the update of old class prototypes. This module dynamically samples pseudo-features in the compensation distribution to correct the old class prototypes with blurred boundaries, focuses on easily confused categories to maintain clear decision boundaries, and optimizes the cross-task and intra-task correction losses to reduce category confusion, effectively alleviating the catastrophic forgetting problem and improving the adaptability and generalization ability of the model.

[0089] As Figure 1 shown, a lifelong object re-identification method based on descriptive text prompt prototype compensation includes the following steps:

[0090] S1: Obtain a pre-trained CLIP image encoder, a CLIP text encoder, and a dataset, introduce visual prompts and perform initialization;

[0091] S2: Add the initialized visual prompts to each layer of the Transformer encoder of the CLIP image encoder;

[0092] S3: Set the training step , and optimize the visual prompts in the CLIP image encoder through the new task loss on the first dataset;

[0093] S4: Use the CLIP image encoder to calculate the visual prototypes and variances of each object identity category in the dataset;

[0094] S5: Optimize the descriptive text of each object identity category in the dataset through the contrastive loss to obtain the category-specific text description of each object identity;

[0095] S6: Input the optimized category-specific text description of each object identity into the CLIP text encoder to obtain the encoded text features of each object identity category as text prototypes;

[0096] S7: Training Step , determine whether the current training step , if so, obtain the text prototype of the current dataset, complete the lifelong target re-identification based on the descriptive text prompt prototype compensation, otherwise go to step S8;

[0097] S8: Optimize the visual prompt in the CLIP image encoder through the new task loss on the current dataset;

[0098] S9: Optimize the compensator through the prompt prototype compensation loss, and use the optimized compensator to update the old-class visual prototypes in the previous dataset;

[0099] S10: Based on the update result, optimize the visual prompt in the CLIP image encoder through the cross-task correction loss, the intra-task correction loss, and the prototype-based distillation loss;

[0100] S11: Based on the optimization result, use the CLIP image encoder to calculate the visual prototypes and variances of each target identity class in the current dataset;

[0101] S12: Optimize the descriptive text of each target identity class in the current dataset through the contrastive loss to obtain the class-specific text description of each target identity;

[0102] S13: Input the optimized class-specific text description of each target identity into the CLIP text encoder to obtain the encoded text features of each target identity class as the text prototype, and return to step S7.

[0103] As Figure 2 shown, the CLIP image encoder in S1 is formed based on the ViT model architecture and is used as the backbone network for feature extraction and representation learning;

[0104] In this embodiment, it is specifically the ViT-B / 16 model architecture, where "B" represents the base model, and 16 means that the model will divide the input image into image patches of 16×16 pixel size. ViT-B / 16 contains a total of 12 Transformer encoder layers, and each image patch is encoded into a 768-dimensional visual feature embedding vector;

[0105] The CLIP text encoder is formed based on the Transformer model architecture and is used to convert the natural language text description related to the target identity into a text feature representation;

[0106] The CLIP text encoder is specifically a 12-layer Transformer model, which uses byte pair encoding to process the text, and the dimension of the text feature embedding vector is 512;

[0107] The input form of the dataset is a set of data streams divided by time steps , the th dataset is expressed as:

[0108]

[0109] wherein, is the target image of the th sample in the training step , is the identity label corresponding to the target image, is the training step and the total number of samples in the training step

[0110] The present invention selects common object re-identification datasets for incremental training, namely Market1501, CUHK-SYSU, DukeMTMC, MSMT17 and CUHK03

[0111] The dimension size of the visual prompt is , which is a set of additional learnable parameter vectors used to be attached to each layer of the image encoder. The role of the visual prompt is to guide the image encoder to continuously learn new tasks through the Prompt Tuning strategy without modifying the backbone network parameters of the ViT model

[0112] wherein, is the number of layers of the ViT model, is the sequence length of the prompt, is the feature encoding dimension

[0113] In S2, for the th layer of the Transformer encoder, the output feature of the previous layer of the th layer of the visual prompt and the Transformer encoder is combined through a concatenation operation to form a new input feature, and the formula is:

[0114]

[0115] wherein, is the output feature of the th layer of the Transformer encoder, is the model parameter of the th layer of the Transformer encoder, represents the concatenation operation;

[0116] After ​After encoding by the layer Transformer encoder, the obtained visual features From the layer output features The category token vector in is extracted, and the formula is:

[0117]

[0118] Among them, Is the encoding of the input image In the image encoder, Is the category token vector of the final output feature.

[0119] The new task loss is:

[0120]

[0121] Among them, Is the new task loss, Is the cross-entropy loss, Is the triplet loss, Is a hyperparameter used to balance the contribution of the triplet loss In the loss calculation;

[0122]

[0123] Cross-entropy loss The classifier is used to measure the difference between the predicted probability of the model for different target identities and the true label;

[0124] Among them, Is the batch size. In the experiment, the batch size is set to 64, including 16 randomly selected target categories, with 4 images for each category, Is the number of categories, Indicates whether the th sample belongs to the th category (0 or 1), Is the probability that the classifier predicts the th sample as the th category, Is the classifier parameter;

[0125]

[0126] Among them, Is the maximum value, Is the Euclidean distance, Is the feature of the anchor sample, Is the feature of the positive sample, Is the feature of the negative sample, Is the pre-set boundary value.

[0127] Triplet loss For deep metric learning, it ensures that samples of the same category are closer in the feature space while pushing the distance between samples of different categories further apart.

[0128] The present invention trains each dataset for 60 epochs, optimizes using the Adam optimizer, and the learning rate linearly increases from to in the first 10 epochs, and then the learning rate decays at the 30th and 50th epochs. Data augmentation is performed on the training images using strategies such as random horizontal flipping, padding, cropping, and erasing.

[0129] The calculation formulas for the visual prototype and variance of each target identity category in the dataset are as follows:

[0130]

[0131] where is the visual prototype of identity , is the number of samples of identity , is the th sample's visual feature, , is the th sample's identity label;

[0132]

[0133] where is the variance of identity , represents the main diagonal elements of the extraction matrix, represents the transpose operation of the matrix.

[0134] Through the above calculation method, the feature distribution of identity can be modeled as a Gaussian distribution .

[0135] The contrastive loss includes image-to-text contrastive loss and text-to-image contrastive loss , and the formula is:

[0136]

[0137] where is to calculate the cosine similarity between the visual feature and the text feature, is the text feature of the th sample extracted using the text encoder, is the text feature of the th sample extracted using the text encoder;

[0138]

[0139] Among them, is the index set of all samples in the batch with the same category, , and represent the identity category to which the / th sample belongs, represents the number of elements in the set, is the th sample in the batch, is the visual feature of the th sample extracted using the image encoder, is the visual feature of the th sample extracted using the image encoder, is the text feature corresponding to the identity category to which the th sample belongs; ;

[0140] In this embodiment, a descriptive text prompt template "A photo of a person." is designed to optimize the target identity-specific text description for each target category. Among them, represents the word embedding text token vector, and the number of learnable text tokens is set to in the experiment.

[0141] By optimizing the descriptive text prompt minimize the image-to-text contrast loss and the text-to-image contrast loss , and the formula is:

[0142]

[0143] Among them, is the sum of the image-to-text contrast loss and the text-to-image contrast loss , is the number of learnable text tokens.

[0144] Specifically, the present invention trains each dataset for 120 epochs, optimizes using the Adam optimizer, and initializes the learning rate to , and then the learning rate is decayed by a cosine scheduler. After training is completed, each target identity has its own class-specific text description.

[0145] To solve the catastrophic forgetting problem caused by feature drift, the present invention learns a compensator to map features from the old feature space to the new feature space. The network structure of the compensator consists of a multi-head cross-attention layer and a feed-forward network layer.

[0146] The compensator in S9 includes a multi-head cross-attention layer and a feed-forward network layer. S9 includes the following sub-steps:

[0147] S91: Calculate the drift value predicted by the old-class visual prototype and the drift value predicted by the old-class text prototype through the multi-head cross-attention layer and the drift value predicted by the old-class text prototype , and the formula is:

[0148]

[0149]

[0150] where represents the multi-head cross-attention, is the visual features extracted from the old model obtained in the previous training step for training samples sampled from the current data set, is the old-class visual prototype calculated in the previous training step, represents the previous task data set the number of target classes included, represents the feature dimension, is the old-class text prototype calculated in the previous training step, is the set of real numbers;

[0151] S92: Combine the drift value predicted by the old-class visual prototype and the drift value predicted by the old-class text prototype to obtain the drift prediction for the current training sample to adapt from the old model to the new model feature space. The formula is:

[0152]

[0153] where is the predicted visual feature at the current time step, is the drift value predicted by the old-class visual or text prototype;

[0154] S93: Optimize the compensator by minimizing the mean square error between the drift prediction value and the true value. The formula is:

[0155]

[0156] Among them, to prompt the prototype to compensate for the loss, is the training step in the th sample's predicted visual feature, is the training step in the th sample's true visual feature;

[0157] This training process keeps all the parameters of the CLIP image encoder frozen to independently train the compensator . Specifically, the present invention trains for 120 epochs for each dataset, optimizes using the Adam optimizer, initializes the learning rate to , and then decays the learning rate through a cosine scheduler. After training is completed, the parameters of the compensator are frozen and only used to update the visual prototypes of the old classes.

[0158] S94: Update the visual prototypes of the old classes in the previous dataset with the optimized compensator , and the formula is:

[0159]

[0160] Among them, is the compensated visual prototype of the old class from training step to training step .

[0161] The cross-task correction loss in the above S10 is:

[0162] After compensating the visual prototypes of the old classes , 's distribution may overlap with the distribution of the new class prototypes . To address the boundary confusion problem caused by prototype updates, the present invention uses the cross-task correction loss to maintain boundary distinctiveness:

[0163] Define the boundary fuzziness through the pairwise cosine similarity between and , and the formula is:

[0164]

[0165] Among them, is the fuzziness between the current task sample feature and the th old class sample, is the The visual prototype features after compensating for the target category are the visual prototype features after compensating for the th target category in the previous dataset, and is the predefined temperature parameter;

[0166] Sample old classes from the blurriness distribution and sample 6 pseudo-features from the Gaussian distribution of each old class prototype for calculating the cross-task correction loss :

[0167]

[0168]

[0169]

[0170] where is the index set of all samples in the batch with the same sample category as the th sample, is the normalization term for positive samples, is the normalization term for pseudo-samples, and is the sampled old-class pseudo-features. During the sampling process, old-class prototypes with higher similarity after compensation are more likely to be sampled for prototype correction, thus better alleviating the forgetting problem of easily confused old-class prototypes.

[0171] To reduce the boundary confusion within the current task, the in-task correction loss is used to correct the inter-class discriminative boundary:

[0172] The in-task correction loss in S10 is:

[0173]

[0174] where is the in-task correction loss, calculates the similarity between the visual features of all samples (including positive and negative samples) and the sample for normalized output.

[0175] To transfer the old-class prototype knowledge to the current task model to reduce the forgetting of old knowledge, calculate the prototype-based distillation loss :

[0176] The prototype-based distillation loss in S10 is:

[0177]

[0178] where is the prototype-based distillation loss, is the Kullback-Leibler divergence, is the Softmax normalization function, is the pairwise similarity matrix between the current task samples and the compensated visual prototypes of the old classes, is sampled from the current dataset and is the visual feature matrix composed of samples. The pairwise similarity matrix of the visual features of samples is calculated, and the pairwise similarity matrix with the old class prototypes as the reference coordinates is calculated. The overall loss of the proposed prompt prototype compensation framework of the present invention is calculated as follows:

[0179] where

[0180]

[0181] is a constant, and and are used to balance the contributions of each loss term in the overall loss calculation. Specifically, the present invention trains for 60 epochs for each dataset, optimizes using the Adam optimizer, initializes the learning rate to and then decays the learning rate at the 30th and 50th epochs. After completing the th training step, increase the value of

[0182] and return to step S7 to continue processing the next dataset until all datasets are processed. It should be noted that the method of the present invention only constructs class prototypes for each target class during the training phase. During the inference phase, the present invention does not need to calculate class prototypes for testing, but directly extracts the visual features of each test sample using the image encoder and performs target matching between the query set and the gallery set features, thereby calculating the mAP and R@1 accuracy metrics. After completing the th training step, increase the value of and return to step S7 to continue processing the next dataset until all datasets are processed. It should be noted that the method of the present invention only constructs class prototypes for each target class during the training phase. During the inference phase, the present invention does not need to calculate class prototypes for testing, but directly extracts the visual features of each test sample using the image encoder and performs target matching between the query set and the gallery set features, thereby calculating the mAP and R@1 accuracy metrics.

[0183] The performance results of the present invention for continuous lifelong learning on 5 target re-identification datasets, namely Market1501, CUHK-SYSU, DukeMTMC, MSMT17, and CUHK03, are shown in Table 1. The comparison algorithms include LwF, AKA, PatchKD, LSTKC, C2R, DKP, and PAEMA.

[0184] Table 1 Comparison of performance results between the method of the present invention and other methods

[0185]

[0186] The experimental results in Table 1 show that the method of the present invention has achieved the optimal results in terms of both the test performance and the overall average performance on 5 object re-identification datasets. Compared with the current state-of-the-art methods, the average mAP / R@1 performance of the present invention has been improved by 3.2% / 3.7%. The experimental results verify that the proposed descriptive text prompt prototype compensation and boundary-aware prototype correction strategies of the present invention can effectively overcome the catastrophic forgetting problem and enhance the adaptability to new tasks.

[0187] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations without departing from the essence of the present invention according to these technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the invention.

Claims

1. A lifelong object re-identification method based on prototype compensation with descriptive text prompts, characterized in that: The following steps are involved: S1: Get the pre-trained CLIP image encoder, CLIP text encoder, and A dataset, introducing visual cues and initialization; S2: Add the initialized visual cues to each layer of the Transformer encoder of the CLIP image encoder; S3: Set up training steps , optimizing visual cues in the CLIP image encoder via a new task loss on the first dataset; S4: Use CLIP image encoder to calculate the visual prototype and variance of each target identity category in the dataset; S5: optimizing the descriptive text of each target identity category in the dataset by contrastive loss to obtain a category-specific text description of each target identity, wherein the contrastive loss includes an image-to-text contrastive loss and a text-to-image contrastive loss; S6: Input the optimized category-specific text description of each target identity into the CLIP text encoder to obtain the encoded text features of each target identity category as the text prototype; S7: Training Step , to determine whether the current training step , if yes, then obtain the text prototype of the current data set and complete the lifelong target re-identification based on the descriptive text prompt prototype compensation, otherwise go to step S8; S8: Optimizing visual cues in the CLIP image encoder via a new task loss on the current dataset; S9: The compensator is optimized by the prompt prototype compensation loss, and the old class visual prototype in the previous data set is updated using the optimized compensator. The prompt prototype supplement loss optimizes the compensator by minimizing the mean square error between the drift prediction value and the true value. The formula is: in, To prompt the prototype to compensate for the loss, For the training step Middle The predicted visual features of samples, For the training step Middle The real visual features of samples, is the batch size; S10: Based on the updated results, the visual cues in the CLIP image encoder are optimized by cross-task correction loss, intra-task correction loss, and prototype-based distillation loss; S11: Based on the optimization results, the CLIP image encoder is used to calculate the visual prototype and variance of each target identity category in the current dataset; S12: Optimize the descriptive text of each target identity category in the current dataset through contrast loss to obtain a category-specific text description of each target identity; S13: Input the optimized category-specific text description of each target identity into the CLIP text encoder, obtain the encoded text features of each target identity category as a text prototype, and return to step S7.

2. The method for lifelong target re-identification based on prototype compensation of descriptive text prompts according to claim 1, characterized in that: The CLIP image encoder in S1 is formed based on the ViT model architecture and serves as a backbone network for feature extraction and representation learning; The CLIP text encoder is formed based on the Transformer model architecture and is used to convert the natural language text description related to the target identity into a text feature representation; The input form of the dataset is a set of data streams divided by time steps , No. Datasets It is expressed as: in, For the training step Middle The target image of samples, is the identity label corresponding to the target image, For the training step The total number of samples in ; The dimension size of the visual cue is ; in, is the number of layers of the ViT model, is the sequence length of the prompt, Encodes the dimension for the feature.

3. The method for lifelong target re-identification based on prototype compensation with descriptive text prompts according to claim 2, characterized in that: In S2, for the The Transformer encoder of the first layer is concatenated Layer Visual Cues And the output features of the previous layer of Transformer encoder Combined to form new input features, the formula is: in, For the The output features of the layer Transformer encoder, For the The model parameters of the layer Transformer encoder, Represents a splicing operation; go through After layer Transformer encoder encoding, the visual features obtained From Layer output features The category token vector in is extracted and the formula is: in, For the input image in the image encoder The encoding, is the category token vector of the final output feature.

4. The method for lifelong target re-identification based on prototype compensation with descriptive text prompts according to claim 3, characterized in that: The new task loss is: in, Loss for new tasks, is the cross entropy loss, is the triplet loss, is a hyperparameter; in, is the batch size, is the number of categories, Marked the Does the sample belong to categories, For the classifier The sample prediction is The probability of the categories, is the classifier parameter; in, is the maximum value, is the Euclidean distance, is the feature of the anchor sample, is the feature of the positive sample, is the feature of negative samples, is a pre-set boundary value.

5. The method for lifelong target re-identification based on prototype compensation with descriptive text prompts according to claim 4, characterized in that: The calculation formula for the visual prototype and variance of each target identity category in the dataset is: in, For identity The visual prototype, For identity The number of samples, For the The visual features of samples, For the The identity labels of the samples; in, For identity The variance of It means extracting the main diagonal elements of the matrix. Represents the transpose operation of a matrix.

6. The method for lifelong target re-identification based on prototype compensation with descriptive text prompts according to claim 5, characterized in that: The contrast loss includes image-to-text contrast loss and text-to-image contrastive loss , the formula is: in, To calculate the cosine similarity between visual features and text features, The first The text features of samples, The first Text features of samples; in, For batch The index set of all samples of the same category, Indicates the number of elements in a collection. For the batch samples, The first The visual features of samples, The first The visual features of samples, For the The text features corresponding to the identity category to which the sample belongs; By optimizing the descriptive text prompt Minimize image-to-text contrastive loss and text-to-image contrastive loss , the formula is: in, is the image-to-text contrast loss and text-to-image contrastive loss the sum of is the number of learnable text tokens.

7. The method for lifelong target re-identification based on prototype compensation with descriptive text prompts according to claim 6, characterized in that: The compensator in S9 includes a multi-head cross attention layer and a feedforward network layer, and S9 includes the following sub-steps: S91: Calculate the drift value predicted by the old visual prototype through a multi-head cross attention layer and the drift value predicted by the old class text prototype , the formula is: in, Indicates multi-head cross attention, is sampled from the current dataset The visual features extracted from the training samples on the old model obtained in the previous training step, is the old class visual prototype calculated in the previous training step, The prototype of the old class text calculated in the previous training step; S92: The drift value predicted by the old visual prototype and the drift value predicted by the old class text prototype Merge to obtain the drift prediction of the current training sample from the old model to the new model feature space. The formula is: in, is the predicted visual feature of the current time step, is the drift value predicted by the old class visual or text prototype; S93: Optimizing the compensator by minimizing the mean square error between the drift prediction value and the true value; S94: The optimized compensator Update the old class visual prototypes in the previous dataset , the formula is: in, From the training step To the training step Compensated old class visual prototype.

8. The method for lifelong target re-identification based on prototype compensation with descriptive text prompts according to claim 7, characterized in that: The cross-task correction loss in S10 is: pass and The pairwise cosine similarity between defines the degree of boundary blurring, and the formula is: in, is the sample feature of the current task and The fuzziness between old class samples, For the previous data set The visual prototype features after target category compensation, For the previous data set The visual prototype features after target category compensation, is a predefined temperature parameter; From the fuzzy distribution Medium Sampling old classes, and from each old class prototype Gaussian distribution The sampled pseudo features are used to calculate the cross-task correction loss : in, For the batch The index set of all samples with the same sample category, is the normalization term of the positive sample, is the normalization term of the pseudo sample, is the old class pseudo feature sampled.

9. The method for lifelong target re-identification based on prototype compensation with descriptive text prompts according to claim 8, characterized in that: The intra-task correction loss in S10 is: in, is the within-task correction loss.

10. The method for lifelong target re-identification based on prototype compensation with descriptive text prompts according to claim 9, characterized in that: The prototype-based distillation loss in S10 is: in, is the prototype-based distillation loss, is the Kullback-Leibler divergence, is the Softmax normalization function, is the pairwise similarity matrix between the current task sample and the compensated old class visual prototype, From the current dataset Sampled in The visual feature matrix consists of samples.

Citation Information

Patent Citations

  • Target re-identification model anti-forgetting training method, target re-identification method and target re-identification device

    CN115439878A

  • Partial prompt learning-based lifelong target re-identification method

    CN118864825A

Cited By

  • Reloading lifelong pedestrian re-identification method and system based on sub-distribution collaborative enhancement

    CN122336860A