Pedestrian Re-identification Methods Based on Similarity Guidance and Mismatch Feature Enhancement

By using similarity-guided and mismatch feature enhancement methods, the problems of insufficient fine-grained interaction and neglect of mismatch features in TI-ReID technology are solved, and more efficient pedestrian re-identification retrieval is achieved.

CN119580302BActive Publication Date: 2025-10-28CHONGQING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411597893.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-10-28
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

Existing TI-ReID technology lacks fine-grained interactive capabilities in image and text matching, ignores the important role of mismatch features, and results in insufficient retrieval accuracy.

Method used

A similarity-guided multimodal interaction module and a mismatch feature emphasis module are adopted. By guiding the selection of masked text features through similarity, cross-modal interaction between images and text is enhanced, and mismatch features are used to improve retrieval accuracy. The loss function is optimized by combining multi-task learning methods.

Benefits of technology

It improves the fine-grained alignment of image and text matching, reduces the false positive rate, enhances the ability to distinguish between similar individuals, and improves retrieval accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580302B_ABST
    Figure CN119580302B_ABST
Patent Text Reader

Abstract

This invention discloses a pedestrian re-identification method based on similarity guidance and mismatch feature enhancement, belonging to the field of intelligent recognition technology. This invention proposes a similarity-guided masking strategy that enhances the use of image information in the masking language modeling process to promote stronger cross-modal interaction. Unlike random masking, this method guides the model's attention to more relevant image-text correspondences, thereby achieving better fine-grained alignment. Furthermore, it introduces a novel mismatch feature enhancement module, which innovatively utilizes mismatch features to improve retrieval accuracy. Previous work mainly focused on matching image-text features; our mismatch feature enhancement module explores the importance of mismatch features, which play a crucial role in distinguishing visually similar individuals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent recognition technology, specifically to a pedestrian re-identification method based on similarity guidance and mismatch feature enhancement. Background Technology

[0002] Text-to-Image Person Re-identification (TI-ReID) aims to retrieve target pedestrians from image databases using text descriptions. This technology was first proposed by Li et al. [1] in 2017, and the first benchmark dataset CUHK-PEDES was released. The core challenge of TI-ReID is how to achieve cross-modal alignment between images and text and overcome modal heterogeneity.

[0003] The following references are cited:

[0004] [1] Li, S., Xiao, T., Li, H., Zhou, B., Yue, D., Wang, X., 2017. Person search with natural language description, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1970–1979.

[0005] [2]Chen, Y., Huang, R., Chang, H., Tan, C., Xue, T., Ma, B., 2021. Cross-modal knowledge adaptation for language-based person search. IEEE Transactions onImage Processing 30, 4057–4069.

[0006] [3] Zheng, Z., Zheng, L., Garrett, M., Yang, Y., Xu, M., Shen, YD, 2020b. Dual-path convolutional image-text embeddings with instance loss. ACM Transactionson Multimedia Computing, Communications, and Applications (TOMM) 16, 1–23.

[0007] [4]Zhang,Y.,Lu,H.,2018.Deep cross-modal projection learning forimage-text matching,in:Proceedings of the European conference on computervision(ECCV),pp.686–701.

[0008] [5]Li,S.,Xiao,T.,Li,H.,Zhou,B.,Yue,D.,Wang,X.,2017.Person search withnatural language description,in:Proceedings of the IEEE conference oncomputer vision and pattern recognition,pp.1970–1979.

[0009] [6]Chen,T.,Xu,C.,Luo,J.,2018b.Improving text-based person search byspatial matching andadaptive threshold,in:2018 IEEE Winter Conference onApplications of Computer Vision(WACV),IEEE.pp.1879–1887.

[0010] [7]Wei,D.,Zhang,S.,Yang,T.,Liu,J.,2023.Calibrating cross-modalfeature for text-based person searching.arXiv preprint arXiv:2304.02278.

[0011] [8]Zha,Z.J.,Liu,J.,Chen,D.,Wu,F.,2020.Adversarial attribute-textembedding for person search with natural language query.IEEE Transactions onMultimedia 22,1836–1846.

[0012] [9]Sarafianos,N.,Xu,X.,Kakadiaris,I.A.,2019.Adversarialrepresentation learning for text-to-image matching,in:Proceedings of theIEEE / CVF international conference on computer vision,pp.5814–5824.

[0013]

[10] Gao,C.,Cai,G.,Jiang,X.,Zheng,F.,Zhang,J.,Gong,Y.,Peng,P.,Guo,X.,Sun,X.,2021.Contextual non-local alignment over full-scale representation fortext-based person search.arXiv preprint arXiv:2101.03036.

[0014]

[11] Ding,Z.,Ding,C.,Shao,Z.,Tao,D.,2021.Semantically self-alignednetwork for text-to-image part-aware person re-identification.arXiv preprintarXiv:2107.12666.

[0015]

[12] Wu,Y.,Yan,Z.,Han,X.,Li,G.,Zou,C.,Cui,S.,2021.Lapscore:language-guided person search via color reasoning,in:Proceedings of the IEEE / CVFInternational Conference on Computer Vision,pp.1624–1633.

[0016]

[13] Wang,Z.,Fang,Z.,Wang,J.,Yang,Y.,2020.Vitaa:Visual-textualattributes alignment in person search by natural language,in:Computer Vision–ECCV 2020:16th European Conference,Glasgow,UK,August 23–28,2020,Proceedings,Part XII 16,Springer.pp.402–420.

[0017]

[14] Aggarwal,S.,Radhakrishnan,V.B.,Chakraborty,A.,2020.Text-basedperson search via attribute-aided matching,in:Proceedings of the IEEE / CVFwinter conference on applications of computer vision,pp.2617–2625.

[0018]

[15] Zhu,A.,Wang,Z.,Li,Y.,Wan,X.,Jin,J.,Wang,T.,Hu,F.,Hua,G.,2021.Dssl:Deep surroundings-person separation learning for text-based personretrieval,in:Proceedings of the 29th ACM International Conference onMultimedia,pp.209–217.

[0019]

[16] Niu,K.,Huang,Y.,Ouyang,W.,Wang,L.,2020.Improving description-based person re-identification by multi-granularity image-textalignments.IEEE Transactions on Image Processing 29,5542–5556.

[0020]

[17] Jing,Y.,Si,C.,Wang,J.,Wang,W.,Wang,L.,Tan,T.,2020.Pose-guidedmulti-granularity attention network for text-based person search,in:Proceedings of the AAAI Conference on Artificial Intelligence,pp.11189–11196.

[0021]

[18] Zheng,K.,Liu,W.,Liu,J.,Zha,Z.J.,Mei,T.,2020a.Hierarchical gumbelattention network for text-basedperson search,in:Proceedings of the 28th ACMInternational Conference on Multimedia,pp.3441–3449.

[0022]

[19] Ji,Z.,Hu,J.,Liu,D.,Wu,L.Y.,Zhao,Y.,2022.Asymmetric cross-scalealignment for text-based person search.IEEE Transactions on Multimedia.

[0023]

[20] Jiang,D.,Ye,M.,2023.Cross-modal implicit relation reasoning andaligning for text-to-image person retrieval,in:Proceedings of the IEEE / CVFConference on Computer Vision and Pattern Recognition,pp.2787–2797.

[0024]

[21] Taylor,W.L.,1953.“cloze procedure”:A new tool for measuringreadability.Journalism quarterly 30,415–433.

[0025]

[22] Shi, D., 2024. Transnext: Robust foveal visual perception for vision transformers, in: Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 17773–17783.

[0026] Existing TI-ReID technology mainly faces the following challenges and limitations:

[0027] 1. Early methods [2-9] mainly focused on global matching of images and text. For example, Zhang et al. [4] proposed cross-modal projection classification loss and cross-modal projection matching loss to align global image and text features. While these methods are intuitive and efficient, they lack the ability to model more fine-grained and complex image-text interactions and may not be able to fully capture the detailed correspondence between visual and text information.

[0028] 2. To address the limitations of global matching, subsequent research [10-19] has shifted to local matching strategies, aiming to capture fine-grained image-text correspondences. For example, Niu et al.

[16] proposed a multi-granularity alignment model that achieves image-text alignment at three different granularities. However, these methods typically rely on region-based comparisons, which incur significant computational overhead.

[0029] 3. Recently, Jiang et al.

[20] introduced the Masked Language Modeling (MLM)

[21] task to implicitly model the image-text correspondence, reducing computational overhead and avoiding additional supervision. However, the random masking strategy in MLM lacks deep involvement of image information, resulting in weak interaction between images and text.

[0030] 4. Existing methods mainly focus on matching image-text features, neglecting the important role of mismatch features in distinguishing visually similar individuals.

[0031] Therefore, a new solution is needed to address the problem arising between problem 3 and problem 4. Summary of the Invention

[0032] The purpose of this invention is to provide a pedestrian re-identification method based on similarity guidance and mismatch feature enhancement. Through a similarity-guided multimodal interaction module, the model is guided to focus on more relevant image-text correspondences, thereby improving the fine-grained image-text alignment effect. Through a mismatch feature enhancement module, mismatch features are utilized in the TI-ReID task to improve the ability to distinguish between similar individuals, thus solving the technical problems mentioned in the background art.

[0033] To achieve the above objectives, the present invention provides the following technical solution: a pedestrian re-identification method based on similarity guidance and mismatch feature enhancement, comprising at least the following steps:

[0034] S1: Image and text feature extraction. Image and text feature maps are extracted using the pre-trained encoder in the contrastive language-image pre-trained model. Specifically, a 12-layer Vision Transformer is used as the image encoder, and a 12-layer Transformer architecture is used as the text encoder. The contrastive language-image pre-trained model is CLIP, and the Vision Transformer is ViT.

[0035] S2: Similarity-guided multimodal interaction utilizes the correlation between image and text features to select text features to be masked. Text features that are highly correlated with image features are more likely to be masked. This strategy encourages the use of image information in the MLM process, thereby enhancing cross-modal interaction between images and text.

[0036] S3: The mismatch feature emphasis module is used to strengthen the mismatch feature, which aims to reduce the false positive rate and improve the overall retrieval accuracy. The mismatch feature emphasis module is MFE.

[0037] S4: Training and inference are performed using a multi-task learning approach, optimizing four loss functions in four parallel components: an SGMI module, an MFE module, an ID loss, and an SDM loss. The SGMI module calculates the MLM loss, the MFE module includes the ITC loss, and the ID loss and SDM loss are calculated independently. During inference, a similarity metric S is calculated for each text query and image candidate pair. This calculation uses the cosine similarity method to quantify the correlation between text and image features. Subsequently, S is used to build a ranked list of individuals in the image library, and the highest-ranked image matching that is most similar to the given text description is retrieved and presented.

[0038] Furthermore, S1 includes at least the following steps:

[0039] S1.1: Image feature extraction, using CLIP pre-trained VIT from input image I∈R H×W×CImage features are obtained from the image, where H, W, and C represent the image height, width, and number of channels, respectively.

[0040] I is divided into N = H × W / P 2 A non-overlapping sequence of patches, where P represents the patch size;

[0041] These patches are mapped to 1D labels using a trainable linear projection.

[0042] F v To represent the feature sequence of the entire image;

[0043] f j v This represents the feature vector of the j-th patch;

[0044] N represents the number of feature sequences;

[0045] After injection location embedding and additional category markers, the marker sequence Input is fed into the transformer block to simulate the relevance of each patch;

[0046] Used as a global image representation, denoted as g v ;

[0047] S1.2: Text feature extraction. For text input T, the text encoder of CLIP is used to obtain the text representation.

[0048] Processing begins with tokenizing the lowercase input text using byte-pair encoding, using a vocabulary of 49,152 tokens;

[0049] The text description is encapsulated within [SOS] and [EOS] tags to indicate the start and end of the sequence;

[0050] Subsequently, the segmented text The data is input into the transformer to analyze the relationships between the markers;

[0051] The highest level [EOS] token of the transformer g is considered a global text representation t .

[0052] Furthermore, S2 includes at least the following steps:

[0053] First, calculate the text features F. t Its corresponding global image feature g v The similarity between them is then normalized using the softmax function, as shown in formula (1):

[0054]

[0055] Where L represents the length of the text tag. <f i t ,g v > indicates f i t and g v The cosine similarity between them, where ρ is the temperature hyperparameter and r i The sampling probability of text tags used as a predetermined proportion for masking; exp is an abbreviation for exponential function, which calculates the exponential value of the input, here mapping each similarity score to a positive value; k is the summation symbol. The index variable in the text represents the exponential sum of the similarity scores between all feature pairs from 1 to L; L represents the number of text features.

[0056] Using the probability mask text features in formula (1), we obtain the text features of the partial mask, denoted as F. mt ;

[0057] F mt and image features F v They are fed together into a multimodal interactive encoder, which consists of a multi-head cross-attention layer and four transformer blocks;

[0058] To effectively integrate image and text information, the text features F of the partial mask are used. mt As the query Q, and the image feature F v As key K and value V in the multi-head cross-attention layer;

[0059] This interaction is implemented as shown in formula (2):

[0060]

[0061] in This represents a masked text representation that incorporates image information. It is a set of mask text marker positions, where d represents the embedding dimension;

[0062] K T In this context, T usually refers to the matrix transpose operation. After transposing K, K is matched with the dimension of the query matrix Q, and the dot product of the two is calculated to generate the attention weight matrix.

[0063] For each mask position The multilayer perceptron classifier predicts the probability distribution of the corresponding original label for the MLP, and is calculated using the following formula (3):

[0064]

[0065] in f represents the predicted probability of the j-th tag in the vocabulary at the i-th mask position; i m This represents the masked text feature vector after incorporating image information, referring to the text feature at mask position i after processing by a multimodal interactive encoder.

[0066] The vocabulary used here is denoted as O, the same as that used in the text encoding process, and consists of 49,152 tags;

[0067] In the final stage of the SGMI module, the MLM loss is calculated based on the prediction of masked text tags. This loss is designed to optimize the model so that text and visual features are aligned at a fine-grained level.

[0068] The MLM loss is the sum of the cross-entropy between the masked text tag and its label, defined by the following formula (4):

[0069]

[0070] in The correct label is marked as 1, and the others are marked as 0.

[0071] Furthermore, the MFE operates in two different modes: mismatched image feature extraction and mismatched text feature extraction, each equipped with a customized information aggregation strategy;

[0072] The mismatched image feature extraction is UIE, the mismatched text feature extraction is UTE, and the information aggregation strategy is IG strategy;

[0073] The IG strategy is a key component of UIE and UTE patterns, used to address the challenge of identifying mismatched features in high-dimensional space;

[0074] The process for UTE is similar to that for UIE. The difference between UTE and UIE is that UTE operates from the text perspective, focusing on text features that do not match any image features. Furthermore, the IG strategy in UTE aggregates image tags in a manner similar to that of UIE.

[0075] Furthermore, the MFE in the UIE mode identifies image features that do not match any text features in a given pair;

[0076] The MFE application in the UIE mode includes at least the following steps:

[0077] After aggregating L text tags into K text tags using the IG strategy, the similarity between each image tag and the K text tags is calculated.

[0078] UIE's IG strategy applies parameterless adaptive average pooling to text tag Ft Aggregate into a more compact representation Where K <L;

[0079] That is: text tags are divided into K groups, and the index of the text features in the j-th group is used. This means that multiple text features within each group are processed to form a single aggregated feature. Recognizing that average pooling leads to significant information loss, a single-layer neural network is used for projection and GELU activation before pooling. After pooling, layer normalization is applied to the output to ensure consistent feature scale.

[0080] The IG strategy is expressed as formula (5):

[0081]

[0082] in Let j represent the set of text features in the j-th group. It is the set of indices of the j-th group, t j These are the features after aggregation;

[0083] Subsequently, text features and visual features F v The similarity is calculated by projecting the image onto the common feature space through a 1x1 convolution, as expressed in formula (6):

[0084] s i,j = <W v f i v W t t j >, i∈[1,N], j∈[1,K] (6)

[0085] Among them W v W t ∈R M×C , <W v f i v W t t j >is W v f i v and W t t j Cosine similarity between them, s i,j This represents the similarity between the i-th image tag and the j-th text tag;

[0086] If the similarity between an image feature and all K text features is less than 0, then it is identified as an image feature that does not match any text feature.

[0087] To highlight the degree of mismatch, the negative similarity scores of this image feature and all text features are added together, expressed as formula (7):

[0088]

[0089] Where s i The summation represents the cumulative mismatch score of the i-th image feature. Only negative similarity values ​​are included in the summation to emphasize the mismatch. The scalar s represents this score. i The s is formed by being replicated M times. i ∈R M , so that it is similar to f i v ∈R M Dimension alignment;

[0090] Subsequently, this mismatch score s i Used with f i v The visual feature v is obtained by multiplying the weights element by element. i ∈R M It emphasizes the degree of mismatch with all text tags, expressed as formula (8):

[0091] v i =s i ⊙f i v (8)

[0092] Using the same method, identify all image tags that do not match any text features, forming... Where G represents the number of such identified image features;

[0093] For each image-text pair, to emphasize the impact of mismatch features, these image features are computed. With text features F t The similarity between them, the average result, is expressed as formula (9):

[0094]

[0095] s m G represents the average similarity score between image and text features when they do not match; G represents the number of image features identified as not matching the text; L is the number of text features. <v i ,f j t > indicates the i-th image feature v that is identified as not matching the text. i With the j-th text feature f j t Cosine similarity;

[0096] Then it is combined with global image-text similarity;

[0097] For B image-text pairs, the intra-pair similarity is calculated by the global feature cosine similarity and s. m The similarity between two images, i.e., between the i-th image and the j-th text, where i ≠ j, is determined solely by the cosine similarity of the global features, as expressed in formulas (10) and (11):

[0098]

[0099] in and Let represent the global features of the i-th image and the j-th text, respectively. The similarity from the i-th image to the j-th text is denoted as . Measuring the correspondence between visual elements and text descriptions, using a similar method... That is, the similarity between the i-th text and the j-th image;

[0100] The resulting hybrid similarity is used to calculate the image-text contrast loss, also known as the ITC loss. The ITC loss is a method for training multimodal models that optimizes the alignment between matching image-text pairs by maximizing the similarity between matching image-text pairs while minimizing the similarity between mismatched pairs.

[0101] In the final step of the MFE module, the similarity scores calculated in previous steps are used to calculate the ITC loss to enhance the effect of mismatched features;

[0102] By emphasizing these contrasting relationships, the ITC loss significantly enhances the model's ability to distinguish visually similar individuals, as detailed in the following formula group (12):

[0103]

[0104] Where γ is the temperature hyperparameter.

[0105] Furthermore, the ID loss is applied as follows: ID loss is used to group individuals based on their identities to ensure matching at the identity level;

[0106] For image-text pairs (I,T), cross-entropy loss is applied to the global features (g). v ,g t );

[0107] ID loss The definition is expressed as formula (13):

[0108]

[0109] Among them Wid y represents the parameters of the fully connected layer used for classification. id This represents the one-hot encoded vector of the real label.

[0110] Furthermore, the SDM loss is applied as follows: The SDM loss is used to minimize the image-text similarity μ. i,j Distribution and matching labels ξ i,j The KL divergence between the normalized distributions is shown in Equations (14) and (15) below:

[0111]

[0112]

[0113] The peak value of the probability distribution is controlled by the temperature hyperparameter τ, y i,j It is a true match tag, y i,j =1 means They are positive pairs from the same identity, while 0 represents a negative pair;

[0114] The SDM loss is expressed as formula (16):

[0115]

[0116] Where ∈ is a decimal used to avoid numerical problems;

[0117] Total training loss It is the weighted sum of these losses:

[0118]

[0119] λ1 and λ2 are weighting factors used to balance the contributions of ITC loss and MLM loss.

[0120] Compared with the prior art, the beneficial effects of the present invention are:

[0121] 1. This invention proposes a similarity-guided masking strategy, which enhances the use of image information in the masking language modeling process to promote stronger cross-modal interaction. Unlike random masking, this method guides the model's attention to more relevant image-text correspondences, thereby achieving better fine-grained alignment.

[0122] 2. This invention introduces a novel mismatch feature emphasis module, which innovatively utilizes mismatch features to improve retrieval accuracy. Previous work mainly focused on matching image-text features; our mismatch feature emphasis module explores the importance of mismatch features, which play a crucial role in distinguishing visually similar individuals. Attached Figure Description

[0123] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0124] Figure 1 This is a schematic diagram of the overall process of the present invention. Detailed Implementation

[0125] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0126] Example 1:

[0127] Given an image-text pair, we first extract image and text feature maps using a pre-trained encoder in a Contrastive Language-Image Pre-training (CLIP) model. Specifically, we utilize a 12-layer Vision Transformer (ViT) as the image encoder and a 12-layer Transformer architecture as the text encoder. Next, the extracted features are computed through four parallel components: a Similarity-Guided Multimodal Interaction (SGMI) module, a Mismatched Feature Emphasis (MFE) module, and an identity identification (ID) loss and a similarity distribution matching (SDM) loss.

[0128] The SGMI module deeply integrates image information into the Masked Language Modeling (MLM) task, implicitly simulating fine-grained image-text correspondences and calculating the MLM loss. The MFE module enhances retrieval performance by emphasizing the role of mismatched features and includes image-text contrast (ITC) loss calculation. Simultaneously, the ID loss is used to optimize identity discrimination representation, while the SDM loss is used to align the similarity distributions between modalities, enhancing the cross-modal consistency of learned features.

[0129] This parallel architecture allows each component to focus on its designated task while processing the same initial features. During backpropagation, the outputs and gradients of these components are aggregated, enabling the model to learn robust and nuanced representations. This approach effectively captures subtle, fine-grained correspondences and key differences between image and text modalities.

[0130] Example 2:

[0131] Based on the above embodiment one, the following technical content is specifically disclosed:

[0132] Please see Figure 1 A pedestrian re-identification method based on similarity guidance and mismatch feature enhancement includes at least the following steps:

[0133] S1: Image and text feature extraction. Image and text feature maps are extracted using the pre-trained encoder in the contrastive language-image pre-trained model. Specifically, a 12-layer Vision Transformer is used as the image encoder, and a 12-layer Transformer architecture is used as the text encoder. The contrastive language-image pre-trained model is CLIP, and the Vision Transformer is ViT.

[0134] S2: Similarity-guided multimodal interaction utilizes the correlation between image and text features to select text features to be masked. Text features that are highly correlated with image features are more likely to be masked. This strategy encourages the use of image information in the MLM process, thereby enhancing cross-modal interaction between images and text.

[0135] It should be noted that the method using MLM to implicitly simulate fine-grained correspondences between images and text exhibits excellent retrieval performance in the TI-ReID task. However, the strategy of random masking text tagging lacks deep involvement in image information, resulting in insufficient cross-modal interaction. Therefore, a similarity-guided multimodal interaction is proposed.

[0136] S3: The MFE module is used to strengthen the mismatch feature, which aims to reduce the false positive rate and improve the overall retrieval accuracy.

[0137] It should be noted that most methods focus on amplifying the beneficial impact of matching image-text features on retrieval, neglecting the influence of mismatch features, which can significantly affect retrieval results. Highlighting the effect of mismatch features can help improve retrieval performance. To this end, we propose the Feature Mismatch Emphasis (MFE) module. While traditional methods focus on using matching features to improve retrieval, MFE takes a complementary approach, emphasizing the discriminative power of mismatches.

[0138] S4: Training and inference are performed using a multi-task learning approach, optimizing four loss functions in four parallel components: the SGMI module, the MFE module, the ID loss, and the SDM loss. The SGMI module calculates the MLM loss, the MFE module includes the ITC loss, and the ID loss and SDM loss are calculated independently. During inference, a similarity metric S is calculated for each text query and image candidate pair. This calculation uses the cosine similarity method to quantify the correlation between text and image features. Subsequently, S is used to build a ranked list of individuals in the image library, and the highest-ranked image matching that is most similar to the given text description is retrieved and presented.

[0139] S1 includes at least the following steps:

[0140] S1.1: Image feature extraction, using CLIP pre-trained VIT from input image I∈R H×W×C Image features are obtained from the image, where H, W, and C represent the image height, width, and number of channels, respectively.

[0141] I is divided into N = H × W / P 2 A non-overlapping sequence of patches, where P represents the patch size;

[0142] These patches are mapped to 1D labels using a trainable linear projection.

[0143] F v f represents the feature sequence of the entire image. j v Let N represent the feature vector of the j-th patch, and N represent the number of feature sequences.

[0144] After injection location embedding and additional class (CLS) markers, the marker sequence Input is fed into the transformer block to simulate the relevance of each patch;

[0145] Used as a global image representation, denoted as g v ;

[0146] S1.2: Text feature extraction. For text input T, the text encoder of CLIP is used to obtain the text representation.

[0147] This encoder is a modified version of the Transformer architecture. Processing begins with tokenizing the lowercase input text using byte-pair encoding (BPE) with a vocabulary of 49,152 tokens;

[0148] The text description is encapsulated within [SOS] and [EOS] tags to indicate the start and end of the sequence;

[0149] Subsequently, the segmented text The data is input into the transformer to analyze the relationships between the markers;

[0150] The [EOS] token at the highest level of the transformer g is considered a global text representation t .

[0151] S2 includes at least the following steps:

[0152] First, calculate the text features F. t Its corresponding global image feature g v The similarity between them is then normalized using the softmax function, as shown in formula (1):

[0153]

[0154] Where L represents the length of the text tag. <f i t ,g v > indicates f i t and g v The cosine similarity between them, where ρ is the temperature hyperparameter and r i The sampling probability of text markers used as a predetermined proportion of the mask;

[0155] exp is short for exponential function; exp calculates the exponential value of the input. Here, exp maps each similarity score to a positive value. k is the summation symbol. The index variable in the table represents the exponential sum of the similarity scores between all feature pairs from 1 to L. L represents the number of text features.

[0156] Using the probability mask text features in formula (1), we obtain the text features of the partial mask, denoted as F. mt ;

[0157] F mt and image features F v They are fed together into a multimodal interactive encoder, which consists of a multi-head cross-attention layer and four transformer blocks;

[0158] To effectively integrate image and text information, the text features F of the partial mask are used. mt As the query Q, and the image feature F v As key K and value V in the multi-head cross-attention layer;

[0159] This interaction is implemented as shown in formula (2):

[0160]

[0161] in This represents a masked text representation that incorporates image information. It is a set of mask text marker positions, where d represents the embedding dimension;

[0162] K T In this context, T usually refers to the matrix transpose operation. Transposing K allows K to match the dimension of the query matrix Q, thereby calculating the dot product of the two and generating the attention weight matrix.

[0163] For each mask position The multilayer perceptron classifier predicts the probability distribution of the corresponding original label for the MLP, and is calculated using the following formula (3):

[0164]

[0165] in This represents the predicted probability of the j-th tag in the vocabulary at the i-th mask position;

[0166] f i m This represents the masked text feature vector after incorporating image information, referring to the text feature at mask position i after processing by a multimodal interactive encoder.

[0167] The vocabulary used here is denoted as O, the same as that used in the text encoding process, and consists of 49,152 tags;

[0168] In the final stage of the SGMI module, the MLM loss is calculated based on the prediction of masked text tags. This loss is designed to optimize the model so that text and visual features are aligned at a fine-grained level.

[0169] The MLM loss is the sum of the cross-entropy between the masked text tag and its label, defined by the following formula (4):

[0170]

[0171] in The correct label is marked as 1, and the others are marked as 0.

[0172] By integrating image information into the MLM task, the model gains a deeper understanding of the complex relationships between the two modalities. The MLM loss is crucial for enhancing cross-modal feature alignment, thereby improving the model's ability to achieve fine-grained cross-modal understanding.

[0173] MFE operates in two different modes: mismatched image feature extraction and mismatched text feature extraction. Each mode is equipped with a customized information aggregation strategy.

[0174] The feature extraction for mismatched images is UIE, the feature extraction for mismatched text is UTE, and the information aggregation strategy is IG strategy.

[0175] The IG strategy is a key component of UIE and UTE patterns, used to address the challenge of identifying mismatched features in high-dimensional spaces;

[0176] The process for UTE is similar to that of UIE. The difference between UTE and UIE is that UTE operates on the text aspect, focusing on text features that do not match any image features. Furthermore, the IG strategy in UTE aggregates image tags in a manner similar to that of UIE.

[0177] MFE in UIE mode identifies image features that do not match any text features in a given pair;

[0178] An MFE application in UIE mode includes at least the following steps:

[0179] After aggregating L text tags into K text tags using the IG strategy, the similarity between each image tag and the K text tags is calculated.

[0180] The probability that an image tag has a similarity of less than zero with all L text tags is extremely low, making it challenging to identify truly mismatched features. To address this issue, UIE's IG strategy applies parameterless adaptive average pooling to the text tags F t Aggregate into a more compact representation Where K <L;

[0181] That is: text tags are divided into K groups, and the index of the text features in the j-th group is used. This means that multiple text features within each group are processed to form a single aggregated feature. Recognizing that average pooling may lead to significant information loss, a single-layer neural network (Linear) is used for projection and GELU activation before pooling. After pooling, layer normalization (LN) is applied to the output to ensure consistent feature scale.

[0182] The IG strategy is expressed as formula (5):

[0183]

[0184] in Let j represent the set of text features in the j-th group. It is the set of indices of the j-th group, t j These are the features after aggregation;

[0185] Subsequently, text features and visual features F v The similarity is calculated by projecting the image onto the common feature space through a 1x1 convolution, as expressed in formula (6):

[0186] s i,j = <W v f i v W t t j >, i∈[1,N], j∈[1,K] (6)

[0187] Among them W v W t ∈R M×C , <W v f i v W t t j >is W v f i v and W t t j Cosine similarity between them, s i,j This represents the similarity between the i-th image tag and the j-th text tag;

[0188] If the similarity between an image feature and all K text features is less than 0, then it is identified as an image feature that does not match any text feature.

[0189] To highlight the degree of mismatch, the negative similarity scores of this image feature and all text features are added together, expressed as formula (7):

[0190]

[0191] Where s i The summation represents the cumulative mismatch score of the i-th image feature. Only negative similarity values ​​are included in the summation to emphasize the mismatch. The scalar s represents this score. i The s is formed by being replicated M times. i ∈R M , so that it is similar to f i v ∈R M Dimension alignment;

[0192] Subsequently, this mismatch score s i Used with f i v The visual feature v is obtained by multiplying the weights element by element. i ∈R M It emphasizes the degree of mismatch with all text tags, expressed as formula (8):

[0193] v i =s i ⊙f i v (8)

[0194] Using the same method, identify all image tags that do not match any text features, forming... Where G represents the number of such identified image features;

[0195] For each image-text pair, to emphasize the impact of mismatch features, these image features are computed. With text features F t The similarity between them, the average result, is expressed as formula (9):

[0196]

[0197] s m G represents the average similarity score between image and text features when they do not match; G represents the number of image features identified as not matching the text; L is the number of text features. <v i ,f j t > indicates the i-th image feature v that is identified as not matching the text. i With the j-th text feature f j t Cosine similarity;

[0198] Then it is combined with global image-text similarity;

[0199] For B image-text pairs, the intra-pair similarity is calculated by the global feature cosine similarity and s. m The similarity between two images, i.e., between the i-th image and the j-th text, where i ≠ j, is determined solely by the cosine similarity of the global features, as expressed in formulas (10) and (11):

[0200]

[0201] in and Let represent the global features of the i-th image and the j-th text, respectively. The similarity from the i-th image to the j-th text is denoted as . Measuring the correspondence between visual elements and text descriptions, using a similar method... That is, the similarity between the i-th text and the j-th image;

[0202] The resulting hybrid similarity is used to calculate the image-text contrast loss, also known as the ITC loss. The ITC loss is a method for training multimodal models that optimizes the alignment between matching image-text pairs by maximizing the similarity between matching image-text pairs while minimizing the similarity between mismatched pairs.

[0203] In the final step of the MFE module, the similarity scores calculated in previous steps are used to calculate the ITC loss to enhance the effect of mismatched features;

[0204] By emphasizing these contrasting relationships, the ITC loss significantly enhances the model's ability to distinguish visually similar individuals, as detailed in the following formula group (12):

[0205]

[0206] Where γ is the temperature hyperparameter.

[0207] The application of ID loss is to group individuals based on their identity to ensure matching at the identity level.

[0208] For image-text pairs (I,T), cross-entropy loss is applied to the global features (g). v ,g t );

[0209] ID loss The definition is expressed as formula (13):

[0210]

[0211] Among them W id y represents the parameters of the fully connected layer used for classification. id This represents the one-hot encoded vector of the real label.

[0212] The application of SDM loss is to minimize the image-text similarity μ. i,j Distribution and matching labels ξ i,j The KL divergence between the normalized distributions is shown in Equations (14) and (15) below:

[0213]

[0214] The peak value of the probability distribution is controlled by the temperature hyperparameter τ, y i,j It is a true match tag, y i,j =1 means They are positive pairs from the same identity, while 0 represents a negative pair;

[0215] The SDM loss is expressed as formula (16):

[0216]

[0217] Where ∈ is a decimal used to avoid numerical problems;

[0218] Total training loss It is the weighted sum of these losses:

[0219]

[0220] λ1 and λ2 are weighting factors used to balance the contributions of ITC loss and MLM loss.

[0221] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

Claims

1. A pedestrian re-identification method based on similarity guidance and mismatch feature enhancement, characterized in that: At least the following steps are included: S1: Image and text feature extraction. The image and text feature maps are extracted using the pre-trained encoder in the contrastive language-image pre-trained model. Specifically, a 12-layer Vision Transformer is used as the image encoder, and a 12-layer Transformer architecture is used as the text encoder. The contrastive language-image pre-trained model is CLIP, and the Vision Transformer is ViT. S2: Similarity-guided multimodal interaction utilizes the correlation between image and text features to select text features to be masked. Text features that are highly correlated with image features are more likely to be masked. This strategy encourages the use of image information in the MLM process, thereby enhancing cross-modal interaction between images and text. S3: The mismatch feature enhancement module is used to enhance the mismatch feature, which aims to reduce the false positive rate and improve the overall retrieval accuracy. The mismatch feature enhancement module is the MFE. The MFE module operates in two different modes: mismatched image feature extraction and mismatched text feature extraction. Each mode is equipped with a customized information aggregation strategy. The mismatched image feature extraction is UIE, the mismatched text feature extraction is UTE, and the information aggregation strategy is IG strategy; The IG strategy is a key component of UIE and UTE patterns, used to address the challenge of identifying mismatched features in high-dimensional space; The process for UTE is similar to that for UIE. The difference between UTE and UIE is that UTE operates from the text perspective, focusing on text features that do not match any image features. Furthermore, the IG strategy in UTE aggregates image tags in a manner similar to that of UIE. The MFE in the UIE mode identifies image features that do not match any text features in a given pair; The MFE application in the UIE mode includes at least the following steps: After aggregating L text tags into K text tags using the IG strategy, the similarity between each image tag and the K text tags is calculated. UIE's IG strategy applies parameterless adaptive average pooling to text tagging. Aggregate into a more compact representation ,in ; That is: text tags are divided into Group, No. Indexing of text features within a group This means that multiple text features within each group are processed to form a single aggregated feature. Recognizing that average pooling leads to significant information loss, a single-layer neural network is used for projection and GELU activation before pooling. After pooling, layer normalization is applied to the output to ensure consistent feature scale. The IG strategy is expressed as formula (5): ; in Indicates the first The set of text features in the group It is the first The set of indexes for a group These are the features after aggregation; Subsequently, text features and visual features The similarity is calculated by projecting the image onto the common feature space using a 1x1 convolution, and then expressed as formula (6): ; in , , yes and Cosine similarity between them Indicates the first The image tag and the first Similarity between text tags; When an image feature is related to all If the similarity of all text features is less than 0, then it is identified as an image feature that does not match any text features. To highlight the degree of mismatch, the negative similarity scores of this image feature and all text features are added together, expressed as formula (7): ; in Indicates the first The cumulative mismatch score of each image feature, summed to include only negative similarity values ​​to emphasize the mismatch, is a scalar. copied Secondary formation , so that it is with Dimensional alignment; Subsequently, this mismatch score Used as with The visual features are obtained by multiplying the weights element-wise. It emphasizes the degree of mismatch with all text tags, expressed as formula (8): ; Using the same method, identify all image tags that do not match any text features, forming... ,in This indicates the number of such identified image features; For each image-text pair, to emphasize the impact of mismatch features, these image features are computed. Text features The similarity between them, the average result, is expressed as formula (9): ; G represents the average similarity score between image and text features when they do not match; G represents the number of image features identified as not matching the text; L is the number of text features. This represents the i-th image feature that is identified as not matching the text. With the j-th text feature Cosine similarity; Then it is combined with global image-text similarity; for For each image-text pair, the intra-pair similarity is calculated by the cosine similarity of global features and... Jointly determined, and the similarity between, i.e., the first The image and the first Between texts, The similarity is determined solely by the cosine similarity of global features, as expressed in formulas (10) and (11): ; ; in and They represent the first The image and the first Global features of the text, from the first The image to the first The similarity of the texts is denoted as . To measure the correspondence between visual elements and text descriptions, a similar method can be used to obtain... That is, the first The text to the first Similarity between images; The resulting hybrid similarity is used to calculate the image-text contrast loss, also known as the ITC loss. The ITC loss is a method for training multimodal models that optimizes the alignment between matching image-text pairs by maximizing the similarity between matching image-text pairs while minimizing the similarity between mismatched pairs. In the final step of the MFE module, the similarity scores calculated in previous steps are used to calculate the ITC loss to enhance the effect of mismatched features; By emphasizing these contrasting relationships, the ITC loss significantly enhances the model's ability to distinguish visually similar individuals. The detailed calculation is shown in the following formula group (12): ; ; ; in It is a temperature hyperparameter; S4: Training and inference are performed using a multi-task learning approach, optimizing four loss functions in four parallel components. These four parallel components include an SGMI module, an MFE module, an ID loss, and an SDM loss. The SGMI module calculates the MLM loss, which is the sum of the cross-entropy between the masked text tokens and their labels. The MFE module includes the ITC loss. Meanwhile, the ID loss and SDM loss are calculated independently. The ID loss is used to group individuals based on their identities, ensuring matching at the identity level. The SDM loss is used to minimize image-text similarity. Distribution and matching tags The KL divergence between the normalized distributions is used to calculate a similarity metric for each text query and image candidate pair during inference. This calculation uses the cosine similarity method to quantify the correlation between text and image features, and then uses... To build a ranked list of individuals in the image library, retrieve and present the highest-ranked image that best matches a given text description.

2. The pedestrian re-identification method based on similarity guidance and mismatch feature enhancement according to claim 1, characterized in that: S1 includes at least the following steps: S1.1: Image feature extraction, using CLIP pre-trained VIT from the input image Image features are obtained from the image, where , and These represent the image's height, width, and number of channels, respectively. Divided into A non-overlapping patch sequence, Indicates the patch size; These patches are mapped to 1D labels using a trainable linear projection. ; To represent the feature sequence of the entire image; Indicates the first The feature vector of each patch; N represents the number of feature sequences; After injection location embedding and additional category markers, the marker sequence Input is fed into the transformer block to simulate the relevance of each patch; Used as a global image representation, denoted as ; S1.2: Text feature extraction, for text input CLIP's text encoder is used to obtain the text representation; Processing begins with tokenizing the lowercase input text using byte-pair encoding, using a vocabulary of 49,152 tokens; The text description is encapsulated within [SOS] and [EOS] tags to indicate the start and end of the sequence; Subsequently, the segmented text The data is input into the transformer to analyze the relationships between the markers; The highest level [EOS] token of the transformer Treated as a global text representation .

3. The pedestrian re-identification method based on similarity guidance and mismatch feature enhancement according to claim 2, characterized in that: S2 includes at least the following steps: First, calculate the text features. Its corresponding global image features The similarity between them is then normalized using the softmax function, as shown in formula (1): ; in Indicates the length of the text tag. express and Cosine similarity between them It's a temperature over-parameter. The sampling probability of text markers used as a predetermined proportion of the mask; It is an abbreviation for exponential function. Calculate the input exponent value, here. Map each similarity score to a positive value; It is the summation symbol. The index variable in the table is used to represent values ​​from 1 to... The sum of the exponential similarity scores between all feature pairs; L represents the number of text features; Using the probability mask text features in formula (1), we obtain the text features of the partial mask, denoted as... ; and image features They are fed together into a multimodal interactive encoder, which consists of a multi-head cross-attention layer and four transformer blocks; To effectively integrate image and text information, partial masking of text features As a query Image features As a key Sum In a multi-headed cross-attention layer; This interaction is implemented as shown in formula (2): ; in This represents a masked text representation that incorporates image information. It is a set of mask text marker positions. Indicates the dimension of the embedding; In this context, T usually refers to the matrix transpose operation, which transforms the matrix into its transpose. After transposing, make With query matrix The dimensions are matched to calculate the dot product of the two and generate the attention weight matrix. For each mask position The multilayer perceptron classifier predicts the probability distribution of the corresponding original label for the MLP, and is calculated as follows (3): ; in Indicates the first in the vocabulary list The first mark in the... The predicted probability of each mask position; This represents the masked text feature vector after incorporating image information, referring to the text feature at mask position i after processing by a multimodal interactive encoder. The vocabulary used here is as follows It is the same as that used in text encoding, consisting of 49,152 tags; In the final stage of the SGMI module, the MLM loss is calculated based on the prediction of masked text tags. This loss is designed to optimize the model so that text and visual features are aligned at a fine-grained level. The MLM loss is the sum of the cross-entropy between the masked text tag and its label, defined by the following formula (4): ; in The correct label is marked as 1, and the others are marked as 0.

4. The pedestrian re-identification method based on similarity guidance and mismatch feature enhancement according to claim 3, characterized in that: The application of the ID loss is as follows: ID loss is used to group individuals based on their identities to ensure matching at the identity level; For image-text pairs Cross-entropy loss is applied to global features. ; ID loss The definition is expressed as formula (13): ; in This represents the parameters of the fully connected layer used for classification. This represents the one-hot encoded vector of the real label.

5. The pedestrian re-identification method based on similarity guidance and mismatch feature enhancement according to claim 4, characterized in that: The application of the SDM loss is as follows: The SDM loss is used to minimize image-text similarity. Distribution and matching tags The KL divergence between the normalized distributions is shown in Equations (14) and (15) below: ; ; The peak value of the probability distribution is determined by the temperature hyperparameter. control, It is a true match tag. express They are positive pairs from the same identity, while 0 represents a negative pair; The SDM loss is expressed as formula (16): ; in It is a decimal number, used to avoid numerical problems; Total training loss It is the weighted sum of these losses: ; in and It is a weighting factor used to balance the contributions of ITC loss and MLM loss.