Text-image pedestrian re-identification method based on sample screening and re-matching

Through cross-modal feature extraction, clean sample screening and noise sample rematch, the problem of noise samples in text-image pedestrian re-identification is solved, and the training efficiency and performance of the model is improved.

CN120339950APending Publication Date: 2025-07-18SICHUAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510430583.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing text-image pedestrian re-identification method fails to effectively process noise samples, resulting in a degradation of model performance and the common loss function fails to rationally utilize data set features.

Method used

Features of images and text are extracted across modal feature extractors, clean and noisy samples are separated using a clean sample filtering module, the noise sample rematch module converts part of the noise samples into clean samples, and trains on clean samples using a weighted contrast loss function.

Benefits of technology

Effectively alleviate the impact of noise samples on the model, expand the number of training samples, and improve the training efficiency and performance of the model in a noisy environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339950A_ABST
    Figure CN120339950A_ABST
Patent Text Reader

Abstract

The invention discloses a text-image pedestrian re-identification method based on sample screening and re-matching, and relates to the field of computer vision and artificial intelligence. The method comprises the following steps: firstly, respectively extracting coarse-grained features and fine-grained features described by an image and a text through a cross-modal feature extractor; secondly, designing a clean sample screening module, and dividing samples in the training data set into clean samples and noise samples through a loss comparison method; in addition, part of noise samples are converted into clean samples through a noise sample rematching module; and finally, through weighted comparison loss, more negative sample information is utilized, and a larger weight is given to a difficult sample, so that the training efficiency is improved in an environment with a noise sample. The method is mainly applied to the fields of pedestrian tracking, public safety and the like, and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a text-image pedestrian re-identification method based on sample screening and re-matching, belonging to the fields of computer vision and multimodal retrieval. Background Art

[0002] Image pedestrian re-identification aims to identify the same target pedestrian under different surveillance camera views, which is a task of great practical significance and challenge. In recent years, pedestrian re-identification has become a research hotspot in the field of computer vision due to its practicality in fields such as pedestrian tracking and public safety. Although deep learning-based image pedestrian re-identification methods have made significant progress in scenarios such as clothing changes and partial occlusions, and have shown performance close to or even exceeding the human level in many public benchmark tests, in some scenarios, people can only give a text description of the target person. Therefore, text-image pedestrian re-identification (TBPR) has gradually received extensive attention.

[0003] In the field of TBPR, due to problems such as low resolution of some images and easy negligence in manual annotation, there are a large number of noise samples in the dataset where the pictures and text descriptions do not match completely. Most existing methods ignore this problem, resulting in a significant decline in model performance. In the field of image-text matching, the problem of noisy descriptions has been studied to some extent, and various solutions have been proposed. For example, the Gaussian Mixture Model (GMM) is often used to model the loss value of samples and classify the samples into clean samples and noise samples. Recently, some researchers have proposed the Beta Mixture Model (BMM), which shows more superior performance in noise sample classification. In addition, some methods screen out clean samples by evaluating the impact of sample training on the model's ability, thus avoiding the model being affected by noise samples.

[0004] Although the existing methods for the problem of noise samples in the field of image-text matching can provide certain references. However, these methods fail to reasonably utilize the characteristics of the dataset in the field of TBPR. In addition, when dealing with negative samples, the commonly used loss functions (such as Triplet Loss and E-LOSS) do not consider the existence of noise samples, further reducing the model performance. Summary of the Invention

[0005] In order to solve the problem of noise samples in the field of TBPR, the present invention proposes a text-image pedestrian re-identification method based on sample screening and re-matching, aiming to alleviate the problem that noise samples in the training dataset affect the model performance.

[0006] The technical solution of the present invention includes the following steps:

[0007] (1) Through a Cross-Modal Feature Extractor (CMFE), token sequences of pictures and text descriptions are extracted respectively, so as to obtain the coarse-grained and fine-grained features of images and texts;

[0008] (2) Through a Clean Sample Selection Module (CSSM), the samples in the training dataset are divided into clean samples and noise samples. Based on the memory effect of a Deep Neural Network (DNN), that is, the loss of clean samples is usually less than that of noise samples, the samples in the training dataset are divided into two categories: clean samples and noise samples by comparing the loss values of similar samples;

[0009] (3) Through a Noisy Sample Rematching Module (NSRM), some of the noise samples filtered out in step (2) are re-converted into clean samples. In the dataset in the TBPR field, the same person contains many pictures and their corresponding text descriptions. The text description can not only describe its original picture, but also describe other pictures of the same person. Therefore, replacing the text description of the noise sample with other more appropriate text descriptions of the same person can reconstitute clean samples, thus expanding the number of clean samples in the training stage;

[0010] (4) Training is carried out on clean samples through a Weighted topK Contrastive Loss (WKCL). This loss function takes into account more negative sample information and assigns a greater weight to difficult samples, thereby effectively improving the model training efficiency in an environment with noise samples.

[0011] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0012] 1. The present invention divides the samples in the training dataset into clean samples and noise samples by introducing a clean sample selection module;

[0013] 2. The present invention effectively expands the number of available training samples by introducing a noisy sample rematching module to re-convert some noise samples into clean samples;

[0014] 3. The present invention improves the training efficiency of the model in an environment with noise samples by introducing a weighted contrast loss, making full use of more negative sample information in a batch and assigning a greater weight to difficult samples. Brief Description of the Drawings

[0015] Figure 1 This is the principle block diagram of the text-image pedestrian re-identification method based on sample screening and re-matching of the present invention;

[0016] Figure 2 This is the schematic diagram of the characteristics of the text-image pedestrian re-identification dataset of the present invention. Detailed Embodiments

[0017] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0018] First, the general composition of the dataset in the TBPR field will be described. Assume that the dataset contains N individuals, and the individual to which each sample belongs is represented as n, where n ∈ {1, 2, 3,..., N}. is a certain sample of individual n, I n is the image of individual n, T n is the text description of individual n, and i ∈ {1, 2, 3,..., N n}, N n represents the total number of images of individual n. j ∈ {1, 2}, because each image has 1 or 2 texts corresponding to it, and the specific number depends on the specific dataset. In the following description, for the sake of more intuitiveness and without affecting understanding, the superscript n of the sample is omitted in some cases in this patent of the present invention. of the superscript n.

[0019] As Figure 1 shown, a text-image pedestrian re-identification method based on sample screening and re-matching includes the following steps:

[0020] (1) For a certain sample (I i , T i,j ), through the cross-modal feature extractor, extract the features of its image and text respectively, and obtain the token sequence of image I i and the token sequence of text T i,j . By processing the token sequences, the coarse-grained and fine-grained features of the image and text are obtained.

[0021] (2) Based on the features obtained in step (1), calculate the loss L i of each sample (I i,j , T wkcl )(I i , T i,j). Input the loss of all samples into the Gaussian mixture model for fitting, and get the probability of each sample being a clean sample. Then select a batch of samples with a high probability of being clean samples to form a clean sample reference set D s (CleanReference Set). The remaining samples constitute set D w , where the samples need to be further judged as clean samples or noise samples.

[0022] (3) For D w Sample (I k ,T k,j ), first calculate I k With D s The similarity of all images in the image is obtained, and the image with the largest similarity and its corresponding text (I m ,T m,j ). Then, through the loss proposed in step (5), we can calculate (I k ,T k,j ) loss L wkcl (I k ,T k,j ) and (I m ,T m,j ) loss L wkcl (I m ,T m,j ), by comparing the losses, we can judge (I k ,T k,j ) is a clean sample or a noise sample. Step (2) and step (3) constitute the CSSM module.

[0023] (4) According to the data set partitioning method in step (3), a portion of the noise samples are converted into clean samples. Figure 2 As shown in the figure, in the dataset of person re-identification, a text description can describe other images of the person to which the corresponding image belongs. The design of the NSRM module is based on this feature of the dataset. For the image of the noise sample, the similarity between the text description of other clean samples of the person to which it belongs and the image of the noise sample is calculated, and the text description with the highest similarity is selected for replacement. This replacement operation is performed according to a certain ratio.

[0024] (5) After step (4), the dynamic reconstruction of the data set is completed. i ,T i,j ) can be used to calculate the loss L wkcl (I i ,T i,j ), the model is obtained by L wkcl (I i ,T i,j ) is trained on clean samples.

[0025] The detailed steps are as follows:

[0026] Step (1): Extract the features of its images and texts through a cross-modal feature extractor.

[0027] Input all image-text pairs into CLIPViT-B / 16, i.e., the cross-modal feature extractor, to obtain the token sequence of the image and the token sequence of the text. Taking the sample (I i ,T i,j ) as an example, the token sequence of the image is where represents the global feature of the image, represents the local feature of the image. The token sequence of the text is where represents the global feature of the text, represents the local feature of the text. Then, calculate the cosine similarity between each token in the local feature token sequence and the global feature. The cosine similarity calculation method can be expressed as:

[0028]

[0029] In the token sequences of the local features of the image and the text, select the 30% of the token sequences with the largest cosine similarity to the global feature to form the effective feature token sequence and Through the processing of and , further obtain the fine-grained features of the image and the text and The specific processing method can be expressed as:

[0030]

[0031] where MP represents the max pooling operation, MLP represents the multi-layer perceptron, and FC represents the fully connected layer.

[0032] Step (2): Establish a clean sample reference set D s .

[0033] For a certain sample (I i ,T i,j ), based on the features obtained in step (1), through the loss function calculation method proposed in step (5), obtain the loss value L of this sample wkcl (I i ,T i,j). Input the loss of all samples into the Gaussian mixture model for fitting, and get the probability that each sample belongs to the clean sample. For example, for sample (I i ,T i,j ), we can get the probability p that the sample is a clean sample i,j , the probability of other samples belonging to clean samples is expressed in a similar way. Next, setting the threshold σ = 0.9, the overall data set D can be divided into D s and D w Among them, D s ={(I i ,T i,j )|p i,j >σ},D w ={(I i ,T i,j )|p i,j ≤σ}. Clean sample reference set D s The probability that the samples contained in are clean samples is greater than 0.9, and D w The probability that the samples in are clean samples is less than 0.9, but they are not all noise samples and need further judgment.

[0034] Step (3): Create a clean sample set D c .

[0035] Based on the division method in step (2), it is considered that For D w A sample of k ,T k,j ), first calculate D s The images and I of all samples in k The cosine similarity between them is used to find the image and its text with the greatest similarity (I m ,T m,j ). By the calculation method proposed in step (5), (I k ,T k,j ) and Calculate the coarse-grained loss Fine-grained features and Calculate the fine-grained loss Similarly, the sample (I m ,T m,j ) and fine-grained loss Next, calculate and They are used to characterize the samples with coarse-grained loss and fine-grained loss (I k ,T k,jWhether it is a clean sample or a noise sample. and are calculated as follows:

[0036]

[0037] where u is a hyperparameter. After experimental verification, the model performance is optimal when u = 0.05. If and are both 1, then the sample (I k , T k,j ) is considered a clean sample; if and are both 0, then the sample (I k , T k,j ) is considered a noise sample; if and one is 1 and the other is 0, the sample (I k , T k,j ) is determined to be a clean sample with a 50% probability. All clean samples form the set D c , and all noise samples form the set D n .

[0038] Step (4): Convert some of the noise samples in D n into clean samples.

[0039] As Figure 2 shown, the datasets in the pedestrian re-identification field have significant characteristics: the same person has multiple images and text descriptions, and each text description can not only describe the original image it belongs to, but also other images of the same person. The NSRM module utilizes this feature of the dataset to convert some of the noise samples into clean samples. For the noise sample of a certain person n, calculate the cosine similarity between the text descriptions of all clean samples of person n and the image of the noise sample, and select the text description with the largest similarity to replace to reconstitute the clean sample . This replacement operation is carried out at a ratio of 15%.

[0040] Step (5): Train the model through WKCL.

[0041] After step (4), the set of clean samples is obtained, and the model will be trained on the clean samples through WKCL. For the clean sample (I i , T i,j ), its loss L wkcl (I i , T i,j ) is calculated as follows:

[0042]

[0043] Among them is the coarse-grained loss of the sample (I i , T i,j ). is the fine-grained loss of the sample (I i , T i,j ). and The specific calculation formulas are as follows:

[0044]

[0045]

[0046] Among them and represent the similarity between positive samples calculated by coarse-grained features,

[0047] and represent the similarity between positive samples calculated by fine-grained features, and represent the weighted similarity of negative samples calculated by coarse-grained features. and represent the weighted similarity of negative samples calculated by fine-grained features. For First, in the same batch, calculate the 80% text descriptions with the highest similarity to I i . These text descriptions and I i constitute hard negative samples. Among these hard negative samples, calculate through the following formula:

[0048]

[0049] Among them, is the coarse-grained feature of image I i , is the coarse-grained feature of text T k,j . and The calculation process and definition of are similar:

[0050]

[0051] To verify the effectiveness of the method of the present invention, the present invention is verified on the commonly used datasets CUHK-PEDES, ICFG-PEDES, and RSTPReid in the field of text-image person re-identification. Four deep learning-based text-image person re-identification methods are selected as comparison methods, specifically:

[0052] Method 1: CFine proposed by Yan et al., reference: "S.Yan, N.Dong, L.Zhang and J.Tang, CLIP-Driven Fine-Grained Text-Image Person Re-Identification, in IEEE Transactions on Image Processing, pp.6032-6046, 2023."

[0053] Method 2: IRRA proposed by Ding et al., reference: "D.Jiang, and M.Ye. Cross-Modal Implicit Relation Reasoning and Aligning for Text-to-Image Person Retrieval, in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp.2787-2797, 2023."

[0054] Method 3: CFAM proposed by Zuo et al., reference: "J.Zuo, H.Zhou, Y.Nie, F.Zhang, T.Guo, N.Sang, UFineBench: Towards Text-based Person Retrieval with Ultra-fine Granularity, in Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp.22010-22019, 2024."

[0055] Method 4: TBPS-CLIP proposed by Cao et al., reference "M. Cao, Y. Bai, Z. Zeng, M. Ye, and M. Zhang. An Empirical Study of CLIP for Text-based Person Search, in Proceedings of the AAAI Conference on Artificial Intelligence, pp. 465-473, 2024."

[0056] As shown in Table 1, the method proposed in the present invention uses Rank1, Rank5, Rank10, and mAP as evaluation indicators. Compared with the other four methods, the performance of the present invention on the three datasets has significant advantages.

[0057] Table 1 Comparison of Rank1, Rank5, Rank10, and mAP indicators with other methods (%)

[0058]

[0059] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A text-image pedestrian re-identification method based on sample screening and re-matching, characterized in that Including the following steps: (1) Through a Cross-Modal Feature Extractor (CMFE), token sequences of pictures and text descriptions are extracted respectively, so as to obtain the coarse-grained and fine-grained features of images and texts; (2) Through a Clean Sample Selection Module (CSSM), the samples in the training dataset are divided into clean samples and noisy samples; (3) Through a Noisy Sample Rematching Module (NSRM), some of the noisy samples filtered out in step (2) are re-converted into clean samples; (4) Training is carried out on clean samples through a Weighted topK Contrastive Loss (WKCL).

2. The text-image pedestrian re-identification method based on sample screening and re-matching according to claim 1, characterized in that, The cross-modal feature extractor CMFE in step (1); Input all image-text pairs into CLIP ViT-B / 16, i.e., CMFE, to obtain the token sequences of the images and the token sequences of the texts; for the sample (I i , T i,j ), the image token sequence is where represents the global feature of the image, represents the local feature of the image; the text token sequence is where represents the global feature of the text, represents the local feature of the text.

3. The cross-modal feature extractor CMFE according to claim 2, wherein Calculate the cosine similarity between each token in the local feature token sequence and the global feature, and the cosine similarity calculation method can be expressed as: Among the token sequences of the local features of the image and text of the sample (I i , T i,j ), select the 30% of the token sequences with the highest cosine similarity to the global feature to form the effective feature token sequence and By processing and , the fine-grained features of the image and text are further obtained and The specific processing method can be expressed as: Where MP represents the max pooling operation, MLP represents the multi-layer perceptron, and FC represents the fully connected layer.

4. A text-image pedestrian re-identification method based on sample screening and re-matching according to claim 1, characterized in that In the step (2), D is established through CSSM c ; for the sample (I i , T i,j ), based on the features obtained in step (1), through the loss function calculation method proposed in step (4), the loss value L wkcl (I i , T i,j ) of this sample is obtained; the losses of all samples are input into the Gaussian mixture model for fitting to obtain the probability that each sample belongs to a clean sample; for the sample (I i , T i,j ), the probability p i,j that this sample is a clean sample can be obtained, and the representation methods of the probabilities that other samples belong to clean samples are similar.

5. The method for establishing D by CSSM according to claim 4 c , characterized in that First, establish a clean sample reference set D s ; By setting the threshold σ = 0.9, the overall dataset D can be divided into D s and D w ; where D s ={(I i , T i,j ) | p i,j > σ}, D w ={(I i , T i,j ) | p i,j ≤ σ}; the probability that the samples included in the clean sample reference set D s are clean samples is greater than 0.9, while the probability that the samples in D w are clean samples is less than 0.9, but not all of them are noise samples and further determination is needed.

6. The method for establishing D according to claim 4 c , characterized in that At D s and D w Based on this, D is established through a loss comparison method c ; For D w For one sample (I k , T k,j ) in s , first calculate the cosine similarity between the images of all samples in D k and I m , T m,j ), and find the image with the maximum similarity and its text (I k , T k,j ); through the calculation method proposed in step (4), from the coarse-grained features of (I and ), calculate the coarse-grained loss Calculate the fine-grained loss from the fine-grained features and Similarly, calculate the coarse-grained loss m and fine-grained loss m,j of the sample (I and Then, calculate and respectively used to characterize whether the sample (I k , T k,j ) is a clean sample or a noise sample on the coarse-grained loss and the fine-grained loss; and The calculation methods are as follows:​ where u is a hyperparameter, and through experimental verification, the model performance is optimal when u = 0.05; if and are both 1, then the sample (I k , T k,j ) is considered a clean sample; if and are both 0, then the sample (I k , T k,j ) is considered a noise sample; if and one is 1 and the other is 0, then with a 50% probability, the sample (I k , T k,j ) is determined to be a clean sample; all clean samples form the set D c , and all noise samples form the set D n .

7. A text-image pedestrian re-identification method based on sample screening and re-matching according to claim 1, characterized in that The noise sample re-matching module in step (3); for the noise samples of a certain person n Calculate the text descriptions of all clean samples of person n and the images of the noise samples of the cosine similarity, and select the text description with the largest similarity Replace Reconstitute the clean samples This replacement operation is carried out at a ratio of 15%.

8. A text-image pedestrian re-identification method based on sample screening and re-matching according to claim 1, characterized in that, The weighted contrastive loss function in step (4); For clean samples (I i , T i,j ), the calculation method of its loss L wkcl (I i , T i,j ) is as follows: Among them is the coarse-grained loss of the sample (I i , T i,j ); is the fine-grained loss of the sample (I i , T i,j ); and The specific calculation formulas are as follows: Among them and represent the similarity between positive samples for coarse-grained feature calculation, and represent the similarity between positive samples for fine-grained feature calculation, and represent the weighted similarity of negative samples for coarse-grained feature calculation; and represent the weighted similarity of negative samples for fine-grained feature calculation; For firstly, within the same batch, calculate the 80% text descriptions with the highest similarity to I i ; These text descriptions and I i constitute hard negative samples; Calculated through the following formula among these hard negative samples: Among them, is the coarse-grained feature of image I i , is the coarse-grained feature of text T k,j ; and The calculation process and definition of are similar to

Citation Information

Cited By

  • Remote sensing image change detection method

    CN121811266A

  • A change detection method of remote sensing image

    CN121811266B