LLM-based cross-modal pedestrian re-identification data enhancement method

Through the cross-modal pedestrian re-identification data enhancement method based on LLM, high-quality text descriptions are generated and feature alignment is performed, which solves the problems of insufficient text annotation and lack of optimization of data enhancement methods, and achieves efficient expansion of the data set and improved model performance.

CN120356032APending Publication Date: 2025-07-22JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510427064.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

There are problems in the lack of high-quality text annotations and short text descriptions that affect model performance in existing text-to-image pedestrian retrieval. The existing data augmentation methods lack optimization strategies for task characteristics, resulting in limited improvement effects.

Method used

A cross-modal pedestrian re-identification data enhancement method based on LLM is adopted, and a new text description is generated by preprocessing the original text description and inputting the LLM. It combines the image filtering mechanism for threshold filtering, uses neural networks for feature learning and projecting to public space, and uses contrast learning loss to achieve image and text feature alignment.

Benefits of technology

It realizes efficient expansion of data sets, generates high-quality text descriptions, solves the LLM generation illusion problem, and improves the accuracy and consistency of cross-modal pedestrian retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356032A_ABST
    Figure CN120356032A_ABST
Patent Text Reader

Abstract

The invention provides a cross-modal pedestrian re-identification data enhancement method based on LLM, and the method comprises the steps: an original image-text pair is included in an existing pedestrian re-identification data set, and an original text description in the preprocessed original image-text pair and a prompt word Prompt are inputted into the LLM to generate a new text description; an IFM is introduced to carry out threshold filtering on the generated text description, the generated text description obtained for the Nmax-th time and the generated text description meeting the threshold requirement are reserved, and the finally stored generated text description is obtained; and carrying out feature learning on the original image-text pair in the data set and the finally stored generated text description by using a neural network, projecting the obtained image global feature and the text description global feature into a public space, and aligning the image global feature and the text description global feature by comparing learning loss. According to the method, the data set can be efficiently expanded with high quality, and the generation illusion problem of the LLM is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pedestrian re-identification methods, and particularly to a cross-modal pedestrian re-identification data augmentation method based on a Large Language Model (LLM). Background Art

[0002] Text-to-image pedestrian retrieval is a technology for querying pedestrian images in an image database based on a text description. It can help users find pedestrian images that match the description by inputting a text description, so as to quickly locate the required images in a large-scale image database. This technology plays an important role in many practical applications.

[0003] In practical applications, text-to-image pedestrian retrieval faces many challenges and problems. One of its main problems is the lack of high-quality text annotations. During the text annotation process, due to its complexity and tediousness, it is inevitable to introduce the subjective bias of annotators, resulting in the generated text description being unable to fully and objectively reflect the true characteristics of the target person. In addition, due to time and cost limitations, the text annotations in existing pedestrian re-identification datasets are usually relatively short, often only capturing partial feature information of the target person. The limitations of this text description directly affect the performance of the model in cross-modal retrieval tasks and cannot comprehensively and accurately match the image of the target person with the text description. Therefore, improving the quality of text annotations and generating more detailed and objective text descriptions is the key to enhancing the effect of cross-modal pedestrian retrieval.

[0004] In the prior art, constructing a large-scale multi-attribute dataset is a common means to improve the model performance. However, simply relying on the way of constructing a large-scale dataset usually requires a large amount of time, computing resources, and manual annotation, and it is difficult to meet the personalized needs of specific tasks. Therefore, data augmentation, as an effective supplementary strategy, has become another method to improve the model training effect. Compared with dataset construction, data augmentation can significantly reduce costs and reduce the dependence on large-scale datasets while expanding the data scale. However, existing data augmentation methods usually focus on simple geometric transformation or noise processing of the original data, lacking strategies optimized for specific tasks, resulting in limited improvement effects in some scenarios. Therefore, how to design an efficient augmentation method for task characteristics while ensuring data diversity is still an important challenge in the current technology. Summary of the Invention

[0005] Aiming at the deficiencies in the prior art, the present invention provides a cross-modal pedestrian re-identification data augmentation method based on LLM.

[0006] The present invention achieves the above technical objectives through the following technical means.

[0007] LLM-based Cross-modal Person Re-identification Data Augmentation Method, including:

[0008] Step 1: There are a large number of original image-text pairs in the existing person re-identification dataset. The original image-text pairs include images and corresponding original text descriptions. Preprocess the original text descriptions, and input the preprocessed original text descriptions and the prompt "Prompt" into the LLM to generate new text descriptions, that is, generate text descriptions;

[0009] Step 2: Introduce an image-based screening mechanism to perform threshold filtering on the generated text descriptions, and regenerate the generated text descriptions that do not meet the threshold requirements according to the method in Step 1, and then perform threshold filtering on the regenerated text descriptions. If after the N_maxth iteration of regeneration, the regenerated text descriptions still do not meet the threshold requirements, then retain the generated text descriptions obtained in the N_maxth time together with the generated text descriptions that have met the threshold requirements to obtain the finally saved generated text descriptions;

[0010] Step 3: Use a neural network to perform feature learning on the original image-text pairs in the dataset in Step 1 and the generated text descriptions finally saved in Step 2 to obtain the global image features and global text description features, project them into a common space, and finally use the contrastive learning loss to align the global image features and global text description features.

[0011] Further, the process of the preprocessing is:

[0012] For the data where each image contains multiple original text descriptions, splice the multiple original text descriptions and insert a separator to form a spliced original text description; for the data where each image contains only one original text description, use this original text description as the spliced original text description.

[0013] Further, the LLM adopts the General Language Model - 4 - 9B - Chat.

[0014] Further, the image-based screening mechanism adopts the BLIP model.

[0015] Furthermore, the process of performing threshold filtering on the generated text descriptions is: First, use the BLIP model to calculate the matching score itm_score and the cosine similarity itc_score for each image and its corresponding original text description in the person re-identification dataset, average the matching scores itm_score and the cosine similarities itc_score of all original image-text pairs respectively to obtain two average scores, and use them as the threshold tuple S t; Then, using the BLIP model, calculate the matching score \(T_{\sim itm\_score}\) and cosine similarity \(T_{\sim itc\_score}\) for the generated text description and the corresponding image in Step 1, and take them as the score tuple \(S\); Finally, compare the score tuple \(S\) with the threshold tuple \(S\) t for comparison. If the values in \(S\) are all greater than the values at the corresponding positions in \(S\) t , the generated text description is considered to meet the threshold requirements; otherwise, it is considered that the generated text description does not meet the threshold requirements.

[0016] Furthermore, the global image features are extracted using the Vision Transformer in the CLIP model, and the global text description features are extracted using the text encoder in the CLIP model.

[0017] Even further, the process of using the contrastive learning loss to align the global image features and global text description features is as follows:

[0018] First, use the contrastive learning loss to calculate the loss \(L\) orig between the image and the original text description, and the loss \(L\) aug between the image and the generated text description; Then, combine the two parts of the losses \(L\) orig and \(L\) aug to obtain the total loss \(L\) total = \(L\) orig + \(L\) aug ; Finally, guide the CLIP model to learn a more accurate matching relationship between the image and the text description through the total loss \(L\) total so as to achieve the alignment of the global image features and global text description features.

[0019] Even further, the calculation formula of the total loss \(L\) total is as follows:

[0020]

[0021] where \(L\) I2T is the loss from the image \(I\) to the original text description \(T\), \(L\) T2I is the loss from the original text description \(T\) to the image \(I\), \(L\) I2T~ is the loss from the image \(I\) to the generated text description \(\widetilde{T}\), and \(L\) T~2I is the loss from the generated text description \(\widetilde{T}\) to the image \(I\).

[0022] Even further, the calculation formulas for the loss \(L\) I2T from the image \(I\) to the original text description \(T\) and the loss \(L\) T2I from the original text description \(T\) to the image \(I\) are:

[0023]

[0024] Among them, N represents the batch size; s(I i ,T i ) represents the similarity between the image I i and the corresponding original text description T i ; v I and v T are the embedding vectors of the image I i and the original text description T i respectively. The dot product v I ·v T represents the similarity between the two, and the denominator is the norm of the embedding vector; exp(.) is the exponential function; τ is the temperature parameter; s(T i ,I i ) represents the similarity between the corresponding original text description T i and the image I i , and the calculation formula is the same as that of s(I i ,T i );

[0025] Furthermore, the loss L aug between the image and the generated text description is calculated in a similar way to the loss L orig between the image and the original text description, with the only difference being that the original text description T i corresponding to the image is replaced by the generated text description T i ~.

[0026] The beneficial effects of the present invention are as follows:

[0027] (1) The dataset expansion is efficient and of high quality: The present invention uses the large language model LLM to regenerate text descriptions, which can increase the diversity of vocabulary and sentence structures while maintaining the original key concepts and semantic information. Compared with traditional dataset construction methods, the present invention can expand the dataset more efficiently and concisely while maintaining high-quality enhanced data.

[0028] (2) Solve the generation hallucination problem of LLM: The present invention introduces an image-based screening mechanism. By calculating the ITC and ITM scores of the generated text and the corresponding original image, the generated text that does not meet the set threshold is screened and filtered out, thus alleviating the generation hallucination problem of LLM and ensuring the quality and consistency of the generated text. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 is a flowchart of the cross-modal pedestrian re-identification data enhancement method based on LLM according to the present invention.

[0030] Figure 2This is a flowchart for generating a new text description from the original text description described in the present invention using an LLM.

[0031] Figure 3 This is a flowchart for filtering the generated text description by the image-based screening mechanism described in the present invention. Detailed implementation manners

[0032] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but the protection scope of the present invention is not limited thereto.

[0033] As Figure 1 shown, the present invention provides a cross-modal pedestrian re-identification data augmentation method based on an LLM, including the following steps:

[0034] Step 1. In the existing pedestrian re-identification dataset, there are a large number of original image-text pairs, and the original image-text pairs include images and corresponding original text descriptions. Preprocess the original text description, and input the preprocessed original text description and the prompt Prompt into the LLM to generate a new text description, that is, generate the text description.

[0035] As a preferred embodiment of the present invention, as Figure 2 shown, Step 1 specifically includes:

[0036] Step 1.1. The preprocessing of the original text description is specifically as follows: (1) For the data where each image contains multiple original text descriptions. First, concatenate the multiple original text descriptions T=(T 1 ,T 2 …) corresponding to each image I together to form a concatenated original text description T ori =Concat(T 1 ,T 2 ,...), where Concat() is the concatenation operation. Secondly, when concatenating the original text descriptions (T 1 ,T 2 …), considering the semantic relationship between the original text descriptions, appropriate connectors need to be inserted to ensure that the concatenated original text description has good readability and coherence. Therefore, the concatenated original text description can be expanded as: T ori =Concat(T 1 ,T 2 ,...)=T 1 +connector+T 2 +..., where the connector is a full stop “.”. (2) For the data where each image I contains only one original text description T, use this original text description as the concatenated original text description, that is, T ori =T.

[0037] Step 1.2. As a specific implementation of the LLM technical route, the General Language Model-4-9B-Chat (GLM-4-9B-Chat) has excellent performance in language understanding and text generation tasks. Therefore, combining with Step 1.1, design the prompt Prompt and use GLM-4-9B-Chat to generate a new text description. The process of generating a new text description is as follows: If an image I contains multiple original text descriptions T = (T 1 , T 2 ,...), obtain the concatenated text description T ori according to Step 1.1, and input the concatenated text description and the prompt Prompt "Integrate the current text into a coherent and logically clear description" into GLM-4-9B-Chat to obtain the generated text description T~; If an image contains only one original text description, directly input the original text description as the concatenated original text description and the prompt Prompt "Rewrite the following text while maintaining the original meaning" into GLM-4-9B-Chat according to Step 1.1 to obtain the generated text description T~. The final formula is as follows:

[0038] T~ = GLM-4-9B-Chat(Concat(T ori , Prompt))

[0039] where T ori is the concatenated original text description; Prompt is a specially designed prompt, the function of which is to instruct GLM-4-9B-Chat to perform a specific operation on T ori ; Concat(T ori , Prompt) is a concatenation operation, that is, concatenate T ori and the prompt Prompt together to form a complete input string, and this string will be used as the input of the GLM-4-9B-Chat model.

[0040] Step 2. Considering that the text descriptions generated by the LLM in Step 1 may deviate from the expectations, a threshold is set, and an Image-based Filtering Mechanism (IFM) is introduced to perform threshold filtering on the generated text descriptions, that is, to evaluate the semantic similarity between the generated text descriptions and the corresponding images, and effectively filter out the generated text descriptions that do not meet the threshold requirements. And regenerate the generated text descriptions that do not meet the threshold requirements according to Step 1, and filter them again. When regenerating, at most N_max iterations of generation are performed. If after the N_maxth iteration, the generated text description still does not meet the threshold requirements, then the generated text description obtained at the N_maxth time and the generated text descriptions that have already met the threshold requirements are retained together to obtain the finally saved generated text descriptions.

[0041] As a preferred embodiment of the present invention, as Figure 3 shown, Step 2 specifically includes:

[0042] Step 2.1. The IFM performs preliminary image and text matching through the Bootstrapping Language-Image Pre-training (BLIP). First, use the BLIP model to calculate the matching score itm_score and cosine similarity itc_score for each image I and its corresponding original text description T in the person re-identification dataset; then, average the matching scores itm_score and cosine similarities itc_score of all original image-text pairs respectively to obtain two average scores, and use them as the threshold tuple S t , which is used for subsequent screening of the generated text descriptions. Also use the BLIP model to calculate the matching score T~_itm_score and cosine similarity T~_itc_score for the generated text description T~ and the corresponding image I in Step 1, and use them as the score tuple S.

[0043] Furthermore, the setting of the threshold S t in Step 2.1 and the calculation of the score S specifically include:

[0044] Step 2.1.1. First, use the image encoder and text encoder of the BLIP model to encode the original image-text pairs in the pedestrian re-identification dataset to obtain image embedding A and text embedding A. Then, input the image embedding A and text embedding A into the cross-attention mechanism to generate a fused multimodal joint feature representation A. Finally, the multimodal joint feature representation A is calculated through the image-text matching loss (Image-Text Matching Loss, ITM) head of the BLIP model to obtain the matching score itm_score between the image and the original text description; the cosine similarity itc_score between the image embedding A and the text embedding A is calculated through the image-text contrast loss (Image-Text Contrastive Loss, ITC) head of the BLIP model. All itm_scores and itc_scores are summarized and averaged to obtain the average matching score itm_average_score and the average cosine similarity itc_average_score. These two averages constitute the tuple S t , and use it as the set threshold for subsequent screening of the generated text description.

[0045] In addition, the matching score itm_score formula is as follows:

[0046] itm_score = E (I,T)~D H(y itm ,p itm (I,T))

[0047] Among them, E (I,T)~D is the expected value on the person re-identification dataset D, which represents the average calculation of the loss for all original image-text pairs (I, T) in the dataset D; H(y, p) is the cross entropy loss function, which measures the difference between the predicted probability p and the true label y; y itm is the true label, represented by a 2-dimensional one-hot vector. If the image I and the original text description T are a matching pair, the label is [1,0]; if they do not match, the label is [0,1]; p itm (I,T) is the predicted probability that the image I and the original text description T are a matching pair calculated by the ITM head.

[0048] The cosine similarity itc_score formula is as follows:

[0049]

[0050] Among them, y i2t (I) is the true label of image I, which is 1 if image I matches the original text description T, otherwise it is 0; t2i(T) is the true label regarding the original text description T, which is 1 if the text T matches the image I, and 0 otherwise; y i2t (I) and y t2i (T) are both represented by one - hot vectors; p i2t (I) is the probability that the predicted image I matches the original text description T, p t2i (T) is the probability that the predicted original text description T matches the image I, both based on the result of softmax normalization. The calculation method of H(y, p) in the itm_score and itc_score formulas is as follows:

[0051]

[0052] Among them, B is the number of classification categories. In the ITM task, B = 2, corresponding to the two categories of "match" and "non - match"; y b is the one - hot encoding of the true label, which is 1 at category b and 0 for the rest; p b is the probability that the model prediction result is the b - th category; log(p b ) is the natural logarithm of the predicted probability p b for calculating the loss.

[0053] Step 2.1.2: First, use the same BLIP model to encode the generated text description T~ and its corresponding image to obtain the image embedding B and the generated text embedding B. Secondly, input the image embedding B and the generated text embedding B into the attention mechanism to generate the fused multi - modal joint feature representation B. Finally, calculate the multi - modal joint feature representation B through the ITM head to obtain the matching score T~_itm_score of the image and the generated text description T~; calculate the cosine similarity T~_itc_score between the image embedding B and the generated text embedding B through the ITC head. Finally, store these two scores in the form of a tuple as S and use it for subsequent screening. In addition, the calculation formula of T~_itm_score can be obtained similarly to the calculation formula of itm_score, and the calculation formula of T~_itc_score can be obtained similarly to the calculation formula of itc_score.

[0054] Step 2.2: Compare the score tuple S calculated in Step 2.1 with the threshold tuple S t If the values in S are all greater than or equal to S tIf the value at the corresponding position in [the relevant data] is such that the generated text description T~ is considered to meet the threshold requirement, then the generated text description T~ is retained; otherwise, it is considered that the generated text description T~ does not meet the threshold requirement. The generated text description T~ that does not meet the threshold requirement is returned to step 1 for regeneration, and the score tuple S is recalculated for the regenerated text description T~. This process is repeated until the generated text description T~ meets the threshold requirement or reaches the preset maximum number of iterations N_max. After reaching the preset maximum number of iterations N_max, if the generated text description T~ still does not meet the threshold requirement, then the text description T~ that does not meet the threshold requirement is retained together with the previously generated text description that meets the threshold requirement and saved as the final result for subsequent training.

[0055] Step 3: Use the corresponding neural network to perform feature learning on the original image-text pairs in the dataset of step 1 and the generated text descriptions finally saved in step 2, project the global features obtained from the two modalities (pedestrian images and text descriptions) into a common space, and finally use the contrastive learning loss to align the global features of the images and the global features of the text descriptions.

[0056] As a preferred embodiment of the present invention, step 3 specifically includes:

[0057] Step 3.1: Feature learning step for pedestrian images: Use the Vision Transformer (ViT) in the Contrastive Language-Image Pretraining (CLIP) model to obtain the global image features. First, split the image I ∈ R H*W*C into M = H*W / P 2 non-overlapping patch sequences of a fixed size, where H represents the height of the image, W represents the width of the image, C represents the number of channels of the image, M represents the number of patches into which the image is split, and P represents the patch size. Then, map the patch sequence to a one-dimensional token sequence through a trainable linear projection By injecting positional embeddings and an additional [CLS] token, the token sequence {f vcls , f v1 , …, f vM} is input into the Transformer to model the correlation of each patch. During this process, the [CLS] token serves as a global token responsible for aggregating the features of the entire image, and f vcls represents the final embedding of the token after passing through the Transformer layer. Finally, use a linear projection to map f vcls to the joint image-text embedding space as the global image representation.

[0058] Step 3.2, Feature Learning Step of Text Description: Use the CLIP model text encoder to extract global text features. First, the text description is tokenized by Byte Pair Encoding (BPE) with a vocabulary of 49,152 lowercase bytes. To clarify the boundaries of the text sequence, the text description is processed through a special token structure {[SOS], text description, [EOS]}, where [SOS] indicates the start of the sequence and [EOS] indicates the end of the sequence. This structure helps the CLIP model correctly identify the boundaries of the text, thus effectively performing feature learning. In particular, the [EOS] token is used to preserve all the information of the text sequence, providing a global representation of the entire text sequence, ensuring that the CLIP model can capture all context information at the end of the sequence. Then, the tokenized text sequence {f tsos , f t1 , …, f tL , f teos} is input into the text encoder for encoding, and the correlation between each token is mined through the masked self-attention mechanism in the encoder. Finally, take the output feature f teos corresponding to the [EOS] token, and map f teos to the image-text joint embedding space through linear projection to obtain the global text representation.

[0059] Step 3.3, First, through Steps 3.1 and 3.2, the global features of the image, the global features of the original text description, and the global features of the generated text description can be obtained, and the three global features are mapped to a unified feature space. Then, using the contrastive learning loss, calculate the loss L orig between the image and the original text description, and the loss L aug between the image and the generated text description. The loss quantifies the matching degree between the image and the text. By optimizing the loss, the CLIP model can adjust the image and text embeddings to make the matching image-text pairs closer in the shared feature space, while the non-matching image-text pairs are pushed away, thus achieving the geometric alignment of cross-modal features. Finally, combine the two parts of the loss L orig and L aug to obtain the total loss L total = L orig + L aug . This total loss guides the CLIP model to learn a more accurate matching relationship between the image and the text description, thus achieving the alignment of the global features of the image and the global features of the text description.

[0060] In addition, the calculation formula of L total is as follows:

[0061]

[0062] Among them, L I2T is the loss from the image to the original text description T, and L T2I is the loss from the original text description T to the image, and L I2T~ is the loss from the image to the generated text description T~, and L T~2I is the loss from the generated text description T~ to the image.

[0063] The calculation formula of L I2T is as follows:

[0064]

[0065] where N represents the batch size, that is, the number of original image-text pairs processed simultaneously in one training batch; s(I i , T i ) represents the similarity between the image I i and the corresponding original text description T i , which is calculated using cosine similarity, that is v I and v T are the embedding vectors of the image I i and the original text description T i respectively. The dot product v I ·v T represents the similarity between the two, and the denominator is the norm of the embedding vector; exp(.) is the exponential function used to ensure the normalization of the similarity value; τ is the temperature parameter used to adjust the smoothness in contrastive learning; is the contrastive loss from the image to the original text description. Therefore, L I2T represents the similarity between the given image I i and the correctly matched original text description T i in the batch N, compared with the sum of the similarities between this image and all other original text descriptions T j . The purpose is to maximize the similarity between I i and T i , while minimizing the similarity between I i and T j ;

[0066] The calculation formula of L T2I is as follows:

[0067]

[0068] where s(T i , I i ) represents the similarity between the corresponding original text description T i and the image I iThe similarity between them is calculated using cosine similarity, and the calculation formula is the same as that of s(I i ,T i ). L T2I The goal is to maximize the similarity between T i and I i , and at the same time minimize the similarity between T i and other images I j ;

[0069] Similarly, the calculation formula for the loss L aug of the image and the generated text description can be obtained. The only difference is that the original text description T i corresponding to the image I i is replaced with the generated text description T i ~.

[0070] The described embodiments are the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Without departing from the essence of the present invention, any obvious improvements, substitutions or variations that those skilled in the art can make all fall within the protection scope of the present invention.

Claims

1. An LLM-based cross-modal person re-identification data augmentation method, characterized in that: Step 1: A large number of original image-text pairs are included in the existing person re-identification dataset. The original image-text pairs include images and corresponding original text descriptions. Preprocess the original text descriptions, and input the preprocessed original text descriptions and the prompt Prompt into the LLM to generate new text descriptions, that is, generate text descriptions; Step 2: Introduce an image-based screening mechanism to perform threshold filtering on the generated text descriptions, and regenerate the generated text descriptions that do not meet the threshold requirements according to the method in Step 1, and then perform threshold filtering on the regenerated text descriptions. If after the N_max-th iteration of regeneration, the regenerated text descriptions still do not meet the threshold requirements, then keep the generated text descriptions obtained in the N_max-th time together with the generated text descriptions that have met the threshold requirements to obtain the finally saved generated text descriptions; Step 3: Use a neural network to perform feature learning on the original image-text pairs in the dataset in Step 1 and the finally saved generated text descriptions in Step 2 to obtain the global image features and the global text description features, project them into a common space, and finally use the contrastive learning loss to achieve the alignment of the global image features and the global text description features.

2. The cross-modal person re-identification data augmentation method based on LLM according to claim 1, wherein The process of the preprocessing is as follows: For the data where each image contains multiple original text descriptions, splice the multiple original text descriptions and insert a separator to form a spliced original text description; for the data where each image contains only one original text description, use this original text description as the spliced original text description.

3. The cross-modal pedestrian re-identification data augmentation method based on LLM according to claim 1, wherein The LLM adopts the General Language Model - 4 - 9B - Chat.

4. The cross-modal person re-identification data augmentation method based on LLM according to claim 1, characterized in that The image-based screening mechanism adopts the BLIP model.

5. The method for enhancing cross-modal pedestrian re-identification data based on LLM according to claim 4, wherein, The process of performing threshold filtering on the generated text description is as follows: First, use the BLIP model to calculate the matching score itm_score and cosine similarity itc_score for each image in the person re-identification dataset and its corresponding original text description. Average the matching scores itm_score and cosine similarities itc_score of all original image-text pairs respectively to obtain two average scores, and use them as the threshold tuple S t ; Then, use the BLIP model to calculate the matching score T~_itm_score and cosine similarity T~_itc_score for the generated text description and the corresponding image in step 1, and use them as the score tuple S; Finally, compare the score tuple S with the threshold tuple S t If the values in S are all greater than or equal to the corresponding values in S t , the generated text description is considered to meet the threshold requirements. Otherwise, the generated text description is considered not to meet the threshold requirements.

6. The cross-modal pedestrian re-identification data augmentation method based on LLM according to claim 1, characterized in that, The global image features are extracted using the Vision Transformer in the CLIP model, and the global text description features are extracted using the text encoder in the CLIP model.

7. The cross-modal person re-identification data augmentation method based on LLM according to claim 6, wherein The process of the contrastive learning loss achieving the alignment of the global image features and the global text description features is as follows: First, using the contrastive learning loss, calculate the loss L between the image and the original text description respectively orig , and the loss L between the image and the generated text description aug ; then, combine the two parts of the loss L orig and L aug to obtain the total loss L total = L orig + L aug ; finally, through the total loss L total guide the CLIP model to learn a more accurate matching relationship between the image and the text description, so as to achieve the alignment of the global features of the image and the global features of the text description.

8. The method for augmenting cross-modal pedestrian re-identification data based on LLM according to claim 7, wherein The total loss L total is calculated as follows: Among them, L I2T is the loss from the image I to the original text description T, and L T2I is the loss from the original text description T to the image I, and L I2T~ is the loss from the image I to the generated text description T~, and L T~2I is the loss from the generated text description T~ to the image I.

9. The cross-modal person re-identification data augmentation method based on LLM according to claim 8, wherein The loss L of the image I to the original text description T I2T and the loss L of the original text description T to the image I T2I are calculated as follows: Among them, N represents the batch size; s(I i ,T i ) represents the similarity between the image I i and the corresponding original text description T i . v I and v T are the embedding vectors of the image I i and the original text description T i respectively. The dot product v I ·v T represents the similarity between the two. The denominator is the norm of the embedding vector; exp(.) is the exponential function; τ is the temperature parameter; s(T i ,I i ) represents the similarity between the corresponding original text description T i and the image I i , and the calculation formula is the same as that of s(I i ,T i ).

10. The method for cross-modal pedestrian re-identification data augmentation based on LLM according to claim 9, wherein The loss L of the described image and the generated text description aug has a calculation formula similar to the loss L of the image and the original text description orig , with the only difference being that the original text description T corresponding to the image i is replaced by the generated text description T i ~.