Unsupervised text-to-image pedestrian re-identification method and system
By constructing image-text matching relationship and multi-level ternary joint learning, the problem of lack of labeled information in unsupervised text-to-image pedestrian re-recognition is solved, and more efficient and accurate pedestrian recognition is achieved.
Patent Information
- Application Number
- CN202510485602.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-22
AI Technical Summary
In the existing unsupervised text-to-image pedestrian re-identification method, the image and text samples lack identity tag information, and the generated text description is unreliable, resulting in difficulty in cross-modal matching. The prior art cannot effectively process text samples, cannot establish an image-text matching relationship, and unreliable pseudo-identity tags hinder model optimization.
By constructing an image-text matching relationship, a large language model is used to generate multiple text descriptions, cluster image features assign pseudo-identity tags, eliminate similarity anomaly samples, and optimize feature distribution using a multi-level ternary joint learning process, including DBSCAN algorithm and interquartile range filtering, combining CLIP model and multi-level ternary loss function to optimize feature matching.
It enhances the reliability of image-text matching, eliminates exception labels, improves the search accuracy and efficiency of unsupervised text-to-image pedestrian re-recognition, and improves the search accuracy by 5.65%.
Smart Images

Figure CN120356242A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of computer vision and artificial intelligence, and particularly to an unsupervised text-to-image pedestrian re-identification method and system. Background Art
[0002] Text-to-image pedestrian re-identification (TIReID) aims to retrieve target pedestrians from a large-scale image library based on specific text descriptions. Different from single-modal pedestrian re-identification methods, TIReID uses different text descriptions as query information to match target images, which is considered a cross-modal task. In recent years, TIReID has received extensive attention due to its potential application value in smart cities and intelligent transportation systems. However, the inherent modality differences make the cross-modal matching process difficult. Some existing technologies use pedestrian identity information to explore invariant features at different granularities to bridge the modality differences. Other methods in some existing technologies generate additional image-text pairs according to identity information to enhance the diversity of training samples. These methods require clear matching relationships and identity information for image-text pairs.
[0003] In practical application scenarios, image samples are easy to obtain and are abundant in number. However, in traditional methods, text descriptions are usually manually annotated. In addition, annotating identity labels for different pedestrian samples takes a lot of time and is obviously infeasible. Unsupervised TIReID aims to mine the internal information of pedestrians from unlabeled samples and has broad application prospects. For unlabeled image and text samples, they not only lack identity label information but also lack the matching relationship between images and texts. This greatly increases the complexity of the task.
[0004] Existing image-based unsupervised learning methods usually perform clustering operations to assign pseudo-identity labels to different pedestrian images. However, they cannot effectively process text samples, cannot establish image-text matching relationships, and are not suitable for unsupervised TIReID tasks. Multimodal large language models (MLLMs) have powerful cross-modal understanding and generation capabilities. Some technicians in this field design different prompts to guide MLLMs to generate text descriptions with pedestrian attributes. However, the reliability of these generated text descriptions is still questionable. In addition, some methods in some existing technologies cluster sample features to generate identity information and use cross-modal matching losses to enhance semantic consistency. When using all pseudo-identity labels, many unreliable identity information hinders the model optimization process. Summary of the Invention
[0005] The present invention provides an unsupervised text-to-image pedestrian re-identification method and system, aiming to solve the technical problems existing in the prior art.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides an unsupervised text-to-image person re-identification method, including:
[0008] Obtain a pedestrian image and a text general template, input them into a large language model, and output multiple text descriptions of the pedestrian image to construct an image-text matching relationship;
[0009] Obtain the image features of the pedestrian image, perform a clustering operation on the image features, assign pseudo identity labels to the pedestrian image and its corresponding text description; and construct an image memory dictionary;
[0010] Obtain the text features of the text description, and according to the image memory dictionary, obtain a similarity set; according to the similarity set, eliminate the pedestrian image with abnormal similarity and its corresponding text description;
[0011] Obtain the image features and text features of the pedestrian image and the text description respectively again, and construct a multi-level triple joint learning process; use the multi-level triple joint learning process to optimize the feature distribution of the pedestrian image and the text description;
[0012] Obtain the text features to be re-identified, perform pedestrian image matching, and obtain a pedestrian recognition result.
[0013] As a further technical solution, the DBSCAN algorithm is used to cluster the image features, and according to the image-text matching relationship, pseudo identity labels are assigned to the pedestrian image and its corresponding text description;
[0014] According to the result after the clustering operation, the constructed image memory dictionary is expressed as:
[0015] where, C[a] represents the image clustering center feature of the a-th identity; l a represents the set of image features with the a-th pedestrian identity; N a represents the number of features.
[0016] As a further technical solution, the specific method for obtaining the similarity set is: calculate the cosine similarity between the text features of the text description and the image clustering center features in the image memory dictionary, and the similarity set is expressed as:
[0017] S a ={sim(t, C[a])|t∈T a}; where, sim() represents the calculation of cosine similarity; T a represents the set of text features with the a-th pedestrian identity;
[0018] The interquartile range filtering algorithm is adopted to eliminate pedestrian images with abnormal similarity and their corresponding text descriptions. The elimination process is expressed as:
[0019] Among them, β represents a hyperparameter to control the filtering intensity; Q1 represents the first quartile of S a after sorting; Q3 represents the third quartile of S a after sorting.
[0020] As a further technical solution, the specific method for obtaining the image features and text features of the pedestrian image and the text description again is as follows: preprocess the pedestrian image and the text description respectively, use the trained CLIP model as the image encoder and the text encoder, and input the preprocessed pedestrian image and text description into the image encoder and the text encoder respectively to obtain the image features and text features again.
[0021] As a further technical solution, the multi-level triplet joint learning process includes a center-level triplet joint learning process and an instance-level triplet joint learning process; the center-level triplet joint learning process is expressed as:
[0022] L center = λ1L ctrl + λ2L cen ; where λ1 and λ2 represent hyperparameters to control the importance of the triplet loss and the extended center loss respectively, L ctrl represents the center-level triplet loss, and L cen represents the extended center loss;
[0023] The center-level triplet loss L ctrl is expressed as:
[0024] Among them, represents the k-th image feature with the r-th identity, represents the most difficult negative instance feature of C[r], m represents the interval of the distance between positive and negative samples, max() represents the maximum value function, and P and K respectively represent the number of pedestrian identities included in a batch and the number of samples for each identity.
[0025] The extended center loss L cen is expressed as:
[0026]
[0027] As a further technical solution, the instance-level triplet joint learning process is expressed as:
[0028] L instance = λ3L inter + λ4Lintra ; where λ3 and λ4 are hyperparameters that respectively control the importance of the inter-modal and intra-modal matching losses, and L inter represents the instance-level inter-modal matching loss, and L intra represents the instance-level intra-modal matching loss;
[0029] The instance-level inter-modal matching loss L inter is expressed as:
[0030]
[0031] where M t and M v respectively represent the positive example text / image feature sets of v i and t i ; and respectively represent the hardest negative example text / image features of v i and t i ; represents the similarity weight from image to text; represents the similarity weight from text to image;
[0032] The instance-level intra-modal matching loss L intra is expressed as:
[0033] where Hv and Ht respectively represent the positive example image / text feature sets of vi and ti; vhnv and thnt respectively represent the hardest negative example image / text features of vi and ti; represents the similarity weight from image to image; represents the similarity weight from text to text.
[0034] As a further technical solution, the similarity weight from image to text is expressed as:
[0035] where τ represents the temperature hyperparameter.
[0036] In a second aspect, the present invention provides an unsupervised text-to-image person re-identification system, including the following modules:
[0037] A graphic-text matching module, configured to: obtain a pedestrian image and a text general template, input them into a large language model, and output multiple text descriptions of the pedestrian image to construct an image-text matching relationship;
[0038] The image memory dictionary construction module is configured to: obtain the image features of pedestrian images, perform clustering operations on the image features, assign pseudo identity labels to pedestrian images and their corresponding text descriptions; and construct an image memory dictionary;
[0039] The abnormal sample elimination module is configured to: obtain the text features of the text descriptions, and obtain a similarity set according to the image memory dictionary; according to the similarity set, eliminate the pedestrian images with abnormal similarities and their corresponding text descriptions;
[0040] The optimized feature distribution module is configured to: respectively obtain the image features and text features of pedestrian images and text descriptions again, and construct a multi-level triple joint learning process; use the multi-level triple joint learning process to optimize the feature distributions of pedestrian images and text descriptions;
[0041] The re-identification module is configured to: obtain the text features to be re-identified, perform pedestrian image matching, and obtain pedestrian recognition results.
[0042] One or more technical solutions of the present invention have the following beneficial effects:
[0043] The present invention proposes a reliable text generation process, constructs an image-text matching relationship, and solves the problems of unreliable text or time-consuming text annotation caused by using a single multi-modal large model or manual annotation to generate text in the prior art; at the same time, the present invention calculates the cosine similarity between the text features of the text descriptions and the image clustering center features in the image memory dictionary, and adopts an interquartile range filtering algorithm to eliminate the pedestrian images with abnormal similarities and their corresponding text descriptions, effectively eliminating the abnormal label samples generated during the clustering process, and enhancing the reliability of identity information; finally, aiming at the problem of insufficient optimization of traditional methods, the present invention proposes a multi-level triple joint learning optimization process from the perspectives of the center and instances, enhances the intra-class compactness and inter-class separability of samples, and greatly improves the retrieval accuracy and efficiency of the unsupervised text-to-image pedestrian re-identification method. Description of the Drawings
[0044] The specification drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0045] Figure 1 It is a schematic diagram of the model structure in the first embodiment of the present invention; Detailed Embodiments
[0046] It should be noted that the following detailed descriptions are all illustrative and are intended to provide further descriptions of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meanings as those commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0047] Example 1
[0048] In this embodiment, an unsupervised text-to-image person re-identification method is provided. As Figure 1 shown, the specific method includes the following steps:
[0049] S1: Obtain a pedestrian image and a text general template, input them into a large language model, and output multiple text descriptions of the pedestrian image to construct an image-text matching relationship;
[0050] S2: Obtain the image features of the pedestrian image, perform a clustering operation on the image features, assign pseudo identity labels to the pedestrian image and its corresponding text description; and construct an image memory dictionary;
[0051] S3: Obtain the text features of the text description, and obtain a similarity set according to the image memory dictionary; according to the similarity set, eliminate the pedestrian image with abnormal similarity and its corresponding text description;
[0052] S4: Obtain the image features and text features of the pedestrian image and the text description respectively again, and construct a multi-level triple joint learning process; use the multi-level triple joint learning process to optimize the feature distribution of the pedestrian image and the text description;
[0053] S5: Obtain the text features to be re-identified, perform pedestrian image matching, and obtain the pedestrian recognition result.
[0054] In step S1, the text general template is expressed as:
[0055] “With [hair description], the [person / woman / man] is wearing [clothing description] and is also carrying [belongings description].”
[0056] “In [clothing description] and [accessory description], the [person / woman / man] also has [hair description].”
[0057] In this embodiment, two different multi-modal large language models (MLLMs) are adopted. The above general template is randomly selected and input into the multi-modal large language model together with the pedestrian image, so as to generate multiple text descriptions of the same pedestrian image. In this embodiment, the two different multi-modal large language models (MLLMs) are Qwen-VL-Chat and Qwen2-VL-7B (both are domestic large models).
[0058] In step S1, the CLIP model parameters are frozen, and the correlation between the pedestrian image and each generated text description is evaluated respectively, and the texts with higher correlation are screened out to construct the image-text matching relationship. Specifically: The trained CLIP model is used to extract the features of the pedestrian image and the text description respectively. The trained CLIP model is a two-tower structure, which includes an image encoder and a text encoder, and extracts image and text features respectively. Then the cosine similarity between the features is calculated to evaluate the correlation between the image and the generated text description, and the texts with high similarity are selected as training samples to construct the image-text matching relationship.
[0059] In step S2, the image features of the pedestrian image are extracted by using the frozen CLIP model, and the DBSCAN algorithm is used to cluster the image features to assign pseudo identity labels to the pedestrian image and its corresponding text description. In this embodiment, the size of the input pedestrian image is 384×128, where 384 is the image height and 128 is the image width.
[0060] In this embodiment, since the image-text matching relationship is established through the MLLM, only the image features need to be clustered, and the corresponding text will also obtain pseudo identity labels, saving the model training time.
[0061] In step S2, according to the result after the clustering operation, an image memory dictionary is constructed to store the central features of each image cluster. The image memory dictionary is expressed as:
[0062] where C[a] represents the central feature of the image cluster of the a-th identity; I a represents the set of image features with the a-th pedestrian identity; N a represents the number of features.
[0063] In this embodiment, during the optimization process of the image memory dictionary, by selecting the most representative difficult sample instances, the central features of each cluster in the image memory dictionary are continuously updated; among them, the most representative difficult sample instance refers to the positive sample instance feature that is farthest (least similar) from the cluster central feature. An example can be used to understand that in each training iteration, there are P identities, where each identity contains K instances, and the batch size is P*K. Since these P identities are pseudo-labels obtained by clustering, they need to be continuously optimized and updated. The P identities correspond to P cluster central features, and the update method is to select one from the K instance features corresponding to the identity to execute the following formula for update. Then there are multiple ways to select the instance, which can be randomly selected or the most representative difficult sample instance can be used. This difficult sample instance is the instance with the lowest similarity to the cluster center among these K instances to update the cluster central feature. The advantage of this update method is to guide the model to learn difficult samples and enhance the feature representation ability.
[0064] The update process is expressed as:
[0065] where η represents the update rate; represents the most difficult positive instance feature of C[a]. In this embodiment, the most difficult sample update strategy is adopted. Specifically, the similarity between the cluster center and each instance is calculated, and the positive example sample with the lowest similarity is selected to update the cluster center; η can be set to 0.1.
[0066] In step S3, the text description corresponding to the abnormal sample cannot well describe the cluster center identity, resulting in a low correlation, so it needs to be filtered out. In this embodiment, an error class filtering module is constructed. The error class filtering module is used to eliminate the abnormal samples generated during the clustering process, thereby enhancing the reliability of the identity label; in the error class filtering module, the cosine similarity between the text feature of the text description and the image cluster center feature in the image memory dictionary is calculated, and the similarity set is expressed as:
[0067] S a ={sim(t, C[a])|t ∈ T a}; where sim() represents the calculation of the cosine similarity; T a represents the text feature set with the a-th pedestrian identity;
[0068] In step S3, the interquartile range filtering algorithm is adopted to eliminate the pedestrian images with abnormal similarity and their corresponding text descriptions. The elimination process is expressed as:
[0069] where β represents the hyperparameter to control the filtering intensity; Q1 represents the first quartile (the 25th percentile) of the sorted S a ; Q3 represents the sorted Sa The third quartile (75th percentile).
[0070] In this embodiment, the abnormal sample filtering process is only for the current batch. The filtered samples will not be included in the training sample set in the current batch. In the next round of training, the filtering algorithm will be re-executed, and the abnormal samples to be filtered will be continuously adjusted according to the clustering results; β can be set to 1.0, and the specific value should be adjusted according to the actual data distribution.
[0071] In step S4, preprocess the pedestrian images, and use data augmentation methods such as random horizontal flipping, random cropping, and random padding. At the same time, perform normalization operations on the pedestrian images;
[0072] In this embodiment, the random horizontal flipping probability can be set to 0.5; the random padding can be set to 10 pixels; the preprocessing of the pedestrian images is only performed during the training process and does not include preprocessing operations in the test phase.
[0073] Preprocess the text description, perform word segmentation and use data augmentation methods such as random masking; in this embodiment, the BPE algorithm can be set for word segmentation, and the random masking probability can be set to 0.15. Similarly, the masking operation on the text description is only performed during the training process.
[0074] In step S4, unfreeze the CLIP model parameters, and extract the enhanced image features and text features again respectively. Use the trained CLIP model as the image encoder and text encoder, input the preprocessed pedestrian images and text descriptions into the image encoder and text encoder respectively, and obtain the image features and text features again. The dimensions of the image features and text features are 512.
[0075] In step S5, construct a multi-level triple joint learning process from the perspectives of the center and instances, and continuously optimize the feature distributions of the pedestrian images and text descriptions.
[0076] In step S5, the multi-level triple joint learning process includes a center-level triple joint learning process and an instance-level triple joint learning process; among them, the center-level triple joint learning process is expressed as:
[0077] L center = λ1L ctrl + λ2L cen ; where λ1 and λ2 represent hyperparameters that respectively control the importance of the triple loss and the extended center loss. L ctrl represents the center-level triple loss, and L cen represents the extended center loss;
[0078] In this embodiment, λ1 and λ2 can be set to 1 and 0.08 after experimental comparison; the center-level loss is generally used to shorten the distance between the instance and the cluster center (identity), which can produce a compact feature distribution.
[0079] The center-level ternary loss L ctrl It is expressed as:
[0080] in, represents the k-th image feature with the r-th identity, represents the most difficult negative instance feature of C[r], m represents the distance interval between positive and negative samples, max() represents the maximum value function, P and K represent the number of pedestrian identities contained in a batch and the number of samples for each identity, respectively.
[0081] In this embodiment, m can be set to 0.1 after experimental comparison.
[0082] The center loss after expansion l cen It is expressed as:
[0083]
[0084] The instance-level ternary joint learning process is expressed as:
[0085] L instance =λ3L inter +λ4L intra ; where λ3 and λ4 represent hyperparameters that control the importance of inter-modal and intra-modal matching losses, respectively, and L inter represents the instance-level modality matching loss, L intra represents the matching loss within the instance-level modality;
[0086] In this embodiment, after experimental comparison, the model performance is best when λ3 and λ4 are both set to 1; the instance-level loss mainly optimizes the cross-modal instance relationship and bridges the gap between modalities.
[0087] Instance-level inter-modality matching loss L inter It is expressed as:
[0088] Where Mt and Mv represent v i and t i A set of positive text / image features; and Respectively represent v i and t i The hardest negative text / image features; Represents the similarity weight from image to text; Represents the similarity weight from text to image;
[0089] In this embodiment, m can be set to 0.1 through experimental comparison; the role of weighting is equivalent to making the model focus on and learn some strongly correlated (high-confidence) image-text pairs, improving the overall learning efficiency.
[0090] The similarity weight from the above-mentioned image to text is expressed as:
[0091] where τ represents the temperature hyperparameter.
[0092] The matching loss L within the instance-level modality intra is expressed as:
[0093] where Hv and Ht respectively represent the set of positive example image / text features of vi and ti; vhnv and thnt respectively represent the most difficult negative example image / text features of vi and ti; represents the similarity weight from image to image; represents the similarity weight from text to text, and their calculation methods are the same as the similarity weight from image to text.
[0094] In this embodiment, the query text used in the test phase is from the test set of the original database, and the text in the training set is not utilized during the training phase. In addition, the identity information is discarded throughout the training process, constituting an unsupervised text-to-image person re-identification method.
[0095] Finally, the performance of the model can be evaluated according to the retrieval accuracy.
[0096] Taking the publicly available large-scale person re-identification database as the test object, for example, when testing on the CUHK-PEDES database, the average precision of this embodiment reaches 46.93%, which is 5.65% higher than the current state-of-the-art unsupervised method.
[0097] Embodiment 2
[0098] This embodiment provides an unsupervised text-to-image person re-identification system, including the following modules:
[0099] An image-text matching module, configured to: obtain a pedestrian image and a text general template, input them into a large language model, and output multiple text descriptions of the pedestrian image to construct an image-text matching relationship;
[0100] An image memory dictionary construction module, configured to: obtain the image features of the pedestrian image, perform a clustering operation on the image features, assign pseudo identity labels to the pedestrian image and its corresponding text description; and construct an image memory dictionary;
[0101] An abnormal sample elimination module, configured to: obtain the text features of the text description, and obtain a similarity set according to the image memory dictionary; eliminate the pedestrian images with abnormal similarity and their corresponding text descriptions according to the similarity set;
[0102] An optimized feature distribution module, configured to: obtain the image features and text features of the pedestrian image and the text description respectively again, and construct a multi-level triple joint learning process; optimize the feature distributions of the pedestrian image and the text description by using the multi-level triple joint learning process;
[0103] A re-identification module, configured to: obtain the text features to be re-identified, perform pedestrian image matching, and obtain a pedestrian recognition result.
[0104] Embodiment III
[0105] The purpose of this embodiment is to provide a computer-readable storage medium for storing a computer program to complete the method described in Embodiment I.
[0106] The method in Embodiment I can be directly implemented by a hardware processor to complete, or by a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.
[0107] Embodiment IV
[0108] The purpose of this embodiment is to provide an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, which can be used to complete the method described in Embodiment I. For the sake of brevity of description, it will not be elaborated here.
[0109] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0110] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0111] For those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An unsupervised text-to-image person re-identification method, characterized in that, include: Obtain pedestrian images and general text templates, input them into a large language model, and output multiple text descriptions of pedestrian images to build image-text matching relationships; Obtain image features of pedestrian images, perform clustering operations on the image features, and assign pseudo identity labels to pedestrian images and their corresponding text descriptions; And build an image memory dictionary; Acquire text features of the text description, and obtain a similarity set according to the image memory dictionary; based on the similarity set, remove pedestrian images and their corresponding text descriptions with abnormal similarity; The image features and text features of the pedestrian image and text description are obtained again, and a multi-level ternary joint learning process is constructed; the feature distribution of the pedestrian image and text description is optimized using the multi-level ternary joint learning process; The text features to be re-identified are obtained, and pedestrian image matching is performed to obtain pedestrian recognition results.
2. The unsupervised text-to-image person re-identification method according to claim 1, characterized in that, The DBSCAN algorithm is used to cluster the image features, and a pseudo identity label is assigned to the pedestrian image and its corresponding text description according to the image-text matching relationship; According to the results of the clustering operation, the constructed image memory dictionary is expressed as: Among them, C[a] represents the image clustering center feature of the a-th identity; I a represents the image feature set with the a-th pedestrian identity; N a represents the number of features.
3. The unsupervised text-to-image person re-identification method according to claim 1, characterized in that, The specific method of obtaining the similarity set is: calculating the cosine similarity between the text features of the text description and the image cluster center features in the image memory dictionary. The similarity set is expressed as: S a = {sim(t, C[a]) | t ∈ T a}; where, sim() represents the cosine similarity calculation; T a represents the text feature set with the a-th pedestrian identity; The interquartile range filtering algorithm is used to remove pedestrian images and their corresponding text descriptions with abnormal similarity. The removal process is expressed as: Among them, β represents a hyperparameter that controls the filtering intensity; Q1 represents the first quartile of S after sorting a ; Q3 represents the third quartile of S after sorting a .
4. The unsupervised text-to-image person re-identification method according to claim 1, characterized in that The specific method for respectively obtaining the image features and text features of the pedestrian image and text description again is: preprocessing the pedestrian image and text description respectively, using the trained CLIP model as the image encoder and text encoder, inputting the preprocessed pedestrian image and text description into the image encoder and text encoder respectively, and obtaining the image features and text features again.
5. The unsupervised text-to-image person re-identification method according to claim 1, wherein The multi-level ternary joint learning process includes a center-level ternary joint learning process and an instance-level ternary joint learning process; the center-level ternary joint learning process is expressed as: L center = λ1L ctrl + λ2L cen ; where λ1 and λ2 are hyperparameters that respectively control the importance of the triplet loss and the extended center loss, and L ctrl represents the triplet loss at the center level, and L cen represents the extended center loss; The triplet loss \(L\) at the center level ctrl is expressed as: Among them, represents the k-th image feature with the r-th identity, represents the most difficult negative instance feature of C[r], m represents the interval of the distance between positive and negative samples, max() represents the maximum value function, and P and K respectively represent the number of pedestrian identities included in a batch and the number of samples for each identity; The extended center loss $L$ cen is expressed as:
6. An unsupervised text-to-image person re-identification method according to claim 5, wherein, The instance-level ternary joint learning process is expressed as: L instance = λ3L inter + λ4K intra ; where λ3 and λ4 are hyperparameters that respectively control the importance of the inter-modal and intra-modal matching losses, and L inter represents the inter-modal matching loss at the instance level, and L intra represents the intra-modal matching loss at the instance level; The matching loss L between instance-level modalities inter is expressed as: Among them, Mt and Mv respectively represent the set of positive example text / image features of v i and t i ; and respectively represent the most difficult negative example text / image features of v i and t i ; represents the similarity weight from image to text; represents the similarity weight from text to image. The matching loss L within the instance-level modality intra is expressed as: Among them, Hv and Ht respectively represent the sets of positive example image / text features of vi and ti; vhnv and thnt respectively represent the most difficult negative example image / text features of vi and ti; represents the similarity weight from image to image; represents the similarity weight from text to text.
7. The unsupervised text-to-image person re-identification method according to claim 6, characterized in that The image-to-text similarity weight is expressed as: Among them, τ represents the temperature hyperparameter.
8. An unsupervised text-to-image person re-identification system, characterized in that, Includes the following modules: The image-text matching module is configured to: obtain a pedestrian image and a general text template, input them into a large language model, and output multiple text descriptions of the pedestrian image to construct an image-text matching relationship; The image memory dictionary building module is configured to: obtain image features of pedestrian images, perform clustering operations on the image features, and assign pseudo identity labels to pedestrian images and their corresponding text descriptions; And build an image memory dictionary; The abnormal sample elimination module is configured to: obtain text features of the text description, and obtain a similarity set according to the image memory dictionary; according to the similarity set, eliminate pedestrian images with abnormal similarity and their corresponding text descriptions; The feature distribution optimization module is configured to: respectively obtain the image features and text features of the pedestrian image and the text description again, and construct a multi-level ternary joint learning process; and optimize the feature distribution of the pedestrian image and the text description using the multi-level ternary joint learning process; The re-identification module is configured to: obtain text features to be re-identified, perform pedestrian image matching, and obtain pedestrian recognition results.
9. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in an unsupervised text-to-image pedestrian re-identification method according to any one of claims 1-7.
10. An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in an unsupervised text-to-image pedestrian re-identification method according to any one of claims 1-7.
Citation Information
Cited By
Unsupervised text-to-image pedestrian re-identification method and system
CN120580739A
An unsupervised pedestrian re-identification method and system from text to image
CN120580739B