Semantic guidance pedestrian re-identification method and system based on text prompt
By adopting a semantic guidance method based on text prompts in pedestrian recognition technology, using the CLIP model and cross attention mechanism to generate personalized language descriptions and image block features, the problem of lack of semantic guidance in the existing technology is solved, and more efficient and accurate pedestrian recognition is achieved.
Patent Information
- Application Number
- CN202410736218.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-07
- Publication Date
- 2025-05-23
AI Technical Summary
The existing pedestrian re-identification technology lacks explicit semantic guidance, resulting in the areas of the model focusing only emphasize specific local discriminant parts, unable to accurately pay attention to all semantic-related areas, and requires additional time-consuming and labor-intensive manual annotation.
The text prompt-based semantic guidance method (PromptSG) is used to generate personalized language descriptions and image block features through the visual encoder and text encoder in the CLIP model, combining the reverse network and the cross-attention mechanism, to realize language-guided image semantic mining.
It improves the search performance of pedestrian re-identification, can generalize to unseen categories, reduces dependence on additional annotated information, and significantly improves the performance and efficiency of the model.
Smart Images

Figure CN120032307A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of information technology, and in particular relates to a semantically guided pedestrian re-identification method and system based on text prompts. Background Art
[0002] With the development of cities and the advancement of science and technology, video surveillance systems have been widely used in public places, commercial areas, transportation hubs, etc. In these scenarios, people often need to identify specific pedestrians, such as tracking criminal suspects, finding missing children, and managing the flow of people in large-scale events. With the rapid growth of surveillance data, this massive data scale has put tremendous pressure on traditional manual surveillance and analysis methods. Therefore, intelligent surveillance systems based on computer vision have begun to emerge. Pedestrian re-identification technology, as an important branch of image retrieval, aims to accurately identify and match the identity of the same pedestrian in cameras at different scenes and time points. With the in-depth mining of semantic information, current methods have made significant progress. These methods focus on extracting parts of the image that are closely related to semantics, such as human body structure, to achieve accurate alignment and matching. However, existing methods only use a single image modality and lack explicit semantic guidance. The area of interest of the model usually only emphasizes specific local discriminant parts, but cannot accurately focus on all semantically related areas. When clear masks or human key points are required as guidance directions, additional, time-consuming and labor-intensive manual annotation is inevitably required. In this paper, a text prompt-based semantic guidance method (PromptSG) is proposed to solve this problem by efficiently utilizing the multimodal large model CLIP.
[0003] In the person re-identification task, a class of techniques based on convolutional neural networks (CNNs) have explored the use of attention mechanisms. For example, AANet generates attribute attention maps by introducing additional attribute labels, guiding the network to focus on extracting perceptually relevant body part features. On the other hand, TransReID, a pioneering work based on visual Transformer, introduced a self-attention-based architecture. However, these methods only apply attention mechanisms to the visual modality and lack explicit language guidance, which may limit their performance to a certain extent and may lead to a degradation in the performance of retrieval models. Most relevant to the research of this paper is CLIP-ReID, which is the first method to utilize the visual-language pre-training model CLIP (Contrastive Language–Image Pre-training) in the pedestrian re-identification task. However, the decoupled use of CLIP-ReID, that is, relying only on visual embeddings during reasoning, makes the learned cues only associated with the identities seen during training, and fails for unseen identities, thus affecting the generalization ability of the model. In addition, the predefined cue learning technique adopted may not be sufficient to fully describe the visual context of a specific pedestrian, which may still lead to a degradation in the retrieval performance of the model. Summary of the invention
[0004] The purpose of the present invention is to use the rich semantics contained in the CLIP model to achieve language-guided image semantic mining in the pedestrian re-identification task where text annotation data is naturally lacking, thereby improving the retrieval performance. Therefore, the present invention first adopts the textual inversion technology, which learns unique tokens to represent the visual context, thereby generating a personalized language description for a given pedestrian. Then, for the generated personalized text prompt, the image blocks output by the image end are optimized by introducing a cross-attention mechanism, guiding the model to focus on the area that is consistent with the semantics of the prompt.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A semantically guided person re-identification method based on textual cues comprises the following steps:
[0007] Input the training image into the visual encoder to obtain the visual embedding;
[0008] Use the inverse network to map the visual embedding to the text space to obtain pseudo-tokens, and integrate the pseudo-tokens into natural language sentences to obtain language hints for the input image;
[0009] Feed the language hint into the text encoder to get text embedding;
[0010] Using visual embedding and text embedding to train a multimodal interaction module, the multimodal interaction module enables interaction between image patches and language cues in a multimodal environment;
[0011] The query image is input into the trained multimodal interaction module to obtain a feature vector that integrates visual and textual information. The feature vector that integrates visual and textual information is used to perform similarity retrieval in the pedestrian image database to obtain the pedestrian re-identification result.
[0012] Furthermore, the visual encoder and the text encoder are the visual encoder and text encoder in the CLIP model.
[0013] Furthermore, the method of mapping the visual embedding to the text space using the inverse network to obtain the pseudo token includes ensuring that the learned pseudo token can effectively convey the context information of the image through the symmetric contrast loss; the symmetric contrast loss is:
[0014]
[0015] Among them, L i2t represents the contrast loss from image to text, L t2i represents the contrast loss from text to image, N represents the number of samples in batch processing, sim represents the cosine similarity function, and v i represents the i-th image in the batch, l p is with v n The corresponding text hint embedding, v p Indicates that p Corresponding visual embeddings, τ represents the temperature hyperparameter used to control the scaling of the similarity score.
[0016] Furthermore, the method of mapping the visual embedding to the text space using the inverse network to obtain the pseudo-token includes encouraging the pseudo-token to capture the visual details belonging to the same identity through a symmetric supervised contrast loss; the symmetric supervised contrast loss is:
[0017]
[0018] Among them, L SupCon represents the symmetric supervised contrast loss, represents the supervised contrastive loss for image to text, represents the supervised contrast loss from text to image, and P(i) represents the loss with v n , l n The set of positive samples with the same label, p + represents the positive sample index, Represents the positive sample of text prompt p + The embedding representation of Represents the visual positive sample p+ The embedded representation of .
[0019] Furthermore, the multimodal interaction module contains a language-guided cross-attention layer that uses text embeddings as queries and the block-wise embeddings of the visual encoder as keys and values; given a pair of image and prompt (x, t p ), input the image x into the visual encoder and get a series of image block embeddings in represents the global visual embedding, v 1 ,...,v M Belong to the local image block embedding; the prompt t p Input the text encoder to get the text embedding l p ; The text embedding is then projected into the query matrix Q, and the image patch embedding is projected into the key matrix K and the value matrix V through three different linear projection layers; the interaction from image patches to text prompts is achieved in the following way:
[0020]
[0021] Where A represents the attention function and d represents the dimension of visual embedding.
[0022] Furthermore, two Transformer blocks are added after the cross-attention layer, and each Transformer block includes a self-attention layer and a feed-forward layer.
[0023] Furthermore, the person re-identification loss used in the training phase includes the ternary loss L Triplet and identity classification loss L ID , where the ternary loss L Triplet The identity classification loss L is used to bring the same pedestrian closer and push different pedestrians away in the embedding space. ID Used to classify a given pedestrian image into its identity category; triple loss L Triplet and identity classification loss L ID The calculation formula is as follows:
[0024]
[0025] L Triplet =max(d p -d n +m,0)
[0026] Among them, y j represents the identity label of the pedestrian, p j represents the category prediction probability of pedestrian images, d p Represents the Euclidean distance of the positive sample pair of images, d nrepresents the Euclidean distance between image negative sample pairs, and m represents a preset threshold or interval.
[0027] A semantically guided person re-identification system based on textual cues, comprising:
[0028] Visual encoder, used to obtain visual embedding based on the input image;
[0029] The reverse network module is used to map the visual embedding to the text space to obtain pseudo-tokens, and integrate the pseudo-tokens into natural language sentences to obtain language hints for the input image;
[0030] Text encoder, used to get text embedding based on language cues;
[0031] The multimodal interaction module is trained using visual embeddings and text embeddings to enable interaction between image patches and language cues in a multimodal environment.
[0032] The recognition module is used to input the query image into the trained multimodal interaction module to obtain a feature vector that integrates visual and textual information, and use the feature vector that integrates visual and textual information to perform similarity retrieval in the pedestrian image database to obtain the pedestrian re-identification result.
[0033] The beneficial effects of the present invention are as follows:
[0034] Compared with the existing methods, this invention proposes a novel semantic guidance method driven by text prompts. Figure 1 As shown, the text hint emphasizes the relevant areas in the image through the cross-attention mechanism, can capture the precise semantic parts, and can also generalize to unseen categories during the reasoning process. In addition, the model can learn the personal tokens of the query image in an end-to-end manner, provide more detailed identity-related guidance, and the design of the present invention does not require additional annotation information such as masks, bounding boxes or precise descriptions. Experiments show that the retrieval performance of the present invention on the existing pedestrian re-identification dataset has been significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a schematic diagram of the core idea of the method of the present invention.
[0036] Figure 2 Schematic diagram of the semantically guided person re-identification method based on text prompts (PromptSG).
[0037] Figure 3 This is an example of an attention map visualization. DETAILED DESCRIPTION
[0038] The present invention is further described in detail below through specific embodiments and drawings.
[0039] like Figure 2 As shown in the figure, a novel semantically guided person re-identification method based on text prompts (PromptSG) proposed in this invention is mainly divided into two parts. In the first stage (training stage), starting from the visual embedding output by the visual encoder of CLIP, an inverse network is used to learn pseudo-tokens, which encapsulate the visual context of the image. After that, a multimodal interaction module is designed, in which the visual encoder uses the attention mechanism to learn semantically faithful representations to form the final re-weighted visual features. In the inference stage, two text input options are provided: one is a simplified prompt oriented for efficiency, and the other is a combined prompt oriented for accuracy. It is worth noting that in the entire framework, the text encoder is frozen, which means that the parameters of the text encoder remain unchanged during the training process and will not be updated by the optimization algorithm.
[0040] (i) Personalized identity-specific cue learning:
[0041] Specifically, the visual encoder V(·) receives an image x as input to obtain a visual embedding v, and then introduces a reverse network, whose core goal is to map the global visual embedding v from the CLIP model visual space V to a pseudo token s in the text space. * This process can be formally expressed as f θ (v) = s * , where s * Located in the text embedding space T * This pseudo-token can then be incorporated into a natural language sentence to produce a linguistic hint for the input image: “A photo of as * The input language prompt will go through a word segmentation process to generate multiple tokens. In the present invention, a token refers to a single word or phrase decomposed from the language prompt, which is the basic unit of text encoding and understanding. The prompt after word segmentation is represented by t p , can be fed into the text encoder of CLIP to obtain the text embedding l p =T(t p ).
[0042] In order to ensure that the learned pseudo tokens can effectively convey the contextual information of the image, the symmetric contrastive loss can be used to achieve the text inverse reconstruction goal. The mathematical expression of the loss function is as follows:
[0043]
[0044] Among them, L i2t represents the contrast loss from image to text, L t2irepresents the contrast loss from text to image, N represents the number of samples in batch processing, sim represents the cosine similarity function, and v i represents the i-th image in the batch, l i represents the i-th text embedding in the batch. p is with v n The corresponding text hint embedding, v p Indicates that p The corresponding visual embedding. τ represents the temperature hyperparameter, which is used to control the scaling of the similarity score.
[0045] However, the present invention believes that images of the same identity should share the same appearance, and the above contrast loss cannot achieve this goal. Therefore, in order to encourage pseudo-tokens to capture visual details belonging to the same identity, the present invention uses symmetric supervised contrast loss, which is specifically formulated as:
[0046]
[0047]
[0048] Among them, L SupCon represents the symmetric supervised contrast loss, represents the supervised contrastive loss for image to text, represents the supervised contrast loss from text to image, and P(i) represents the loss with v n , l n The set of positive samples with the same label, p + represents the positive sample index, Represents the positive sample of text prompt p + The embedding representation of Represents the visual positive sample p + The embedded representation of .
[0049] (II) Semantic guidance based on prompts:
[0050] The above generates personalized language cues by integrating identity-related pseudo-tokens, thereby enhancing its ability to convey a more specific visual context of the image. Then the core idea of the present invention is to finely guide image features through language and explicitly determine which area of the image is aligned with the language cues. Intuitively, the present invention believes that image blocks that are closely related to the semantics of "people" should play a more important role in the distinction and recognition process. Based on this, the present invention designs a multimodal interaction module, which is specifically used to realize the interaction between image blocks and language cues in a multimodal environment.
[0051] Specifically, the present invention adopts a language-guided cross-attention layer, such as Figure 2As shown in Figure 2, it uses text embeddings as queries and the block-wise embeddings of the visual encoder as keys and values. Given a pair of image and prompt (x, t p ), first input the image x into the visual encoder to get a series of image block embeddings here, represents the global visual embedding, and the rest of v i , i∈[1,M] belongs to the local image patch embedding. Similarly, the prompt is fed into the text encoder to obtain the text embedding l p Subsequently, the text embeddings are projected to a query matrix Q, while the image patch embeddings are projected to a key matrix K and a value matrix V through two different linear projection layers. Therefore, the interaction of image patches to text cues can be achieved as follows:
[0052]
[0053] Where A represents the attention function and d represents the dimension of visual embedding. Figure 2 The Z in the equation represents the output representation of the attention function A(), that is,
[0054] This interaction aggregates the attention map and highlights the areas with high semantic response. Drawing on the multimodal fusion method, this paper adds two Transformer blocks after the cross attention layer ( Figure 2 2x is used to represent the number of Transformer blocks), each Transformer block contains a self-attention layer and a feed-forward layer to further refine and integrate the information of high semantic response areas to obtain the final representation.
[0055] Finally, the present invention adopts the standard pedestrian re-identification loss, namely the ternary loss L Triplet and identity classification loss L ID , to optimize the framework of the present invention. Among them, the ternary loss L Triplet The identity classification loss L is used to bring the same pedestrian closer and push different pedestrians away in the embedding space. ID Used to classify a given pedestrian image into the identity category it belongs to.
[0056]
[0057] L Triplet =max(d p -d n +m,0) (8)
[0058] Where N represents the number of batch samples, y j represents the identity label of the pedestrian, p j represents the category prediction probability of pedestrian images, dp Represents the Euclidean distance of the positive sample pair of images, d n represents the Euclidean distance between image negative sample pairs, and m represents a preset threshold or interval.
[0059] Effects of the present invention:
[0060] Datasets: Experiments were conducted on four basic datasets: Market1501, DukeMTMC-ReID, MSMT17, and CUHK03-NP. Market1501 contains 1,501 people and 32,668 images from 6 cameras, and the DukeMTMC-ReID dataset contains 1,404 pedestrians and 36,411 images from 8 cameras. The MSMT17 dataset contains 4,101 pedestrians and 126,441 images from 15 cameras. The CUHK03-NP dataset contains 1,467 pedestrians and 13,164 images from 2 cameras.
[0061] Evaluation indicators: The rank-1 accuracy (Rank-1) and the average precision (mAP) of all classes are used to evaluate the model of the present invention. The larger the mAP and Rank-1, the stronger the retrieval performance of the model.
[0062] Experimental Results: To fully validate the proposed method, PromptSG is benchmarked against current state-of-the-art methods, which can generally be divided into three categories: CNN-based methods, ViT-based methods, and CLIP-based methods.
[0063] As shown in Table 1, it is observed that the proposed PromptSG achieves the best results and establishes a new state-of-the-art performance. Among ViT-based methods, the seminal work TransReID sets a strong benchmark for ViT-based methods by leveraging the potential of Transformer. On this basis, PHA further enhances the preservation of key high-frequency elements of the image. Unlike existing ViT methods that only capture block-by-block unimodal information, the proposed PromptSG method shows that the interaction of different modalities can improve the performance of each modality. Compared with the CLIP-based method CLIP-ReID, when using ViT-B / 16 as the visual backbone, the proposed PromptSG outperforms it by 5.0% / 1.5% and 13.8% / 3.9% in mAP / Rank-1 on the Market-1501 and MSMT17 datasets. A key difference between CLIP-ReID and the proposed method is the combination of query-specific pseudo-labels. The results of the proposed method further emphasize that integrating textual information during inference can also improve performance. Compared with CNN-based methods, to ensure a fair comparison, the proposed method also implements PromptSG using the ResNet-50 backbone network. With the exception of the LTReID method that utilizes higher resolution images, our method consistently outperforms the other methods by a considerable margin, highlighting the robustness and superiority of our method on a variety of architectures.
[0064] Ablation experiment of the whole module: As shown in Table 2, comparing rows b) and c) with a), a similar conclusion is reached, that is, removing the text-to-image or image-to-text contrast loss leads to performance degradation on both datasets. Further comparing rows a) and d), it is observed that the reduction caused by removing semantic information is greater than the reduction caused by removing only ID-specific appearance information. The complete model PromptSG of the present invention utilizes language supervision of semantics and appearance during training, achieving significant performance improvement.
[0065] Ablation exploration experiment of personalized prompts: In order to better understand the learned pseudo-labels * Can it provide more fine-grained guidance for learning visual embeddings? We train a strong baseline model. During training and testing, the text prompts are not correlated with the s * The results in Table 3 show that the combination s * has a significant impact on the overall performance. When s is removed from the training process *The performance indicator mAP drops by 1.9% to 2.6% when . Although the present invention focuses on the unimodal re-identification task, the above paradigm also has the potential to be applied to multimodal test sets, for example, by combining image features with text to achieve better alignment for text-to-image person retrieval.
[0066] Comparison of training efficiency: As shown in Table 4, the present invention conducts a comparative analysis of the single-stage PromptSG and the two-stage CLIP-ReID methods, focusing on the number of learnable parameters and training speed. In terms of training parameters, CLIP-ReID introduces additional parameters through ID-wise learnable prompts on the basis of CLIP, while the method of the present invention is mainly extended through a fixed-size mapping network and an interaction module. Although CLIP-ReID has 2%-4% fewer parameters than the present invention on both datasets, it may experience continuous growth in parameters when the number of categories is large or dynamically changing. In contrast, PromptSG shows stronger robustness in the number of parameters and achieves about two times the training speed acceleration.
[0067] Visualization analysis: In order to intuitively understand and verify the effectiveness of the method of the present invention, a qualitative analysis was conducted and Figure 3 The visualization of the attention map is presented in Figure 1. Specifically, examples from the Market-1501 and MSMT17 datasets are shown, each containing two training images and two gallery images. Some challenging examples are selected, including images with complex backgrounds or showing multiple individuals. d) PromptSG is compared in depth with b) CLIP-ReID and c) PromptSG without image combination training. The results show that the method of the present invention has significant advantages in both the accuracy of capturing semantic information and the delicacy of focusing on appearance details. In the first row of examples from the Market-1501 dataset, the attention map of CLIP-ReID is often disturbed by background elements such as "vehicles" and it is difficult to focus on the target pedestrian. In contrast, although PromptSG without image combination training tends to emphasize semantic information related to "people", it mainly focuses on general positions such as the head, arms and legs, and the capture of appearance features is still insufficient. The method of the present invention not only accurately captures these key parts, but also deeply explores subtle appearance features such as hats and backpacks, so as to more accurately identify different individuals. In addition, in the example of the MSMT17 dataset, when multiple pedestrians appear in the image, the method of the present invention shows excellent filtering ability and can effectively eliminate unnecessary pedestrian interference. In this regard, CLIP-ReID cannot achieve this filtering effect.
[0068] Table 1 Protection effects of different methods on different models
[0069]
[0070]
[0071] Table 2 Ablation experiment
[0072]
[0073] Table 3 Personalized prompt ablation exploration
[0074]
[0075] Table 4 Comparison of training efficiency
[0076]
[0077] Another embodiment of the present invention provides a semantically guided person re-identification system based on text prompts, which includes:
[0078] Visual encoder, used to obtain visual embedding based on the input image;
[0079] The reverse network module is used to map the visual embedding to the text space to obtain pseudo-tokens, and integrate the pseudo-tokens into natural language sentences to obtain language hints for the input image;
[0080] Text encoder, used to get text embedding based on language cues;
[0081] The multimodal interaction module is trained using visual embeddings and text embeddings to enable interaction between image patches and language cues in a multimodal environment.
[0082] The recognition module is used to input the query image into the trained multimodal interaction module to obtain a feature vector that integrates visual and textual information, and use the feature vector that integrates visual and textual information to perform similarity retrieval in the pedestrian image database to obtain the pedestrian re-identification result.
[0083] The specific implementation process of each module refers to the above description of the method of the present invention.
[0084] Another embodiment of the present invention provides a computer device (computer, server, smart phone, etc.), which includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in the method of the present invention.
[0085] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, magnetic disk, optical disk), wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the steps of the method of the present invention are implemented.
[0086] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. It can be understood by those skilled in the art that various replacements, changes and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the contents disclosed in the embodiments of this specification, and the scope of protection of the present invention shall be subject to the scope defined in the claims.
Claims
1. A semantically guided person re-identification method based on textual cues, characterized in that: The following steps are involved: Input the training image into the visual encoder to obtain the visual embedding; Use the inverse network to map the visual embedding to the text space to obtain pseudo tokens, and integrate the pseudo tokens into natural language sentences to obtain language hints for the input image; Feed the language hint into the text encoder to get the text embedding; Using visual embedding and text embedding to train a multimodal interaction module, the multimodal interaction module enables interaction between image patches and language cues in a multimodal environment; The query image is input into the trained multimodal interaction module to obtain a feature vector that integrates visual and textual information. The feature vector that integrates visual and textual information is used to perform similarity retrieval in the pedestrian image database to obtain the pedestrian re-identification result.
2. The method according to claim 1, characterized in that The visual encoder and the text encoder are the visual encoder and text encoder in the CLIP model.
3. The method according to claim 1, characterized in that The method uses an inverse network to map the visual embedding to the text space to obtain a pseudo token, including ensuring that the learned pseudo token can effectively convey the context information of the image through a symmetric contrast loss; the symmetric contrast loss is: Among them, L i2t represents the contrast loss from image to text, L t2i represents the contrast loss from text to image, N represents the number of samples in batch processing, sim represents the cosine similarity function, and v i represents the i-th image in the batch, l p is with v n The corresponding text hint embedding, v p Indicates that p Corresponding visual embeddings, τ represents the temperature hyperparameter used to control the scaling of the similarity score.
4. The method according to claim 3, characterized in that The method maps the visual embedding to the text space using an inverse network to obtain pseudo tokens, including encouraging the pseudo tokens to capture visual details belonging to the same identity through a symmetric supervised contrast loss; the symmetric supervised contrast loss is: Among them, L Supcon represents the symmetric supervised contrast loss, represents the supervised contrastive loss for image to text, represents the supervised contrast loss from text to image, and P(i) represents the loss with v n , l n The set of positive samples with the same label, p + represents the positive sample index, Represents the positive sample of text prompt p + The embedding representation of Represents the visual positive sample p + The embedded representation of .
5. The method according to claim 1, characterized in that The multimodal interaction module contains a language-guided cross-attention layer that uses text embeddings as queries and the block-wise embeddings of the visual encoder as keys and values; Given a pair of images and hints (x, t p ), input the image x into the visual encoder and get a series of image blocks embedded in represents the global visual embedding, v1,...,v M Belong to the local image block embedding; the prompt t p Input the text encoder to get the text embedding l p ; The text embeddings are then projected into the query matrix Q, and the image patch embeddings are projected into the key matrix K and the value matrix V through two different linear projection layers; the interaction from image patches to text prompts is achieved in the following way: Where A represents the attention function and d represents the dimension of visual embedding.
6. The method according to claim 5, characterized in that Two Transformer blocks are added after the cross-attention layer, each Transformer block contains a self-attention layer and a feed-forward layer.
7. The method according to claim 1, characterized in that The pedestrian re-identification loss used in the training stage includes the triple loss L Triplet and identity classification loss L ID , where the ternary loss L rriplet The identity classification loss L is used to bring the same pedestrian closer and push different pedestrians away in the embedding space. ID Used to classify a given pedestrian image into its identity category; triple loss L Triplet and identity classification loss L ID The calculation formula is as follows: L Triplet =max(d p -d n +m,0) Where N represents the number of batch samples, y j represents the identity label of the pedestrian, p j represents the category prediction probability of pedestrian images, d p Represents the Euclidean distance of the positive sample pair of images, d n represents the Euclidean distance between image negative sample pairs, and m represents a preset threshold or interval.
8. A semantically guided person re-identification system based on textual cues, characterized in that: include: Visual encoder, used to obtain visual embedding based on the input image; The reverse network module is used to map the visual embedding to the text space to obtain pseudo-tokens, and integrate the pseudo-tokens into natural language sentences to obtain language hints for the input image; Text encoder, used to get text embedding based on language cues; The multimodal interaction module is trained using visual embeddings and text embeddings to enable interaction between image patches and language cues in a multimodal environment. The recognition module is used to input the query image into the trained multimodal interaction module to obtain a feature vector that integrates visual and textual information, and use the feature vector that integrates visual and textual information to perform similarity retrieval in the pedestrian image database to obtain the pedestrian re-identification result.
9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Multi-modal panoramic image blind quality evaluation method and system based on AI generation description
CN120356071A
Image cross-modal pedestrian re-identification method and system based on text clue alignment
CN120429469A
Method and device for identifying forest fire hidden danger of power transmission line based on multi-modal large model
CN120852793A
Fire hazard identification method and device for power transmission line based on multi-modal large model
CN120852793B
Face anti-counterfeiting method and system based on domain guide prompt distribution learning
CN121033948A