A text-to-image cross-modal person re-identification method and system
By using overall optimization loss and local optimization loss in text-to-image cross-modal pedestrian re-identification, combined with text mask prediction loss, and optimizing training image and text feature encoder, the problems of inconsistent cross-modal global matching relationship and poor local feature correlation are solved, significantly improving the recognition accuracy.
Patent Information
- Application Number
- CN202411318602.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-20
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2044-09-20
AI Technical Summary
In the prior art, due to inconsistent cross-modal global matching relationships caused by different text descriptions, and due to the complexity of the scene and the diversity of text descriptions, there is redundant information between the image and text description, resulting in no good correlation between local features, which in turn leads to lower cross-modal recognition accuracy of text-to-image.
Through the proposed overall optimization loss and local optimization loss, combined with text mask prediction loss, the optimization training loss function of image and text feature encoder is constructed, the encoder is trained, and the cross-modal granularity correlation invariant features is explored, which significantly enhances the ability of deep models to extract key pedestrian information.
It effectively solves the problem of inconsistent global matching relationships, and through iterative optimization of fine-grained information, the accuracy of text-to-image pedestrian recognition search is improved.
Smart Images

Figure CN119295779B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a text-to-image cross-modal pedestrian re-identification method and system. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Text-to-image person re-identification (TI ReID) is a cross-modal retrieval task that aims to identify and locate pedestrian images from a large-scale image library based on a given text description query. Unlike traditional image-based person re-identification tasks, TI ReID can effectively utilize different text descriptions to complete target person matching, thereby making up for the shortcomings of the target image. However, since text and images belong to different modalities, text descriptions mainly focus on abstract concepts, while images mainly provide intuitive visual expressions. This modality heterogeneity makes it difficult to align different sample features, significantly increasing the difficulty of retrieval.
[0004] At present, some methods use visual and language encoders to extract image / text features, and then use contrastive loss to align the entire feature distribution; in addition, considering that the pre-trained visual-language model has a certain generalization ability, some researchers fine-tune the pre-trained model to mine cross-modal invariant pedestrian information and achieve good recognition performance. However, these methods only consider the global matching relationship equivalently, without paying attention to the inconsistency of the matching relationship between samples. In actual scenarios, text annotations are usually manually annotated or machine-generated, which inevitably introduces a certain degree of subjectivity and randomness.
[0005] Generally speaking, detailed text descriptions tend to produce higher similarities, while rough text descriptions tend to produce lower similarities. Therefore, the global matching relationship between images and texts shows significant inconsistency and should not be treated equally. Existing methods mainly use global information to achieve feature alignment and ignore fine-grained features. In this regard, some studies use part-of-speech tagging methods to extract important noun phrases in texts, and use auxiliary techniques such as pose estimation and image segmentation to locate key areas in images, explicitly establish connections between important text and image local features, and then reduce cross-modal feature differences from a local perspective. The recognition performance has been significantly improved, which proves the importance of local features. However, due to the complexity of the scene and the diversity of text descriptions, there is redundant information between image and text descriptions. General local learning methods find it difficult to mine important local features, resulting in poor correlation between local features. Summary of the invention
[0006] The present invention provides a text-to-image cross-modal pedestrian re-identification method and system to solve the problems in existing methods, such as inconsistent cross-modal global matching relationships caused by different text descriptions, and redundant information between image and text descriptions due to the complexity of the scene and the diversity of text descriptions, resulting in a lack of good correlation between local features, which in turn leads to low cross-modal recognition accuracy from text to image.
[0007] According to a first aspect of an embodiment of the present invention, a text-to-image cross-modal person re-identification method is provided, comprising:
[0008] Get the text description to be queried and its corresponding image library;
[0009] For the text description and image library, the pre-trained image and text feature encoders are used respectively to obtain text features and image features of image samples in the image library; wherein the training of the image and text feature encoders is specifically as follows: for the input training samples, the image feature encoder is used to extract the image attention map and image features of the training samples; the text feature encoder is used to extract the text attention map and text features of the training samples; based on the obtained image features and text features, the overall matching loss is calculated; based on the obtained image attention map and text attention map, feature extraction is performed to obtain image local features and text local features, and based on the image local features and text local features, the local matching loss is calculated; based on the overall matching loss and local matching loss, combined with the text mask prediction loss, an optimized training loss function of the image and text feature encoder is constructed, and the encoder is trained based on the optimized training loss function;
[0010] Based on the obtained text features and the image features of image samples in the image library, the image corresponding to the text description to be queried is determined through similarity calculation, thus realizing text-to-image cross-modal pedestrian re-identification.
[0011] Furthermore, the overall matching loss specifically includes a global matching loss from image to text and a global matching loss from text to image, wherein the global matching loss from image to text represents the KL divergence between the matching probability from image to text and the true matching probability; the global matching loss from text to image represents the KL divergence between the matching probability from text to image and the true matching probability.
[0012] Furthermore, the local matching loss is specifically expressed as follows:
[0013] L LJLS =ηL trl +(1-η)L id
[0014] Where η represents the decay coefficient, which decays from 1 to 0 as the number of model training rounds increases; Ltrl represents the triple loss function, L id represents the identity loss function.
[0015] Furthermore, the text mask prediction loss is obtained based on a pre-built cross encoder, wherein the cross encoder includes a multi-head cross attention layer and several standard Transformer blocks.
[0016] Furthermore, the text mask prediction loss is specifically expressed as follows:
[0017]
[0018] Where M represents the set of masked text tokens; W represents the size of the vocabulary; Indicates the predicted probability that the wth token is a real word; represents the predicted probability that the wth token is the lth word; y n Represents the ground truth word at the masked position.
[0019] Furthermore, the image and text feature encoders both adopt a CLIP model, which includes a plurality of self-attention layers and a feedforward neural network.
[0020] Furthermore, the feature extraction is performed based on the obtained image attention map and text attention map to obtain image local features and text local features, specifically: obtaining the image attention map and text attention map of the last layer in the image feature encoder and the text feature encoder, and selecting image local features and text local features from the image attention map and the text attention map according to a preset strategy.
[0021] According to a second aspect of an embodiment of the present invention, a text-to-image cross-modal person re-identification system is provided, comprising:
[0022] A data acquisition unit, which is used to acquire a text description to be queried and its corresponding image library;
[0023] A feature extraction unit, which is used to obtain text features and image features of image samples in the image library using pre-trained image and text feature encoders for the text description and image library respectively; wherein the training of the image and text feature encoders is specifically as follows: for the input training samples, the image feature encoder is used to extract the image attention map and image features of the training samples; the text feature encoder is used to extract the text attention map and text features of the training samples; based on the obtained image features and text features, the overall matching loss is calculated; based on the obtained image attention map and text attention map, feature extraction is performed to obtain image local features and text local features, and based on the image local features and text local features, the local matching loss is calculated; based on the overall matching loss and local matching loss, combined with the text mask prediction loss, an optimized training loss function of the image and text feature encoder is constructed, and the encoder is trained based on the optimized training loss function;
[0024] The pedestrian re-identification unit is used to determine the image corresponding to the text description to be queried based on the obtained text features and the image features of the image samples in the image library through similarity calculation, so as to realize text-to-image cross-modal pedestrian re-identification.
[0025] According to a third aspect of an embodiment of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored and running on the memory, wherein when the processor executes the program, the text-to-image cross-modal pedestrian re-identification method is implemented.
[0026] According to a fourth aspect of an embodiment of the present invention, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the text-to-image cross-modal pedestrian re-identification method is implemented.
[0027] Compared with the prior art, the present invention has the following beneficial effects:
[0028] The scheme described in the present invention provides a text-to-image cross-modal pedestrian re-identification method and system. The scheme effectively solves the problem of inconsistent global matching relationships caused by different text descriptions through the proposed overall optimization loss. Through the proposed local optimization loss, fine-grained information is iteratively optimized from the perspectives of metric learning and representation learning. At the same time, combined with the text mask prediction loss, optimized training of text and image feature encoders is achieved, cross-modal granularity-related invariant features are explored, the ability of deep models to extract key pedestrian information is significantly enhanced, and the accuracy of text-to-image pedestrian re-identification retrieval is improved.
[0029] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0031] Figure 1 A flowchart of a text-to-image cross-modal person re-identification method described in an embodiment of the present invention;
[0032] Figure 2 It is a schematic diagram of the overall model structure described in an embodiment of the present invention. DETAILED DESCRIPTION
[0033] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0034] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.
[0035] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0036] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.
[0037] In order to solve the problems existing in the prior art, such as Figure 1 As shown, this embodiment provides a text-to-image cross-modal pedestrian re-identification method, including:
[0038] Step 1: Get the text description to be queried and its corresponding image library;
[0039] Step 2: For the text description and the image library, use pre-trained image and text feature encoders to obtain text features and image features of image samples in the image library;
[0040] The training of the image and text feature encoder is specifically as follows: for the input training sample, the image feature encoder is used to extract the image attention map and image features of the training sample; the text feature encoder is used to extract the text attention map and text features of the training sample; based on the obtained image features and text features, the overall matching loss is calculated; based on the obtained image attention map and text attention map, feature extraction is performed to obtain image local features and text local features, and based on the image local features and text local features, the local matching loss is calculated; based on the overall matching loss and local matching loss, combined with the text mask prediction loss, an optimized training loss function of the image and text feature encoder is constructed, and the encoder is trained based on the optimized training loss function;
[0041] In a specific implementation, the overall matching loss specifically includes a global matching loss from image to text and a global matching loss from text to image, wherein the global matching loss from image to text represents the KL divergence between the matching probability from image to text and the true matching probability; the global matching loss from text to image represents the KL divergence between the matching probability from text to image and the true matching probability.
[0042] The text mask prediction loss is obtained based on a pre-built cross encoder, where the cross encoder includes a multi-head cross attention layer and several standard Transformer blocks.
[0043] The image and text feature encoders both adopt the CLIP model, which includes several self-attention layers and a feedforward neural network.
[0044] It should be noted here that the reason for choosing the CLIP model is:
[0045] (1) As a pre-trained model of visual language, CLIP already has a large amount of prior knowledge about image-text matching. Fine-tuning on downstream tasks can quickly capture the rich semantic relationship between visual and text samples and improve the overall recognition accuracy.
[0046] (2) The number of parameters of CLIP is relatively small compared to some other visual language models such as ALBEF and BLIP, and the space occupied is very considerable. Generally, a 3090 GPU can meet the training process of the entire model, which is a very good choice when computing power is insufficient.
[0047] Furthermore, regarding the use of cross-encoders, most of the existing works use a dual-tower structure, that is, only two encoders for images and texts are trained, and matching loss is used for feature alignment. It is obvious that the features of the two modalities lack interactive behavior. The use of cross-encoders can achieve full interaction between local information of the two modalities, and use MLM classification tasks to improve the model's ability to extract detailed information.
[0048] Specifically, during the training process of the image and text feature encoder, the following processing is specifically performed:
[0049] (1) Training sample preprocessing
[0050] Before each round of model optimization, the images and texts are preprocessed, and then the image and text feature encoders based on the CLIP model are used to extract the feature representations of the images and texts respectively.
[0051] In one or more embodiments, the size of the input image is 384×128, where 384 is the image height and 128 is the image width.
[0052] In a specific implementation, the image preprocessing includes: performing random horizontal flipping, random cropping, random erasing, and random filling data enhancement methods on the pedestrian image, and performing normalization operations on the pedestrian image.
[0053] In one or more embodiments, the random horizontal flip probability may be set to 0.5; the random padding may be set to 10 pixels; and the random cropped image size may be set to 384×128.
[0054] In a specific implementation, the preprocessing of the text includes: word segmentation of the text and data enhancement methods of random masking.
[0055] In one or more implementations, word segmentation may be performed using a BPE algorithm, the probability of a random mask may be set to 0.15, and after feature extraction, the dimension of each image feature and text feature is 512.
[0056] (2) Loss function construction
[0057] 1) Overall matching optimization loss
[0058] Based on the extracted sample features, an overall matching optimization scheme is proposed to constrain the global feature distribution of image features and text features; the overall matching optimization loss is expressed as:
[0059]
[0060] in, represents the global matching loss from image to text; represents the global matching loss for text to image.
[0061] Among them, the global matching loss from image to text It is expressed as:
[0062]
[0063] in, Represents a dynamic constraint factor; ∈ represents a minimum value to prevent the denominator from being zero; represents the matching probability between the i-th image and the j-th text feature; q i,j represents the true matching probability; N represents the total number of samples.
[0064] Optionally, the matching probability p i,j The calculation is expressed as:
[0065]
[0066] Where τ represents the temperature hyperparameter; Represents the cosine similarity between the i-th image feature and the j-th text feature.
[0067] Among them, the dynamic constraint factor It is expressed as:
[0068]
[0069] Among them, α represents a hyperparameter; Indicates calculating the cosine similarity between the features of the nth image and text pair.
[0070] In one or more implementations, α may be set to 1.45.
[0071] 2) Local features and local optimization losses
[0072] Based on the attention map in the CLIP model and the extracted sample features, important local features are selected; then the selected local features are processed by a maximum pooling operator to obtain a final local representation;
[0073] Among them, for the local features of the selected images and texts and The corresponding maximum pooling process is expressed as:
[0074]
[0075] Among them, MLP represents multi-layer perceptron; FC represents fully connected layer; Maxpool represents maximum pooling function.
[0076] In one or more embodiments, the local features of the image and text and The corresponding selection method includes: first, select the attention map of the last layer in the CLIP model. The attention map reflects the dependency between different local features. According to experience, 30% of the local features with the highest weight are selected. and Then perform the selected local features Normalize to obtain regularized features and
[0077] In the specific implementation, based on the pooled local features, a local joint learning strategy is proposed to iteratively optimize the cross-modal fine-grained information;
[0078] Among them, for the local feature v after pooling local and t local The corresponding local joint learning strategy is expressed as:
[0079]
[0080] Among them, η represents the decay coefficient, and its value decays from 1 to 0 as the number of model training rounds increases; represents the triplet loss function; represents the identity loss function.
[0081] Among them, the triple loss function It is expressed as:
[0082]
[0083] Among them, m represents the hyperparameter that controls the interval between positive and negative samples; Indicates that The most difficult local features of text; Indicates that The most difficult local features of an image.
[0084] In one or more embodiments, m may be set to 0.1.
[0085] Among them, the local features of the i-th image The corresponding local features of the most difficult negative text The selection method includes: firstly calculating the local features of the i-th image With all local features of text t local The cosine similarity between them is calculated, and all similarities are sorted from high to low; except The corresponding positive text features The feature with the highest similarity other than
[0086] Among them, the identity loss function It is expressed as:
[0087]
[0088] in, and represents the predicted probability of the true identity; and Represents the regularized global features of images and text.
[0089] 3) Final loss representation
[0090] Combine global and local constraints to jointly optimize the deep learning network and gradually explore cross-modal granularity-related invariant features;
[0091] Among them, the final loss for optimizing the deep learning model is expressed as:
[0092]
[0093] in, represents the text mask prediction loss, represents the overall matching optimization loss, represents the local joint optimization loss.
[0094] Among them, the text mask prediction loss It is expressed as:
[0095]
[0096] Where M represents the set of masked text tokens; W represents the size of the vocabulary; Indicates the predicted probability that the wth token is a real word; represents the predicted probability that the wth token is the lth word; y n Represents the ground truth word at the masked position.
[0097] Step 3: Based on the obtained text features and the image features of the image samples in the image library, the image corresponding to the text description to be queried is determined through similarity calculation to achieve text-to-image cross-modal pedestrian re-identification.
[0098] Based on the optimized deep learning network, the text query features of pedestrians are extracted, pedestrian image matching is performed, and pedestrian recognition results are obtained.
[0099] Finally, in order to prove the effectiveness of the scheme described in this embodiment, this embodiment uses a large-scale public pedestrian re-identification database as a test object. For example, when tested on the CUHK-PEDES database, the correct search rate of the first matching image returned by the scheme described in this embodiment reaches 74.03%, and the average accuracy reaches 66.27%. The pedestrian re-identification method of this embodiment effectively explores the difficult sample problem of pedestrian images and greatly improves the correct search rate of pedestrian re-identification, which shows the effectiveness of the method of the present invention.
[0100] In one or more embodiments, a text-to-image cross-modal person re-identification system is provided corresponding to the above method, including:
[0101] A data acquisition unit, which is used to acquire a text description to be queried and its corresponding image library;
[0102] A feature extraction unit, which is used to obtain text features and image features of image samples in the image library using pre-trained image and text feature encoders for the text description and image library respectively; wherein the training of the image and text feature encoders is specifically as follows: for the input training samples, the image feature encoder is used to extract the image attention map and image features of the training samples; the text feature encoder is used to extract the text attention map and text features of the training samples; based on the obtained image features and text features, the overall matching loss is calculated; based on the obtained image attention map and text attention map, feature extraction is performed to obtain image local features and text local features, and based on the image local features and text local features, the local matching loss is calculated; based on the overall matching loss and local matching loss, combined with the text mask prediction loss, an optimized training loss function of the image and text feature encoder is constructed, and the encoder is trained based on the optimized training loss function;
[0103] The pedestrian re-identification unit is used to determine the image corresponding to the text description to be queried based on the obtained text features and the image features of the image samples in the image library through similarity calculation, so as to realize text-to-image cross-modal pedestrian re-identification.
[0104] It can be understood that the system described in this embodiment corresponds to the method in the above embodiment, and its technical details are described in detail in Embodiment 1, so they are not repeated here.
[0105] In further embodiments, there is also provided:
[0106] An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed on the processor, wherein when the computer instructions are executed by the processor, the method described in Embodiment 1 is performed. For the sake of brevity, no further description is given here.
[0107] It should be understood that in this embodiment, the processor may be a central processing unit CPU, and the processor may also be other general-purpose processors, digital signal processors DSP, application-specific integrated circuits ASIC, off-the-shelf programmable gate arrays FPGA or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0108] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A portion of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.
[0109] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the method described in embodiment 1 is completed.
[0110] The method in the first embodiment can be directly embodied as a hardware processor, or a combination of hardware and software modules in the processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.
[0111] Those skilled in the art will appreciate that the units, i.e., algorithm steps, of the various examples described in the present embodiment can be implemented in electronic hardware or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0112] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A text-to-image cross-modal person re-identification method, characterized in that: include: Get the text description to be queried and its corresponding image library; For the text description and image library, the pre-trained image and text feature encoders are used respectively to obtain text features and image features of image samples in the image library; wherein the training of the image and text feature encoders is specifically as follows: for the input training samples, the image feature encoder is used to extract the image attention map and image features of the training samples; the text feature encoder is used to extract the text attention map and text features of the training samples; based on the obtained image features and text features, the overall matching loss is calculated; based on the obtained image attention map and text attention map, feature extraction is performed to obtain image local features and text local features, and based on the image local features and text local features, the local matching loss is calculated; based on the overall matching loss and local matching loss, combined with the text mask prediction loss, an optimized training loss function of the image and text feature encoder is constructed, and the encoder is trained based on the optimized training loss function; Based on the obtained text features and the image features of image samples in the image library, the image corresponding to the text description to be queried is determined through similarity calculation, thus realizing text-to-image cross-modal pedestrian re-identification.
2. The text-to-image cross-modal person re-identification method according to claim 1, characterized in that: The overall matching loss specifically includes a global matching loss from image to text and a global matching loss from text to image, wherein the global matching loss from image to text represents the KL divergence between the matching probability from image to text and the true matching probability; the global matching loss from text to image represents the KL divergence between the matching probability from text to image and the true matching probability.
3. The text-to-image cross-modal person re-identification method according to claim 1, characterized in that: The local matching loss is specifically expressed as follows: Where η represents the decay coefficient, which decays from 1 to 0 as the number of model training rounds increases; represents the triplet loss function, represents the identity loss function.
4. The text-to-image cross-modal person re-identification method according to claim 1, characterized in that: The text mask prediction loss is obtained based on a pre-built cross encoder, where the cross encoder includes a multi-head cross attention layer and several standard Transformer blocks.
5. The text-to-image cross-modal person re-identification method according to claim 1, characterized in that: The text mask prediction loss is specifically expressed as follows: Where M represents the set of masked text tokens; W represents the size of the vocabulary; Indicates the predicted probability that the wth token is a real word; represents the predicted probability that the wth token is the lth word; y n Represents the ground truth word at the masked position.
6. The text-to-image cross-modal person re-identification method according to claim 1, characterized in that: The image and text feature encoders both adopt the CLIP model, which includes several self-attention layers and a feedforward neural network.
7. The text-to-image cross-modal person re-identification method according to claim 1, characterized in that: The method performs feature extraction based on the obtained image attention map and text attention map to obtain image local features and text local features, specifically comprising: obtaining the image attention map and text attention map of the last layer in the image feature encoder and the text feature encoder, and selecting image local features and text local features from the image attention map and the text attention map according to a preset strategy.
8. A text-to-image cross-modal person re-identification system, characterized in that: include: A data acquisition unit, which is used to acquire a text description to be queried and its corresponding image library; A feature extraction unit, which is used to obtain text features and image features of image samples in the image library using pre-trained image and text feature encoders for the text description and image library respectively; wherein the training of the image and text feature encoders is specifically as follows: for the input training samples, the image feature encoder is used to extract the image attention map and image features of the training samples; the text feature encoder is used to extract the text attention map and text features of the training samples; based on the obtained image features and text features, the overall matching loss is calculated; based on the obtained image attention map and text attention map, feature extraction is performed to obtain image local features and text local features, and based on the image local features and text local features, the local matching loss is calculated; based on the overall matching loss and local matching loss, combined with the text mask prediction loss, an optimized training loss function of the image and text feature encoder is constructed, and the encoder is trained based on the optimized training loss function; The pedestrian re-identification unit is used to determine the image corresponding to the text description to be queried based on the obtained text features and the image features of the image samples in the image library through similarity calculation, so as to realize text-to-image cross-modal pedestrian re-identification.
9. An electronic device comprising a memory, a processor and a computer program stored and running on the memory, characterized in that: When the processor executes the program, the text-to-image cross-modal pedestrian re-identification method as described in any one of claims 1-7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, a text-to-image cross-modal pedestrian re-identification method as described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Cross-modal retrieval method based on multilevel feature representation alignment
CN113792207A
Multi-modal pedestrian re-identification method based on multi-level cross-modal difference harmonic
CN116682144A