A cross-modal group alignment method for integrated multimodal person re-identification

Through the cross-modal group alignment method, combined with local feature filtering and hyperplane constraints, the problem of insufficient feature integration in multimodal pedestrian re-identification is solved, accurate recognition is achieved under conditions of lighting changes or occlusion, and the matching accuracy and model performance of pedestrian re-identification are improved.

CN119832599BActive Publication Date: 2025-09-26XIAMEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411938930.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-09-26
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

Existing person re-identification technologies fail to effectively integrate multimodal data in cross-modal retrieval, especially in cases of illumination changes or occlusion. Traditional methods cannot accurately match the fine-grained features of pedestrians, and local learning strategies ignore global relationships and fail to fully utilize the complementary information between images and text.

Method used

A cross-modal group alignment method is adopted. The RGB image and sketch features are extracted by sharing the image feature extractor, and the text features are extracted by combining the text feature extractor. Feature fusion is performed through the fusion feature extractor. Fine-grained feature deep fusion is performed using local feature filtering and cross-modal domain contrast learning modules. The global feature distribution is constrained to the hyperplane by combining the hyperplane constraint module to achieve alignment of features from different modalities.

Benefits of technology

It achieves accurate recognition under conditions of lighting changes or occlusion, enhances the matching accuracy of multimodal pedestrian re-identification, and improves the stability of feature optimization and model performance through cross-modal grouping domain contrast learning and hyperplane constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832599B_ABST
    Figure CN119832599B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-modal grouping alignment method for integrated multimodal pedestrian re-identification, comprising the following steps: S1, using a shared image feature extractor to extract features from RGB images and sketches, and using a text feature extractor to extract features from texts; S2, using a fusion feature extractor to fuse features of sketches and texts; S3, filtering out redundant features from local features through filtering processing, and then deeply fusing fine-grained features between modalities through a cross-modal domain contrast learning module to achieve fine-grained feature alignment; S4, using a hyperplane constraint module to constrain the distribution of global features of three modalities of the same pedestrian ID in a shared space to a hyperplane; S5, aligning the three modalities through contrast learning of the global features in the same hyperplane, and finally achieving text retrieval of RGB images, sketch retrieval of RGB images, and text fusion sketch retrieval of RGB images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and computer vision, and in particular to a cross-modal grouping alignment method for integrated multimodal person re-identification. Background Art

[0002] Person re-identification (ReID) aims to identify the same individual across different times, locations, and camera views, matching a given pedestrian image to a large-scale image or video database. It has a wide range of applications, including intelligent surveillance, public safety, traffic management, and smart retail. ReID can be categorized into single-modality and cross-modal retrieval.

[0003] While unimodal ReID only matches RGB images, cross-modal ReID allows retrieval of RGB images across multiple modalities, including infrared, text, and sketches. Text-based retrieval allows users to find pedestrian images based on natural language descriptions, while sketches provide more intuitive visual details. However, traditional models are typically trained on only a single modality, limiting the flexible use of multimodal data. In a typical urban surveillance scenario, a surveillance system searches for a missing woman based on eyewitness descriptions. Traditional feature grouping strategies may separately analyze the video image features from the camera and the text or sketch features extracted from the eyewitness description, failing to achieve fine-grained cross-modal feature fusion. When the details of the eyewitness description are not obvious in the image, such as the text description of a woman's scarf being red, but the scarf appears darker and no longer red due to changes in ambient light in the actual surveillance scene, traditional algorithms may fail to correctly match the pedestrian.

[0004] UNIReID integrates different retrieval methods (such as text, sketches, and RGB images) to meet diverse practical needs. However, UNIReID relies solely on global features for cross-modal alignment, potentially overlooking critical local details. Existing local learning paradigms typically employ hard segmentation strategies to extract local information, ignoring the global relationships between features. This leads to incomplete understanding, fails to fully exploit the potential complementary information between images and text, and fails to fully leverage the advantages of fusing features from different modalities. Summary of the Invention

[0005] To solve the above problems, the present invention proposes a cross-modal grouping alignment method for integrated multimodal person re-identification.

[0006] To achieve the above object, the present invention adopts the following technical solutions:

[0007] A cross-modal grouping alignment method for integrated multimodal person re-identification is characterized by comprising the following steps:

[0008] S1, use the same shared image feature extractor to extract features from RGB images and sketches, and use the text feature extractor to extract features from texts;

[0009] S2, using a fusion feature extractor to fuse the features of the sketch and text, to support the retrieval method of sketch fusion text;

[0010] S3: The local features are filtered to remove redundant features, and then the fine-grained features between modalities are deeply integrated through the cross-modal contrast learning module to achieve fine-grained feature alignment;

[0011] S4. Global features are constrained to a hyperplane through the hyperplane constraint module. The distribution of the three modal global features of the same pedestrian ID in the shared space is constrained to a hyperplane.

[0012] The specific process of step S4 is:

[0013] S41. The distance between features belonging to different categories but the same modality is smaller than a threshold, that is, the difference between the distance between the RGB modal features of different pedestrian categories and the distance between the sketch modality, text modality, and fusion modality within the same modality is smaller than the threshold. This can be formally expressed as the difference between the four edges of the three-dimensional figure being smaller than the threshold.

[0014] S42. For any two different pedestrian IDs, the distance between the RGB modal features of pedestrian ID1 and the sketch, text, and fused modal features of pedestrian ID2 is equal to the distance between the sketch, text, and fused modal features of pedestrian ID1 and the RGB modal features of pedestrian ID2. This is formally represented as the lengths of the diagonals on two planes and one diagonal plane of a solid figure being equal.

[0015] S43. Through the hyperplane constraint, the distribution of the features of the four modalities of the same pedestrian ID, namely RGB modality, sketch modality, text modality and fusion modality, in the shared space will present a hyperplane;

[0016] S5. The global features in the same hyperplane align the three modalities through contrastive learning, and finally realize text retrieval of RGB images, sketch retrieval of RGB images, and text fusion sketch retrieval of RGB images.

[0017] Preferably, the specific process of step S1 is:

[0018] S11. During the training process, the training data is defined as , where N represents the number of samples within a batch, I represents the RGB modality, S represents the sketch modality, T represents the text modality, and k represents the sample number;

[0019] S12. Input the data of the RGB modality and the sketch modality into the same image feature extractor, and input the data of the text modality into the text feature extractor, to obtain global features of the RGB modality, the sketch modality, and the text modality;

[0020] S13. Perform feature extraction on the patch tokens of the RGB modality and the sketch modality to obtain RGB local features and sketch local features, respectively; perform feature extraction on the vocabulary tokens of the text modality to obtain text local features.

[0021] Preferably, the specific process of step S2 is: splicing the global features of the sketch and text, adding a CLS identifier, and then passing it through a feature extractor, taking the CLS identifier and passing it through the output corresponding to the feature extractor. As a fusion feature of sketch and text.

[0022] Preferably, the specific process of the filtering processing in step S3 is: calculating the similarity scores of the global features and local features within the modality, then arranging them in descending order, selecting and retaining the top-ranked local features, and eliminating the remaining local features.

[0023] Preferably, the specific process of deep fusion of fine-grained features between modalities in step S3 is:

[0024] S31. For the image domain, the filtered local features of the text are mapped to the image domain through a linear layer. The calculation formula is: , where Represents the text features after mapping, Indicates linear mapping of the filtered text local features. Represents the feature dimension after mapping;

[0025] S32. For the text domain, the filtered RGB local features and sketch local features are mapped to the text domain through a linear layer. The calculation formula is: , , where Represents the mapped RGB features, Indicates linear mapping of the local features of the filtered RGB image. Represents the feature dimension after mapping; Represents the sketch feature after mapping, Indicates linear mapping of the filtered sketch local features. Represents the feature dimension after mapping;

[0026] S33, the characteristics of the image domain are , the characteristics of the text domain are , the features of the image domain and the text domain are fused through a cross-modal grouping strategy; for the image domain, the local features of each modality in the image domain are first randomly shuffled and divided into n equal parts, and then the local features of the three modalities are combined to generate n groups, each group contains local features of three different modalities; for the text domain, the local features of each modality in the text domain are first randomly shuffled and divided into n equal parts, and then the local features of the three modalities are combined to generate n groups, each group contains local features of three different modalities;

[0027] S34, extracting the global features within each group of the image domain and the text domain through an image feature extractor and a text feature extractor, and the global features within the group of the pth group in the image domain are , the global feature within the group p of the text domain is, , and then stack the global features of the same group number with different pedestrian IDs to form a feature matrix. The p-th feature matrix of the image domain and text domain are: , , where represents the feature matrix of the pth stack in the image domain, represents the global feature set within the p-th group of all pedestrians in the image domain, represents the feature matrix of the p-th stack of the text domain, represents the global feature set within the p-th group of all pedestrians in the text domain, Represents the dimensions of the stacked feature matrix;

[0028] All feature matrices in the image domain and text domain are expressed as: , , where Represents the feature matrix of all stacks in the image domain, Represents the collection of n stacked feature matrices in the image domain, Represents the collection of all stacked feature matrices of the text field, Represents the collection of n stacked feature matrices of the text domain, Represents the dimension of the collection of all stacked feature matrices in the image domain; All feature matrices in the pairwise calculation contrast loss, there will be loss components, and finally the loss in the image domain is obtained; All feature matrices in the pairwise comparison loss are calculated, and there will be loss components, and finally the loss of the text domain is obtained.

[0029] Preferably, the specific process of step S5 is: after the hyperplane constraint, contrastive learning is used for the three retrieval methods of text retrieval RGB image, sketch retrieval RGB image, and text fusion sketch retrieval RGB image, respectively, to maximize the cosine similarity between true sample pairs and minimize the cosine similarity between error sample pairs.

[0030] By adopting the above technical solution, the present invention has the following beneficial effects: It randomly mixes and groups local features from pedestrian images, sketches, and text, breaking down modal barriers and achieving deep, fine-grained cross-modal feature integration. The present invention's Cross-Modal Grouping Alignment (CMGA) method for integrated multimodal person re-identification aligns fine-grained features from different modalities, enhancing matching accuracy. This is motivated by the fact that by grouping red scarf features in images and sketches with the corresponding text description "red scarf," cross-modal information interaction can accurately identify scarf color despite lighting variations or occlusion. To this end, the present invention designs Cross-Modal Grouping In-Domain Contrastive Learning (CGIC), mapping RGB, sketch, and text features to the image and text domains for in-domain cross-modal grouping to achieve fine-grained feature fusion. Furthermore, the present invention employs contrastive learning to optimize the cross-modal embedding space. In addition, traditional methods often lead to chaotic feature distribution during the initial optimization process, resulting in unstable optimization. The present invention further introduces hyperplane constrained (HC) contrastive learning to ensure that different modal features of the same pedestrian ID are distributed on a common hyperplane in the shared space, which significantly improves feature optimization and model performance, and enhances the performance in multimodal pedestrian re-identification tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a flow chart of the present invention;

[0032] Figure 2 It is a framework structure diagram of the present invention;

[0033] Figure 3 It is a structural schematic diagram of the filtration process of the present invention;

[0034] Figure 4 Schematic diagram of the structure of the hyperplane constraint of the present invention. DETAILED DESCRIPTION

[0035] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0036] like Figures 1 to 4 As shown in FIG, a cross-modal grouping alignment method for integrated multimodal person re-identification includes the following steps:

[0037] S1, use the same shared image feature extractor to extract features from RGB images and sketches, and use the text feature extractor to extract features from texts;

[0038] The specific process of step S1 is:

[0039] S11. During the training process, the training data is defined as , where N represents the number of samples within a batch, I represents the RGB modality, S represents the sketch modality, T represents the text modality, and k represents the sample number;

[0040] S12. Input the data of the RGB modality and the sketch modality into the same image feature extractor, and input the data of the text modality into the text feature extractor, to obtain global features of the RGB modality, the sketch modality, and the text modality;

[0041] S13, performing feature extraction on the patch tokens of the RGB modality and the sketch modality to obtain RGB local features and sketch local features respectively; performing feature extraction on the vocabulary tokens of the text modality to obtain text local features;

[0042] S2, using a fusion feature extractor to fuse the features of the sketch and text, to support the retrieval method of sketch fusion text;

[0043] The specific process of step S2 is: splicing the global features of the sketch and text, adding a CLS identifier, and then passing it through a feature extractor, taking the CLS identifier and passing it through the output corresponding to the feature extractor as a fusion of sketch and text;

[0044] S3: The local features are filtered to remove redundant features, and then the fine-grained features between modalities are deeply integrated through the cross-modal contrast learning module to achieve fine-grained feature alignment;

[0045] The specific process of the filtering process in step S3 is: calculating the similarity scores of the global features and local features within the modality, then sorting them in descending order, selecting the top-ranked local features, and eliminating the remaining local features;

[0046] The specific process of deep fusion of fine-grained features between modalities in step S3 is as follows:

[0047] S31. For the image domain, the filtered local features of the text are mapped to the image domain through a linear layer. The calculation formula is: , where Represents the text features after mapping, Indicates linear mapping of the filtered text local features. Represents the feature dimension after mapping;

[0048] S32. For the text domain, the filtered RGB local features and sketch local features are mapped to the text domain through a linear layer. The calculation formula is: , , where Represents the mapped RGB features, Indicates linear mapping of the local features of the filtered RGB image. Represents the feature dimension after mapping; Represents the sketch feature after mapping, Indicates linear mapping of the filtered sketch local features. Represents the feature dimension after mapping;

[0049] S33, the characteristics of the image domain are , the characteristics of the text domain are , the features of the image domain and the text domain are fused through a cross-modal grouping strategy; for the image domain, the local features of each modality in the image domain are first randomly shuffled and divided into n equal parts, and then the local features of the three modalities are combined to generate n groups, each group contains local features of three different modalities; for the text domain, the local features of each modality in the text domain are first randomly shuffled and divided into n equal parts, and then the local features of the three modalities are combined to generate n groups, each group contains local features of three different modalities;

[0050] S34, extracting the global features within each group of the image domain and the text domain through an image feature extractor and a text feature extractor, and the global features within the group of the pth group in the image domain are , the global feature within the group p of the text domain is, , and then stack the global features of the same group number with different pedestrian IDs to form a feature matrix. The p-th feature matrix of the image domain and text domain are: , , where represents the feature matrix of the pth stack in the image domain, represents the global feature set within the p-th group of all pedestrians in the image domain, represents the feature matrix of the p-th stack of the text domain, represents the global feature set within the p-th group of all pedestrians in the text domain, Represents the dimensions of the stacked feature matrix;

[0051] All feature matrices in the image domain and text domain are expressed as: , , where Represents the feature matrix of all stacks in the image domain, Represents the collection of n stacked feature matrices in the image domain, Represents the collection of all stacked feature matrices of the text field, Represents the collection of n stacked feature matrices of the text domain, Represents the dimension of the collection of all stacked feature matrices in the image domain; All feature matrices in the pairwise comparison loss are calculated, and there will be loss components, and finally the loss in the image domain is obtained; All feature matrices in the pairwise comparison loss are calculated, and there will be loss components, and finally get the loss of the text domain;

[0052] S4. Global features are constrained to a hyperplane through the hyperplane constraint module. The distribution of the three modal global features of the same pedestrian ID in the shared space is constrained to a hyperplane.

[0053] The specific process of step S4 is:

[0054] S41. The distance between features belonging to different categories but the same modality is smaller than a threshold, that is, the difference between the distance between the RGB modal features of different pedestrian categories and the distance between the sketch modality, text modality, and fusion modality within the same modality is smaller than the threshold. This can be formally expressed as the difference between the four edges of the three-dimensional figure being smaller than the threshold.

[0055] S42. For any two different pedestrian IDs, the distance between the RGB modal features of pedestrian ID1 and the sketch, text, and fused modal features of pedestrian ID2 is equal to the distance between the sketch, text, and fused modal features of pedestrian ID1 and the RGB modal features of pedestrian ID2. This is formally represented as the lengths of the diagonals on two planes and one diagonal plane of a solid figure being equal.

[0056] S43. Through the hyperplane constraint, the distribution of the features of the four modalities of the same pedestrian ID, namely RGB modality, sketch modality, text modality and fusion modality, in the shared space will present a hyperplane;

[0057] S5. Global features in the same hyperplane are used to align the three modalities through contrastive learning, ultimately achieving text retrieval from RGB images, sketch retrieval from RGB images, and text-sketch fusion retrieval from RGB images.

[0058] The specific process of step S5 is as follows: after the hyperplane constraint, contrastive learning is used for the three retrieval methods of text retrieval RGB image, sketch retrieval RGB image, and text fusion sketch retrieval RGB image, respectively, to maximize the cosine similarity between true sample pairs and minimize the cosine similarity between error sample pairs.

[0059] Performance testing:

[0060] This test used three multimodal datasets constructed from the paper "Towards Modality-Agnostic Person Re-identification with Descriptive Query" by Cuiqun Chen (2023) at CVPR 2023: Tri-CUHK-PEDES, Tri-ICFG-PEDES, and Tri-RSTPReid. These datasets cover three modalities: RGB images (visible light images), text descriptions, and hand-drawn sketches (sketched images). The sketches were synthesized using the Meitu API, which is suitable for low-resolution pedestrian images. This method generates hand-drawn sketches of pedestrians based on the background-removed RGB images. Detailed dataset statistics are shown in Table 1. Multimodal data not only enriches the information sources for person re-identification but also provides an experimental basis for verifying the effectiveness and robustness of the proposed method.

[0061] Table 1: Statistics of the dataset

[0062]

[0063] For the Tri-CUHK-PEDES dataset, the training set contains 34,054 RGB and sketch images with 11,003 pedestrian identities; the validation set contains 3,078 RGB and sketch images with 1,000 pedestrian identities; and the test set contains 3,074 RGB and sketch images with 1,000 pedestrian identities. For the Tri-ICFG-PEDES dataset, the training set contains 34,674 RGB and sketch images with 3,102 identities; the test set contains 19,848 RGB and sketch images with 1,000 pedestrian identities. For the Tri-RSTPReid dataset, the training set contains 18,505 RGB and sketch images with 3,701 pedestrian identities; the validation set contains 1,000 RGB and sketch images with 200 pedestrian identities; and the test set contains 1,000 RGB and sketch images with 200 pedestrian identities.

[0064] The proposed cross-modal group alignment method (CMGA) for integrated multimodal person re-identification was compared with other existing methods. To ensure a fair comparison with existing methods, the proposed cross-modal group alignment method (CMGA) for integrated multimodal person re-identification was tested on three retrieval tasks using the original image collections of three datasets. The test results are shown in Table 2.

[0065] Table 2: Test results of CMGA of the present invention and other methods

[0066]

[0067] Results show that the proposed dual encoder for learning visual and textual modal features is more effective than most CNN-based methods. CMGA employs a cross-modal grouping and hyperplane constraint strategy to achieve fine-grained cross-modal feature integration and a unified initial feature optimization direction, resulting in superior retrieval performance. As shown in Table 2, the Rank 1 accuracy of the text-fused sketch retrieval RGB image task on the three datasets reached 88.42%, 84.13%, and 76.20%, respectively, which are 2.13%, 1.96%, and 3% higher than UNIReid, respectively. This significant performance improvement demonstrates the effectiveness and potential of the proposed cross-modal grouping alignment method (CMGA) for integrated multimodal person re-identification in multimodal retrieval tasks.

[0068] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A cross-modal grouping alignment method for integrated multimodal person re-identification, characterized by: The following steps are involved: S1, use the same shared image feature extractor to extract features from RGB images and sketches, and use the text feature extractor to extract features from texts; S2, using a fusion feature extractor to fuse the features of the sketch and text, to support the retrieval method of sketch fusion text; S3: The local features are filtered to remove redundant features, and then the fine-grained features between modalities are deeply integrated through the cross-modal contrast learning module to achieve fine-grained feature alignment; S4. Global features are constrained to a hyperplane through the hyperplane constraint module. The distribution of the three modal global features of the same pedestrian ID in the shared space is constrained to a hyperplane. The specific process of step S4 is: S41. The distance between features belonging to different categories but the same modality is smaller than a threshold, that is, the difference between the distance between the RGB modal features of different pedestrian categories and the distance between the sketch modality, text modality, and fusion modality within the same modality is smaller than the threshold. This can be formally expressed as the difference between the four edges of the three-dimensional figure being smaller than the threshold. S42. For any two different pedestrian IDs, the distance between the RGB modal features of pedestrian ID1 and the sketch, text, and fused modal features of pedestrian ID2 is equal to the distance between the sketch, text, and fused modal features of pedestrian ID1 and the RGB modal features of pedestrian ID2. This is formally represented as the lengths of the diagonals on two planes and one diagonal plane of a solid figure being equal. S43. Through the hyperplane constraint, the distribution of the features of the four modalities of the same pedestrian ID, namely RGB modality, sketch modality, text modality and fusion modality, in the shared space will present a hyperplane; S5. The global features in the same hyperplane align the three modalities through contrastive learning, and finally realize text retrieval of RGB images, sketch retrieval of RGB images, and text fusion sketch retrieval of RGB images.

2. The cross-modal grouping alignment method for integrated multimodal person re-identification according to claim 1, characterized in that: The specific process of step S1 is: S11. During the training process, the training data is defined as , where N represents the number of samples within a batch, I represents the RGB modality, S represents the sketch modality, T represents the text modality, and k represents the sample number; S12. Input the data of the RGB modality and the sketch modality into the same image feature extractor, and input the data of the text modality into the text feature extractor, to obtain global features of the RGB modality, the sketch modality, and the text modality; S13. Perform feature extraction on the patch tokens of the RGB modality and the sketch modality to obtain RGB local features and sketch local features, respectively; perform feature extraction on the vocabulary tokens of the text modality to obtain text local features.

3. The cross-modal grouping alignment method for integrated multimodal person re-identification according to claim 1, characterized in that: The specific process of step S2 is: splicing the global features of the sketch and text, adding a CLS identifier, and then passing it through a feature extractor, taking the CLS identifier and passing it through the output corresponding to the feature extractor As a fusion feature of sketch and text.

4. The cross-modal grouping alignment method for integrated multimodal person re-identification according to claim 1, characterized in that: The specific process of the filtering process in step S3 is: calculating the similarity scores of the global features and local features within the modality, then sorting them in descending order, selecting the top-ranked local features, and eliminating the remaining local features.

5. The cross-modal grouping alignment method for integrated multimodal person re-identification according to claim 1, characterized in that: The specific process of deep fusion of fine-grained features between modalities in step S3 is as follows: S31. For the image domain, the filtered local features of the text are mapped to the image domain through a linear layer. The calculation formula is: , where Represents the text features after mapping, Indicates linear mapping of the filtered text local features. Represents the feature dimension after mapping; S32. For the text domain, the filtered RGB local features and sketch local features are mapped to the text domain through a linear layer. The calculation formula is: , , where Represents the mapped RGB features, Indicates linear mapping of the local features of the filtered RGB image. Represents the feature dimension after mapping; Represents the sketch feature after mapping, Indicates linear mapping of the filtered sketch local features. Represents the feature dimension after mapping; S33, the characteristics of the image domain are , the characteristics of the text domain are , the features of the image domain and the text domain are fused through a cross-modal grouping strategy; for the image domain, the local features of each modality in the image domain are first randomly shuffled and divided into n equal parts, and then the local features of the three modalities are combined to generate n groups, each group contains local features of three different modalities; for the text domain, the local features of each modality in the text domain are first randomly shuffled and divided into n equal parts, and then the local features of the three modalities are combined to generate n groups, each group contains local features of three different modalities; S34, extracting the global features within each group of the image domain and the text domain through an image feature extractor and a text feature extractor, and the global features within the group of the pth group in the image domain are , the global feature within the group p of the text domain is, , and then stack the global features of the same group number with different pedestrian IDs to form a feature matrix. The p-th feature matrix of the image domain and text domain are: , , where represents the feature matrix of the pth stack in the image domain, represents the global feature set within the p-th group of all pedestrians in the image domain, represents the feature matrix of the p-th stack of the text domain, represents the global feature set within the p-th group of all pedestrians in the text domain, Represents the dimensions of the stacked feature matrix; All feature matrices in the image domain and text domain are expressed as: , , where Represents the feature matrix of all stacks in the image domain, Represents the collection of n stacked feature matrices in the image domain, Represents the collection of all stacked feature matrices of the text field, Represents the collection of n stacked feature matrices of the text domain, Represents the dimension of the collection of all stacked feature matrices in the image domain; All feature matrices in the pairwise comparison loss are calculated, and there will be loss components, and finally the loss in the image domain is obtained; All feature matrices in the pairwise comparison loss are calculated, and there will be loss components, and finally the loss of the text domain is obtained.

6. The cross-modal grouping alignment method for integrated multimodal person re-identification according to claim 1, characterized in that: The specific process of step S5 is as follows: after the hyperplane constraint, contrastive learning is used for the three retrieval methods of text retrieval RGB image, sketch retrieval RGB image, and text fusion sketch retrieval RGB image, respectively, to maximize the cosine similarity between true sample pairs and minimize the cosine similarity between error sample pairs.

Citation Information

Patent Citations

  • Identity recognition method based on face-fingerprint cooperation

    CN102622590A

  • Multi-modal identity authentication method based on vein similar image knowledge migration network

    CN112241680A