Cross-modal pedestrian re-identification method and system based on feature enhancement and association optimization reasoning
Through the method of fine-grained feature enhancement and association optimization inference, the problem of large modal differences and insufficient utilization of association information in cross-modal pedestrian re-identification is solved, and efficient identification under different light conditions is achieved.
Patent Information
- Application Number
- CN202510283188.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-29
AI Technical Summary
In cross-modal pedestrian re-identification, multimodal image feature style differences are large and difficult to align. The existing inference strategy does not use the associated information sufficiently, resulting in poor recognition effect in environments with large light changes.
A method based on feature enhancement and association optimization inference is adopted, through fine-grained feature enhancement and association perception inference, a channel-level modal alignment loss function is designed, and a two-stage inference mechanism is used to extract and correct features using modal shared features and modal specific features to improve the accuracy of similarity calculation.
Effectively alleviate modal differences, improve the robustness and accuracy of cross-modal retrieval, and is suitable for complex security monitoring scenarios.
Smart Images

Figure CN120388392A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of robot image processing, and particularly relates to a cross-modal pedestrian re-identification method and system based on feature enhancement and correlation optimization inference. Background Art
[0002] With the wide popularization of monitoring devices and the continuous reduction of network transmission costs, video surveillance has become one of the key technologies for ensuring social public safety. As an important part of the security field, video surveillance can realize functions such as video acquisition, processing, control, and emergency command of the monitored scene, and is widely used in multiple fields such as public security, transportation, water conservancy, and finance. However, with the continuous growth of public security requirements and the rapid progress of hardware technology, the resolution of video cameras in the security field has been continuously improved, and the deployment scenarios have become increasingly diversified, resulting in an explosive growth in the data scale, which poses new challenges to the efficient analysis and processing of data. In recent years, the rapid development of deep learning technology has brought new opportunities to the field of video surveillance. Modern intelligent security systems can efficiently acquire, store, and analyze massive amounts of monitored video data by means of deep learning methods. By performing feature recognition and extraction on people, vehicles, objects, etc. in the video, the system can achieve identity verification and target recognition, enabling the machine to "understand" the world and actively make predictions and responses. In addition, the application of complex deep learning algorithms in video analysis not only improves the intelligent level of the system, but also brings significant economic benefits and business growth to related fields.
[0003] Person re-identification (PIR) utilizes computer vision techniques to determine whether a specific person exists within an image or video sequence. Specifically, it involves detecting the presence of a target person at a specific time using camera-captured data. Currently, most research is based on the most common visible light image data, leveraging discriminative attributes such as appearance and texture to enable cross-camera person retrieval. With the advancement of deep learning, neural network-based feature extraction methods have eliminated the traditional complex manual extraction process, enabling more diverse and generalizable feature representations and achieving significant success in PIR. However, in practical applications, the image quality of standard visible light cameras degrades significantly in low-light or highly variable environments, making it difficult for PIR models based on visible light images to perform effectively and often failing to meet practical application requirements. With the advancement of hardware, more video surveillance systems are integrating visible light and infrared cameras. Leveraging the infrared cameras' superior imaging capabilities at night and in low-light conditions, they enable 24-hour, uninterrupted video monitoring. However, due to the significant style differences and data heterogeneity between the two modalities, there is a need to conduct research on pedestrian image retrieval tasks across visible light and infrared modalities. This cross-modal pedestrian re-identification technology is also of great significance in ensuring public safety.
[0004] The main challenge in cross-modal person re-identification lies in the "cross-modal" problem. How to model the modality well to reduce the differences between the two modality images and learn the robust features shared between the two modalities is the key to current research. Currently, the research on cross-modal person re-identification can be roughly divided into the following three categories. The first category is to use metric learning to force the feature distributions extracted from different modality images to be more consistent. The second category is to learn a neural network that can map different modality images to a shared feature space by sharing some deep networks. The third category is to perform some transformation processing at the image level, such as using a generative adversarial network to transform the images of the two modalities into a consistent modality, and then using the same network to extract features for retrieval. One of the most intuitive ways to solve cross-modal tasks is to reduce modality differences at the data level. One type of existing method is generative-based. It mainly applies generative adversarial networks (GANs) to generate missing modality information or perform mutual image transformation between the two modalities. Similar to such GAN-based methods, many scholars have tried to use deep neural networks to generate samples of the intermediate modality to assist model training. Limited by the lack of color information in infrared images, such methods focus more on the generation of infrared images. However, such methods need to rely on complex image generation networks to achieve data augmentation, which brings a very large computational overhead. More methods focus on optimizing the network, such as attention mechanisms, adversarial mechanisms, etc., to make the model pay more attention to discriminative information related to the person and irrelevant to the style. The instance normalization (IN) method has been proven to be able to well eliminate modality differences and has also been used multiple times in cross-modal person re-identification tasks. Recently, many local feature-based methods have been applied to the field of cross-modal person re-identification. By extracting more refined features, it is more convenient to achieve semantic alignment of features and at the same time pay attention to more discriminative features. Such local feature-based methods usually can bring better results, but compared with global feature-based methods, the splicing of local features greatly increases the dimension of the feature vector during inference and also brings higher retrieval overhead. Summary of the Invention
[0005] Aiming at the problems in the cross-modal person re-identification system, such as the large differences in the feature styles of multi-modal images being difficult to align and the existing inference strategies not making full use of associated information, the present invention proposes a cross-modal person re-identification method and system based on feature enhancement and association optimization inference. Through fine-grained feature enhancement and association-aware inference, modality differences can be alleviated and the associated information in the image library can be fully utilized.
[0006] To achieve the above object, the technical solution of the present invention includes the following content.
[0007] A cross-modal pedestrian re-identification method based on feature enhancement and correlation optimization reasoning, the method comprising:
[0008] Calculate the initial similarity between the query image and the gallery images, and select a gallery image as a proxy image according to the initial similarity; wherein, the query image and the gallery images are of different modalities, and all gallery images are in the same modality;
[0009] Calculate the similarity between the proxy image and the gallery images;
[0010] Based on the similarity between the proxy image and the gallery images, correct the initial similarity to obtain the final similarity, and obtain the re-identification image sequence of the query image according to the final similarity.
[0011] Further, calculating the initial similarity between the query image and the gallery images includes:
[0012] Construct a modality-shared feature extraction network; wherein, the structure of the modality-shared feature extraction network includes: a residual network, a global pooling layer, and a BNNeck structure;
[0013] Train the modality-shared feature extraction network on the training set;
[0014] Based on the trained modality-shared feature extraction network, obtain the query image feature and the modality-shared feature of the gallery images respectively;
[0015] According to the cosine similarity between the query image feature and the modality-shared feature of the gallery images, obtain the initial similarity between the query image and the gallery images.
[0016] Further, training the modality-shared feature extraction network on the training set includes:
[0017] Input the training sample images into the residual network to obtain feature maps;
[0018] Use global pooling to reduce the dimension of the feature maps to obtain the dimension-reduced feature maps;
[0019] Utilize the BNNeck structure to adjust the feature distribution of the dimension-reduced feature maps to obtain the training sample image features;
[0020] Use a classifier to classify the training sample image features to obtain the classification result of the training sample images;
[0021] Based on the channel data, calculate the channel-level modality alignment loss L fea ;
[0022] Based on the dimension-reduced feature maps, calculate the triplet loss;
[0023] Calculate the identity classification loss based on the classification result of the training sample image;
[0024] Integrate the fine-grained channel-level modality alignment loss, the triplet loss, and the identity classification loss to obtain the first total loss, and perform backpropagation based on the first total loss.
[0025] Furthermore, the channel-level modality alignment loss where Diag represents taking the elements on the diagonal, and the elements in the correlation matrix R refers to the i-th channel vector of the shared features of the visible light image modality, refers to the j-th channel vector of the shared features of the infrared image modality, and N c represents the total number of channels.
[0026] Furthermore, calculating the similarity between the proxy image and the gallery images includes:
[0027] Construct a modality-specific feature extraction network; wherein, the structure of the modality-specific feature extraction network includes: a residual network, a global pooling layer, and a BNNeck structure;
[0028] Train the modality-specific feature extraction network on the training set using images from the same modality as the gallery;
[0029] Based on the trained modality-specific feature extraction network, obtain the modality-specific features of the proxy image and the gallery images respectively;
[0030] According to the cosine similarity between the modality-specific features of the proxy image and the gallery images, obtain the similarity between the proxy image and the gallery images.
[0031] Furthermore, training the modality-specific feature extraction network on the training set includes:
[0032] Input the training sample image into the residual network to obtain a feature map;
[0033] Use global pooling to reduce the dimension of the feature map to obtain a dimension-reduced feature map;
[0034] Utilize the BNNeck structure to adjust the feature distribution of the dimension-reduced feature map to obtain the training sample image features;
[0035] Use a classifier to classify the training sample image features to obtain the classification result of the training sample image;
[0036] Based on the dimension-reduced feature map, calculate the triplet loss;
[0037] Based on the classification results of the training sample images, calculate the identity classification loss;
[0038] Integrate the triplet loss and the identity classification loss to obtain the second total loss, and perform backpropagation based on the second total loss.
[0039] A cross-modal pedestrian re-identification system based on feature enhancement and association optimization reasoning, the system includes:
[0040] A first calculation module, configured to calculate the initial similarity between the query image and the gallery images, and select a gallery image as a proxy image according to the initial similarity; wherein, the query image and the gallery images are in different modalities, and all gallery images are in the same modality;
[0041] A second calculation module, configured to calculate the similarity between the proxy image and the gallery images;
[0042] A re-identification module, configured to correct the initial similarity based on the similarity between the proxy image and the gallery images to obtain the final similarity, and obtain the re-identification image sequence of the query image according to the final similarity.
[0043] An electronic device, characterized in that the electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the cross-modal pedestrian re-identification method according to any one of the above is implemented.
[0044] A computer-readable storage medium, characterized in that computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the cross-modal pedestrian re-identification method according to any one of the above is implemented.
[0045] A computer program product, characterized in that when the computer program product runs on a computer device, the computer device is enabled to execute the cross-modal pedestrian re-identification method according to any one of the above.
[0046] Compared with the prior art, the present invention has at least the following beneficial effects.
[0047] 1) The present invention designs a channel-level modality alignment loss function to constrain the channel distribution similarity of different modality features at a fine-grained level, reduce the modality gap and improve the feature information capacity.
[0048] 2) For the retrieval process, this paper proposes a two-stage inference mechanism: the first stage screens candidate proxy images by calculating cosine similarity based on shared features; the second stage combines the modality-specific features of the image library extracted by an independent residual network to explore the intra-modal correlations of samples in the library, optimize similarity ranking through feature complementarity, and accurately locate the target. By decoupling shared features from modality-specific features, this paper balances cross-modal alignment and intra-modal detail expression, significantly improving the robustness and accuracy of cross-modal retrieval in complex scenarios such as security monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Overall flow chart of the method of the present invention. DETAILED DESCRIPTION
[0050] In order to describe the method of the present invention more clearly and completely, the present invention will be further described below in conjunction with specific embodiments of the present invention and the accompanying drawings.
[0051] The modal person re-identification method of the present invention performs a two-stage reasoning process, aiming to deeply explore the modality-specific information within the library samples. In the first stage, a residual network (Residual Network, ResNet) with shared parameters is used to model the modal shared features of the query image (Query) and the image gallery image (Galley), and the proxy image most similar to the query image is screened from the image gallery according to the cosine similarity between the features; wherein, a fine-grained feature enhancement method is designed in the modeling process of the modal shared features, and the channel-level modal alignment loss used guides the extraction process of the modal shared features to promote modal alignment while increasing the information load of the features. In the second stage, another residual network is introduced and the modality-specific features of the Gallery image are modeled with the help of proxy images. By obtaining more accurate correlation information, the information within the modality is further integrated, and the initial similarity of the first stage is corrected to obtain more accurate retrieval results.
[0052] like Figure 1 As shown, the modal pedestrian re-identification method of the present invention mainly includes the following steps.
[0053] Step 1: Calculate the initial similarity between the query image and the image library images, and select an image library image as a proxy image based on the initial similarity.
[0054] Step 1.1: Modality-shared feature extraction.
[0055] In order to better learn the common feature expression space of the two modalities, a parameter-shared feature extractor is often used. In this invention, we also refer to many existing works and select the pre-trained ResNet50 as the basic framework. The input of the network is a pedestrian image of uniform size captured by an infrared camera and a visible light camera respectively, and a feature map is obtained after passing through the residual network. Then global pooling is used to reduce its dimension. It is then sent to the BNNeck structure to adjust the feature distribution. BNNeck refers to adding a batch normalization layer at the end of the network for feature normalization to solve the problem of inconsistent feature distribution, and calculating the triplet loss before the normalization layer, and then further calculating the identity classification loss through the classifier. Such a structure is more conducive to constraining features under different distribution conditions, promoting inter-class similarity and intra-class differences. The two modalities use a parameter-shared feature extractor to more effectively learn the common representation space of visible light and infrared modalities.
[0056] In one embodiment, in addition to the common identity classification loss and triplet loss mentioned above, the present invention also designs a fine-grained feature enhancement method and proposes a fine-grained channel-level modal alignment loss. This method uses fine-grained modal alignment loss to bring the feature similarity between different modalities closer for data of two modalities. By performing modal alignment at the channel level, the model can maintain higher feature consistency and feature carrying capacity during cross-modal learning. This fine-grained alignment strategy can effectively reduce the differences between modalities, improve the complementarity between modalities, and thereby enhance the final model's ability to comprehensively utilize multimodal information. Specifically, the present invention is based on the fine-grained modal alignment loss to promote the consistency of the distribution of data in the same channel, and suppress the expression of correlation for different channels. In this way, not only can the two modalities express the same fine-grained attributes more consistently, but they can also carry more effective information in a limited feature expression space.
[0057]
[0058] Among them, c represents the data of each channel of each modal feature, v represents the visible light modality, r represents the infrared modality, L fea represents the channel-level modality alignment loss, R represents the correlation matrix between different channels, which is obtained by calculating the cosine similarity between channels of different modalities, and N c The number of channels is represented by Diag, which represents the diagonal elements. Elements on the main diagonal represent the correlation of corresponding channels and need to be promoted. Elements in other positions represent the correlation of non-corresponding channels and their expression should be suppressed to obtain higher information carrying capacity.
[0059] Step 1.2: Initial similarity calculation.
[0060] First, the present invention still calculates the cosine similarity between the query features and the library image features according to the general method. The magnitude of this value can characterize the similarity degree between images. The higher the similarity, the greater the probability that the people in the corresponding two images belong to the same identity. Therefore, the image most similar to the image to be retrieved can be found through this cosine similarity, that is, there are similar features. In such a case, it is still somewhat optimized to retrieve the database again based on the found closest image to correct the similarity score. Such an operation can effectively improve the position of the positive sample in the list, facilitating the obtaining of the correct image better, and at the same time not introducing too much computational complexity. Since the samples in the Gallery come from the same modality, there are relatively smaller differences in style representation compared to cross-modal. Selecting reliable samples for retrieval within the Gallery may be a more reasonable method. Considering that the initially calculated cosine similarity can perform a preliminary ranking of the sample similarities, the present invention selects the Rank1 proxy selected by the cosine similarity.
[0061]
[0062] Step 1.3: Proxy selection.
[0063] For each query image, find the column index j corresponding to the element with the maximum cosine similarity value in the i-th row of the obtained cosine similarity matrix S qg and then select the image x g,j indexed as j in the gallery set into the proxy set G p where x g,j represents the image sample indexed as j in the gallery image library, S qg represents the cosine similarity matrix between the query image features and the library image features, and N q represents the number of query images, which is the same as the number of rows of the cosine similarity matrix S qg .
[0064] Step 2: Calculate the similarity between the proxy image and the library images.
[0065] The mainstream paradigm uses a parameter-sharing network to process two modalities, with the focus on extracting modality-shared information. However, during the training process, some modality-specific information, such as color and grayscale distribution information, will inevitably be lost. But this modality-specific information is crucial for mining reliable relevant information within the image gallery. It is very reasonable to explore more accurate correlations between samples using the modality-specific features of the gallery. Therefore, considering that all gallery images come from the same modality, mining the associated information within the gallery is equivalent to performing association analysis within the modality. At this time, using modality-specific features that retain part of the style information will better serve the association mining. The present invention designs a more reliable association information mining strategy by introducing an additional residual network to train an additional modality-specific feature encoder for the gallery. The structure of this encoder is the same as that of the modality-shared feature encoder. It also uses the pre-trained ResNet50 on ImageNet plus the BNN neck structure as the basic framework, and uses triplet loss and identity classification loss as the optimization objectives. Finally, based on this modality-specific feature encoder, a similarity matrix S between the proxy image and the gallery images can be obtained. gg 。
[0066] Step 3: Correct the initial similarity based on the similarity between the proxy image and the gallery images to obtain the final similarity, and obtain the re-identification image sequence of the query image according to the final similarity.
[0067] The present invention corrects the initial similarity according to the similarity matrix S gg and then obtains the image sequence accordingly.
[0068] S = S qg + S gg
[0069] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the spirit and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.
Claims
1. A cross-modal pedestrian re-identification method based on feature enhancement and correlation optimization inference, characterized in that The method includes: Calculating an initial similarity between a query image and images in an image library, and selecting an image in the image library as a proxy image according to the initial similarity; wherein, the query image and the images in the image library are in different modalities, and all images in the image library are in the same modality; Calculating the similarity between the proxy image and the images in the image library; Based on the similarity between the proxy image and the images in the image library, correcting the initial similarity to obtain a final similarity, and obtaining a re-identification image sequence of the query image according to the final similarity.
2. The method according to claim 1, wherein Calculating the initial similarity between the query image and the images in the image library includes: Constructing a modality-shared feature extraction network; wherein, the structure of the modality-shared feature extraction network includes: a residual network, a global pooling layer, and a BNNeck structure; Training the modality-shared feature extraction network on a training set; Based on the trained modality-shared feature extraction network, respectively obtaining the query image features and the modality-shared features of the images in the image library; According to the cosine similarity between the query image features and the modality-shared features of the images in the image library, obtaining the initial similarity between the query image and the images in the image library.
3. The method according to claim 2, characterized in that Training the modality-shared feature extraction network on a training set includes: Inputting a training sample image into the residual network to obtain a feature map; Using global pooling to reduce the dimension of the feature map to obtain a dimension-reduced feature map; Using the BNNeck structure to adjust the feature distribution of the dimension-reduced feature map to obtain training sample image features; Using a classifier to classify the training sample image features to obtain the classification result of the training sample image; Based on the channel data, calculate the channel-level modal alignment loss L fea ; Based on the dimension-reduced feature map, calculating a triplet loss; Based on the classification result of the training sample image, calculating an identity classification loss; Integrating the fine-grained channel-level modality alignment loss, the triplet loss, and the identity classification loss to obtain a first total loss, and performing backpropagation based on the first total loss.
4. The method according to claim 3, wherein The channel-level modal alignment loss where Diag means taking the elements on the diagonal, and the elements in the correlation matrix R refers to the i-th channel vector of the shared features of the visible light image modality, refers to the j-th channel vector of the shared features of the infrared image modality, and N c represents the total number of channels.
5. The method according to claim 1, characterized in that, Calculating the similarity between the proxy image and the images in the image library includes: Constructing a modality-specific feature extraction network; wherein, the structure of the modality-specific feature extraction network includes: a residual network, a global pooling layer, and a BNNeck structure; Using images from the same modality as the image library to train the modality-specific feature extraction network on a training set; Based on the trained modality-specific feature extraction network, respectively obtaining the modality-specific features of the proxy image and the images in the image library; According to the cosine similarity between the modality-specific features of the proxy image and the images in the image library, obtaining the similarity between the proxy image and the images in the image library.
6. The method according to claim 5, wherein Training the modality-specific feature extraction network on a training set includes: Inputting a training sample image into the residual network to obtain a feature map; Using global pooling to reduce the dimension of the feature map to obtain a dimension-reduced feature map; Using the BNNeck structure to adjust the feature distribution of the dimension-reduced feature map to obtain training sample image features; Using a classifier to classify the training sample image features to obtain the classification result of the training sample image; Based on the dimension-reduced feature map, calculating a triplet loss; Based on the classification results of the training sample images, calculate the identity classification loss; Integrate the triplet loss and the identity classification loss to obtain the second total loss, and perform backpropagation based on the second total loss.
7. A cross-modal pedestrian re-identification system based on feature enhancement and association optimization reasoning, characterized in that The system includes: A first calculation module, configured to calculate an initial similarity between a query image and an image library image, and select an image library image as a proxy image according to the initial similarity; wherein, the query image and the image library image are in different modalities, and all the image library images are in the same modality; A second calculation module, configured to calculate the similarity between the proxy image and the image library images; A re-identification module, configured to correct the initial similarity based on the similarity between the proxy image and the image library images to obtain a final similarity, and obtain a re-identification image sequence of the query image according to the final similarity.
8. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the cross-modal pedestrian re-identification method based on feature enhancement and association optimization inference as described in any one of claims 1-6 is implemented.
9. A computer-readable storage medium, characterized in that, Computer program instructions are stored on the computer-readable storage medium, and when the computer program instructions are executed by a processor, the cross-modal pedestrian re-identification method based on feature enhancement and association optimization inference as described in any one of claims 1-6 is implemented.
10. A computer program product, characterized in that, When the computer program product runs on a computer device, the computer device is caused to execute the cross-modal pedestrian re-identification method based on feature enhancement and association optimization inference as described in any one of claims 1-6.