Cross-modal pedestrian re-identification method based on cross-scale information interaction

By introducing a cross-modal obfuscator and a cross-scale information interaction module in the cross-modal pedestrian re-identification model, the identification problem between visible and infrared light images is solved using contrast loss constraints, and high-precision cross-modal pedestrian re-identification is achieved, which is suitable for security monitoring and other fields.

CN120071384AActive Publication Date: 2025-05-30DALIAN MARITIME UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411905192.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-30
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

The prior art is difficult to achieve efficient cross-modal pedestrian re-identification between visible and infrared images. Due to factors such as modal differences, pedestrian attitude changes and background interference, the recognition accuracy is difficult to meet actual needs.

Method used

A cross-modal pedestrian re-identification model based on cross-scale information interaction is adopted. By introducing a cross-modal obfuscator and a cross-scale information interaction module, the modal difference between visible light and infrared light images is reduced, and the model's robustness to pedestrian pose changes and background interference is enhanced.

Benefits of technology

High-precision cross-modal pedestrian re-identification is realized, which can effectively reduce the impact of modal differences, posture changes and background interference, and provides more accurate pedestrian identity recognition results, which are suitable for security monitoring and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071384A_ABST
    Figure CN120071384A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal pedestrian re-identification method based on cross-scale information interaction, which belongs to the technical field of computer vision, and mainly comprises the following steps of: sending a visible light pedestrian image and an infrared pedestrian image into a cross-modal obfuscator, pre-processing the two input images with different modalities by using a mature image processing technology, and obtaining a cross-modal pedestrian re-identification result; an input object meeting the model requirement is obtained; then, processing an input object by adopting a parameter separation feature extraction network to respectively obtain visible light image features and infrared image features; on the basis, a parameter sharing feature extraction network is utilized, and a cross-scale interaction module is combined, so that modal sharing features are further extracted; meanwhile, for two features with different scales output by the feature extraction network, respectively extracting global feature information and local region feature information from the features; and finally, optimizing the model by adopting a multi-loss joint optimization strategy according to the obtained features, and finally determining a query result based on the similarity between the query image and the image library image features. According to the method, the accuracy and efficiency of visible light infrared pedestrian re-identification can be remarkably improved, and the method has wide application prospects in the fields of security monitoring and the like needing pedestrian identity identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision image retrieval, and particularly relates to a cross-modal pedestrian re-identification model, method and application based on cross-scale information interaction. Background Art

[0002] Pedestrian re-identification aims to retrieve images of the same pedestrian under different cameras from a database, and is widely used in fields such as intelligent security and intelligent life. With the development of deep learning, its accuracy has been significantly improved compared with traditional manual feature extraction methods, but there is still room for improvement. Modern cameras are often equipped with two modes of visible light and infrared light to ensure night work. However, most mainstream pedestrian re-identification models rely on visible light image recognition and are difficult to effectively recognize infrared light images. Therefore, the cross-modal pedestrian re-identification task appears. Its goal is to achieve mutual retrieval of pedestrian pictures in different modalities (visible light and infrared light). Different from visible light cameras, infrared cameras can work in low-light or dark environments. The imaging principles of the two are different. The former depends on light sources and reflected light when imaging, and the image contains rich information such as colors and details. The latter depends on object thermal radiation imaging, and the image does not contain information such as colors. This makes the pictures taken by the two have modal differences. In addition to facing modal differences between cameras with different spectra, it is also affected by many factors such as pedestrian pose changes, occlusions, and lighting changes, resulting in the recognition accuracy of existing models being difficult to meet actual requirements. Summary of the Invention

[0003] In view of the deficiencies of the prior art, the present invention provides a cross-modal pedestrian re-identification model and method based on cross-scale information interaction. In order to fully mine modality-invariant information in a large range, we introduce a cross-modal confuser and a cross-scale information interaction module, and are constrained by a contrastive loss based on auxiliary information. It overcomes the influence factors such as the bimodal differences between visible light images and infrared light images, pedestrian pose changes, and background interference, provides high-precision cross-modal pedestrian re-identification prediction results, and realizes cross-modal pedestrian re-identification.

[0004] To achieve the above object, the present technical solution provides a cross-modal pedestrian re-identification method based on cross-scale information interaction, including the following steps:

[0005] S1. Obtain paired visible light images and infrared images, input the visible light images therein into a cross-modal confuser for preprocessing to generate pseudo-infrared images, and after expanding the pseudo-infrared images and the infrared images, form a training set;

[0006] S2. Construct a cross-modal pedestrian re-identification network model based on cross-scale information interaction. The cross-modal pedestrian re-identification network model based on cross-scale information interaction adopts an improved Resnet50 network framework, and includes a separated-parameter feature extraction network and a modality-sharing feature extraction network with shared parameters.

[0007] The separation parameter feature extraction network uses the first-stage feature extraction networks in two parallel Resnet50 frameworks to extract visible light image features and infrared image features respectively.

[0008] The modality-sharing feature extraction network with parameter sharing includes the second to fifth-stage feature extraction networks in the Resnet50 framework and a cross-scale information interaction module, and the cross-scale information interaction module is inserted into the second to fourth-stage feature extraction networks respectively.

[0009] S3. Training the cross-modal person re-identification network model based on cross-scale information interaction using the training data in the training dataset, including:

[0010] Inputting the visible light image features and infrared image features into the modality-sharing feature extraction network with parameter sharing to obtain a multi-scale modality-sharing feature representation.

[0011] Based on the two-scale features output by the fourth to fifth-stage feature extraction networks obtained from the modality-sharing feature extraction network with parameter sharing, global features and local feature representations of the two modalities of visible light images and infrared images are obtained through normalization and pooling.

[0012] According to the multi-scale global features and local features obtained by the model, the weighted sum of the cross-entropy classification loss, triplet loss, and cross-modal contrast loss is used for backpropagation and the model parameters of the cross-modal person re-identification network model are updated. The model parameters include the parallel first-stage feature extraction network in the Resnet50 framework, the shared second to fifth-stage feature extraction networks, and the parameters in the cross-scale information interaction module.

[0013] S4. Obtaining a query image and a gallery image, inputting them into the trained cross-modal person re-identification network model, and obtaining a retrieval result based on the similarity between the two.

[0014] Furthermore, inputting the visible light image into a cross-modal confuser for preprocessing to generate a pseudo-infrared image, including preprocessing according to the following process:

[0015]

[0016] where x ir (x s ) represents the original infrared image, represents the infrared image after channel alignment processing with the visible light image, transfer means repeating a single channel three times to align with the three-channel format of the visible light image, and x vis represents the original visible light image. Denote the pseudo-infrared image generated after being processed by the cross-modal confuser, where H and W represent the height and width of the image, and T cc Denote the image transformation operations adopted by the cross-modal confuser. The image transformation operations include weighted gray-scale transformation, cross-channel information confusion transformation, and spectral jitter. For the above three transformations, the execution probabilities corresponding to each are p 1 , p 2 , p 3 , and the detailed processes of the three transformations are as follows:

[0017] Weighted gray-scale transformation: The visible light image is divided into three channels, and then they are fused using the following formula:

[0018]

[0019] where x r , x g , x b respectively represent the information of the red, green, and blue channels of the visible light image, and α 1 , α 2 , α 3 are weight factors randomly generated in the range of [0, 1] and with a sum of 1, denote the visible light picture after weighted gray-scale transformation;

[0020] Perform cross-channel information confusion transformation on the visible light image according to the following formula:

[0021]

[0022] where, respectively represent randomly selecting two of the three channels separated from the visible light image as the foreground and background respectively, and then by cutting a random rectangle from the foreground and pasting it into the background to finally form a single image

[0023] The transformation formula for spectral jitter is as follows:

[0024]

[0025]

[0026] where, denote randomly selecting one channel and repeating it 3 times to obtain a degenerate spectral image, and x ch denote the single-channel information randomly selected from the three channels of the visible light image, and then replacing the information of the other channels with the selected x ch to obtain the degenerate spectral image β 1Denote the weight factor, which is randomly taken in the range of [0, 1], β 2 = 1 - β 1 When β 1 is 0, it represents the degenerate spectral image. When β 1 is 1, x vis represents the original visible light image, and represents the visible light image after spectral transformation.

[0027] Furthermore, the separation parameter feature extraction network uses the first-stage networks in two parallel Resnet50 frameworks to extract the visible light image features and infrared image features respectively. The formula is:

[0028]

[0029] ResnetStage1 represents the feature extraction network in the first stage of the Resnet50 framework, which respectively represent the visible light image features and infrared image features obtained through the specific modality feature extraction network.

[0030] Furthermore, the cross-scale information interaction module generates Query by using the feature map with channel compression, generates Key and Value by using the features with channel compression and spatial compression, realizes cross-scale interaction through multi-head self-attention, and restores the spatial and channel dimensions through the fully connected layer and bilinear interpolation; this module includes three components: spatial channel compression, cross-scale interaction, and spatial channel restoration, and is used to represent the image features of ResnetStagei passing through the modality sharing network as where i represents the output of the i-th stage of the Resnet framework. Here, the value of i is 2, 3, 4, and includes the visible light image features and infrared image features obtained by the separation parameter feature extraction network and concatenated in the batch dimension. Cross-scale information interaction is realized according to the following formula:

[0031]

[0032] where represents the image features after batch normalization, where Conv represents the channel dimension compression operation with a channel compression ratio of r implemented by using a 1×1 convolution kernel cc Pooling represents the spatial dimension compression operation with a channel compression ratio of r implemented by using average pooling with a pooling kernel size of r sc to generate sc and and The features are flattened and then input into the cross-scale information interaction component, which is implemented using multi-head attention. Among them, the flattened acts as the query, while the flattened acts as the key and value. The output features of the multi-head attention are represented as The obtained features are input into the spatial channel restoration component, where FC represents the fully connected layer, which is used to restore the channel dimension. Then, after the Sigmoid and Upsample operations, the spatial dimension is restored to obtain the Mask, and the Mask and the result of batch normalization are multiplied element-wise to obtain the final output of this module and the obtained is input into the (i + 1)-th stage of Resnet.

[0033] Furthermore, based on the two-scale features output by the fourth to fifth stage feature extraction networks of the modality-shared feature extraction network with parameter sharing, the global features and local feature representations of the visible light image and the infrared image are obtained through normalization and pooling, including:

[0034] The two-scale features output by the fourth to fifth stage feature extraction networks are respectively defined as For the global feature branch is respectively used to perform generalized average pooling on it to obtain the global feature, and the local feature branch evenly divides it into multiple equal local blocks along the horizontal direction, and performs generalized average pooling on each evenly divided local block to obtain the local region feature. The specific formula is as follows:

[0035]

[0036] where H, W, and C represent the height, width, and number of channels of the feature map, x i,j,c represents the element at the coordinate (i, j) on the c-th channel in the feature map, where i = 1, 2,..., H; j = 1, 2,..., W; c = 1, 2,..., C, p is an adjustable parameter, Gem() represents the operation of performing generalized average pooling on the feature, and Partition represents the horizontal division operation to obtain equal local feature maps.

[0037] Among them represents the global feature and N local features obtained by performing generalized average pooling on the output of the fifth stage of the Resnet framework, represents the global feature and M local features obtained by performing generalized average pooling on the output after the fourth stage of the Resnet framework and the inserted cross-scale information interaction module.

[0038] Furthermore, according to the multi-scale global features and local features obtained by the model, the weighted sum of the cross-entropy classification loss, triplet loss, and cross-modal contrast loss is used for backpropagation to update the model parameters of the cross-modal person re-identification network model. The calculation processes of each loss are as follows:

[0039] First, the features are obtained through a multi-layer perceptron For A cross-modal contrast loss is established for constraint, and this loss is represented by and its formula is expressed as follows:

[0040]

[0041] Among them, W 1 , W 2 represents an affine transformation, BN represents batch normalization, and ReLU is an activation function. In vis , VI represents that the anchor represents visible light features, while other samples involved in the loss correspond to infrared light features. I infra (i) represents the set of the top K most dissimilar infrared samples with the same ID as the visible sample i. z is the feature obtained after applying L2 normalization to ; z p refers to the feature z corresponding to the sample extracted from the set P infra (i); while z a represents the feature corresponding to the sample extracted from the set A(i), where A(i) represents the set of all samples in the infrared modality with different IDs from i, as well as all samples in P infra (i). Similarly, for the case where the anchor is an infrared feature, the corresponding loss is The final constraint of this loss is and is the mean of

[0042] For and respectively, the cross-entropy classification loss is calculated for constraint, and its formula is expressed as follows:

[0043]

[0044] Among them, p() is the probability that the classifier correctly predicts the feature, and E represents the expectation.

[0045] For the global feature the cross-entropy classification loss and triplet loss are calculated respectively for it. For the local feature the cross-entropy classification loss is calculated respectively for it, and its formula is expressed as follows:

[0046]

[0047] where p() is the probability that the classifier correctly predicts the feature, and E represents the expectation; represents the triplet loss value, d(a, p) is the distance between the anchor sample and the positive sample, d(a, n) is the distance between the anchor sample and the negative sample, and margin is a manually set interval parameter to ensure a sufficient distance gap between the positive and negative samples;

[0048] The final overall multi-loss joint constraint function is expressed as follows:

[0049]

[0050] where λ 1 and λ 2 are hyperparameters used to balance the contributions of each loss function.

[0051] Compared with the prior art, the present invention has the following advantages:

[0052] The present invention proposes a cross-modal pedestrian re-identification method based on cross-scale information interaction. This method can fully exploit the modality-invariant information in a large range. We introduce a cross-modal confuser and a cross-scale information interaction module, and a contrastive loss constraint based on auxiliary information. It effectively reduces the influencing factors such as the modality difference between visible light images and infrared light images, pedestrian pose changes, and background interference, provides high-precision cross-modal pedestrian re-identification prediction results, and has broad application prospects in many fields such as security monitoring that require pedestrian identity recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0054] Figure 1 It is a schematic flowchart of a cross-modal pedestrian re-identification method based on cross-scale information interaction in an embodiment of the present invention.

[0055] Figure 2 It is an architecture diagram of a cross-modal pedestrian re-identification model based on cross-scale information interaction in an embodiment of the present invention.

[0056] Figure 3 It is a schematic diagram of the principle of the cross-scale information interaction module in an embodiment of the present invention.

[0057] Figure 4Schematic diagram of the transformation effect of the cross-modal obfuscator in the embodiments of the present invention. Detailed implementation manners

[0058] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0059] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0060] As Figure 1 shown, the present invention provides a cross-modal pedestrian re-identification method based on cross-scale information interaction, which mainly includes the following steps:

[0061] S1. Obtain paired visible light images and infrared images, input the paired visible light images and infrared images into a cross-modal obfuscator for preprocessing, convert the visible light images into pseudo-infrared images close to infrared images to promote the model to mine modality-invariant information of the same identity, and at the same time make full use of the multi-channel color information existing in the original visible light images, and expand the transformed multiple pairs of pseudo-infrared images and infrared images to form a training set. The specific expansion steps include random flipping, padding, and Resize, etc.

[0062] Specifically, multiple groups of image sets composed of corresponding visible light images and infrared light images marked with the same human category are obtained as the training set. To implement the training and testing of a cross-modal pedestrian re-identification model based on cross-scale information interaction, multiple groups of image sets composed of corresponding visible light images and infrared light images marked with the same human category can also be obtained as the validation set and the testing set. In this solution, image processing is performed on the visible light images and infrared light images in the training set to uniformly adjust the scale to 394*192. Since there are significant modal differences between visible light images and infrared images, before inputting the data into the model, a cross-modal mixer is first used to preprocess the visible light images, which can not only increase the diversity of samples but also enable the model to learn more robust feature representations. This solution inputs the processed training set into the cross-modal pedestrian re-identification framework for training. This solution divides the training set into 80 batches to train the cross-modal pedestrian re-identification framework. The batch size of each batch of the training set is 48. Specifically, each batch of the training set corresponds to 6 different pedestrians, and each pedestrian corresponds to 4 visible light images and 4 infrared light images.

[0063] Input the visible light images therein into the cross-modal mixer for preprocessing to generate pseudo-infrared images, including preprocessing according to the following process:

[0064]

[0065] Among them, x ir (x s ) represents the original infrared image, represents the infrared image after channel alignment processing with the visible light image, transfer means repeating a single channel three times to align with the three-channel format of the visible light image, x vis represents the original visible light image, represents the pseudo-infrared image generated after being processed by the cross-modal mixer, H, W represent the height and width of the image, T cc represents the image transformation operation adopted by the cross-modal mixer. The image transformation operation includes weighted grayscale transformation, cross-channel information confusion transformation, and spectral jitter. For the above three transformations, the execution probability corresponding to each is p 1 , p 2 , p 3 , and the detailed processes of the three transformations are as Figure 4 shown, including:

[0066] Weighted grayscale transformation: The visible light image is divided into three channels, and then they are fused using the following formula:

[0067]

[0068] Among them, x r 、xg , x b respectively represent the information of the red, green, and blue channels of the visible light image. α 1 , α 2 , α 3 are weight factors randomly generated in the range of [0, 1] with a sum of 1. represents the visible light picture after weighted gray-scale transformation;

[0069] Perform cross-channel information confusion transformation on the visible light image according to the following formula:

[0070]

[0071] where respectively represent randomly selecting two channels from the three channels separated from the visible light image as the foreground and background respectively, and then by cutting a random rectangle from the foreground and pasting it into the background to finally form a single image

[0072] The transformation formula for spectral jitter is as follows:

[0073]

[0074] where represents randomly selecting one channel and repeating it 3 times to obtain a degenerate spectral image, and x ch represents the single-channel information randomly selected from the three channels of the visible light image, and then replacing the information of other channels with the selected x ch to obtain a degenerate spectral image β 1 represents a weight factor, randomly taking values in the range of [0, 1], and β 2 = 1 - β 1 , when β 1 is 0, it represents the degenerate spectral image. When β 1 is 1, x vis represents the original visible light image, represents the visible light image after spectral transformation.

[0075] S2. Construct a cross-modal pedestrian re-identification network model based on cross-scale information interaction. The cross-modal pedestrian re-identification network model based on cross-scale information interaction includes a cross-modal confuser, a feature extraction and embedding network, and a multi-loss joint optimization strategy. The cross-modal confuser uses technical means such as weighted gray scale, channel mixing, and spectral jitter to process paired visible light images and infrared light images. The feature extraction and embedding network is used to extract the image features of the two input modalities.

[0076] In this application, the feature extraction and embedding network includes: the backbone network of the cross-modal person re-identification network based on cross-scale information interaction adopts a deep residual network based on ResNet50, an infrared light feature extraction network, a visible light feature extraction network, and a fully connected layer. Among them, the feature extraction network in the first stage of Resnet50 with parameter separation is used to extract specific modal features, and the last four stages with shared parameters in ResNet50 are used as the modal shared feature extraction network. Cross-scale information interaction modules are inserted in the second, third, and fourth stages to fully mine the modal invariant information in a large range and obtain features at two scales. The global features and local features are obtained through the global feature branch and the local feature branch respectively. The pooling method used is generalized mean pooling. The acquisition method of local features is realized by dividing the features into equal local feature blocks. The pooling method used is generalized mean pooling. Loss constraints are performed according to the global features and local features at the two scales obtained, and the model weights are updated according to the loss value through backpropagation.

[0077] S3. Train the cross-modal person re-identification network model based on the training data in the training dataset, including: input the visible light image features and infrared image features into the modal shared feature extraction network with shared parameters to obtain multi-scale modal shared feature representations. Based on the two-scale features output by the fourth to fifth stage feature extraction networks obtained from the modal shared feature extraction network with shared parameters, the global features and local feature representations of the two modalities of visible light images and infrared images are obtained through normalization and pooling. According to the multi-scale global features and local features obtained by the model, the weighted sum of the cross-entropy classification loss, triplet loss, and cross-modal contrast loss is used for backpropagation and the model parameters of the cross-modal person re-identification network model are updated, including the parameters in the first stage feature extraction network in parallel in the Resnet50 framework, the shared second to fifth stage feature extraction networks, and the cross-scale information interaction module.

[0078] In the present invention, the image feature extraction and embedding network needs to be trained. Regarding the training method, the stochastic gradient descent method (SGD) is used as the optimizer in this application, with a momentum of 0.9. The warm-up method is used for preheating in the first 10 rounds, and the learning rate linearly increases from 0.01 to 0.1. Then, the learning rate is reduced to one-tenth of the original value at the 50th round and the 80th round respectively.

[0079] The following specifically describes the training process of the cross-modal person re-identification network model:

[0080] S301. Construct the model input, and input the paired visible light image and infrared image into the cross-modal mixer to obtain the model input.

[0081] Specifically, obtain multiple groups of image sets composed of corresponding visible light images and infrared light images that label the same category of humans as the training set. Specifically, in this solution, multiple groups of image sets composed of corresponding visible light images and infrared light images that label the same pedestrian are obtained from existing cross-modal pedestrian re-identification data sets as the training set.

[0082] To implement the training and testing of a cross-modal pedestrian re-identification model based on cross-scale information interaction, multiple groups of image sets composed of corresponding visible light images and infrared light images that label the same category of humans can also be obtained as the validation set and the testing set.

[0083] S302. Preprocess the visible light images in the training set, and the processing method can be referred to in formulas (1)-(5).

[0084] At the same time, common transformations such as random flipping and random erasing are respectively applied to the visible light images and infrared images. A set of visible light image features and a set of infrared image features obtained through the above series of transformations are respectively represented by and where it contains 4 visible light pictures and 4 infrared pictures of 6 specific identity pedestrians, and and constitute a training batch.

[0085] S303. Based on the obtained model inputs, input them into a specific modality feature extractor with separated parameters to respectively obtain visible light image features and infrared image features. Construct a specific modality feature extraction network, and adopt the feature extraction network in the first stage of the Resnet50 framework. Input the visible light images into the visible light image feature extraction network to obtain visible light image features, and input the infrared images into the visible light image feature extraction network to obtain visible light image features:

[0086]

[0087] ResnetStage1 represents the feature extraction network in the first stage of the Resnet50 framework, respectively represent the visible light image features and infrared image features obtained through the specific modality feature extraction network.

[0088] S304. Based on the obtained modality-specific features, input them into a feature extraction network with shared parameters and multiple cross-scale information interaction modules to obtain modality-shared features.

[0089] The obtained visible light image features and infrared images are concatenated in the batch dimension to obtain the feature

[0090] The obtained The feature extraction network of the (i + 1)-th stage of the input Resnet50 framework obtains new new where i performs an operation of adding 1 based on the previous i.

[0091] Determine whether the current stage is the second stage, the third stage, or the fourth stage. If so, input it into the cross-scale information interaction module.

[0092] The cross-scale information interaction module generates Query using the feature map with channel compression, generates Key and Value using the features with channel compression and spatial compression, realizes cross-scale interaction through multi-head self-attention, and restores the spatial and channel dimensions through a fully connected layer and bilinear interpolation. This module is composed of three components in series: spatial channel compression, cross-scale interaction, and spatial channel restoration, and is used to represent the image features of ResnetStagei passing through the modality sharing network as where i represents the output of the i-th stage of the Resnet framework. Here, the value of i is 2, 3, 4, and includes the visible light image features obtained by the separable parameter feature extraction network and the infrared image features concatenated in the batch dimension. First, the obtained is input into the spatial channel compression component:

[0093]

[0094] where represents the image features after batch normalization, where Conv represents the channel dimension compression operation with a channel compression ratio of r implemented using a 1×1 convolution kernel cc and Pooling represents the spatial dimension compression operation with a channel compression ratio of r implemented using average pooling with a pooling kernel size of r sc ; sc

[0095] The obtained features are input into the cross-scale information interaction component:

[0096]

[0097] The previously generated and features are flattened and then input into the cross-scale information interaction component, which is implemented using multi-head attention. Among them, the flattened acts as the query (Query), and the flattened acts as the key (Key) and value (Value). The output features of the multi-head attention are represented as

[0098] Input the obtained features into the spatial channel restoration component:

[0099]

[0100] Among them, FC represents the fully connected layer, which is used to restore the channel dimension. Then, after Sigmoid and Upsample operations, the spatial dimension is restored to obtain the Mask, and the Mask and the result of batch normalization are multiplied element-wise to obtain the final output of this module Reuse to represent.

[0101] After passing through the second to fourth stages and the cross-scale interaction modules after each stage, the obtained is input into the fifth stage of the Resnet50 framework to obtain

[0102] S305. Obtain the global features and regional features of the two-modal images based on the features of the last two scales obtained from the shared feature extraction network.

[0103] The obtains the global feature through generalized average pooling, and then the is horizontally divided into M local blocks of equal size along the height dimension. Generalized average pooling is performed on each local block to obtain M local features:

[0104]

[0105] Among them, H, W, and C represent the height, width, and number of channels of the feature map, and x i,j,c represents the element at coordinates (i, j) on channel c in the feature map, where (i = 1, 2, …, H; j = 1, 2, …, W; c = 1, 2, …, C), p is an adjustable parameter, which is set to 3 in this example. Partition represents the horizontal division operation to obtain a total of N local blocks, and Gem represents performing the pooling operation in formula (10) on each local feature to obtain a total of N pooled local features.

[0106] Similarly, the obtains the global feature through generalized average pooling, and then the is horizontally divided into N local blocks of equal size along the height dimension. Generalized average pooling is performed on each local block to obtain N local features:

[0107]

[0108] Among them, Partition represents the horizontal division operation to obtain There are N local blocks in total. Gem represents performing the pooling operation in formula (11) on each local feature to obtain N pooled local features in total.

[0109] S306. According to the obtained global features and regional features, optimize the model using a multi-loss optimization strategy. Input the query image and the gallery image into the trained model, and finally output the retrieval result according to the similarity between the two.

[0110] First, obtain features through a multi-layer perceptron For Establish a cross-modal contrastive loss for constraint. This loss is represented by and its formula is expressed as follows:

[0111]

[0112]

[0113] Among them, W 1 , W 2 represents an affine transformation, BN represents batch normalization, ReLU is an activation function, In vis , VI represents that the anchor represents visible light features, while other samples involved in the loss correspond to infrared light features. I infra (i) represents the set of the top K most dissimilar infrared samples with the same ID as the visible sample i. z is the feature obtained after applying L2 normalization to ; z p refers to the feature z corresponding to the sample extracted from the set P infra (i); while z a represents the feature corresponding to the sample extracted from the set A(i). A(i) represents the set of all samples in the infrared modality with different IDs from i, as well as all samples in P infra (i). Similarly, for the case where the anchor is an infrared feature, the corresponding loss is The final constraint of this item of loss is and the mean of

[0114] For and respectively calculate the cross-entropy classification loss for constraint, and its formula is expressed as follows:

[0115]

[0116] Among them, p() is the probability that the classifier correctly predicts the feature, and E represents the expectation;

[0117] For global features Calculate the cross-entropy classification loss and the triplet loss for them respectively. For local features Calculate the cross-entropy classification loss for them respectively. The formula is expressed as follows:

[0118]

[0119] Where p() is the probability that the classifier correctly predicts the feature, and E represents the expectation; represents the triplet loss value, d(a, p) is the distance between the anchor sample and the positive sample, d(a, n) is the distance between the anchor sample and the negative sample, and margin is a manually set interval parameter to ensure there is a sufficient distance gap between the positive and negative samples;

[0120] The final overall multi-loss joint constraint function is expressed as follows:

[0121]

[0122] Where λ 1 and λ 2 are hyperparameters used to balance the contributions of each loss function.

[0123] S4. Obtain the query image and the gallery image and input them into the trained cross-modal person re-identification network model, and obtain the retrieval result according to the similarity between the two.

[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A cross-modal person re-identification method based on cross-scale information interaction, characterized in that: The following steps are involved: S1, obtaining a pair of visible light images and infrared images, inputting the visible light images into a cross-modal obfuscator for preprocessing to generate pseudo infrared images, and expanding the pseudo infrared images and infrared images to form a training set; S2. Construct a cross-modal person re-identification network model based on cross-scale information interaction. The cross-modal person re-identification network model based on cross-scale information interaction adopts an improved Resnet50 network framework, including a separated parameter feature extraction network and a parameter-sharing modality-shared feature extraction network. The separation parameter feature extraction network uses two parallel Resnet50 frameworks in the first stage feature extraction network to extract visible light image features and infrared image features respectively. The parameter-sharing modality-sharing feature extraction network includes the second to fifth stage feature extraction networks and the cross-scale information interaction module in the Resnet50 framework, and the cross-scale information interaction module is respectively inserted into the second to fourth stage feature extraction networks; S3. Training a cross-modal person re-identification network model based on cross-scale information interaction based on the training data in the training dataset, including: The visible light image features and infrared image features are input into the parameter-sharing modality-shared feature extraction network to obtain multi-scale modality-shared feature representation. The two scale features output by the fourth to fifth stage feature extraction network obtained by the modality sharing feature extraction network based on parameter sharing are normalized and pooled to obtain the global and local feature representations of the visible light image and infrared image. According to the multi-scale global features and local features obtained by the model, the weighted sum of the cross entropy classification loss, the triplet loss and the cross-modal contrast loss is used for back propagation and the model parameters of the cross-modal person re-identification network model are updated, wherein the model parameters include the first-stage feature extraction network in parallel in the Resnet50 framework, the shared second to fifth-stage feature extraction networks and the parameters in the cross-scale information interaction module; S4. Obtain the query image and the gallery image as inputs into the trained cross-modal person re-identification network model, and obtain the retrieval result based on the similarity between the two.

2. A cross-modal person re-identification method based on cross-scale information interaction according to claim 1, characterized in that: The visible light image is input into the cross-modal obfuscator for preprocessing to generate a pseudo infrared image, including preprocessing according to the following process: Among them, x ir (x s ) represents the original infrared image, represents the infrared image after channel alignment with the visible light image. transfer means repeating a single channel three times to align it with the three-channel format of the visible light image. vis represents the original visible light image, represents the pseudo infrared image generated after the cross-modal obfuscator processing, H, W represent the height and width of the image, T cc It represents the image transformation operation adopted by the cross-modal obfuscator, which includes weighted grayscale transformation, cross-channel information obfuscation transformation and spectral jittering. The execution probability of each of the above three transformations is p1, p2, and p3 respectively. The detailed process of the three transformations is as follows: Weighted grayscale transformation: The visible light image is divided into three channels and then fused using the following formula: Among them, x r 、x g 、x b They represent the information of the red, green and blue channels of the visible light image respectively. α1, α2 and α3 are randomly generated weight factors in the range of [0,1] with a sum of 1. Represents a visible light image after weighted grayscale transformation; The visible light image is transformed through cross-channel information confusion according to the following formula: in, Respectively, two channels are randomly selected from the three channels separated from the visible light image as the foreground and background, and then Cut random rectangles from the image and paste them to the background Finally, a single image is formed The transformation formula of spectral jitter is as follows: in, represents randomly selecting a channel and repeating it 3 times to obtain a degenerate spectral image, x ch Represents the single channel information randomly selected from the three channels of the visible light image, and then replaces the information of other channels with the selected x ch , thus obtaining a degenerate spectral image β1 represents the weight factor, which takes a random value in the range [0,1], β2 = 1-β1, when β1 is 0, That is, it represents a degenerate spectral image. When β1 is 1, x vis represents the original visible light image, Represents the visible light image after spectral transformation.

3. The cross-modal person re-identification method based on cross-scale information interaction according to claim 1, characterized in that: The separation parameter feature extraction network uses two parallel Resnet50 frameworks in the first stage to extract visible light image features and infrared image features respectively. The formula is: ResnetStage1 represents the first stage feature extraction network in the Resnet50 framework. They respectively represent the visible light image features and infrared image features obtained through the specific modality feature extraction network.

4. The cross-modal person re-identification method based on cross-scale information interaction according to claim 1, characterized in that: The cross-scale information interaction module uses channel compressed feature maps to generate queries, uses channel compressed and spatial compressed features to generate keys and values, realizes cross-scale interaction through multi-head self-attention, and restores spatial and channel dimensions through fully connected layers and bilinear interpolation; the cross-scale information interaction module includes three components: spatial channel compression, cross-scale interaction, and spatial channel restoration, which are used to represent the image features of ResnetStagei through the modality sharing network as Where i represents the output of the i-th stage of the Resnet framework, where the value of i is 2, 3, 4, and Includes visible light image features obtained by the separation parameter feature extraction network and infrared image features It is spliced ​​in the batch dimension, and cross-scale information interaction is achieved according to the following formula: in Represents the batch-normalized image features, where Conv represents the channel compression ratio r achieved using a 1×1 convolution kernel cc The channel dimension compression operation, Pooling means using a pooling kernel size of r sc The average pooling achieves a channel compression ratio of r sc The spatial dimension compression operation produces and The features are flattened and then input into the cross-scale information interaction component, which is implemented using multi-head attention. Acting as a query, while flattened Acting as keys and values, the output features of multi-head attention are represented as The features obtained Input spatial channel recovery component, where FC represents the fully connected layer, which is used to restore the channel dimension. Then, the spatial dimension is restored through Sigmoid and Upsample operations to obtain Mask, and the Mask and batch normalization results are combined. Multiply element by element to get the final output of this module and will obtain Enter Resnet stage i+1.

5. The cross-modal person re-identification method based on cross-scale information interaction according to claim 1, characterized in that: The two scale features output by the fourth-fifth stage feature extraction network obtained by the modality sharing feature extraction network based on parameter sharing are normalized and pooled to obtain the global and local feature representations of the visible light image and infrared image modal images, including: The two scale features output by the feature extraction network in the fourth and fifth stages are defined as for The global feature branch is used to perform generalized average pooling to obtain the global feature. The local feature branch divides it into multiple equally divided local blocks along the horizontal direction. Generalized average pooling is performed on each equally divided local block to obtain the local area feature. The specific formula is as follows: Among them, H, W, C represent The height, width and number of channels of the feature map, x i,j,c Represents the element with coordinates (i, j) on channel c in the feature map, where i = 1, 2, ..., H; j = 1, 2, ..., W; c = 1, 2, ..., C, p is an adjustable parameter, Gem() represents the generalized average pooling operation on the feature, Partition represents the horizontal partitioning operation to obtain an equal local feature map, It represents the global features and N local features obtained by generalized average pooling of the output of the fifth stage. It represents the output after the fourth stage of the Resnet framework and the inserted cross-scale information interaction module, and then the global features and M local features obtained by generalized average pooling.

6. The cross-modal person re-identification method based on cross-scale information interaction according to claim 1, characterized in that: According to the multi-scale global features and local features obtained by the model, the weighted sum of the cross entropy classification loss, the triplet loss and the cross-modal contrast loss is back-propagated and the model parameters of the cross-modal person re-identification network model are updated. The calculation process of each loss is as follows: First, the features are obtained through a multi-layer perceptron right Establish a cross-modal contrast loss for constraint, which is used It is expressed as follows: Among them, W1, W2 represent affine transformation, BN represents batch normalization, and ReLU is the activation function. The VI in denoted that the anchor represents the visible light feature, while the other samples involved in the loss correspond to the infrared light features. vis represents the set of visible light image samples in the current training batch, P infra (i) represents the set of the top K least similar infrared samples with the same ID as the visible sample i, and z is Features obtained after applying L2 normalization; z p It means from the set P infra (i) The feature z corresponding to the sample extracted; and z a represents the features corresponding to the samples extracted from the set A(i), A(i) represents the set of all samples in the infrared modality with IDs different from i, and P infra (i) All samples. Similarly, when the anchor is an infrared feature, the corresponding loss is The final loss constraint is and The mean of against L, and The cross entropy classification loss is calculated separately for constraints, and the formula is expressed as follows: Where p() is the probability that the classifier correctly predicts the feature, and E represents the expectation; For global features The cross entropy classification loss and triple loss are calculated for each of them, and the local features are The cross entropy classification loss is calculated respectively, and the formula is expressed as follows: Where p() is the probability that the classifier correctly predicts the feature, and E represents the expectation; represents the triple loss value, d(a,p) is the distance between the anchor sample and the positive sample, d(a,n) is the distance between the anchor sample and the negative sample, and margin is an artificially set interval parameter to ensure that there is enough distance between the positive and negative samples; The final total multi-loss joint constraint function is expressed as follows: where λ1 and λ2 are hyperparameters used to balance the contribution of each loss function.

Citation Information

Patent Citations

  • Cross-modal pedestrian re-identification method based on image generation and shared learning network

    CN114241517A

  • Cross-modal pedestrian re-identification method and system based on multi-feature learning

    CN114495010A

  • Cross-modal pedestrian re-identification method based on data enhancement and graph matching

    CN119169522A