A Cross-modal Person Re-identification Method Based on Specific Modal Feature Compensation
By generating adversarial networks to convert cross-modal image styles, combined with the feature fusion method of attention mechanism and joint constraint strategy, the problem of ignoring specific modal information in the existing technology is solved, and the accuracy of cross-modal pedestrian re-identification is significantly improved.
Patent Information
- Application Number
- CN202210401883.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-18
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-18
AI Technical Summary
The existing cross-modal pedestrian re-identification technology ignores specific modal information when eliminating modal differences, resulting in limited pedestrian characteristic representation ability and low accuracy.
Generative adversarial networks are used to convert the style of visible light images and infrared images, generate cross-modal paired pedestrian images, and improve the discrimination of fusion features through feature fusion methods and joint constraint strategies based on attention mechanism.
Through style conversion and feature fusion methods, high-quality images are generated, the performance of pedestrian recognition is improved, and the accuracy of cross-modal pedestrian recognition is improved.
Smart Images

Figure CN115171148B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of image processing and pattern recognition, and particularly relates to a cross-modal pedestrian re-identification method based on specific modal feature compensation. Background Art
[0002] Pedestrian re-identification technology can find target pedestrians with the same identity under the fields of view of different cameras. With the establishment of smart cities and safe cities, video surveillance has been widely popularized, and pedestrian re-identification technology is widely applied in fields such as intelligent video surveillance, security, and criminal investigation. It is a popular research topic in the current field of computer vision. Existing pedestrian re-identification technologies mainly focus on pedestrian re-identification under visible light. However, visible light cameras cannot capture effective pedestrian information in the dark. Therefore, many new cameras will automatically switch to infrared cameras at night to capture effective pedestrian information. In this case, cross-modal pedestrian re-identification technology has been proposed, aiming to find pedestrians with the same identity by matching visible light images and infrared images under different cameras, and realizing cross-modal pedestrian re-identification.
[0003] Cross-modal pedestrian re-identification is not only affected by factors such as illumination changes, pedestrian pose changes, shooting perspective changes, and external occlusions, resulting in large appearance differences of the same pedestrian under different lenses. In addition, due to different imaging principles, there are serious modal differences between visible light images and infrared images. Therefore, eliminating modal differences is an important challenge faced by cross-modal pedestrian re-identification.
[0004] Existing methods for eliminating modal differences are mainly based on methods of shared modal feature learning. That is, a shared network is used to extract modal-irrelevant features of visible light images and infrared images for cross-modal pedestrian matching. However, specific modal information has important value for pedestrian re-identification. Only using modal-irrelevant features ignores specific modal information, which will limit the representation ability of pedestrian features and thus hinder the performance of cross-modal pedestrian re-identification. Summary of the Invention
[0005] In order to overcome the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a cross-modal pedestrian re-identification method based on specific modal feature compensation to solve the problem of low accuracy of cross-modal pedestrian re-identification.
[0006] In order to achieve the above purpose, the technical solution adopted by the present invention is:
[0007] A cross-modal pedestrian re-identification method based on specific modal feature compensation, comprising:
[0008] Collect visible light pedestrian images as visible light domain training images, and collect infrared pedestrian images as infrared domain training images;
[0009] Use a generative adversarial network to perform style transfer on pedestrian images in the visible light domain and the infrared domain, and generate cross-modal paired pedestrian images;
[0010] Obtain the fusion features between paired pedestrian images as the representation features of pedestrian images for pedestrian re-identification.
[0011] In one embodiment, the style transfer is implemented by a generative network and a discriminative network based on style transfer, including:
[0012] Input the pedestrian image in the visible light domain into the generative network, and output the corresponding pedestrian image in the infrared domain;
[0013] Input the pedestrian image in the infrared domain into the generative network, and output the corresponding pedestrian image in the visible light domain.
[0014] Compared with the prior art, the beneficial effects of the present invention are:
[0015] The present invention provides a cross-modal pedestrian re-identification method based on specific modal feature compensation, which uses a generative adversarial network for multi-modal image style transfer to achieve style transfer between visible light images and infrared images, thereby generating high-quality images; constructing a paired image feature fusion method based on the attention mechanism can enable the network to focus on the complementary information and redundant information between different modal paired images to improve the performance of pedestrian re-identification; constructing a joint constraint strategy can obtain more robust and discriminative fusion features, further improving the accuracy of cross-modal pedestrian re-identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 It is a flowchart of a cross-modal pedestrian re-identification method based on specific modal feature compensation disclosed by the present invention;
[0018] Figure 2 It is an algorithm network block diagram of a cross-modal pedestrian re-identification method proposed by the present invention. Among them, the upper half of the dashed box is the style transfer sub-network, and the lower half of the dashed box is the pedestrian re-identification sub-network;
[0019] Figure 3 It is a schematic diagram of the paired image feature fusion framework proposed by the present invention;
[0020] Figure 4 It is a schematic diagram of the joint constraint strategy framework proposed by the present invention. Detailed implementation manners
[0021] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0022] The terms "comprising" and "having" in the description of the embodiments, claims and drawings of the present invention, and any variations thereof, are intended to cover non-exclusive inclusion. For example, a series of steps or units are included.
[0023] The technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0024] The present invention is a cross-modal pedestrian re-identification method based on specific modal feature compensation. The generative adversarial network is used to perform style conversion on pedestrian images in two domains to generate cross-modal paired pedestrian images, and the fusion features between the paired images are used to improve the performance of cross-modal pedestrian re-identification; and based on the paired image fusion method and joint constraint strategy of the attention mechanism, the discriminability of the fusion features is enhanced, and the performance of cross-modal pedestrian re-identification is further improved.
[0025] As Figure 1 shown, the present invention specifically includes the following steps:
[0026] (1) Collect and preprocess the cross-modal pedestrian re-identification dataset to obtain training samples. Among them, visible light pedestrian images are used as visible light domain training images, and infrared pedestrian images are used as infrared domain training images;
[0027] In this embodiment, the same preprocessing operations are performed on visible light and infrared pictures: add pixel points with a width of l and a value of 0 to each side of the input image, use random cropping to obtain the same picture size, and randomly horizontally flip the picture. In this embodiment, the value of l is 10, and the obtained picture size is 288*144.
[0028] In order to eliminate the influence of color information, the visible light pedestrian images can be grayscale processed.
[0029] (2) Construct a generative network and a discriminative network based on style conversion. This model uses the generative adversarial idea to realize the style conversion between pedestrian images in the visible light domain and pedestrian images in the infrared domain to generate cross-modal paired pedestrian images.
[0030] The style conversion of the present invention refers to:
[0031] A pedestrian image input generation network in the visible light domain outputs the corresponding pedestrian image in the infrared domain;
[0032] A pedestrian image input generation network in the infrared domain outputs the corresponding pedestrian image in the visible light domain.
[0033] That is, when the original image is a visible light pedestrian image, the infrared pedestrian image is the generated image; when the original image is an infrared pedestrian image, the visible light pedestrian image is the generated image.
[0034] In this embodiment, as Figure 2 shown in the upper part of G2I a style conversion branch B from visible light to infrared and I2G a style conversion branch B from infrared to visible light are included in the generation network and the discriminator network, and each branch includes a generator and a discriminator, satisfying:
[0035]
[0036]
[0037]
[0038] Among them, X G is the visible light pedestrian image, and X I is the infrared pedestrian image;
[0039] represents the adversarial loss function between the infrared pedestrian image and the generated infrared pedestrian image;
[0040] represents the adversarial loss function between the visible light pedestrian image and the generated visible light pedestrian image;
[0041] D I* (X I ) represents the discrimination result of the discriminator on the real infrared pedestrian image;
[0042] D G* (X G ) represents the discrimination result of the discriminator on the real visible light pedestrian image;
[0043] G G2I represents that the generator takes the visible light pedestrian image as the input and then obtains a new infrared pedestrian image;
[0044] G I2G represents that the generator takes the infrared pedestrian image as the input and then obtains a new visible light pedestrian image;
[0045] D I* [G G2I (X G) represents the discrimination result of the discriminator for the generated infrared pedestrian image;
[0046] D G* [G I2G (X I ) represents the discrimination result of the discriminator for the generated visible-light pedestrian image;
[0047] L GAN represents and the sum of the adversarial losses;
[0048] The generator network and the discriminator network are trained using the following loss functions:
[0049] L rec o ns = ||X G - G I2G (X G )||1 + ||X I - G G2I (X I )||1
[0050] L cyc = ||X G - G I2G [G G2I (X G )]||1 + ||X I - G G2I [G I2G (X I )]||1
[0051]
[0052] where L recons is the reconstruction loss function that defines the visible-light pedestrian image or the infrared pedestrian image and the generated visible-light pedestrian image G I2G (X G ) or the infrared pedestrian image G G2I (X I );
[0053] L cyc is the cycle-consistency loss function that defines the visible-light pedestrian image or the infrared pedestrian image and the generated visible-light pedestrian image G I2G [G G2I (X G )] or G G2I [G I2G (X I );
[0054] and The identity loss functions for visible light pedestrian images and infrared pedestrian images are denoted as \(L\) ID denotes and the sum of the identity losses;
[0055] and use the cross - entropy loss function as the identity loss function for visible light pedestrian images and infrared pedestrian images respectively, where and are the predicted scores of visible light pedestrian images and infrared pedestrian images respectively, and \(y\) is the true pedestrian identity label;
[0056] \(\|\cdot\|_1\) represents the \(L1\) norm;
[0057] The objective functions of the generation network and discriminant network based on style transfer are:
[0058] \(L1 = L\) ID +\(\lambda_1L\) recons +\(\lambda_2L\) cyc +\(\lambda_3L\) gan
[0059] where \(L1\) represents the objective function of the generation network and discriminant network based on style transfer;
[0060] \(\lambda_1\), \(\lambda_2\) and \(\lambda_3\) are weighting coefficients.
[0061] (3) Construct a paired image feature fusion method based on the attention mechanism to obtain the fusion features between paired pedestrian images, that is, the fusion features of the original image features and the generated image features, as the representation features of pedestrian images for pedestrian re - identification
[0062] In this embodiment, as Figure 3 shown, the paired image feature fusion method based on the attention mechanism includes the following steps:
[0063] (31) Use four independent ResNet50s to extract four different types of features \(F\) V , \(F\) I* , \(F\) I and \(F\) G* , representing visible light pedestrian image features, generated infrared pedestrian image features, infrared pedestrian image features and generated visible light pedestrian image features respectively. In this embodiment, only the first four convolutional blocks of ResNet50 are used;
[0064] (32) Take the modality compensation of visible light pedestrian images as an example, that is, when the original image is a visible light pedestrian image, \(F\) V and \(F\) I* first pass through two channel attention modules;
[0065] EF V = CAM(F V ) = w SV * F V , EF I* = CAM(F I* ) = w SI* * F I*
[0066] w SV = σ(GAP(F V ) + GMP(F V ))
[0067] (33) For the EF V and EF I* respectively pass through two convolutional blocks and then pass through two channel attention modules;
[0068] CF V = ConvB(EF V , θ1), CF I* = ConvB(EF I* , θ2)
[0069] F SV = CAM(CF V ), F SI* = CAM(CF I* )
[0070] (34) Perform an average operation on F SV and F SI* to obtain the final pedestrian image fusion feature;
[0071] F VI* = Mean(F SV , F SI* ) = (F SV + F SI* ) / 2
[0072] Among them, EF V and EF I* represent the enhanced visible light pedestrian image feature and the generated infrared pedestrian image feature;
[0073] CAM(·) represents the channel attention module, w (·) represents the channel weight map, GAP(·) and GMP(·) respectively represent global average pooling and global maximum pooling;
[0074] CF V and CF I* represent the convolutional visible light pedestrian image feature and the generated infrared pedestrian image feature;
[0075] FSV and F SI* represent the finally enhanced visible-light pedestrian image features and the generated infrared pedestrian image features;
[0076] F VI* represents the fusion feature of the visible-light pedestrian image and the generated infrared pedestrian image;
[0077] When the original image is a visible-light pedestrian image, replace F V and F I* with F I and F G* , and perform steps (32) to (34) to obtain the finally enhanced infrared pedestrian image feature F SI and the generated visible-light pedestrian image feature F SG* as well as the pedestrian fusion feature F IG* of the infrared pedestrian image and the generated visible-light pedestrian image.
[0078] (4) Construct a joint constraint strategy, and use a loss function to jointly constrain the original image features, the generated image features, and the fusion features between paired pedestrian images to further improve the robustness and discriminability of the fusion features, and obtain a trained cross-modal pedestrian re-identification network based on specific modal feature compensation;
[0079] In this embodiment, as Figure 4 shown, constructing a joint constraint strategy includes the following steps:
[0080] (51) As Figure 2 the lower half and Figure 4 shown, through the cross-modal pedestrian re-identification network, six different types of features F SV , F SI , F SI* , F SG* , F VI* and F IG* are finally obtained;
[0081] (52) Taking F VI* and F IG* as an example, first perform a block operation on the two groups of features respectively to obtain block P1, and
[0082] (53) For each feature block and use global average pooling operation to obtain a global feature vector, and then send it to a fully connected layer to obtain pedestrian features and
[0083]
[0084]
[0085]
[0086]
[0087] (54)Finally, send the feature blocks of each pedestrian into the pedestrian identity classifier to predict the identity of each pedestrian;
[0088]
[0089]
[0090] Specifically, the Euclidean distance of the pedestrian image features can be calculated, and the matching results of different pedestrian images can be obtained according to the Euclidean distance.
[0091] (55) The joint constraint strategy is trained using the following loss function:
[0092] ξ ID (P id ,P gt )=-P gt log(P id )
[0093]
[0094]
[0095]
[0096] L2=L id +λ4L hc
[0097] Among them, represents the fusion feature of the visible light pedestrian image and the generated infrared pedestrian image after partitioning, represents the fusion feature of the infrared pedestrian image and the generated visible light pedestrian image after partitioning;
[0098] Part(·) represents the partitioning strategy, GAP(·) represents the global average pooling operation, and FC(·) represents the fully connected layer;
[0099] and respectively represent the predicted pedestrian identity scores;
[0100] P id and P gt respectively represent the predicted pedestrian identity score and the true pedestrian identity;
[0101] M represents M visible light pedestrian images, and their corresponding feature is Fvisible where \(N\) represents \(N\) infrared pedestrian images, and their corresponding features are \(F\). infrared ;
[0102] c visible and \(c\) infrared represent the centers of feature distributions of visible-light pedestrian images and infrared pedestrian images, respectively;
[0103] F visible,m and \(F\) infrared,n represent the features of the \(m\)-th visible image and the features of the \(n\)-th infrared pedestrian image, respectively;
[0104] \(\|\cdot\|_2\) represents the \(L_2\) norm;
[0105] L id represents the pedestrian identity loss function;
[0106] L hc represents the metric loss function;
[0107] \(\lambda_4\) represents the weighting coefficient;
[0108] \(L_2\) represents training the cross-modal pedestrian re-identification network with a joint constraint strategy;
[0109] (5) Verify the effectiveness of the proposed cross-modal pedestrian re-identification method based on specific modal feature compensation, and test the trained cross-modal pedestrian re-identification network using a publicly available dataset to obtain corresponding results.
[0110] In this embodiment, to verify the effectiveness of the proposed pedestrian re-identification method, metric performance evaluation is carried out on the publicly available datasets SYSU-MM01 and RegDB.
[0111] The technical effects of the present invention are further described below in conjunction with simulation experiments:
[0112] 1. Simulation conditions: All simulation experiments are implemented on an operating system of Ubuntu 16.04.5, a hardware environment of GPU Nvidia GeForce GTX 2080Ti, and using the PyTorch deep learning framework;
[0113] 2. Simulation content and result analysis:
[0114] The results obtained by experimenting with the present invention and the existing cross-modal pedestrian re-identification method based on shared modal feature learning on two public cross-modal pedestrian re-identification datasets SYSU-MM01 and RegDB are objectively evaluated using recognized evaluation metrics. The evaluation simulation results are shown in Tables 1 and 2:
[0115] Table 1 Experimental results on the SYSU-MM01 dataset
[0116]
[0117] Table 2 Experimental Results on the RegDB Dataset
[0118]
[0119]
[0120] Table 3 Experimental Results on the SYSU-MM01 Dataset
[0121] Methods Rank-1 Rank-10 Rank-20 mAP Baseline 48.03 88.74 95.12 46.83 Baseline+PwIF 57.00 92.17 97.41 54.51 Baseline+PwIF+IAI 64.23 95.19 98.73 61.21
[0122] Wherein:
[0123] Rank-1, Rank-10, Rank-20, and mAP respectively represent the Top-1 pedestrian image recognition accuracy, Top-10 pedestrian image recognition accuracy, Top-20 pedestrian image recognition accuracy, and mean average precision;
[0124] All-Search represents pedestrian re-identification in the panoramic mode, including indoor and outdoor camera scenarios;
[0125] Indoor-Search represents pedestrian re-identification in the indoor mode;
[0126] Single-shot means that only one image is selected for each pedestrian identity in the gallery;
[0127] Baseline, Baseline+PwIF, and Baseline+PwIF+IAI respectively represent the basic network, the basic network plus the paired image feature fusion method, and the basic network plus the paired image fusion method and the joint constraint strategy.
[0128] The higher the Rank-1, Rank-10, Rank-20, and mAP, the better. As can be seen from Table 1 and Table 2, on the two public datasets, the present invention achieves the best performance in each index, and significantly improves the cross-modal pedestrian re-identification performance. As can be seen from Table 3, the paired image feature fusion method and the joint constraint strategy of the present invention jointly improve the accuracy of the cross-modal pedestrian re-identification task, further enhancing the performance of the basic network, fully demonstrating the effectiveness and superiority of the method of the present invention.
[0129] The above has described the embodiments of the present invention in detail. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
[0130] In the above specific embodiments, the object, technical solution and beneficial effects of the present invention have been further described in detail. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A cross-modal pedestrian re-identification method based on specific modal feature compensation, characterized in that Including: Collect visible-light pedestrian images as visible-light domain training images, and collect infrared pedestrian images as infrared domain training images; Use a generative adversarial network to perform style conversion on pedestrian images in the visible-light domain and the infrared domain, and generate cross-modal paired pedestrian images; Obtain the fusion features between paired pedestrian images as the representation features of pedestrian images for pedestrian re-identification; Among them, the fusion features are the fusion features of the original image and the generated image, and are obtained through a paired image feature fusion method based on an attention mechanism. The method is as follows: (1) Four independent ResNet50s are used to extract four different types of features F V , F I* , F I and represent visible-light pedestrian image features, generated infrared pedestrian image features, infrared pedestrian image features, and generated visible-light pedestrian image features respectively; the original image is a visible-light pedestrian image or an infrared pedestrian image, and the generated image is an infrared pedestrian image or a visible-light pedestrian image; (2) When the original image is a visible light pedestrian image, F V and first pass through two channel attention modules; w SV = σ(GAP(F V ) + GMP(F V )) (3) EF V and respectively pass through two convolutional blocks and then pass through two channel attention modules; CF V = ConvB(EF V , θ1), F SV = CAM(CF V ), (4) Average operation on F SV and to obtain the final pedestrian image fusion feature; Among them, EF V and represent the enhanced visible light pedestrian image features and the generated infrared pedestrian image features; CAM(·) represents the channel attention module, and w (·) represents the channel weight map. GAP(·) and GMP(·) represent global average pooling and global max pooling respectively; CF V and represent the visible-light pedestrian image features and the generated infrared pedestrian image features after convolution; F SV and represent the finally enhanced visible-light pedestrian image features and the generated infrared pedestrian image features; Represents the fusion feature of the visible light pedestrian image and the generated infrared pedestrian image; When the original image is a visible-light pedestrian image, replace F V and with F I and Execute steps (2) to (4) to obtain the finally enhanced infrared pedestrian image feature F SI and generate the visible-light pedestrian image feature as well as the pedestrian fusion feature of the infrared pedestrian image and the generated visible-light pedestrian image 2. The cross-modal pedestrian re-identification method based on specific modal feature compensation according to claim 1, characterized in that Perform the same preprocessing operations on visible-light and infrared images: add pixel points with a width of l and a value of 0 to each side of the input image, and use random cropping to obtain the same image size, and then randomly flip the image horizontally.
3. The cross-modal pedestrian re-identification method based on specific modal feature compensation according to claim 2, wherein The preprocessing operation also includes: performing grayscale processing on visible-light images.
4. The cross-modal pedestrian re-identification method based on specific modal feature compensation according to claim 1, wherein The style conversion is realized through a generative network and a discriminative network based on style conversion, including: Input the pedestrian image in the visible-light domain into the generative network, and output the corresponding pedestrian image in the infrared domain; Input the pedestrian image in the infrared domain into the generative network, and output the corresponding pedestrian image in the visible-light domain.
5. The cross-modal pedestrian re-identification method based on specific modal feature compensation according to claim 4, wherein The generation network and the discriminant network adopt the generative adversarial idea and include a style conversion branch B from the visible light domain to the infrared domain G2I and a style conversion branch B from the infrared domain to the visible light domain I2G , each branch includes a generator and a discriminator, and satisfies: Among them, X G is a visible light pedestrian image, and X I is an infrared pedestrian image; Denotes the adversarial loss function between the infrared pedestrian image and the generated infrared pedestrian image; Indicates the adversarial loss function between the visible-light pedestrian image and the generated visible-light pedestrian image; Indicates the discrimination result of the discriminator for real infrared pedestrian images; Indicates the discriminant result of the discriminator for real visible-light pedestrian images; G G2I Indicates that the generator takes a visible-light pedestrian image as input and then obtains a new infrared pedestrian image; G I2G Indicates that the generator takes the infrared pedestrian image as input and then obtains a new visible-light pedestrian image; Indicates the discrimination result of the discriminator on the generated infrared pedestrian image; Indicates the discrimination result of the discriminator on the generated visible light pedestrian image; L GAN denotes and the sum of the adversarial losses; The generative network and the discriminative network use the following loss function for training: L recons = ||X G - G I2G (X G )||1 + ||X I - G G2I (X I )||1 L cyc = ||X G - G I2G [G G2I (X G )]||1 + ||X I - G G2I [G I2G (X I )]||1 Among them, L recons is the reconstruction loss function that defines the visible light pedestrian image or the infrared pedestrian image and the generated visible light pedestrian image G I2G (X G ) or the infrared pedestrian image G G2I (X I ); L cyc is a cyclic consistency loss function defined between the visible light pedestrian image or the infrared pedestrian image and the generated visible light pedestrian image G I2G [G G2I (X G )] or G G2I [G I2G (X I )]; and respectively represent the identity loss functions of visible light pedestrian images and infrared pedestrian images, L ID denotes and the sum of the identity losses; and respectively use the cross-entropy loss function as the identity loss function for visible-light pedestrian images and infrared pedestrian images. Among them, and are the predicted scores of visible-light pedestrian images and infrared pedestrian images respectively, and y is the true pedestrian identity label; ||·||1 represents the L1 norm; The objective function L1 of the generative network and the discriminative network based on style conversion is: L1 = L ID + λ1L recons + λ2L cyc + λ3L gan Among them, λ1, λ2, and λ3 are weighting coefficients.
6. The cross-modal pedestrian re-identification method based on specific modal feature compensation according to claim 5, wherein The generated cross-modal paired pedestrian images are and Among them, represents the visible light pedestrian image and its corresponding generated infrared pedestrian image, represents the infrared pedestrian image and its corresponding generated visible light pedestrian image.
7. The cross-modal pedestrian re-identification method based on specific modal feature compensation according to claim 1, wherein Construct a joint constraint strategy, and use the loss function to jointly constrain the original image features, the generated image features, and the fusion features between paired pedestrian images to improve the robustness and discriminability of the fusion features, and obtain a trained cross-modal pedestrian re-identification network based on specific modal feature compensation.
8. The cross-modal pedestrian re-identification method based on specific modal feature compensation according to claim 7, wherein The construction of the joint constraint strategy includes the following steps: (1) Obtain F through a cross-modal pedestrian re-identification network SV , F SI , and (2) For and First, perform a chunking operation on each of them separately to obtain block P1, and (3) For each feature block and use the global average pooling operation to obtain a global feature vector, and then send it to a fully connected layer to obtain the pedestrian feature and where p1 = 1,..., P1; (4) Finally, send the feature blocks of each pedestrian into the pedestrian identity classifier to predict the identity of each pedestrian; (5) The joint constraint strategy uses the following loss function for training: ξ ID (P id ,P gt )=-P gt log(P id ) L2 = L id + λ4L hc Among them, represents the fusion feature of the segmented visible-light pedestrian image and the generated infrared pedestrian image, represents the fusion feature of the segmented infrared pedestrian image and the generated visible-light pedestrian image; Part(·) represents the chunking strategy, GAP(·) represents the global average pooling operation, and FC(·) represents the fully connected layer; and respectively represent the predicted pedestrian identity scores; P id and P gt respectively represent the predicted pedestrian identity score and the true pedestrian identity; M represents M visible light pedestrian images, and their corresponding feature is F visible , N represents N infrared pedestrian images, and their corresponding feature is F infrared ; c visible and c infrared respectively represent the feature distribution centers of visible light pedestrian images and infrared pedestrian images; F visible,m and F infrared,n respectively represent the features of the m-th visible image and the features of the n-th infrared pedestrian image; ||·||2 represents the L2 norm; L id Indicates the pedestrian identity loss function; L hc represents a metric loss function; λ4 represents the weighting coefficient; L2 represents the joint constraint strategy for training the pedestrian re-identification network.
9. The cross-modal pedestrian re-identification method based on specific modal feature compensation according to claim 1, characterized in that Use a publicly available dataset to test the trained cross-modal pedestrian re-identification network and obtain the corresponding results.
Citation Information
Patent Citations
Infrared and visible light image fusion method based on detail enhancement channel attention
CN111161201A
Cross-modal pedestrian re-identification method based on multi-modal image style conversion
CN111539255A