A full-time multimodal person re-identification method based on simulation augmentation and prototype learning

Through the method of simulating lighting augmentation and multimodal prototype learning, the feature extraction of the multimodal pedestrian re-identification model in the whole period is optimized, which solves the problems of light changes and data loss, and improves the recognition performance and robustness.

CN118799919BActive Publication Date: 2025-05-09WUHU SIMBA NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410972222.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2025-05-09
Estimated Expiration
2044-07-19

AI Technical Summary

Technical Problem

The existing multimodal pedestrian re-identification method for full-time multimodal pedestrian re-identification is difficult to provide effective discrimination information when facing light changes and data loss, resulting in a degradation of identification performance.

Method used

Using a method based on simulation augmentation and prototype learning, diversified data are generated through simulated lighting augmentation, and interactive learning of multimodal prototypes and instances is combined to optimize initial feature extraction to ensure the stability and accuracy of the model under the conditions of variable lighting and missing modes.

Benefits of technology

The model's discrimination ability in the multimodal pedestrian re-identification task in the full-time period is improved, the robustness of light changes and data loss is enhanced, and the recognition performance is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118799919B_ABST
    Figure CN118799919B_ABST
Patent Text Reader

Abstract

The present invention discloses a full-time multimodal pedestrian re-identification method based on simulation augmentation and prototype learning. A multi-branch network is designed using an illumination simulation augmentation module and multimodal prototype learning to improve the accuracy of the multimodal pedestrian re-identification model in various variable illumination scenes. The idea of ​​data augmentation and prototype learning is combined with subspace feature constraints to augment modal images susceptible to illumination changes, and interact prototype and instance features, thereby improving the robustness of the model to illumination changes. The present invention generates training data under a variety of illumination conditions by training an illumination simulation augmentation module, and designs a multimodal prototype for feature learning and interaction. It can cope with possible missing situations, and can also enable the model to stably re-identify the same pedestrian. The present invention has achieved good results on a full-time multimodal pedestrian re-identification dataset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer image processing technology, and in particular to a full-time multimodal pedestrian re-identification method based on simulation augmentation and prototype learning. Background Art

[0002] Pedestrian re-identification technology is a very important direction in computer vision research and has attracted great attention from industry and academia. This technology plays a vital role in many research and application aspects such as trajectory tracking, crime prevention, traffic safety, emergency management, smart cities, etc. The pedestrian re-identification (single-modal) task is to use computer vision and deep learning technology to identify pedestrians in surveillance videos, and to re-identify the same pedestrian in different camera perspectives, different times and places. Full-time multimodal pedestrian re-identification is to introduce multiple time periods and multiple complementary modal data on the basis of traditional re-identification tasks to adapt to harsh and complex real-world scenarios.

[0003] Existing person re-identification methods can be divided into three categories: global feature extraction, local information mining, and feature metric learning. Global feature extraction is mainly a basic framework that focuses on obtaining the overall feature representation of pedestrians from the entire image for subsequent matching and recognition processes. This type of method has strong adaptability to re-identification tasks in almost all scenarios. Due to the lack of fine-grained features, these methods based on global feature extraction may find it difficult to capture sufficient discriminative information when faced with occlusion or posture changes, resulting in reduced recognition performance. Local information mining is an important part of the field of person re-identification. This type of fine-grained information utilization method focuses on analyzing and identifying detailed features of various parts of the pedestrian's body, such as clothing, hairstyle, posture, etc., and focuses on mining detailed information that is beneficial to specific tasks to enhance the accuracy of individual recognition. If the key local features are not obvious or lost, this type of method may be negatively affected and cannot achieve good recognition results. Feature metric learning is mainly a special loss design other than the basic classification loss. Its purpose is to learn an effective distance or similarity metric to accurately measure the differences between different pedestrian features, thereby improving the recognition accuracy.

[0004] In order to address the challenges of illumination changes and data missing in the full-time multimodal pedestrian re-identification task, some cross-modal re-identification methods propose to narrow the gap between modalities by aligning different modalities or generating auxiliary modalities, while multimodal methods improve performance by integrating stable modal information into the final features. However, all of the above methods still have some problems. For example, existing methods mainly focus on reducing the heterogeneity between modalities or utilizing modal information that is not affected by illumination, and often ignore the impact of illumination changes within a single modality; generation-based methods require training of additional modules; category-based prototypes may be difficult to work because the training and test identities of the re-identification task do not overlap. Therefore, in real-world environments with variable lighting conditions and unknown test situations, existing methods often do not achieve good results. Summary of the invention

[0005] Purpose of the invention: The purpose of the present invention is to solve the deficiencies in the prior art and provide a full-time multimodal pedestrian re-identification method based on simulation augmentation and prototype learning.

[0006] The present invention aims to guide the optimization of the extracted initial features through data augmentation based on illumination changes and interactive learning of multimodal prototypes and instances, so that the model can ultimately provide favorable discriminant information when dealing with full-time multimodal pedestrian re-identification tasks.

[0007] Technical solution: A full-time multimodal pedestrian re-identification method based on simulation augmentation and prototype learning of the present invention comprises the following steps:

[0008] Step 1: Obtain three modal original images of the same pedestrian scene, including visible light image I R , Near infrared image I N and thermal infrared image I T , the original image sizes of the three modalities are all 3×256×128 (number of channels, height, width); the corresponding modal prototypes are denoted as P R , P N and P T ,The modal prototype sizes of the three modes are 1×768 (number of channels and length);

[0009] Step 2: The original visible light image I R The augmented images I are obtained for the morning, afternoon and night periods by simulating the illumination augmentation module. Rm ,I Ra and I Re ;

[0010] Step 3: The augmented images I of the three time periods Rm ,I Ra and I Re, the original near infrared image I N and thermal infrared image I T Send it to the feature extractor to get the corresponding feature f Rm ,f Ra ,f Re ,f N ,f T ; Then the three features f of the augmented image Rm ,f Ra ,f Re After averaging, we get the feature f R ;

[0011] Step 4: Simultaneously transform feature f R and model p R Send IM RGB Interaction module, feature f N and modality P N Send IM NIR Interaction modules, and features f T and modality P T Send IM TIR The interaction module realizes the information interaction of the three feature prototype instance features, and obtains and

[0012] Step 5: To preserve the features of all original images and augmented images, and Concatenate by channel to get the final feature f for classifier training final ; The classifier consists of a convolutional layer, a ReLU activation layer, and a fully connected layer; the classifier is used to reduce the feature dimension to a fixed dimension for subsequent loss calculation;

[0013] Step 6: For feature f Rm ,f Ra ,f Re 、f R ,f N ,f T and f final To perform subspace constraints, the method is:

[0014] For the augmented multiple visible light features [f Rm ,f Ra ,f Re ]Use orthogonal loss l ort Constrain them to make the differences as large as possible, which can ensure the increased diversity;

[0015] For different modal features after interaction [f R ,f N ,f T ]Use consistent loss Lclo To constrain the model so that it can maintain the unity of multi-modality;

[0016] For the final feature f final Use cross entropy classification loss L ce With triplet loss L tri Conduct training.

[0017] Furthermore, after completing the subspace feature constraints in step 6, due to the particularity of the re-identification task, the features required in the test phase and the training phase are inconsistent. To deal with the missing and non-missing situations in multimodality, two test strategies are designed:

[0018] (1) When there is no missing mode, after inputting the test data, the features of each mode are obtained after feature extraction and prototype instance interaction, and then the channels are spliced ​​to form f final ,This feature will serve as the final identifier of the pedestrian;

[0019] (2) When the modality is missing, after the test data is input, the image of the existing modality will be extracted and the prototype instance interaction will be used to obtain the corresponding modality features. The features of the missing modality will be restored by the incomplete recovery module and then spliced ​​together with other modality features to form f final as the final identifier.

[0020] Furthermore, the step 2 simulates the illumination augmentation module to adjust the original visible light image I R The light intensity of the simulated light augmentation module is equipped with a brightness adjustment function based on the visible light image I R The shooting time is adjusted by the brightness adjustment function to adjust the visible light image I R Adjustment is performed, and the shooting time is divided into three time periods: morning, afternoon, and night. Finally, the augmented image I is obtained. Rm ,I Ra ,I Re , the augmented image has the same size as the original visible light image.

[0021] Furthermore, the specific processing method of the simulated illumination augmentation module is:

[0022] Step A: Obtaining light image I R , collection time label y time , the lower bound alpha of the augmented range and the upper bound beta of the augmented range;

[0023] Step B: According to y tRme The value of visible light image I R The collection time period: If y time =0, then the light image I is judged R The collection time is in the morning, if y time=1, then the light image I is judged R The collection time is afternoon, if y time =2, then the light image I is judged R The collection time is night;

[0024] Step C: Based on the light image I R The acquisition time is used to adjust the brightness of the original visible light image in different ways;

[0025] If the acquisition time is in the morning, the original visible light image is kept unchanged, and the two sub-functions min_dim_trans and max_dim_trans that reduce the brightness are used to adjust it, so as to obtain the simulated dark afternoon scene image, the simulated extremely dark night scene image and the original morning visible light image;

[0026] If the acquisition time is in the afternoon, the sub-function min_brighten_trans that increases the brightness is first used to adjust the original visible light image to simulate the morning scene, and then the original image is kept unchanged, and the sub-function min_dim_trans that reduces the brightness is used to adjust it to simulate the night scene, so that the simulated morning scene image, the original afternoon visible light image, and the simulated night scene image are obtained;

[0027] If the acquisition time is at night, the sub-functions min_brighten_trans and max_brighten_trans for improving the brightness are first used to adjust the original visible light image, thereby obtaining a simulated morning scene image, a simulated afternoon scene image and an original night visible light image.

[0028] Furthermore, the feature encoder in step 3 includes a light-sensitive encoder Light-Insensitive Encoder Light sensitive encoder Light-Insensitive Encoder Both are based on the visual transformer network structure. The illumination-sensitive encoder extracts the features of visible light images, and the illumination-insensitive encoder extracts the features of near-infrared images and thermal infrared images. The expressions are as follows:

[0029]

[0030] Finally, the features of the three visible light image instances are averaged to obtain the visible light center feature f R .

[0031] Furthermore, the processing method of the three interaction modules in step 4 is:

[0032] First, the cosine similarity between the current instance feature and the corresponding modal prototype is calculated, and then the modal attention value A is calculated based on the cosine similarity. Finally, the attention value is multiplied by the modal prototype, and the result is added to the original modal instance feature. Finally, the instance feature after interaction is obtained for subsequent constraints and training.

[0033] To ensure that the features after interaction can maintain both modal consistency and information richness, step 6 constrains different modal features in the same subspace. The specific method is as follows:

[0034] First, to ensure the information richness of the augmented features, the augmented visible light features should be as different as possible; unit vectors have low similarity when they are orthogonal, so through L ort To constrain all visible light features to remain orthogonal in the subspace. This goal can be expressed as:

[0035]

[0036] Here, · represents the inner product of two vectors; and because the unit orthogonal vector can be represented in space as two vectors with an angle of 90 degrees. The final orthogonal loss L ort Defined as:

[0037]

[0038] Among them, θ() is the cosine similarity of two unit vectors, as shown below:

[0039]

[0040] Here, ||·|| represents the norm of the vector, and p=1e-8 can avoid division by zero.

[0041] Secondly, to ensure that the visible light modality has the same illumination stability as the near infrared / thermal infrared modality, the differences between different multimodal features should be as small as possible; since the unit vectors are close in the subspace, it means that the vectors are highly similar, so let the multimodal features have a consistent loss L clo Under the supervision of is constrained to be close to the state, the expression is as follows:

[0042]

[0043] Among them, L2 represents the mean square error loss.

[0044] In order to better train, in addition to using L ort and L clo In addition, the cross entropy loss L is also used ce and triplet loss L tri ; Step 2 The total loss function of the entire training process is:

[0045] L=L ce +L tr +L ort +L clo ,

[0046] Among them, L ce ,L tri ,L ort ,L clo Represent classification loss, triplet loss, orthogonal loss and consistent loss respectively; through L ce ,L tri Constrain the final pedestrian features by L ort ,L clo The distribution of multiple modal features in the constrained subspace is controlled; at the same time, the stochastic gradient descent algorithm (SGD) is used to update the network parameters.

[0047] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0048] (1) The present invention uses simulated illumination augmentation to change the brightness of input samples, and can generate images under complex illumination conditions for training, thereby using data augmentation to improve the robustness of the model.

[0049] (2) The present invention adopts subspace constraints to ensure the diversity of feature augmentation and the stability of multimodal features to illumination.

[0050] (3) The present invention designs a multimodal prototype to learn modal domain information and uses prototype-instance interaction to embed relevant information into instance features.

[0051] (4) The present invention utilizes incomplete information recovery to combine instance features of existing modalities with pre-trained modal prototypes to recover unique information for each missing sample. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 It is a schematic diagram of the overall process of the present invention;

[0053] Figure 2 Schematic diagram of the network model designed specifically;

[0054] Figure 3 The main structure of the interactive module for the invented prototype instance;

[0055] Figure 4 A comparison chart of the visualization results of actual retrieval in different scenarios between the present invention and the benchmark method. DETAILED DESCRIPTION

[0056] The technical solution of the present invention is described in detail below, but the protection scope of the present invention is not limited to the embodiments.

[0057] like Figure 1 As shown, the full-time multimodal pedestrian re-identification method based on simulation augmentation and prototype learning of the present invention includes the following steps:

[0058] Step 1: Obtain three modal original images of the same pedestrian scene, including visible light image I R , Near infrared image I N and thermal infrared image I T , the original image size of the three modalities is 3×256×128 (number of channels, height, width); in order to obtain multimodal domain information that is independent of identity, the corresponding modal prototypes are recorded as P R , P N and P T ,The modal prototype sizes of the three modes are 1×768 (number of channels and length);

[0059] Step 2: The original visible light image I R The augmented images I are obtained for the morning, afternoon and night periods by simulating the illumination augmentation module. Rm ,I Ra and I Re ;

[0060] Step 3: The augmented images I of the three time periods Rm ,I Ra and I Re , the original near infrared image i N and thermal infrared image I T Send it to the feature extractor to get the corresponding feature f Rm ,f Ra ,f Re ,f N ,f T ; Then the three features f of the augmented image Rm ,f Ra ,f Re After averaging, we get the feature f R ;

[0061] Step 4: Simultaneously transform feature f R and model p R Send IM RGB Interaction module, feature f N and modality P N Send IM NIR Interaction modules, and features f T and modality P T Send IM TIR The interaction module realizes the information interaction of the three feature prototype instance features, and obtains and

[0062] Step 5: To preserve the features of all original images and augmented images, and Concatenate by channel to get the final feature f for classifier training final ; The classifier consists of a convolutional layer, a ReLU activation layer, and a fully connected layer; the classifier is used to reduce the feature dimension to a fixed dimension for subsequent loss calculation;

[0063] Step 6: For feature f Rm ,f Ra ,f Re 、f R ,f N ,f T and f final To perform subspace constraints, the method is:

[0064] For the augmented multiple visible light features [f Rm ,f Ra ,f Re ]Use the orthogonal loss L ort Constrain them to make the differences as large as possible, which can ensure the increased diversity;

[0065] For different modal features after interaction [f R ,f N ,f T ]Use consistent loss L clo To constrain the model so that it can maintain the unity of multi-modality;

[0066] For the final feature f final Use cross entropy classification loss L ce With triplet loss L tri Conduct training.

[0067] If there is no missing mode, the pedestrian feature f can be obtained directly according to the training process final If there is a missing mode, the incomplete recovery module is needed to obtain the missing mode features, and then splice them together with other existing mode features for the final test. Figure 2 As shown in part (c) of .

[0068] like Figure 2 As shown, in the network model of the present invention, the simulated illumination augmentation module, the three prototype instance interaction modules and the subspace feature constraint are the three core parts. Data augmentation can improve the stability of the model, and the prototype can retain the characteristics of modal specific information. Combining multiple feature constraints to ensure network learning, the model can learn modal features containing stable and diverse information, thereby performing better pedestrian identification.

[0069] In order to alleviate the data gap between the training set and the test set and improve the model's ability to process images across time periods, step 2 simulates the illumination augmentation module to adjust the original visible light image I R The light intensity of the simulated light augmentation module is equipped with a brightness adjustment function based on the visible light image I R The shooting time is adjusted by the brightness adjustment function to adjust the visible light image R R Adjustment is performed, and the shooting time is divided into three time periods: morning, afternoon, and night. Finally, the augmented image I is obtained. Rm ,I Ra ,I Re , the augmented image has the same size as the original visible light image. The augmentation function definitions include: small brightness enhancement function min_brighten_trans, large brightness enhancement function max_brighten_trans, small brightness reduction function min_dim_trans, and large brightness reduction function max_dim_trans. In addition to the above four augmentation functions, there are other basic augmentation operations performed on the original visible light image: scaling, random horizontal flipping, zero padding, normalization, and random erasing.

[0070] The specific processing method of the simulated illumination augmentation module of this embodiment is as follows:

[0071] Step A: Obtaining light image I R , Collection time label Y time , the lower bound alpha of the augmented range and the upper bound beta of the augmented range;

[0072] Step B: According to y time The value of visible light image I R The collection time period: If y time =0, then the light image I is judged R The collection time is in the morning, if y time =1, then the light image I is judged R The collection time is afternoon, if y time =2, then the light image I is judged R The collection time is night;

[0073] Step C: Based on the light image I R The acquisition time is used to adjust the brightness of the original visible light image in different ways;

[0074] If the acquisition time is in the morning, the original visible light image is kept unchanged, and the two sub-functions min_dim_trans and max_dim_trans that reduce the brightness are used to adjust it, so as to obtain the simulated dark afternoon scene image, the simulated extremely dark night scene image and the original morning visible light image;

[0075] If the acquisition time is in the afternoon, the sub-function min_brighten_trans that increases the brightness is first used to adjust the original visible light image to simulate the morning scene, and then the original image is kept unchanged, and the sub-function min_dim_trans that reduces the brightness is used to adjust it to simulate the night scene, so that the simulated morning scene image, the original afternoon visible light image, and the simulated night scene image are obtained;

[0076] If the acquisition time is at night, the sub-functions min_brighten_trans and max_brighten_trans for improving the brightness are first used to adjust the original visible light image, thereby obtaining a simulated morning scene image, a simulated afternoon scene image and an original night visible light image.

[0077] In summary, the simulated illumination augmentation module can easily convert the current image into different brightness levels according to the image shooting time label, and finally generate an image list containing two simulated illumination images and one original image [I Rm ,I Ra ,I Re ,I N ,I T ].

[0078] After the original visible light image is augmented by simulated illumination, it is necessary to extract multimodal instance features to obtain their identity information and interact the instance features with the modal prototypes to absorb the domain knowledge in the interactive module. Due to differences in imaging principles, only the visible light modality is affected by changes in illumination, while near infrared and thermal infrared are almost unaffected by ambient illumination. Therefore, this embodiment uses two independent encoders with the same structure to extract features of modalities that are more sensitive to illumination (augmented visible light) and modalities that are not affected by illumination (near infrared and thermal infrared). Specifically, the feature encoder in step 3 of this embodiment includes an illumination-sensitive encoder. Light-Insensitive Encoder Light sensitive encoder Light-Insensitive Encoder Both are based on the visual transformer network structure. The illumination-sensitive encoder extracts the features of visible light images, and the illumination-insensitive encoder extracts the features of near-infrared images and thermal infrared images. The expressions are as follows:

[0079]

[0080] Finally, the features of the three visible light image instances are averaged to obtain the visible light center feature f R .

[0081] like Figure 3As shown, the processing method of the three interaction modules in step 4 of this embodiment is:

[0082] First, the cosine similarity between the current instance feature and the corresponding modal prototype is calculated, and then the modal attention value A is calculated based on the cosine similarity. Finally, the attention value is multiplied by the modal prototype, and the result is added to the original modal instance feature. Finally, the instance feature after interaction is obtained for subsequent constraints and training.

[0083] For example, for a visible light image, the attention value A between the visible light center feature and the visible light modal prototype is first calculated by cosine similarity. RR , and then integrate the modal information in the prototype into the instance features through a two-step operation of multiplication and addition. R The process is formulated as:

[0084] f i p =f i +proj(p R )×A RR ,i∈[R m ,R a ,R e ]

[0085] in Represents cosine similarity; proj is composed of a linear layer, a batch normalization layer, and a ReLU layer.

[0086] For near infrared and thermal infrared modalities, the interactive module IM N and IM N The overall process and IM R Similar; similarity calculation is performed first, and then weighted interaction is performed, which can be expressed as:

[0087]

[0088] After completing the interaction between instance features and corresponding modal prototypes, features It contains both identity-independent modal information and identity-related instance information.

[0089] Furthermore, the detailed method of the subspace feature constraint in step 6 is:

[0090] First, through the orthogonal loss function L ort To constrain all visible light features to remain orthogonal in the subspace, it is expressed as follows:

[0091]

[0092] Where "·" represents the inner product of two vectors; unit orthogonal vectors are represented in space as two vectors with an angle of 90 degrees; θ() is the cosine similarity of two unit vectors, as shown below:

[0093]

[0094] ||·|| represents the norm of a vector;

[0095] Second, use a consistent loss function L clo Supervise the three modal features so that they are constrained to be close to each other, and the expression is:

[0096]

[0097] L2 represents the mean square error loss.

[0098] At this point, the total loss function of the entire training process is:

[0099] L=L ce +L tri +L ort +L clo ,

[0100] Among them, L ce ,L tri ,L ort ,L clo Represent classification loss, triplet loss, orthogonal loss and consistent loss respectively; through L ce ,L tri Constrain the final pedestrian features by L ort ,L clo The distribution of multiple modal features in the constrained subspace is controlled; at the same time, the stochastic gradient descent algorithm (SGD) is used to update the network parameters.

[0101] Embodiment 1:

[0102] Step 1: This embodiment uses the AllDay 843 and AllDay 843-G datasets as experimental data. Both datasets select a large number of multi-time period and multi-modal image groups as training and testing data.

[0103] Among them, the AllDay843 dataset: contains 91,371 images of 843 identities from 6 different perspectives with diverse lighting, weather, and background changes. The dataset is divided into two subsets: 785 identities for training and 58 identities for testing. AllDay843-G dataset: In order to compare with existing single-modal and multi-modal methods, cycleGAN is used to complete the missing data and AllDay843 is extended to AllDay843-G. Compared with AllDay843, AllDay843-G provides aligned images of three modalities for each sample. The comparison results are shown in Table 1.

[0104] Table 1 Quantitative experimental data on the AllDay843 dataset

[0105]

[0106] Step 2. This embodiment uses PyTorch of GeForce RTX 3090 GPU as the experimental platform. Two visual transformers are used as feature encoders, which are parameter-independent and pre-trained on ImageNet. After basic data augmentation and the proposed simulated lighting augmentation, the images are resized to a uniform size of 256×128. The batch size of the training process is set to 32 and the maximum number of iterations is set to 60. The initial learning rate is set to 0.001. The dimension of the modal prototype is 1×768 dimensions, which is the same as the instance feature dimension extracted by the encoder. The final feature scale of each pedestrian is 768×5=3840 dimensions.

[0107] Step 3: To facilitate quantitative evaluation, this embodiment uses two widely used indicators: CMC and mAP. CMC is a cumulative probability indicator used to evaluate the expected correct match in a given candidate list, while mAP comprehensively considers the precision and recall of all relevant samples.

[0108] (1) The horizontal axis of the CMC (Cumulative Matching Characteristics) curve is the rank, and the vertical axis is the recognition rate (the recognition rate refers to the probability of finding the correct match in the first n results). At different ranking positions, the curve will show the probability of finding the correct match before that point. Specifically, for each query, the algorithm will find the top N images in the gallery that are most similar to it. If the correct match appears in this list, the ranking is considered to be successfully hit. The probability of finding the correct match in the top k results among all queries is calculated, which is the Rank-k accuracy. Suppose there are N queries, and each query is sorted according to similarity in the candidate set. Let Acc kIt indicates the proportion of queries that contain correct matches in the first k candidate samples, that is:

[0109]

[0110] (2) mAP (mean Average Precision) is a commonly used evaluation metric in multi-target retrieval tasks. It takes into account the precision and recall of all relevant samples. For each query, the area under the precision-recall curve is calculated to obtain the AP (Average Precision) of the query, and then the AP of all queries is averaged to obtain mAP. For a query, its precision (P) and recall (R) can be defined as:

[0111]

[0112] Where TP is the true positives (the number of correct matches), FP is the false positives (the number of incorrect matches), and FN is the false negatives (the number of correct matches not found).

[0113] Qualitative evaluation:

[0114] like Figure 4 As shown, this embodiment performs the search results on the AllDay843 dataset compared with other technical solutions. Figure 4 It can be seen from the above that the present invention has a good effect. Figure 4 Judging from the results in (a), the lighting conditions are better in the morning, but the comparison method still has incorrect matches in most cases, as shown in the red box. Figure 4 (b) shows the afternoon period. As the lighting conditions change, we can see that the number of matching errors increases. This may be due to the difference in light intensity, which affects the ability of the contrast method to identify pedestrians. Figure 4 The night time in (c) is also challenging. Due to the significant decrease in lighting conditions, the number of correct matches is also small, indicating that the comparison model's ability to capture pedestrian features is reduced. Compared with the comparison method, the method of the present invention has better retrieval results in different time periods. Except for a small number of samples with incorrect matches, in most cases, the identity of pedestrians in images collected across time periods can be accurately identified.

Claims

1. A full-time multimodal person re-identification method based on simulation augmentation and prototype learning, characterized in that: The following steps are involved: Step 1: Obtain three modal original images of the same pedestrian scene, including visible light image I R , Near infrared image I N and thermal infrared image I T , the original image sizes of the three modalities are all 3×256×128; the corresponding modal prototypes are denoted as P R , P N and P T , the modal prototype sizes of the three modes are all 1×768; Step 2: The original visible light image I R The augmented images I are obtained for the morning, afternoon and night periods by simulating the illumination augmentation module. Rm ,I Ra and I Re ; Step 3: The augmented images I of the three time periods Rm ,I Ra and I Re , the original near infrared image I N and thermal infrared image I T Send it to the feature extractor to get the corresponding feature f Rm ,f Ra ,f Re ,f N ,f T ; Then the three features f of the augmented image Rm ,f Ra ,f Re After averaging, we get the feature f R ; Step 4: Simultaneously transform feature f R and model p R Send IM RGB Interaction module, feature f N and modality P N Send IM NIR Interaction modules, and features f T and modality P T Send IM TIR The interaction module realizes the information interaction of the three feature prototype instance features, and obtains and Step 5: To preserve the features of all original images and augmented images, and Concatenate by channel to get the final feature f for classifier training final ; The classifier consists of a convolutional layer, a ReLU activation layer, and a fully connected layer; the classifier is used to reduce the feature dimension to a fixed dimension for subsequent loss calculation; Step 6: For feature f Rm ,f Ra ,f Re 、f R ,f N ,f T and f final To perform subspace constraints, the method is: For the augmented multiple visible light features [f Rm ,f Ra ,f Re ]Use the orthogonal loss L ort Constrain them to make the differences as large as possible, which can ensure the increased diversity; For different modal features after interaction [f R ,f N ,f T ]Use consistent loss L clo To constrain the model so that it can maintain the unity of multi-modality; For the final feature f final Use cross entropy classification loss L ce With triplet loss L tri Conduct training.

2. The full-time multimodal person re-identification method based on simulation augmentation and prototype learning according to claim 1 is characterized in that: After completing the subspace feature constraints in step 6, due to the particularity of the re-identification task, the features required in the test phase and the training phase are inconsistent. To deal with the missing and non-missing situations in multimodality, two test strategies are designed: (1) When there is no missing mode, after inputting the test data, the features of each mode are obtained after feature extraction and prototype instance interaction, and then the channels are spliced ​​to form f final ,This feature will serve as the final identifier of the pedestrian; (2) When the modality is missing, after the test data is input, the image of the existing modality will be extracted and the prototype instance interaction will be used to obtain the corresponding modality features. The features of the missing modality will be restored by the incomplete recovery module and then spliced ​​together with other modality features to form f final as the final identifier.

3. The full-time multimodal person re-identification method based on simulation augmentation and prototype learning according to claim 1 is characterized in that: The step 2 simulates the illumination augmentation module to adjust the original visible light image I R The illumination intensity of the image is simulated, and a brightness adjustment function is set in the illumination augmentation module. R The shooting time is adjusted by the brightness adjustment function to adjust the visible light image I R Adjustment is performed, and the shooting time is divided into three time periods: morning, afternoon, and night. Finally, the augmented image I is obtained. Rm ,I Ra ,I Re , the augmented image has the same size as the original visible light image.

4. According to claim 1 or 3, the full-time multimodal pedestrian re-identification method based on simulated augmentation and prototype learning, the specific processing method of the simulated illumination augmentation module is: Step A: Obtaining visible light image i R , collection time label y time , the lower bound alpha of the augmented range and the upper bound beta of the augmented range; Step B: According to y time The value of visible light image I R The collection time period: If y time =0, then the visible light image I R The collection time is in the morning, if y time =1, then the visible light image I R The collection time is afternoon, if y time =2, then the visible light image I R The collection time is night; Step C: Based on the visible light image I R The acquisition time is used to make different brightness adjustments to the original visible light image; If the acquisition time is in the morning, the original visible light image is kept unchanged, and the two sub-functions min_dim_trans and max_dim_trans that reduce the brightness are used to adjust it, so as to obtain the simulated dark afternoon scene image, the simulated extremely dark night scene image and the original morning visible light image; If the acquisition time is in the afternoon, the sub-function min_brighten_trans that increases the brightness is first used to adjust the original visible light image to simulate the morning scene, and then the original image is kept unchanged, and the sub-function min_dim_trans that reduces the brightness is used to adjust it to simulate the night scene, so that the simulated morning scene image, the original afternoon visible light image, and the simulated night scene image are obtained; If the acquisition time is at night, the sub-functions min_brighten_trans and max_brighten_trans for improving the brightness are first used to adjust the original visible light image, thereby obtaining a simulated morning scene image, a simulated afternoon scene image and an original night visible light image.

5. The full-time multimodal person re-identification method based on simulation augmentation and prototype learning according to claim 1, characterized in that: The feature encoder in step 3 includes a light-sensitive encoder Light-Insensitive Encoder Light sensitive encoder Light-Insensitive Encoder Both are based on the visual transformer network structure. The illumination-sensitive encoder extracts the features of visible light images, and the illumination-insensitive encoder extracts the features of near-infrared images and thermal infrared images. The expressions are as follows: Finally, the features of the three visible light image instances are averaged to obtain the visible light center feature f R .

6. The full-time multimodal person re-identification method based on simulation augmentation and prototype learning according to claim 1, characterized in that: The processing method of the three interaction modules in step 4 is: First, the cosine similarity between the current instance feature and the corresponding modal prototype is calculated, and then the modal attention value A is calculated based on the cosine similarity. Finally, the attention value is multiplied by the modal prototype, and the result is added to the original modal instance feature. Finally, the instance feature after interaction is obtained for subsequent constraints and training.

7. The full-time multimodal person re-identification method based on simulation augmentation and prototype learning according to claim 1, characterized in that: The detailed method of the subspace feature constraint in step 6 is: First, through the orthogonal loss function l ort To constrain all visible light features to remain orthogonal in the subspace, it is expressed as follows: Among them, "·" represents the inner product of two vectors; unit orthogonal vectors are represented in space as two vectors with an angle of 90 degrees; θ( ) is the cosine similarity of two unit vectors, as shown below: ||·|| represents the norm of a vector; Second, use a consistent loss function L clo Supervise the three modal features so that they are constrained to be close to each other, and the expression is: L2 represents the mean square error loss.

8. The full-time multimodal person re-identification method based on simulation augmentation and prototype learning according to claim 1, characterized in that: The total loss function of the entire training process in step 2 is: L=L ce +L tri +L ort +L clo , Among them, L ce ,L tri ,L ort ,L clo Represent classification loss, triplet loss, orthogonal loss and consistent loss respectively; through L ce ,L tri Constrain the final pedestrian features by L ort ,L clo The distribution of multiple modal features in the constrained subspace is controlled; at the same time, the stochastic gradient descent algorithm (SGD) is used to update the network parameters.

Citation Information

Patent Citations

  • Visible light infrared pedestrian re-identification method based on multi-modal relation aggregation

    CN114511878A

  • Cross-modal pedestrian re-identification method based on spectrum perception and attention mechanism

    CN116798070A