Multi-mode pedestrian recognition system and application
By introducing an adaptive modal aggregation module and a multimodal Transformer into the multimodal pedestrian recognition system, the problem of quality differences between modals affecting the recognition performance is solved, and more efficient multimodal pedestrian recognition is achieved.
Patent Information
- Application Number
- CN202510303409.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-27
AI Technical Summary
The existing multimodal pedestrian recognition technology has the recognition performance under modal fusion and is affected by the differences in image quality of different modalities, resulting in poor recognition performance.
A multimodal pedestrian recognition system is designed. Through the task sharing learning module and the task specific learning module, the adaptive modal aggregation module and the multimodal Transformer are used to mine the relationship between images of different modalities to form modal complementarity, thereby improving the multimodal pedestrian recognition performance.
Through modal complementarity and commonal extraction, the accuracy and robustness of multimodal pedestrian recognition are improved, and are suitable for fields such as intelligent monitoring systems and public safety.
Smart Images

Figure CN120220186A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-modal pedestrian recognition technology, belonging to the field of computer vision technology. Background Art
[0002] Pedestrian recognition is an important task in computer vision, aiming to match pedestrians among different images. The most common pedestrian recognition is RGB image pedestrian recognition, that is, to match whether two RGB images from different perspectives or scenes are of the same pedestrian. Since RGB images are greatly affected by factors such as illumination, which limits their applications, the multi-modal pedestrian recognition technology combining visible light (RGB) modality, infrared (Infrared, IR) modality, and depth (Depth, D) modality has more applications. This technology can give full play to the roles of images in different modalities and has a broader application scenario, such as cross-modal pedestrian recognition (judging whether the pedestrians in two modal images are the same) and modality fusion pedestrian recognition (fusing images of multiple modalities to recognize pedestrians). This technology has extensive applications in many fields such as intelligent monitoring systems and public security.
[0003] However, in practical applications, the available high-quality modalities may be unknown. For example, during the day, high-quality RGB images can be captured. At night, if the person is far from the camera, the infrared image may have higher quality, and when the temperature difference is small, the depth image may show higher quality. There are disagreements between different modal images, which affects the recognition performance under modality fusion. Summary of the Invention
[0004] Object of the Invention: The technical problem to be solved by the present invention is to provide a multi-modal pedestrian recognition system aiming at the deficiencies of the prior art. This system can explore the relationships between different modal images, form modal complementarity to improve the multi-modal pedestrian recognition performance.
[0005] The present invention discloses a multi-modal pedestrian recognition system, including a task sharing learning module 100, a task specific learning module 200, and an identification module 300;
[0006] The task sharing learning module 100 includes an RGB branch 110, an infrared branch 120, and a depth branch 130;
[0007] The RGB branch 110 is used to extract the global feature D of the RGB image R and divide the global feature D of the RGB image R into N feature blocks, perform feature enhancement on each feature block to obtain the RGB enhanced feature block
[0008] The infrared branch 120 is used to extract the global feature D of the infrared imageI and globally feature D of the infrared image I is divided into N feature blocks, and feature enhancement is performed on each feature block to obtain infrared enhanced feature blocks
[0009] The depth branch 130 is used to extract the global feature D of the depth image D and globally feature D of the depth image D is divided into N feature blocks, and feature enhancement is performed on each feature block to obtain depth enhanced feature blocks
[0010] The task specific learning module 200 includes an adaptive modality aggregation module 210 and a single modality recognition module 220;
[0011] The adaptive modality aggregation module 210 includes a multi-modal sequence generation module 211, a multi-modal Transformer 212, and a multi-modal fusion feature correction module 213; the multi-modal sequence generation module 211 integrates the feature block In token , N RGB enhanced feature blocks N infrared enhanced feature blocks N depth enhanced feature blocks to form a multi-modal sequence Seq M , and the elements in the multi-modal sequence are the integrated feature block In token , feature blocks with the number of elements being 3N + 1; the size of the integrated feature block In token is the same as that of and is randomly initialized;
[0012] The multi-modal Transformer 212 includes K parallel attention branches and a splicing unit; each attention branch consists of an attention model and a feature block calculation unit; each attention model is used to obtain the mutual attention between the elements in the multi-modal sequence Seq M , and the feature block calculation unit is used to calculate a new feature block according to the multi-modal sequence and the mutual attention, and the calculation method is:
[0013]
[0014] where F u,k is the new feature block calculated by the k-th attention branch according to the u-th element F M in the multi-modal sequence Seq u ; k = 1, 2,..., K, u = 1, 2,..., 3N + 1; p u,j,k is the mutual attention of the multi-modal sequence Seq MThe uth element F u With the jth element F j The concatenation unit concatenates the K new feature blocks obtained by the K attention branches, and then uses a convolutional layer to reduce the size of the concatenated feature blocks to the same size as the original. The sizes of are consistent to form a fusion feature block; 3N+1 fusion feature blocks are combined into a preliminary multimodal fusion sequence
[0015] The multimodal fusion feature correction module 213 is used to correct the initial multimodal fusion sequence To make corrections, the specific steps are:
[0016] The preliminary multimodal fusion sequence Input the fully connected layer to obtain the intermediate multimodal fusion features Calculate the relationship value between the integrated feature block in the preliminary multimodal fusion sequence and each element in the intermediate multimodal fusion feature, and normalize the relationship value; correct the intermediate multimodal fusion feature to obtain the corrected multimodal fusion feature, and the correction method is:
[0017]
[0018] Among them G u ′ is the u-th element of the modified multimodal fusion feature, G u is the u-th element of the intermediate multimodal fusion feature; For the initial multimodal fusion feature The integrated feature block in With the jth element G of the intermediate multimodal fusion feature j The normalized relationship value of ; the calculation steps are:
[0019] Calculate preliminary multimodal fusion features The integrated feature block in With the jth element G of the intermediate multimodal fusion feature j The relationship value q j , j=1,2,…,3N+1;
[0020] Q j Normalize and get the normalized relationship value
[0021] G u ′ constitutes the modified multimodal fusion feature G;
[0022] The unimodal recognition module 220 includes an intra-modal Transformer 221 and a unimodal loss calculation module 222; the intra-modal Transformer 221 is used to extract RGB enhanced feature blocks. Infrared enhancement feature block Depth enhancement feature block The in-modal features are obtained to get the RGB in-modal feature block Infrared in-modal feature block Depth in-modal feature block The single-modal loss calculation module 222 includes three softmax loss functions, which are respectively used to calculate the RGB single-modal loss function L pr , infrared single-modal loss function L pi and depth single-modal loss function L pd ;
[0023] The recognition module 300 is used to calculate the Euclidean distance between the multi-modal fusion features of the pedestrian to be recognized and the candidate set pedestrians, and select the candidate image corresponding to the minimum Euclidean distance as the recognition result.
[0024] Furthermore, the RGB branch 110 includes an RGB shallow feature extractor 111, an RGB deep feature extractor 112, an RGB feature segmentation module 113, and an RGB attention pooling layer 114; the RGB shallow feature extractor 111 is used to extract the shallow features S of the RGB image R , and the RGB deep feature extractor 112 is used to extract the global features D of the RGB image according to the shallow features S of the RGB image R , the RGB feature segmentation module 113 is used to divide the global features D of the RGB image R into N feature blocks R , and the RGB attention pooling layer 114 is used to obtain N RGB enhancement feature blocks according to the N global RGB image feature blocks
[0025] The infrared branch 120 includes an infrared shallow feature extractor 121, an infrared deep feature extractor 122, an infrared feature segmentation module 123, and an infrared attention pooling layer 124; the infrared shallow feature extractor 121 is used to extract the shallow features S of the infrared image I , and the infrared deep feature extractor 122 is used to extract the global features D of the infrared image according to the shallow features S of the infrared image I , the infrared feature segmentation module 123 is used to divide the global features D of the infrared image I into N feature blocks R , and the infrared attention pooling layer 124 is used to obtain N infrared enhancement feature blocks according to the N infrared global feature blocks
[0026] The depth branch 130 includes a depth shallow feature extractor 131, a depth deep feature extractor 132, a depth feature segmentation module 133, and a depth attention pooling layer 134; the depth shallow feature extractor 131 is used to extract the shallow feature S of the depth image D , and the depth deep feature extractor 132 is used to extract the global feature D of the depth image according to the shallow feature S of the depth image D D , and the depth feature segmentation module 133 is used to divide the global feature D of the depth image D into N feature blocks The depth attention pooling layer 134 is used to obtain N depth enhanced feature blocks according to the N depth global feature blocks
[0027] The training of the parameters in the above multi-modal pedestrian recognition system includes the following steps:
[0028] S11. Input the pedestrian images in the training set into the RGB branch 110, the infrared branch 120, and the depth branch 130 according to the modal types respectively, and obtain the global feature D of the RGB image R , the global feature D of the infrared image I , and the global feature D of the depth image D ; input D R , D I , and D D into the softmax loss function respectively, and obtain the RGB global loss L gr , the infrared global loss L gi , and the depth global loss L gd ;
[0029] S12. Input the multi-modal fusion feature G into the softmax loss function to obtain the multi-modal consistency constraint L f ;
[0030] S13. Perform iterative training, and the goal of iterative training is to minimize the first overall loss function L total :
[0031] L total = L g + L p + L f
[0032] where L g = L gr + L gi + L gd is the global loss; L p = L pr + L pi + L pd is the single-modal loss.
[0033] To explore the commonalities of image features in different modalities and narrow the modality gap, the specific task learning module 200 further includes a multi-modal alignment learning module 230; the multi-modal alignment learning module 230 includes a multi-modal common feature extraction module 231 and a multi-modal alignment learning loss calculation module 232;
[0034] The multi-modal common feature extraction module 231 is used to extract the RGB modality common feature C R , the infrared modality common feature C I and the depth modality common feature C D ;
[0035] The multi-modal alignment learning loss calculation module 232 is used to calculate the multi-modal alignment learning loss L a , and the steps are as follows:
[0036] Collect the nth type of feature from the infrared modality common feature C I Collect W negative samples m represents the negative sample category, m≠n, w is the negative sample serial number, w∈{1,2…,W}; collect the nth type of feature from the RGB modality common feature C m represents the negative sample category, m≠n, w is the negative sample serial number, w∈{1,2…,W}; collect the nth type of feature from the RGB modality common feature C R Collect the nth type of feature
[0037] For Force L2 normalization of the feature embedding, that is
[0038] Calculate the loss function of infrared modality alignment learning:
[0039]
[0040] where τ is a hyperparameter that controls the data distribution level;
[0041] Collect the nth type of feature from the depth modality common feature C D Collect the nth type of feature Collect W negative samples
[0042] For Force L2 normalization of the feature embedding;
[0043] Calculate the loss function of depth modality alignment learning:
[0044]
[0045] The multi-modal alignment learning loss L a is calculated as follows:
[0046] L a = 0.5×LIR +0.5×L D 。
[0047] Furthermore, the steps for the multi-modal common feature extraction module 231 to extract the RGB-modal common feature C R , the infrared-modal common feature C I and the depth-modal common feature C D are as follows:
[0048] Connect N RGB enhanced feature blocks N infrared enhanced feature blocks N depth enhanced feature blocks to form the RGB-modal enhanced feature F R , the infrared-modal enhanced feature F I , and the depth-modal enhanced feature F D ;
[0049] Respectively apply average pooling layers to the RGB-modal enhanced feature F R , the infrared-modal enhanced feature F I , and the depth-modal enhanced feature F D to obtain the RGB-modal average enhanced feature the infrared-modal average enhanced feature and the depth-modal average enhanced feature
[0050] Respectively apply a 1×1 convolutional layer to the RGB-modal average enhanced feature the infrared-modal average enhanced feature and the depth-modal average enhanced feature to obtain the RGB-modal common feature C R , the infrared-modal common feature C I and the depth-modal common feature C D .
[0051] The training of the parameters in the above multi-modal pedestrian recognition system includes the steps of:
[0052] S21. Input the pedestrian images in the training set into the RGB branch 110, the infrared branch 120, and the depth branch 130 according to the modal types respectively to obtain the RGB image global feature D R , the infrared image global feature D I and the depth image global feature D D ; Input D R , D I and D D into the softmax loss function respectively to obtain the RGB global loss L gr , the infrared global loss L gi and the depth global loss L gd;
[0053] S22. Input the multi-modal fusion feature G into the softmax loss function to obtain the multi-modal consistency constraint L. f ;
[0054] S23. Conduct iterative training, and the goal of iterative training is to minimize the second overall loss function L'. total :
[0055] L' total = L g + L p + L f + L a
[0056] where L g = L gr + L gi + L gd , is the global loss; L p = L pr + L pi + L pd , is the single-modal loss.
[0057] Furthermore, the multi-modal fusion feature correction module 213 calculates the relationship value q of the integrated feature block in the preliminary multi-modal fusion feature j and the j-th element G j of the intermediate multi-modal fusion feature, and the calculation formula is:
[0058]
[0059] where "·" represents the dot product operator, and j = 1, 2,..., 3N + 1.
[0060] The present invention also discloses a multi-modal pedestrian recognition method, including:
[0061] Input the RGB, infrared, and depth three-modal images of the pedestrian to be recognized into the above multi-modal pedestrian recognition system, and obtain the recognition result according to the output of the recognition module.
[0062] The present invention also discloses a cross-modal pedestrian recognition method, including:
[0063] Input the pedestrian image to be recognized into the corresponding modal branch of the above multi-modal pedestrian recognition system according to the modal type, and obtain the global feature of the corresponding modal.
[0064] Input the pedestrian images in the candidate set into the corresponding modal branch of the above multi-modal pedestrian recognition system according to the modal type, and obtain the global features of the corresponding modal.
[0065] Calculate the Euclidean distance between the global features of each candidate set of pedestrian images and select the candidate pedestrian image corresponding to the minimum Euclidean distance as the recognition result.
[0066] The present invention also discloses a computer-readable storage medium, on which computer instructions are stored, and when the computer instructions run, they execute the above-mentioned multi-modal pedestrian recognition method or cross-modal pedestrian recognition method.
[0067] Beneficial effects:
[0068] The multi-modal pedestrian recognition system disclosed by the present invention has the following advantages:
[0069] 1. By using the adaptive modality aggregation module to explore the relationships between different modality images, overcome the differences between modalities to form modality complementarity and generate multi-modal fusion features. The adaptive modality aggregation module constructs multi-modal fusion features in two stages. In the first stage, a multi-modal Transformer is used to generate a preliminary multi-modal fusion sequence. The elements in the preliminary multi-modal fusion sequence embed the relationships between the intra-modal feature blocks and also introduce cross-modal relationships, enriching the feature representation. In the second stage, the importance of each modality element is re-evaluated based on the updated integrated feature blocks, and the feature blocks of the three modalities are re-integrated according to the importance. The finally generated multi-modal fusion features make full use of the heterogeneous relationships between modalities, thereby enhancing the representation ability of the fusion modality features to improve the performance of modality fusion recognition;
[0070] 2. By using the multi-modal alignment learning module to explore the commonalities between modalities to narrow the modality gap. This module reduces the distance between the IR and depth maps and the RGB image in a single direction through contrastive learning, mapping the three modalities to the same space to reduce the differences between multiple modalities, thereby achieving cross-modal recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0072] Figure 1 is a schematic structural diagram of the multi-modal pedestrian recognition system in Embodiment 1;
[0073] Figure 2 is a schematic diagram of the components of each branch in the task sharing learning module;
[0074] Figure 3 is a schematic diagram of the components of the adaptive modality aggregation module;
[0075] Figure 4Schematic diagram of the composition of the multi-modal Transformer;
[0076] Figure 5 Schematic diagram of the composition of the single-modal recognition module;
[0077] Figure 6 Schematic diagram of the structure of the multi-modal pedestrian recognition system in Embodiment 2;
[0078] Figure 7 Schematic diagram of the composition of the multi-modal alignment learning module. Detailed implementation manner
[0079] Embodiment 1
[0080] A multi-modal pedestrian recognition system, as Figure 1 shown, includes a task sharing learning module 100, a task specific learning module 200, and a recognition module 300; wherein the task sharing learning module 100 includes an RGB branch 110, an infrared branch 120, and a depth branch 130; the RGB branch 110 is used to extract the global RGB image feature D R , and divide the global RGB image feature D R into N feature blocks, perform feature enhancement on each feature block to obtain the RGB enhanced feature block The infrared branch 120 is used to extract the global infrared image feature D I , and divide the global infrared image feature D I into N feature blocks, perform feature enhancement on each feature block to obtain the infrared enhanced feature block The depth branch 130 is used to extract the global depth image feature D D , and divide the global depth image feature D D into N feature blocks, perform feature enhancement on each feature block to obtain the depth enhanced feature block
[0081] The extraction of the global image feature can be carried out in various ways. In the present invention, the shallow features of the image are first extracted, the deep features are extracted from the shallow features, and the extracted deep features are used as the global features of the image. As Figure 2 (a) shows, the RGB branch 110 includes an RGB shallow feature extractor 111, an RGB deep feature extractor 112, an RGB feature segmentation module 113, and an RGB attention pooling layer 114; the RGB shallow feature extractor 111 is used to extract the shallow feature S R of the RGB image, and the RGB deep feature extractor 112 is used to extract the global RGB image feature D R according to the shallow RGB image feature S R , and the RGB feature segmentation module 113 is used to divide the global RGB image feature DR Divided into N feature blocks The RGB attention pooling layer 114 is used to obtain N RGB enhanced feature blocks according to N RGB image global feature blocks Each feature block represents the local spatial feature of the RGB modality;
[0082] Such as Figure 2 (b), similar to the RGB branch 110, the infrared branch 120 includes an infrared shallow feature extractor 121, an infrared deep feature extractor 122, an infrared feature segmentation module 123, and an infrared attention pooling layer 124; the infrared shallow feature extractor 121 is used to extract the shallow feature S of the infrared image I and the infrared deep feature extractor 122 is used to extract the infrared image global feature D according to the infrared image shallow feature S I Extract the infrared image global feature D I and the infrared feature segmentation module 123 is used to divide the infrared image global feature D R Divided into N feature blocks The infrared attention pooling layer 124 is used to obtain N infrared enhanced feature blocks according to N infrared global feature blocks Each feature block represents the local spatial feature of the infrared modality;
[0083] Such as Figure 2 (c), the depth branch 130 includes a depth shallow feature extractor 131, a depth deep feature extractor 132, a depth feature segmentation module 133, and a depth attention pooling layer 134; the depth shallow feature extractor 131 is used to extract the shallow feature S of the depth image D and the depth deep feature extractor 132 is used to extract the depth image global feature D according to the depth image shallow feature S D Extract the depth image global feature D D and the depth feature segmentation module 133 is used to divide the depth image global feature D D Divided into N feature blocks The depth attention pooling layer 134 is used to obtain N depth enhanced feature blocks according to N depth global feature blocks Each feature block represents the local spatial feature of the depth modality.
[0084] In this embodiment, the RGB shallow feature extractor 111, the infrared shallow feature extractor 121, and the depth shallow feature extractor 131 adopt shallow feature extractors with non-shared parameters; the RGB deep feature extractor 112, the infrared deep feature extractor 122, and the depth deep feature extractor 132 adopt deep feature extractors with shared parameters in the three modalities; the RGB attention pooling layer 114, the infrared attention pooling layer 124, and the depth attention pooling layer 134 all adopt the attention pooling layer described in the literature: Wu J, Jiang J, Qi M, et al. An end-to-end heterogeneous restraint network for RGB-D cross-modal person re-identification[J]. ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM), 2022, 18(4): 1-22.
[0085] The task-specific learning module 200 includes an adaptive modality aggregation module 210 and a single-modality recognition module 220; the adaptive modality aggregation module 210 adaptively aggregates the enhanced features of the three modalities to achieve modality fusion pedestrian recognition. As Figure 3 shown, the adaptive modality aggregation module 210 includes a multi-modal sequence generation module 211, a multi-modal Transformer 212, and a multi-modal fusion feature correction module 213; the multi-modal sequence generation module 211 combines the integrated feature block In token , N RGB enhanced feature blocks N infrared enhanced feature blocks N depth enhanced feature blocks into a multi-modal sequence Seq M , that is, the elements in the multi-modal sequence Seq M are the integrated feature block In token , the feature blocks and the number of elements is 3N + 1; the size of the integrated feature block In token is the same as that of and is randomly initialized;
[0086] As Figure 4 (a) shown, the multi-modal Transformer 212 includes K parallel attention branches and a splicing unit; each attention branch consists of an attention model and a feature block calculation unit, as Figure 4 (b) shown; each attention model is used to obtain the multi-modal sequence Seq MThe mutual attention among elements in the middle, and the feature block calculation unit is used to calculate new feature blocks according to the multi-modal sequence and the mutual attention. The calculation method is as follows:
[0087]
[0088] Where F u,k is the new feature block calculated by the k-th attention branch according to the u-th element F M in the multi-modal sequence Seq u ; k = 1, 2, …, K, u = 1, 2, …, 3N + 1; p u,j,k is the mutual attention between the u-th element F M and the j-th element F u in the multi-modal sequence Seq j calculated by the attention model in the k-th attention branch.
[0089] The splicing unit splices the K new feature blocks obtained from the K attention branches together, and then uses a convolutional layer to reduce the size of the spliced feature blocks to be the same as that of to form a fused feature block; the 3N + 1 fused feature blocks are combined into a preliminary multi-modal fusion sequence
[0090] Since the multi-modal sequence Seq M includes feature blocks of three modalities, the mutual attention obtained by the attention model in each attention branch includes both intra-modal attention and inter-modal attention. For each element in Seq M , all elements are weighted and summed according to their mutual attention values with all elements to obtain a new feature block. In this way, the new feature block not only embeds the relationship between the intra-modal feature blocks but also introduces cross-modal relationships, enriching the feature expression of the preliminary multi-modal fusion sequence . In addition, since the integrated feature block In M in the multi-modal sequence Seq token is randomly initialized, calculating its relationship with each modal element will not be biased towards a certain modality. The integrated feature block in the preliminary multi-modal fusion feature synthesizes the influence of each element of each modality and corrects In token . In this embodiment, the multi-modal Transformer 212 includes 5 parallel attention branches and a splicing unit.
[0091] The multi-modal fusion feature correction module 213 is used to correct the preliminary multi-modal fusion sequence, aiming to correct according to the integrated feature block in the preliminary multi-modal fusion feature Re-evaluate the importance of each modal element, and re-integrate the feature blocks of the three modalities according to the importance to achieve a correction of modal fusion. The specific steps are as follows:
[0092] Input the preliminary multi-modal fusion sequence into the fully connected layer to obtain the intermediate multi-modal fusion features Calculate the relationship value between the integrated feature block in the preliminary multi-modal fusion sequence and each element in the intermediate multi-modal fusion features. The larger the relationship value, the more relevant the element is to the entire pedestrian, and a larger correction weight is set for this element. The present invention uses the normalized relationship value as the correction weight to correct the elements of the preliminary multi-modal fusion sequence . In the present invention, the dot product is used as the relationship value between the feature blocks, that is: the integrated feature block in the preliminary multi-modal fusion features and the j-th element Gj j of the intermediate multi-modal fusion features, and the relationship value q j is:
[0093] where "·" represents the dot product operator, and j = 1, 2,..., 3N + 1.
[0094] Normalize q j to obtain the normalized relationship value
[0095] Correct the intermediate multi-modal fusion features to obtain the corrected multi-modal fusion features. The correction method is:
[0096]
[0097] where Gu u ′ is the u-th element of the corrected multi-modal fusion features, Gu u is the u-th element of the intermediate multi-modal fusion features; Gu u σ constitutes the corrected multi-modal fusion features G.
[0098] The adaptive modal aggregation module 210 can make full use of the heterogeneous relationships between modalities, thereby enhancing the representation ability of the fused modal features to improve the performance of modal fusion recognition.
[0099] As Figure 5 shown, the single-modal recognition module 220 includes an intra-modal Transformer 221 and a single-modal loss calculation module 222; the intra-modal Transformer 221 is used to extract the intra-modal features of the RGB enhanced feature block infrared enhanced feature block depth enhanced feature block to obtain the RGB intra-modal feature block Infrared modality internal feature block Depth modality internal feature block In the present invention, the Transformer network in the literature: A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Kaiser, and I. Polosukhin, "Attention is all you need," in Advances in Neural Information Processing Systems, 2017, pp. 5998–6008 is adopted to extract the modality internal feature blocks to enrich the internal spatial representation of the modality, so as to cope with the situation of partial pedestrian absence. The single-modal loss calculation module 222 includes three softmax loss functions, which are respectively used to calculate the RGB single-modal loss function L pr 、infrared single-modal loss function L pi and depth single-modal loss function L pd .
[0100] The recognition module 300 is used to calculate the Euclidean distance between the pedestrian to be recognized and the multi-modal fusion features of the candidate set pedestrians, and select the candidate image corresponding to the minimum Euclidean distance as the recognition result.
[0101] In this embodiment, the parameters of the multi-modal pedestrian recognition system are trained according to the following steps:
[0102] S11. Input the pedestrian images in the training set into the RGB branch 110, infrared branch 120, and depth branch 130 respectively according to the modality type to obtain the RGB image global feature D R 、infrared image global feature D I and depth image global feature D D ; Input D R 、D I and D D into the softmax loss function respectively to obtain the RGB global loss L gr 、infrared global loss L gi and depth global loss L gd ;
[0103] S12. Input the multi-modal fusion feature G into the softmax loss function to obtain the multi-modal consistency constraint L f ;
[0104] S13. Perform iterative training, and the goal of iterative training is to minimize the first overall loss function L total :
[0105] L total = L g + L p + L f
[0106] where L g = L gr + L gi + L gd , is the global loss; L p = L pr + L pi + L pd , is the single-modal loss. Adding the multi-modal consistency constraint L f to the overall loss function helps to constrain different modalities to optimize in the same direction, thereby reducing the differences between modalities to improve the accuracy of cross-modal recognition. Therefore, the use of the adaptive modality aggregation module can balance different tasks.
[0107] To verify the effectiveness of the multi-modal pedestrian recognition system disclosed in the present invention, in this embodiment, a multi-modal pedestrian dataset is constructed. Each sample in this dataset contains RGB modality, infrared modality, and depth modality images of pedestrians. Using this dataset as the candidate set, a multi-modal pedestrian recognition method is carried out, including the steps:
[0108] Each sample in the candidate set is input into the task-sharing learning module of the trained multi-modal pedestrian recognition system, and the adaptive modality aggregation module outputs the multi-modal fusion features corresponding to the candidate set samples;
[0109] The RGB, infrared, and depth modality images of the pedestrian to be recognized are input into the trained multi-modal pedestrian recognition system, and the adaptive modality aggregation module outputs the multi-modal fusion features corresponding to the pedestrian to be recognized;
[0110] The recognition module calculates the Euclidean distance between the multi-modal fusion features of the pedestrian to be recognized and the pedestrians in the candidate set, and selects the candidate image corresponding to the minimum Euclidean distance as the recognition result and outputs it;
[0111] Obtain the recognition result according to the output of the recognition module.
[0112] In this embodiment, the recognition results of the multi-modal pedestrian recognition system in this embodiment and the multi-modal pedestrian recognition system without the adaptive modality aggregation module are compared. The results are shown in Table 1:
[0113] Table 1: Experimental results of Embodiment 1
[0114] Experimental system Rank-1 Rank-5 mAP System A 85.5 91.9 66.9 System B 92.2 96.1 75.1
[0115] System A, without the adaptive modality aggregation module, the recognition module will use N RGB enhanced feature blocks N infrared enhancement feature blocks N depth enhancement feature blocks Stitch them together as the recognition feature, calculate the Euclidean distance between the recognition features of the pedestrian to be recognized and the candidate set of pedestrians, and select the candidate image corresponding to the minimum value of the Euclidean distance as the recognition result and output. System B is the multi-modal pedestrian recognition system in this embodiment. As can be seen from Table 1, in this embodiment, by adding an adaptive modality aggregation module, the recognition performance is effectively improved.
[0116] Embodiment 2
[0117] In this embodiment, on the basis of Embodiment 1, multi-modal alignment learning is added to map each modality to the same mapping space to narrow the relationship between the three modalities, so as to realize different cross-modal pedestrian recognition. As Figure 6 shown, the task-specific learning module 200 further includes a multi-modal alignment learning module 230; as Figure 7 shown, the multi-modal alignment learning module 230 includes a multi-modal common feature extraction module 231 and a multi-modal alignment learning loss calculation module 232.
[0118] The multi-modal common feature extraction module 231 is used to extract the RGB modality common feature C R , the infrared modality common feature C I and the depth modality common feature C D , and the specific steps include:
[0119] Connect the N RGB enhancement feature blocks N infrared enhancement feature blocks N depth enhancement feature blocks F i D to form the RGB modality enhancement feature F R , the infrared modality enhancement feature F I , and the depth modality enhancement feature F D ;
[0120] Use the average pooling layer for the RGB modality enhancement feature F R , the infrared modality enhancement feature F I , and the depth modality enhancement feature F D respectively to obtain the RGB modality average enhancement feature the infrared modality average enhancement feature the depth modality average enhancement feature
[0121] Use the 1×1 convolutional layer for the RGB modality average enhancement feature the infrared modality average enhancement feature the depth modality average enhancement feature respectively to obtain the RGB modality common feature CR 、Common feature C of the infrared modality I and common feature C of the depth modality D .
[0122] The multi-modal alignment learning loss calculation module 232 is used to calculate the multi-modal alignment learning loss L a , and the steps are as follows:
[0123] Collect the nth type of feature from the common feature C of the infrared modality I Collect W negative samples m represents the negative sample category, m≠n, w is the negative sample serial number, w∈{1,2…,W};
[0124] Collect the nth type of feature from the common feature C of the RGB modality R
[0125] For Force L2 normalization of the feature embedding, that is
[0126] Calculate the loss function of infrared modality alignment learning:
[0127]
[0128] τ is a hyperparameter that controls the data distribution level. The larger the τ value, the softer the probability distribution. In the present invention, it is set to 0.2.
[0129] Collect the nth type of feature from the common feature C of the depth modality D Collect W negative samples
[0130] For Force L2 normalization of the feature embedding;
[0131] Calculate the loss function of depth modality alignment learning:
[0132]
[0133] The multi-modal alignment learning loss L a is calculated as follows:
[0134] L a = 0.5×L IR + 0.5×L D .
[0135] Compared with Embodiment 1, the training objective of the parameters in the multi-modal pedestrian recognition system in this embodiment is to minimize the second overall loss function L′ total :
[0136] L′ total = L g + L p + L f + L a
[0137] That is, the multi-modal alignment learning loss L is increased a 。L a includes the loss functions of infrared modal alignment learning and depth modal alignment learning, which can align both the infrared modality and the depth modality to the RGB modality, reduce the distance between the infrared and depth images and the RGB image, map the three modalities to the same space, reduce the differences between different modalities, and thus help improve the performance of multiple cross-modal recognitions.
[0138] In this embodiment, the trained multi-modal pedestrian recognition system is used for cross-modal pedestrian recognition to determine whether the pedestrians in two different-modal images are the same pedestrian, so as to verify the cross-modal performance. The cross-modal pedestrian recognition method includes the following steps:
[0139] Input the pedestrian image to be recognized into the corresponding modal branch of the trained multi-modal pedestrian recognition system according to the modal type, and obtain the global feature of the corresponding modality
[0140] Input the pedestrian images in the candidate set into the corresponding modal branch of the trained multi-modal pedestrian recognition system according to the modal type, and obtain the global feature of the corresponding modality
[0141] Calculate the Euclidean distance between the global feature of each candidate set pedestrian image and select the candidate pedestrian image corresponding to the minimum Euclidean distance as the recognition result.
[0142] In the experiment of this embodiment, three recognition systems are used: System A, removing the adaptive modal aggregation module; System B, the multi-modal pedestrian recognition system in Embodiment 1; System C, the multi-modal pedestrian recognition system in Embodiment 2. The experiment is divided into three groups. The first group RGB-IR: the pedestrian image to be recognized is in the RGB modality, M1 = R, and the pedestrian images in the candidate set are in the infrared modality, M2 = I; the second group IR-D: the pedestrian image to be recognized is in the IR modality, M1 = I, and the pedestrian images in the candidate set are in the depth modality, M2 = D; the third group D-RGB: the pedestrian image to be recognized is in the depth modality, M1 = D, and the pedestrian images in the candidate set are in the RGB modality, M2 = R. The experimental results are shown in Table 2:
[0143] Table 2: Experimental Results of Embodiment 2
[0144] Experimental system Mode Rank-1 mAP System A RGB-IR 32.0 13.7 System B RGB-IR 35.4 15.8 System C RGB-IR 36.9 16.0 System A IR-D 39.5 18.7 System B IR-D 39.0 20.3 System C IR-D 40.1 21.9 System A D-RGB 25.4 17.8 System B D-RGB 30.3 20.0 System C D-RGB 31.1 20.6
[0145] As can be seen from the above experimental results, the addition of the adaptive modality aggregation module and multi-modal alignment learning in the present invention can effectively improve the cross-modal recognition performance of the network.
[0146] The present invention also discloses a computer-readable storage medium, on which computer instructions are stored, and when the computer instructions run, they execute the above multi-modal pedestrian recognition method or cross-modal pedestrian recognition method.
[0147] The present invention also discloses an electronic device, including a processor and a storage medium, where the storage medium is the above computer-readable storage medium; the processor loads and executes the instructions and data in the storage medium to implement the above multi-modal pedestrian recognition method or cross-modal pedestrian recognition method.
[0148] The present invention provides an idea and method for multi-modal pedestrian recognition. There are many methods and ways to specifically implement this technical solution. The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by existing technologies.
Claims
1. A multimodal pedestrian recognition system, characterized in that: It includes a task sharing learning module (100), a task specific learning module (200) and an identification module (300); The task sharing learning module (100) includes an RGB branch (110), an infrared branch (120) and a depth branch (130); The RGB branch (110) is used to extract the global feature D of the RGB image. R , and the RGB image global feature D R Divide into N feature blocks, enhance each feature block, and obtain RGB enhanced feature blocks The infrared branch (120) is used to extract the global feature D of the infrared image. I , and the infrared image global feature D I Divide into N feature blocks, enhance each feature block, and obtain infrared enhanced feature blocks The depth branch (130) is used to extract the global feature D of the depth image. D , and the global feature D of the deep image D Divide into N feature blocks, enhance each feature block, and obtain a deep enhanced feature block The task-specific learning module (200) includes an adaptive modality aggregation module (210) and a single modality recognition module (220); The adaptive modality aggregation module (210) includes a multimodal sequence generation module (211), a multimodal transformer (212) and a multimodal fusion feature correction module (213); the multimodal sequence generation module (211) integrates the feature block In token , N RGB enhanced feature blocks N infrared enhancement feature blocks N depth-enhanced feature blocks Combined into a multimodal sequence Seq M , the elements in the multimodal sequence are integrated feature blocks In token , feature block The number of elements is 3N+1; the integrated feature block In token The size and The sizes of are consistent and randomly initialized; The multimodal Transformer (212) includes K parallel attention branches and a splicing unit; each attention branch is composed of an attention model and a feature block calculation unit; each attention model is used to obtain a multimodal sequence Seq M The mutual attention between the elements in the feature block calculation unit is used to calculate the new feature block according to the multimodal sequence and the mutual attention. The calculation method is: Among them, F u,k For the kth attention branch, according to the multimodal sequence Seq M The uth element F u The new feature block calculated; k = 1, 2, ..., K, u = 1, 2, ..., 3N + 1; p u,j,k The multimodal sequence Seq computed for the attention model in the kth attention branch M The uth element F u With the jth element F j The concatenation unit concatenates the K new feature blocks obtained by the K attention branches, and then uses a convolutional layer to reduce the size of the concatenated feature blocks to the same size as the original. The sizes of are consistent to form a fusion feature block; 3N+1 fusion feature blocks are combined into a preliminary multimodal fusion sequence The multimodal fusion feature correction module (213) is used to correct the initial multimodal fusion sequence To make corrections, the specific steps are: The preliminary multimodal fusion sequence Input the fully connected layer to obtain the intermediate multimodal fusion features Calculate the relationship value between the integrated feature block in the preliminary multimodal fusion sequence and each element in the intermediate multimodal fusion feature, and normalize the relationship value; correct the intermediate multimodal fusion feature to obtain the corrected multimodal fusion feature, and the correction method is: Among them G u ′ is the u-th element of the modified multimodal fusion feature, G u is the u-th element of the intermediate multimodal fusion feature; For the initial multimodal fusion feature The integrated feature block in With the jth element G of the intermediate multimodal fusion feature j The normalized relationship value of The calculation steps are: Calculate preliminary multimodal fusion features The integrated feature block in With the jth element G of the intermediate multimodal fusion feature j The relationship value q j , j=1,2,…,3N+1; Q j Normalize and get the normalized relationship value G′ u Composing the modified multimodal fusion feature G; The unimodal recognition module (220) comprises an intra-modal Transformer (221) and a unimodal loss calculation module (222); the intra-modal Transformer (221) is used to extract RGB enhanced feature blocks. Infrared Enhanced Feature Block Depth Enhancement Feature Block The intra-modal features of RGB modality are obtained. Infrared modality feature block Deep Intra-Modal Feature Blocks The unimodal loss calculation module (222) includes three softmax loss functions, which are used to calculate the RGB unimodal loss function L pr , infrared single modal loss function L pi and the deep unimodal loss function L pd ; The recognition module (300) is used to calculate the Euclidean distance between the multimodal fusion features of the pedestrian to be recognized and the pedestrians in the candidate set, and select the candidate image corresponding to the minimum value of the Euclidean distance as the recognition result.
2. The multimodal pedestrian recognition system according to claim 1, characterized in that: The RGB branch (110) includes an RGB shallow feature extractor (111), an RGB deep feature extractor (112), an RGB feature segmentation module (113) and an RGB attention pooling layer (114); the RGB shallow feature extractor (111) is used to extract shallow features S of the RGB image. R The RGB deep feature extractor (112) is used to extract the shallow features S of the RGB image according to the R Extract global features D from RGB images R The RGB feature segmentation module (113) is used to segment the RGB image global feature D R Divided into N feature blocks The RGB attention pooling layer (114) is used to obtain N RGB enhanced feature blocks according to the N RGB image global feature blocks. The infrared branch (120) comprises an infrared shallow feature extractor (121), an infrared deep feature extractor (122), an infrared feature segmentation module (123) and an infrared attention pooling layer (124); the infrared shallow feature extractor (121) is used to extract shallow features S of the infrared image. I The infrared deep feature extractor (122) is used to extract the infrared image shallow feature S I Extract global features of infrared images D I The infrared feature segmentation module (123) is used to segment the infrared image global feature D R Divided into N feature blocks The infrared attention pooling layer (124) is used to obtain N infrared enhanced feature blocks according to the N infrared global feature blocks. The depth branch (130) includes a depth shallow feature extractor (131), a depth deep feature extractor (132), a depth feature segmentation module (133) and a deep attention pooling layer (134); the depth shallow feature extractor (131) is used to extract shallow features S of the depth image. D The deep feature extractor (132) is used to extract the shallow features S of the depth image according to the depth image. D Extract the global feature D of the deep image D The depth feature segmentation module (133) is used to segment the depth image global feature D D Divided into N feature blocks The deep attention pooling layer (134) is used to obtain N deep enhancement feature blocks according to the N deep global feature blocks.
3. The multimodal pedestrian recognition system according to claim 1, characterized in that: The training of parameters in the multimodal pedestrian recognition system comprises the following steps: S11, input the pedestrian images in the training set into the RGB branch (110), infrared branch (120) and depth branch (130) according to the modality type, and obtain the RGB image global feature D R , infrared image global features D I and the global feature D of the deep image D ; D R , D I and D D Input the softmax loss function separately to get the RGB global loss L gr , infrared global loss L gi and the deep global loss L gd ; S12. Input the multimodal fusion feature G into the softmax loss function to obtain the multimodal consistency constraint L f ; S13, perform iterative training, the goal of which is to minimize the first overall loss function L total : L total =L g +L p +L f Where L g =L gr +L gi +L gd , is the global loss; L p =L pr +L pi +L pd , is a single-mode loss.
4. The multimodal pedestrian recognition system according to claim 1, characterized in that: The task-specific learning module (200) further includes a multimodal alignment learning module (230); the multimodal alignment learning module (230) includes a multimodal common feature extraction module (231) and a multimodal alignment learning loss calculation module (232); The multi-modal common feature extraction module (231) is used to extract the RGB modal common feature C R 、Common features of infrared modalities C I Shared feature C with the deep modality D ; The multimodal alignment learning loss calculation module (232) is used to calculate the multimodal alignment learning loss L a , the steps are as follows: From the infrared modality, the common feature C I Collect the nth type of features Collect W negative samples m represents the negative sample category, m≠n, w is the negative sample number, w∈{1,2…,W}; from the RGB modality common feature C R Collect the nth type of features right Force L2 normalized feature embedding, that is Calculate the loss function for infrared modality alignment learning: Where τ is a hyperparameter that controls the data distribution level; From the deep modality common features C D Collect the nth type of features Collect W negative samples right Force L2 normalized feature embedding; Calculate the loss function for deep modality alignment learning: Multimodal alignment learning loss L a The calculation is as follows: L a =0.5×L IR +0.5×L D 。 5. The multimodal pedestrian recognition system according to claim 4, characterized in that: The multi-modal common feature extraction module (231) extracts the RGB modal common feature C R 、Common features of infrared modalities C I Shared feature C with the deep modality D The steps are: N RGB enhanced feature blocks are N infrared enhancement feature blocks N depth-enhanced feature blocks Connect to RGB modality enhanced feature F R 、Infrared mode enhancement feature F I , deep modality enhancement feature F D ; Enhance the features F for RGB mode respectively R 、Infrared mode enhancement feature F I , deep modality enhancement feature F D Use the average pooling layer to obtain the average enhanced features of the RGB modality Infrared modal average enhancement feature Deep modal average enhancement feature Average enhanced features for RGB modalities respectively Infrared modal average enhancement feature Deep modal average enhancement feature Use a 1×1 convolution layer to obtain the common feature C of the RGB modality R 、Common features of infrared modalities C I Shared feature C with the deep modality D .
6. The multimodal pedestrian recognition system according to claim 4, characterized in that: The training of parameters in the multimodal pedestrian recognition system comprises the following steps: S21, input the pedestrian images in the training set into the RGB branch (110), infrared branch (120) and depth branch (130) according to the modality type, and obtain the RGB image global feature D R , infrared image global features D I and the global feature D of the deep image D ; D R , D I and D D Input the softmax loss function separately to get the RGB global loss L gr , infrared global loss L gi and the deep global loss L gd ; S22. Input the multimodal fusion feature G into the softmax loss function to obtain the multimodal consistency constraint L f ; S23, perform iterative training, the goal of which is to minimize the second overall loss function L' total : L' total =L g +L p +L f +L a Where L g =L gr +L gi +L gd , is the global loss; L p =L pr +L pi +L pd , is a single-mode loss.
7. The multimodal pedestrian recognition system according to claim 1, characterized in that: The multimodal fusion feature correction module (213) calculates the preliminary multimodal fusion feature The integrated feature block in With the jth element G of the intermediate multimodal fusion feature j The relationship value q j The calculation formula is: Wherein "·" represents the dot product operator, j = 1, 2, ..., 3N + 1.
8. A multimodal pedestrian recognition method, characterized in that: include: The RGB, infrared and depth images of the pedestrian to be identified are input into the multimodal pedestrian recognition system described in any one of claims 1 to 7, and the recognition result is obtained according to the output of the recognition module.
9. A cross-modal pedestrian recognition method, characterized in that: include: Inputting the pedestrian image to be identified into the corresponding modal branch of the multimodal pedestrian recognition system according to any one of claims 1 to 7 according to the modal type, and obtaining the global features of the corresponding modality Input the pedestrian images in the candidate set into the corresponding modal branch of the multimodal pedestrian recognition system according to any one of claims 1 to 7 according to the modal type, and obtain the global features of the corresponding modality calculate The global features of each candidate pedestrian image The Euclidean distance between them is used to select the candidate pedestrian image corresponding to the minimum value of the Euclidean distance as the recognition result.
10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed, the multimodal pedestrian recognition method according to claim 8 or the cross-modal pedestrian recognition method according to claim 9 is executed.