The present application relates to the technical field of
pedestrian re-identification, and discloses an RGB-D cross-
modal pedestrian re-identification method,
system and medium based on joint cross attention, which first constructs joint RGB-D
semantic representation and alignment
processing to realize effective interaction between
modes; a multi-stage cross attention learning strategy is introduced to gradually learn the features of each mode at
multiple stages, and the interaction between
modes at each stage is only realized through cross attention, and only in the last stage, the depth features of the visible light image and the visible light features of the depth image are connected, the intra-
modal and inter-
modal relationships are gradually modeled, and therefore more robust fused multi-modal
pedestrian feature representation is obtained. Finally, the mean-max feature representation of the RGB-D mode is used to construct a joint representation at the semantic level, guide the obtained fine-grained cross-modal interaction features, and realize alignment at the semantic level. Furthermore, the heterogeneity of the cross-modal features is fitted, and more robust multi-modal pedestrian feature representation is obtained.