Cross-modal feature extraction method based on visible light and infrared pedestrian re-identification

By extracting features from visible light and infrared video sequences using the LPCG and CCIA modules, the problem of modeling local structure and spatial location information in cross-modal pedestrian re-identification was solved. This enabled refined alignment and fusion of cross-modal features, improving recognition accuracy and stability.

CN122049583APending Publication Date: 2026-05-15NANTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANTONG UNIV
Filing Date
2026-01-15
Publication Date
2026-05-15

Smart Images

  • Figure CN122049583A_ABST
    Figure CN122049583A_ABST
Patent Text Reader

Abstract

The invention provides a cross-modal feature extraction method based on visible light and infrared pedestrian re-identification, and the method comprises the steps: carrying out the frame-level processing of a visible light video sequence and an infrared video sequence, and obtaining the frame-level spatial features; based on the frame-level spatial features, introducing a cross-branch context interpolation attention mechanism, executing controllable interpolation operation between the self-attention features and context features from another modal branch, and obtaining branch enhancement features; and performing end-to-end optimization on the feature extraction network by using a joint loss function combining cross entropy loss and triple loss to obtain a cross-modal feature extraction model. According to the invention, through joint modeling of local area structures, position information and channel responses in visible light and infrared branches, refined alignment and effective fusion of cross-modal features are realized.
Need to check novelty before this filing date? Find Prior Art