Cross-modal feature extraction method based on visible light and infrared pedestrian re-identification
By extracting features from visible light and infrared video sequences using the LPCG and CCIA modules, the problem of modeling local structure and spatial location information in cross-modal pedestrian re-identification was solved. This enabled refined alignment and fusion of cross-modal features, improving recognition accuracy and stability.
CN122049583APending Publication Date: 2026-05-15NANTONG UNIV
View PDF 0 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANTONG UNIV
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-15
Smart Images

Figure CN122049583A_ABST
Abstract
The invention provides a cross-modal feature extraction method based on visible light and infrared pedestrian re-identification, and the method comprises the steps: carrying out the frame-level processing of a visible light video sequence and an infrared video sequence, and obtaining the frame-level spatial features; based on the frame-level spatial features, introducing a cross-branch context interpolation attention mechanism, executing controllable interpolation operation between the self-attention features and context features from another modal branch, and obtaining branch enhancement features; and performing end-to-end optimization on the feature extraction network by using a joint loss function combining cross entropy loss and triple loss to obtain a cross-modal feature extraction model. According to the invention, through joint modeling of local area structures, position information and channel responses in visible light and infrared branches, refined alignment and effective fusion of cross-modal features are realized.
Need to check novelty before this filing date? Find Prior Art