A multi-modal image feature matching method based on window local and global attention
By combining the FPN architecture with window local and global attention, the problems of low efficiency and high computational burden of the ViT model in multimodal image feature matching are solved, achieving efficient and accurate feature matching, which is suitable for multimodal image datasets.
CN117649541BActive Publication Date: 2026-07-24YUNNAN UNIV +1
View PDF 0 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YUNNAN UNIV
- Filing Date
- 2023-11-30
- Publication Date
- 2026-07-24
Smart Images

Figure CN117649541B_ABST
Abstract
The application relates to the technical field of image processing, and discloses a window local and global attention-based multi-modal image feature matching method, which uses an FPN architecture to preliminarily extract a group of features after image data enhancement, uses window attention to interact local features, selects important features in each window, interacts global information of all features, completes final feature extraction, uses a bidirectional softmax function to process features after attention interaction, trains the model, and realizes feature matching under multi-modal images. The window local and global attention-based multi-modal image feature matching method interacts local information through window attention, significantly reduces the calculation amount of global information interaction on the basis of excellent local attention interaction, has excellent matching ability and matching accuracy, has very good generalization on various multi-modal data sets, and has very high practical value.
Need to check novelty before this filing date? Find Prior Art