A multi-modal image feature matching method based on window local and global attention

By combining the FPN architecture with window local and global attention, the problems of low efficiency and high computational burden of the ViT model in multimodal image feature matching are solved, achieving efficient and accurate feature matching, which is suitable for multimodal image datasets.

CN117649541BActive Publication Date: 2026-07-24YUNNAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YUNNAN UNIV
Filing Date
2023-11-30
Publication Date
2026-07-24

Smart Images

  • Figure CN117649541B_ABST
    Figure CN117649541B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of image processing, and discloses a window local and global attention-based multi-modal image feature matching method, which uses an FPN architecture to preliminarily extract a group of features after image data enhancement, uses window attention to interact local features, selects important features in each window, interacts global information of all features, completes final feature extraction, uses a bidirectional softmax function to process features after attention interaction, trains the model, and realizes feature matching under multi-modal images. The window local and global attention-based multi-modal image feature matching method interacts local information through window attention, significantly reduces the calculation amount of global information interaction on the basis of excellent local attention interaction, has excellent matching ability and matching accuracy, has very good generalization on various multi-modal data sets, and has very high practical value.
Need to check novelty before this filing date? Find Prior Art