An image-text matching method based on semantic selection and hierarchical alignment

By using a semantic selection and hierarchical alignment approach, and leveraging gated attention and adaptive weight similarity calculation, the feature representation of the Transformer encoder is optimized, thus solving the bottleneck problems of computational cost and speed in existing models and achieving more efficient image-text matching.

CN116450877BActive Publication Date: 2026-07-21NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTHEASTERN UNIV CHINA
Filing Date
2023-04-26
Publication Date
2026-07-21

Smart Images

  • Figure CN116450877B_ABST
    Figure CN116450877B_ABST
Patent Text Reader

Abstract

The application provides an image-text matching method based on semantic selection and hierarchical alignment, and relates to the technical field of multi-modal data processing. In view of the requirement of the retrieval speed of the model, the application adopts an alignment-based model as a basic framework; in view of the complexity of the corresponding relationship of the modal characteristics and the redundancy of the information in the modal, the application uses a gating attention unit and an adaptive weight fine-grained calculation method to select and filter the characteristics; meanwhile, in order to excavate the performance of the Transformer encoder, the application designs a cross-modal hierarchical alignment method, optimizes the encoder, and obtains high-quality modal characteristic representation. The method provided by the application aims to help the alignment model to understand the heterogeneous modal information, and obtain higher retrieval performance without increasing additional interaction.
Need to check novelty before this filing date? Find Prior Art