An image-text matching method based on semantic selection and hierarchical alignment
By using a semantic selection and hierarchical alignment approach, and leveraging gated attention and adaptive weight similarity calculation, the feature representation of the Transformer encoder is optimized, thus solving the bottleneck problems of computational cost and speed in existing models and achieving more efficient image-text matching.
CN116450877BActive Publication Date: 2026-07-21NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NORTHEASTERN UNIV CHINA
- Filing Date
- 2023-04-26
- Publication Date
- 2026-07-21
Smart Images

Figure CN116450877B_ABST
Abstract
The application provides an image-text matching method based on semantic selection and hierarchical alignment, and relates to the technical field of multi-modal data processing. In view of the requirement of the retrieval speed of the model, the application adopts an alignment-based model as a basic framework; in view of the complexity of the corresponding relationship of the modal characteristics and the redundancy of the information in the modal, the application uses a gating attention unit and an adaptive weight fine-grained calculation method to select and filter the characteristics; meanwhile, in order to excavate the performance of the Transformer encoder, the application designs a cross-modal hierarchical alignment method, optimizes the encoder, and obtains high-quality modal characteristic representation. The method provided by the application aims to help the alignment model to understand the heterogeneous modal information, and obtain higher retrieval performance without increasing additional interaction.
Need to check novelty before this filing date? Find Prior Art