多模态知识增强的跨模态表示学习与检索方法及相关设备
By collecting local fine-grained and global coarse-grained features of images and text, and using a multimodal graph attention network for cross-modal retrieval, the problem of low efficiency in cross-modal representation learning and retrieval in existing technologies is solved, and more efficient cross-modal semantic association and retrieval are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2023-08-22
- Publication Date
- 2026-07-17
AI Technical Summary
In existing technologies, deep learning-based cross-modal visual-semantic embedding methods have failed to fully exploit cross-modal semantic knowledge between images and text, resulting in low efficiency in cross-modal representation learning and retrieval of multimodal data.
By collecting local fine-grained features and global coarse-grained features from images and text, a multimodal graph attention network is used to perform implicit fine-grained semantic association reasoning within and between modalities, generating an efficient unified hash representation across modalities, and performing hash mapping for cross-modal retrieval.
It improves the accuracy and efficiency of cross-modal retrieval of image and text data, and can be better applied to real-world scenarios without label supervision, learning richer and more comprehensive cross-modal semantic associations.
Smart Images

Figure CN117349454B_ABST