多模态知识增强的跨模态表示学习与检索方法及相关设备

By collecting local fine-grained and global coarse-grained features of images and text, and using a multimodal graph attention network for cross-modal retrieval, the problem of low efficiency in cross-modal representation learning and retrieval in existing technologies is solved, and more efficient cross-modal semantic association and retrieval are achieved.

CN117349454BActive Publication Date: 2026-07-17BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2023-08-22
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In existing technologies, deep learning-based cross-modal visual-semantic embedding methods have failed to fully exploit cross-modal semantic knowledge between images and text, resulting in low efficiency in cross-modal representation learning and retrieval of multimodal data.

Method used

By collecting local fine-grained features and global coarse-grained features from images and text, a multimodal graph attention network is used to perform implicit fine-grained semantic association reasoning within and between modalities, generating an efficient unified hash representation across modalities, and performing hash mapping for cross-modal retrieval.

Benefits of technology

It improves the accuracy and efficiency of cross-modal retrieval of image and text data, and can be better applied to real-world scenarios without label supervision, learning richer and more comprehensive cross-modal semantic associations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117349454B_ABST
    Figure CN117349454B_ABST
Patent Text Reader

Abstract

本公开提供一种多模态知识增强的跨模态表示学习与检索方法及相关设备,包括:获取数据信息集,其中所述数据信息集包括图像数据以及文本数据;采集所述数据信息集的局部特征,并基于所述局部特征确定所述数据信息集的细粒度特征;采集所述数据信息集的全局特征,并基于所述全局特征确定所述数据信息集的粗粒度特征;基于所述细粒度特征以及所述粗粒度特征,对所述数据信息集进行跨模态检索。本公开中,通过构建的多模态知识图谱,并基于多模态图注意力网络对模态内和模态间的隐含细粒度语义关联进行了推理,之后对推理得到的结果进行哈希映射并生成跨模态高效统一哈希表示,最终基于所生成的哈希表示进行跨模态检索。
Need to check novelty before this filing date? Find Prior Art