一种采用多模态融合技术实现的媒资检索方法及系统

By utilizing knowledge graphs and cross-attention models for cross-modal semantic completion and deep fusion in media asset retrieval, the problems of fragile cross-modal alignment logic and information damage in media asset retrieval technology are solved, achieving high-precision deep semantic retrieval and improving the stability and accuracy of media asset retrieval.

CN122045441BActive Publication Date: 2026-07-17JIANGSU BROADCASTING CORPORATION

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGSU BROADCASTING CORPORATION
Filing Date
2026-04-20
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing media asset retrieval technologies suffer from weak cross-modal alignment logic, lack of completion capabilities when information is damaged, lack of knowledge guidance in the fusion process, and insufficient generalization ability of retrieval intent. As a result, when processing complex or noisy media asset data, the alignment accuracy between modalities decreases and semantic shift occurs severely, making it difficult to meet the needs of deep semantic retrieval.

Method used

By acquiring the original features of video, audio, and text from high-quality media asset datasets, we extract features using a visual Transformer model, an acoustic convolutional neural network, and a pre-trained language model. We then identify target entities and map them to a knowledge graph. Cross-modal semantic completion and weight reshaping are performed, and a cross-attention model is used for deep fusion to generate a fused feature vector. Finally, we construct a semantic index for retrieval.

Benefits of technology

It achieves cross-modal semantic alignment, solves the semantic offset and false association problems in traditional multimodal fusion, improves the discrimination accuracy and retrieval efficiency under complex retrieval tasks, and ensures the stability and accuracy of retrieval under extreme conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122045441B_ABST
    Figure CN122045441B_ABST
Patent Text Reader

Abstract

本发明公开了一种采用多模态融合技术实现的媒资检索方法及系统,涉及知识图谱技术领域,包括,获取媒资高质量数据集中的视频、音频及文本原始特征,并从中识别目标实体;将目标实体作为语义锚点映射至知识图谱,获取节点间的逻辑关联关系并对所述原始特征进行跨模态语义补全;将所述逻辑关联关系作为偏置参数注入交叉注意力模型,对对齐后的多模态特征进行权重重塑与深度融合,生成融合特征向量;基于融合特征向量构建语义索引,响应检索指令并输出经语义泛化匹配后的检索结果。本发明方法通过引入知识图谱驱动的语义锚点捕捉,有效解决了媒资数据在多模态融合过程中的逻辑断层与检索精度瓶颈。
Need to check novelty before this filing date? Find Prior Art