一种采用多模态融合技术实现的媒资检索方法及系统
By utilizing knowledge graphs and cross-attention models for cross-modal semantic completion and deep fusion in media asset retrieval, the problems of fragile cross-modal alignment logic and information damage in media asset retrieval technology are solved, achieving high-precision deep semantic retrieval and improving the stability and accuracy of media asset retrieval.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIANGSU BROADCASTING CORPORATION
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-17
AI Technical Summary
Existing media asset retrieval technologies suffer from weak cross-modal alignment logic, lack of completion capabilities when information is damaged, lack of knowledge guidance in the fusion process, and insufficient generalization ability of retrieval intent. As a result, when processing complex or noisy media asset data, the alignment accuracy between modalities decreases and semantic shift occurs severely, making it difficult to meet the needs of deep semantic retrieval.
By acquiring the original features of video, audio, and text from high-quality media asset datasets, we extract features using a visual Transformer model, an acoustic convolutional neural network, and a pre-trained language model. We then identify target entities and map them to a knowledge graph. Cross-modal semantic completion and weight reshaping are performed, and a cross-attention model is used for deep fusion to generate a fused feature vector. Finally, we construct a semantic index for retrieval.
It achieves cross-modal semantic alignment, solves the semantic offset and false association problems in traditional multimodal fusion, improves the discrimination accuracy and retrieval efficiency under complex retrieval tasks, and ensures the stability and accuracy of retrieval under extreme conditions.
Smart Images

Figure CN122045441B_ABST