A multimodal
entity linking method based on dual encoders and a
hybrid expert mechanism is proposed. This invention relates to multimodal
entity linking technology at the intersection of
natural language processing and
computer vision. Addressing the problems of low
inference efficiency, insufficient cross-
modal interaction, and shallow
modal fusion in existing methods, this invention proposes a multimodal
entity linking method based on dual encoders and a
hybrid expert mechanism. A dual-
tower architecture is used to independently
encode mentions and entities. Entity embeddings can be pre-computed offline and indexed, and linking is completed during
inference through fast vector retrieval. A
hybrid expert mechanism is introduced to achieve adaptive
feature transformation of samples, and a gating network dynamically selects expert combinations. Bidirectional cross-
modal attention is used to establish fine-grained alignment at the word-image block
granularity. A channel attention mechanism dynamically balances the contributions of textual and visual modalities. The model is jointly optimized by multiple constraints, including load balancing loss. This invention achieves efficient retrieval while maintaining deep
inference capabilities, simplifies inference
time complexity, and is suitable for scenarios such as
knowledge graph construction and intelligent
question answering systems.