一种基于空间-语义对齐的跨模态文档信息抽取方法
By aligning spatial features and semantic information in two directions and extracting hierarchical cross-modal information, the problem of insufficient accuracy and robustness in information extraction from complex documents is solved, achieving efficient information extraction in low-resource scenarios and improving the generalization ability of information extraction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING INST OF COMP TECH & APPL
- Filing Date
- 2025-10-29
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies suffer from limitations in single-modal features, insufficient interaction between spatial features and semantic information, and constraints in multi-level information modeling when processing complex or structured documents. This results in insufficient accuracy and robustness of information extraction, especially in low-resource scenarios where generalization ability is inadequate.
By aligning spatial features and semantic information in two directions and extracting hierarchical cross-modal information, a two-way interactive attention mechanism is used to achieve dynamic adjustment and collaborative modeling of spatial layout and text semantics, and multi-level information extraction is performed at the global, regional and entity levels.
It improves the accuracy and robustness of information extraction from complex documents, enhances the generalization ability in low-resource environments, and provides high-quality structured information support.
Smart Images

Figure CN121210686B_ABST