一种基于空间-语义对齐的跨模态文档信息抽取方法

By aligning spatial features and semantic information in two directions and extracting hierarchical cross-modal information, the problem of insufficient accuracy and robustness in information extraction from complex documents is solved, achieving efficient information extraction in low-resource scenarios and improving the generalization ability of information extraction.

CN121210686BActive Publication Date: 2026-07-17BEIJING INST OF COMP TECH & APPL

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING INST OF COMP TECH & APPL
Filing Date
2025-10-29
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies suffer from limitations in single-modal features, insufficient interaction between spatial features and semantic information, and constraints in multi-level information modeling when processing complex or structured documents. This results in insufficient accuracy and robustness of information extraction, especially in low-resource scenarios where generalization ability is inadequate.

Method used

By aligning spatial features and semantic information in two directions and extracting hierarchical cross-modal information, a two-way interactive attention mechanism is used to achieve dynamic adjustment and collaborative modeling of spatial layout and text semantics, and multi-level information extraction is performed at the global, regional and entity levels.

Benefits of technology

It improves the accuracy and robustness of information extraction from complex documents, enhances the generalization ability in low-resource environments, and provides high-quality structured information support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121210686B_ABST
    Figure CN121210686B_ABST
Patent Text Reader

Abstract

本发明涉及一种基于空间‑语义对齐的跨模态文档信息抽取方法,属于人工智能、计算机视觉、自然语言处理领域。本发明设计了空间特征与语义信息双向对齐模型,通过构建空间特征与语义特征的双向对齐模型,使文档的布局信息能够动态调节文本语义特征的关注分布,同时语义信息也反向优化空间特征,实现空间布局与语义信息的协同建模,从而提升复杂文档信息抽取的准确性和鲁棒性。本发明设计了层级化跨模态信息抽取模型,通过层级化跨模态信息抽取模型,在全局层面识别文档整体结构,在区域层面聚焦局部关键内容,并在实体层面对细粒度文本与视觉元素进行精细建模,实现跨模态实体及其语义关系的精准识别,增强信息抽取的泛化能力和适用性。
Need to check novelty before this filing date? Find Prior Art