一种轻量化多模态融合的文档信息结构化提取方法及系统
By employing a lightweight multimodal fusion method, utilizing MobileNetV3, a feature pyramid network, and a multimodal encoder, the problem of high computational resources and low accuracy in document structured information extraction is solved, achieving efficient and low-cost structured information extraction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUXI YIMAIDE TECH CO LTD
- Filing Date
- 2025-12-09
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies for extracting structured information from documents suffer from problems such as unstructured output, easy failure of preset rules, high computational resource requirements, low computational accuracy, and high implementation difficulty, making it difficult to meet the needs of small and medium-sized enterprises and resource-constrained scenarios.
A lightweight multimodal fusion method is adopted, which extracts text features through MobileNetV3, detects text regions by combining a feature pyramid network and a dual-branch Tokenized MLP module, detects table regions by using SLANet_plus, and extracts structured document information through a multimodal encoder, thereby reducing the consumption of computing resources.
It enables structured information output, reduces the complexity of processing procedures and the consumption of computing resources, improves recognition accuracy and generalization ability, and lowers the deployment threshold.
Smart Images

Figure CN121640483B_ABST
Abstract
Citation Information
Patent Citations
Multi-mode irony detection method and device, computer equipment and storage medium
CN117891940A
Scene text recognition method and device, equipment and medium
CN118314564A