一种轻量化多模态融合的文档信息结构化提取方法及系统

By employing a lightweight multimodal fusion method, utilizing MobileNetV3, a feature pyramid network, and a multimodal encoder, the problem of high computational resources and low accuracy in document structured information extraction is solved, achieving efficient and low-cost structured information extraction.

CN121640483BActive Publication Date: 2026-07-17WUXI YIMAIDE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUXI YIMAIDE TECH CO LTD
Filing Date
2025-12-09
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing technologies for extracting structured information from documents suffer from problems such as unstructured output, easy failure of preset rules, high computational resource requirements, low computational accuracy, and high implementation difficulty, making it difficult to meet the needs of small and medium-sized enterprises and resource-constrained scenarios.

Method used

A lightweight multimodal fusion method is adopted, which extracts text features through MobileNetV3, detects text regions by combining a feature pyramid network and a dual-branch Tokenized MLP module, detects table regions by using SLANet_plus, and extracts structured document information through a multimodal encoder, thereby reducing the consumption of computing resources.

Benefits of technology

It enables structured information output, reduces the complexity of processing procedures and the consumption of computing resources, improves recognition accuracy and generalization ability, and lowers the deployment threshold.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121640483B_ABST
    Figure CN121640483B_ABST
Patent Text Reader

Abstract

本发明提供了一种轻量化多模态融合的文档信息结构化提取方法及系统,涉及数据处理技术领域,方法包括:获取文档图像;对文档图像进行预处理,得到优化图像;通过MobileNetV3,提取优化图像的文本特征;对多种文本特征进行多尺度特征融合,得到融合特征图;通过双分支Tokenized MLP模块,提取融合特征图的位置特征与空间位置间的上下文关系特征;根据融合特征图的位置特征与空间位置间的上下文关系特征,检测优化图像的文本区域;通过LPRNet,对优化图像的文本区域进行文本识别;通过SLANet_plus,检测优化图像的表格区域;通过多模态编码器,结合文本区域和表格区域,提取结构化文档信息。
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Multi-mode irony detection method and device, computer equipment and storage medium

    CN117891940A

  • Scene text recognition method and device, equipment and medium

    CN118314564A