A multi-modal large model certificate identification method based on adaptive region of interest enhancement
The multimodal large-model document recognition method with adaptive region of interest enhancement solves the problem of insufficient information carrying capacity in small target areas, realizes deep fusion of local details and global context, improves recognition accuracy and optimizes inference latency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG ZHIGANGTONG TECH CO LTD
- Filing Date
- 2026-05-08
- Publication Date
- 2026-07-10
AI Technical Summary
Existing document recognition methods based on multimodal large language models have insufficient information carrying capacity when processing small target regions. The super-resolution enhancement strategy lacks closed-loop linkage, cannot achieve deep fusion of local details and global context, and is difficult to achieve an adaptive balance between recognition accuracy and inference latency.
The information carrying capacity in the visual token space is quantified by the first round of inference. An adaptive region of interest enhancement method is adopted, and the semantic information output by the first round of VLM inference is used for selective super-resolution enhancement. A multi-granularity fusion strategy is used to achieve deep fusion of local high-density tokens and global tokens.
It improves the information carrying capacity of small target areas, enhances recognition accuracy, optimizes inference latency without increasing computing resources, and achieves deep fusion of local details and global context.
Smart Images

Figure CN122365092A_ABST