A multi-modal large model certificate identification method based on adaptive region of interest enhancement

The multimodal large-model document recognition method with adaptive region of interest enhancement solves the problem of insufficient information carrying capacity in small target areas, realizes deep fusion of local details and global context, improves recognition accuracy and optimizes inference latency.

CN122365092APending Publication Date: 2026-07-10ZHEJIANG ZHIGANGTONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG ZHIGANGTONG TECH CO LTD
Filing Date
2026-05-08
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing document recognition methods based on multimodal large language models have insufficient information carrying capacity when processing small target regions. The super-resolution enhancement strategy lacks closed-loop linkage, cannot achieve deep fusion of local details and global context, and is difficult to achieve an adaptive balance between recognition accuracy and inference latency.

Method used

The information carrying capacity in the visual token space is quantified by the first round of inference. An adaptive region of interest enhancement method is adopted, and the semantic information output by the first round of VLM inference is used for selective super-resolution enhancement. A multi-granularity fusion strategy is used to achieve deep fusion of local high-density tokens and global tokens.

Benefits of technology

It improves the information carrying capacity of small target areas, enhances recognition accuracy, optimizes inference latency without increasing computing resources, and achieves deep fusion of local details and global context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122365092A_ABST
    Figure CN122365092A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on adaptive region of interest enhancement multi-modal big model certificate identification method, this method includes: input certificate image, carry out VLM first round inference, output the boundary box of each field, recognized text and confidence; The patch coverage of each field is calculated with VLM semantic confidence, judge whether there is the field to be enhanced that patch coverage is insufficient or confidence is too low;If there is, then according to patch coverage gap calculates adaptive super-resolution ratio, crop and expand ROI region, mosaic picture is spliced after batch, and once again is sent into super-resolution model reconstruction, restores high-frequency stroke details;Original image and super-resolution ROI image are regarded as independent input, and are encoded into global token and local enhancement token respectively, and the final recognition result is output after multi-granularity fusion by language decoder. Through closed-loop cascade architecture and visual token information density redistribution, the small target character recognition accuracy is improved without modifying model parameters, and the on-demand allocation of computing resources is realized.
Need to check novelty before this filing date? Find Prior Art