Certificate information extraction and quality evaluation method, device, equipment and medium

By employing a multimodal large-scale model parallel inference and active learning strategy, a document information extraction and quality assessment model is constructed. This solves the problem of poor user experience caused by single judgment in traditional document image processing, and achieves efficient and accurate document information extraction and quality assessment.

CN122369012APending Publication Date: 2026-07-10TONGDUN NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN Β· China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-03
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In traditional document image processing, the serial processing method based on optical character recognition results in a single pass or rejection judgment for the entire image, which cannot effectively distinguish images with clearly identifiable key fields, leading to a poor user experience.

Method used

A multimodal large-scale model parallel inference and active learning strategy is adopted. By synthesizing training data and expanding the training set with high-confidence pseudo-labels, an instruction fine-tuning dataset is constructed to train the document information extraction and quality assessment model, so as to realize character recognition and quality assessment simultaneously.

Benefits of technology

It alleviated the data annotation bottleneck, improved the system's reliability and processing efficiency, effectively suppressed recognition illusions under low-quality images, and improved the accuracy of character recognition and the fine granularity of image quality assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122369012A_ABST
    Figure CN122369012A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and medium for document information extraction and quality assessment. The method includes: extracting synthetic training data from each character region in annotated sample document images; using the synthetic training data to perform preliminary training on multiple multimodal large models with different backbone networks to obtain multiple candidate models; inputting unannotated real document images into the multiple candidate models for parallel inference to obtain the character recognition results and quality assessment results output by each candidate model; adding characters with the same recognition results and quality assessment results output by the multiple candidate models, along with their corresponding character recognition results and quality assessment results, as high-confidence pseudo-labels to an expanded training set; marking characters with different recognition results and quality assessment results as difficult sample data for manual correction; uniformly formatting the synthetic training data, expanded training set, and manually corrected difficult sample data to construct an instruction fine-tuning dataset; and using the instruction fine-tuning dataset to train the multimodal large model to obtain the document information extraction and quality assessment model.
Need to check novelty before this filing date? Find Prior Art