A multi-task OCR certificate text dataset collaborative generation and division method

By generating OCR document text datasets from a single information source through a unified processing flow, the problems of fragmented and inconsistent data generation processes are solved, improving data preparation efficiency and model integration effects, especially demonstrating excellent performance in multi-line text recognition.

CN122290143APending Publication Date: 2026-06-26SICHUAN KAFAN NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SICHUAN KAFAN NETWORK TECH CO LTD
Filing Date
2026-05-27
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing OCR technologies suffer from several problems in recognizing structured documents, including fragmented data generation processes, inconsistent dataset structures, a disconnect between annotation strategies and model requirements, and difficulty in supporting semantically consistent annotation of multi-line text. These issues result in inefficient data preparation and poor model integration.

Method used

A collaborative generation and segmentation method for multi-task OCR document text datasets is adopted. Through a unified processing flow, an adapted text detection, text recognition, and semantic entity recognition dataset is generated from a single information source. Parallel derivation and parallel segmentation ensure the consistency and collaboration of the datasets and support semantic consistency annotation of multi-line text.

Benefits of technology

It improves data preparation efficiency, ensures the fairness and accuracy of multi-model evaluation, enhances the performance of semantic entity recognition tasks, lowers the technical threshold, and generates higher-quality datasets that are suitable for real-world application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122290143A_ABST
    Figure CN122290143A_ABST
Patent Text Reader

Abstract

This invention discloses a collaborative generation and partitioning method for multi-task OCR document text datasets, belonging to the fields of computer vision and deep learning technology. Addressing the problems of fragmented processes, low efficiency, and poor data consistency in existing technologies that prepare separate datasets for text detection, text recognition, and semantic entity recognition tasks, this invention includes: obtaining an original annotation file containing text box positions, text content, and semantic category labels; based on this file, generating three structurally consistent datasets for text detection, text recognition, and semantic entity recognition in parallel through a unified processing flow; and finally, simultaneously partitioning the three datasets according to the same strategy and proportions. This invention achieves "one-time annotation, multi-directional generation," significantly improving data preparation efficiency and quality, ensuring the consistency of multi-task training data, and providing a reliable data foundation for training high-precision end-to-end structured information extraction models.
Need to check novelty before this filing date? Find Prior Art