A document image synthesis and data set automatic generation method and system based on Doctags language

By using a reverse rendering method based on the Doctags language, logically consistent document images and annotation information are generated, solving the problems of high cost and noise in existing technologies. This enables the efficient and automated construction of training datasets for document understanding models, improving the model's recognition and understanding capabilities.

CN122135386AActive Publication Date: 2026-06-02南京通达海软件有限公司

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
南京通达海软件有限公司
Filing Date
2026-05-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing document dataset construction techniques rely on manual annotation or automated tool generation, resulting in high costs, low efficiency, and noise, making it difficult to generate high-quality training data. In particular, when dealing with documents with complex layout structures, it is impossible to achieve absolute consistency between images and annotation information, which limits the development of document intelligent processing technology.

Method used

We employ a reverse rendering method based on the Doctags language to generate synthetic document images by parsing the Doctags information of the documents. We utilize techniques such as adaptive font size calculation, vector drawing, and water-filling algorithms to ensure the logical consistency and spatial correspondence between the images and the annotation information at the pixel level, thereby constructing a high-precision training dataset.

Benefits of technology

It achieves low-cost, automated generation of training data for large-scale document understanding models. The generated dataset can support high-quality training of multimodal large models, solves the problem of handling complex page elements, and improves the model's recognition and understanding capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135386A_ABST
    Figure CN122135386A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for document image synthesis and automatic dataset generation based on the Doctags language. The invention first acquires the original document and parses its Doctags information. Then, based on the Doctags information, it generates a synthesized document image through reverse rendering. Reverse rendering involves taking the Doctags information as input and drawing visual elements aligned with the Doctags information on a canvas according to semantic categories and physical bounding box coordinates, thus obtaining the synthesized document image. Finally, the synthesized document image and its corresponding Doctags information are associated and stored to form image-tag pairs to construct the dataset. This invention employs reverse thinking, ensuring that the generated images maintain absolute logical consistency and spatial correspondence with the annotation information at the pixel level, eliminating mismatch errors between images and annotations, and enabling the automated construction of large-scale, high-precision document understanding model training data.
Need to check novelty before this filing date? Find Prior Art