A document image synthesis and data set automatic generation method and system based on Doctags language
By using a reverse rendering method based on the Doctags language, logically consistent document images and annotation information are generated, solving the problems of high cost and noise in existing technologies. This enables the efficient and automated construction of training datasets for document understanding models, improving the model's recognition and understanding capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 南京通达海软件有限公司
- Filing Date
- 2026-05-06
- Publication Date
- 2026-06-02
AI Technical Summary
Existing document dataset construction techniques rely on manual annotation or automated tool generation, resulting in high costs, low efficiency, and noise, making it difficult to generate high-quality training data. In particular, when dealing with documents with complex layout structures, it is impossible to achieve absolute consistency between images and annotation information, which limits the development of document intelligent processing technology.
We employ a reverse rendering method based on the Doctags language to generate synthetic document images by parsing the Doctags information of the documents. We utilize techniques such as adaptive font size calculation, vector drawing, and water-filling algorithms to ensure the logical consistency and spatial correspondence between the images and the annotation information at the pixel level, thereby constructing a high-precision training dataset.
It achieves low-cost, automated generation of training data for large-scale document understanding models. The generated dataset can support high-quality training of multimodal large models, solves the problem of handling complex page elements, and improves the model's recognition and understanding capabilities.
Smart Images

Figure CN122135386A_ABST