Unified solution for non-text object analysis and understanding in rich vision document
Through the UNTOA-VRD model combined with the fine-tuned large language model and unified multi-task algorithm, a unified visual-rich document was uniformly analyzed, which solved the problem of incomplete recognition and interpretation of non-text objects in the existing technology, achieved efficient multi-task analysis and understanding, and improved the accuracy and efficiency of document processing.
Patent Information
- Application Number
- CN202510051564.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-13
AI Technical Summary
When the prior art deals with non-text objects in visually rich documents, there are problems of insufficient identification and interpretation, and multi-stage modeling strategies lead to complex model maintenance and updates, reduced efficiency, and cannot meet the needs of large-scale document processing.
Provide a unified solution to analyze visually rich documents through UNTOA-VRD model, combining fine-tuned large language model module and unified multi-task algorithm module to achieve unified analysis and understanding of multiple tasks.
This method can efficiently analyze and understand non-text objects in visually rich documents, simplify the modeling process, improve the overall accuracy of document understanding, and enhance synergy between multi-task analysis.
Smart Images

Figure CN119992575A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and natural language processing, and in particular to a unified solution for analyzing and understanding non-text objects in visually rich documents. Background Art
[0002] The widespread use of the Internet and mobile devices has spawned a large number of digital documents, such as academic papers, business reports, and technical manuals. These rich visual documents not only contain rich text information, but also integrate non-text elements such as formulas, tables, and charts.
[0003] In the field of rich visual document analysis, there are significant challenges in efficiently and accurately processing various non-text objects. Although natural language processing and computer vision technologies have made significant progress in text extraction and semantic analysis, such as the application of optical character recognition and deep learning models, there are still many technical difficulties in the field of multimodal integration. Existing methods mainly rely on automated systems to extract semantic information, but there are two main problems: first, the recognition and interpretation of non-text elements such as formulas, tables, and charts are still imperfect; second, existing models are usually limited to single-task analysis when processing non-text objects in rich visual documents. This limitation requires a multi-stage modeling strategy, which makes model maintenance and updating complicated, reduces efficiency, and cannot meet the needs of large-scale document processing. In addition, the diversity of document types and formats limits the adaptability and generalization ability of single-task models, and increases the cost of system deployment and scalability.
[0004] In this context, the industry urgently needs a unified multi-task analysis and understanding method that can handle multiple tasks simultaneously in a single model. Summary of the invention
[0005] In view of this, an object of the present invention is to provide a unified solution for analyzing and understanding non-text objects in visually rich documents.
[0006] The technical solution adopted by the present invention to solve the technical problem is to provide a unified solution for analyzing and understanding non-text objects in rich visual documents, including the steps of:
[0007] S1. Input the rich visual document to the UNTOA-VRD model. The UNTOA-VRD model performs layout analysis P on the rich visual document. The user inputs a command to form a user command C. The UNTOA-VRD model forms a recognition task T according to the user command C.
[0008] S2, UNTOA-VRD model analysis identifies whether task T only includes layout analysis P, if so, outputs analysis result R, R = P;
[0009] If not, proceed to step S3;
[0010] S3. The UNTOA-VRD model analyzes the recognition task T to obtain several recognition tasks t. The UNTOA-VRD model locates the area of all recognition tasks t in the rich visual document and performs the corresponding recognition task t in the area. The UNTOA-VRD model analyzes whether the recognition task T contains the layout analysis P. If so, the analysis result R is output, R = P∪{rt|t∈T};
[0011] If not, output the analysis result R, R = {rt|t∈T}.
[0012] As a further improvement of the present invention, the UNTOA-VRD model includes a fine-tuning large language model module and a unified multi-task algorithm module.
[0013] As a further improvement of the present invention, step S0 is also included before step S1: multiple labeled data sets are shuffled and integrated to form a large data set to input into the large language model, and the large language model is fine-tuned with all parameters to form a fine-tuned large language model module.
[0014] As a further improvement of the present invention, the UNTOA-VRD model includes the InternViT-300M-448pX visual encoder.
[0015] As a further improvement of the present invention, the recognition task t in step S3 includes formula recognition, table recognition and chart recognition.
[0016] As a further improvement of the present invention, step S0 is also included before step S1: fine-tuning the large language model module using the PubLayNet dataset to perform positioning tasks and layout analysis tasks.
[0017] As a further improvement of the present invention, step S0 is also included before step S1: fine-tuning the large language model module using the PubTabNet-HTML dataset to perform table recognition and parsing tasks.
[0018] The beneficial effects of the present invention are at least as follows: the UNTOA-VRD model utilizes a fine-tuned lightweight large language model to analyze and understand non-text objects in rich visual documents, and at the same time integrates the designed unified multi-task algorithm module so that the model can perform unified analysis on multiple tasks, which not only simplifies the modeling process, but also improves the overall accuracy of rich visual document understanding while enhancing the synergy between multi-task analyses. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic diagram of the steps of the present invention. DETAILED DESCRIPTION
[0020] The technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings.
[0021] Reference Figure 1 The present invention proposes a unified solution for analyzing and understanding non-text objects in rich visual documents, comprising the steps of:
[0022] S1. Input the rich visual document to the UNTOA-VRD model. The UNTOA-VRD model performs layout analysis P on the rich visual document. The user inputs a command to form a user command C. The UNTOA-VRD model forms a recognition task T according to the user command C.
[0023] S2, UNTOA-VRD model analysis identifies whether task T only includes layout analysis P, if so, outputs analysis result R, R = P;
[0024] If not, proceed to step S3;
[0025] S3. The UNTOA-VRD model analyzes the recognition task T to obtain several recognition tasks t. The UNTOA-VRD model locates the area of all recognition tasks t in the rich visual document and performs the corresponding recognition task t in the area. The UNTOA-VRD model analyzes whether the recognition task T contains the layout analysis P. If so, the analysis result R is output, R = P∪{rt|t∈T};
[0026] If not, output the analysis result R, R = {rt|t∈T}.
[0027] The user input instructions are analyzed through natural language processing technology to form user instructions C, and the user instructions C are formed into recognition tasks T, and it is determined whether the analysis and recognition task T only includes layout analysis P, etc. Since natural language processing technology is relatively mature, as mentioned in the invention patent with authorization announcement number CN118193765B, this application will not go into details.
[0028] As a further improvement of the present invention, the UNTOA-VRD model includes a fine-tuning large language model module and a unified multi-task algorithm module.
[0029] Specifically, for fine-tuning the large language model module, the present invention selected PubLayNet for layout analysis, and annotated 360,000 document images. UniMER is used for formula recognition, containing more than 1 million LaTeX-image pairs. PubTabNet-HTML is used for table recognition, providing table images and HTML tag annotations. MMC is used for chart recognition, providing a benchmark for chart reasoning capabilities. First, training tests of a single task are performed to evaluate the feasibility of unified multi-task analysis and understanding. After testing, the single task has achieved different degrees of indicator advantages. The present invention randomly shuffles multiple labeled data sets into a single data set and feeds it to the model for full parameter fine-tuning for unified multi-task training.
[0030] In combination with the fine-tuned large language model module, the present invention designs a unified multi-task algorithm module, and the workflow is as follows: First, input a rich visual document image. Regardless of the user's specific instructions, the model always performs layout analysis first to ensure a comprehensive understanding of the document structure and provide a basis for the positioning and identification of subsequent tasks. Subsequently, according to the user's instructions, the model adopts an appropriate response strategy. If the user's instructions only include layout analysis, the model directly outputs the layout analysis result R=P. If the instructions include a combination of layout analysis and other tasks, after completing the layout analysis, the model locates the area At related to other tasks based on the analysis results and user instructions, and performs the recognition task of task t, and finally outputs R=P∪{rt∣t∈T}. In the case where the user's instructions do not include layout analysis, the model still first performs layout analysis to assist in positioning, and then performs the specified task and outputs R={rt∣t∈T}.
[0031] As a further improvement of the present invention, step S0 is also included before step S1: multiple labeled data sets are shuffled and integrated to form a large data set to input into the large language model, and the large language model is fine-tuned with all parameters to form a fine-tuned large language model module.
[0032] Specialized annotation labels are introduced for various non-text objects (such as formulas, tables, and charts) as well as layout information. These labels are embedded into the training data as explicit information, thus providing the model with prior indications of content type and structural information.
[0033] The method of full parameter fine-tuning is adopted to train all parameters of the model end-to-end to fully utilize its potential in multimodal data processing. Through this method, the model can fully adjust the internal weights to better adapt to specific document types and domain knowledge, and realize accurate recognition and understanding of non-text elements such as formulas, tables and charts. By fine-tuning all parameters of multiple labeled data sets, the present invention effectively enhances the ability of lightweight large language models in multimodal non-text information processing. The data sets used cover multiple tasks such as layout analysis, formula recognition, table recognition and chart understanding. The model can effectively migrate between multiple tasks and maintain high performance, greatly improving its generalization ability in a multi-task environment. The unified multi-task algorithm module can flexibly respond to different user instructions. Whether it is a single task layout analysis or a combination of layout analysis and other tasks, UNTOA-VRD can be completed efficiently and has strong scalability.
[0034] As a further improvement of the present invention, the UNTOA-VRD model includes the InternViT-300M-448pX visual encoder.
[0035] As a further improvement of the present invention, the recognition task t in step S3 includes formula recognition, table recognition and chart recognition.
[0036] As a further improvement of the present invention, step S0 is also included before step S1: fine-tuning the large language model module using the PubLayNet dataset to perform positioning tasks and layout analysis tasks.
[0037] As a further improvement of the present invention, step S0 is also included before step S1: fine-tuning the large language model module using the PubTabNet-HTML dataset to perform table recognition and parsing tasks.
Claims
1. A unified solution for analyzing and understanding non-text objects in visually rich documents, characterized in that: Includes steps: S1. Input the rich visual document to the UNTOA-VRD model. The UNTOA-VRD model performs layout analysis P on the rich visual document. The user inputs a command to form a user command C. The UNTOA-VRD model forms a recognition task T according to the user command C. S2, UNTOA-VRD model analysis identifies whether task T only includes layout analysis P, if so, outputs analysis result R, R = P; If not, proceed to step S3; S3. The UNTOA-VRD model analyzes the recognition task T to obtain several recognition tasks t. The UNTOA-VRD model locates the area of all recognition tasks t in the rich visual document and performs the corresponding recognition task t in the area. The UNTOA-VRD model analyzes whether the recognition task T contains the layout analysis P. If so, the analysis result R is output, R = P∪{rt|t∈T}; If not, output the analysis result R, R = {rt|t∈T}.
2. The unified solution for analyzing and understanding non-text objects in rich visual documents according to claim 1, characterized in that: The UNTOA-VRD model includes a fine-tuned large language model module and a unified multi-task algorithm module.
3. The unified solution for analyzing and understanding non-text objects in rich visual documents according to claim 2, characterized in that: Before step S1, step S0 is also included: shuffling and integrating multiple labeled data sets to form a large data set to input into the large language model, and performing full-parameter fine-tuning training on the large language model to form a fine-tuned large language model module.
4. The unified solution for analyzing and understanding non-text objects in rich visual documents according to claim 1, characterized in that: The UNTOA-VRD model includes the InternViT-300M-448pX vision encoder.
5. The unified solution for analyzing and understanding non-text objects in rich visual documents according to claim 1, characterized in that: The recognition task t in step S3 includes formula recognition, table recognition and chart recognition.
6. The unified solution for analyzing and understanding non-text objects in rich visual documents according to claim 2, characterized in that: Before step S1, step S0 is also included: fine-tuning the large language model module using the PubLayNet dataset to perform positioning tasks and layout analysis tasks.
7. The unified solution for analyzing and understanding non-text objects in rich visual documents according to claim 2, characterized in that: Before step S1, step S0 is also included: fine-tuning the large language model module using the PubTabNet-HTML dataset to perform table recognition and parsing tasks.
Citation Information
Patent Citations
Method of integrating event reminders and growth memory
CN118193765B