Visual document content extraction and analysis system and method
Through the multi-task unified modeling module and end-to-end decoding framework, multimodal content in complex documents is extracted and understood, and the problems of non-text object recognition and task separation in the existing technology are solved, and efficient and unified document analysis and understanding are achieved.
Patent Information
- Application Number
- CN202510537626.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-27
AI Technical Summary
Existing automatic reading comprehension systems of documents are difficult to effectively identify and understand complex non-text objects, such as formulas, tables, charts, etc., and task separation and model fragmentation lead to parameter redundancy and complex maintenance.
The multi-task unified modeling module is adopted to extract multi-modal features through visual language mask prediction technology and contrast learning method, generate unified multi-modal representations, and achieve end-to-end content extraction and information extraction decoding through a shared decoder and a joint loss function of dynamic weights.
It realizes unified processing of multimodal tasks in complex documents, improves the coordination efficiency between different tasks and the generalization performance of models, and solves the efficiency loss and consistency problems caused by task separation in traditional methods.
Smart Images

Figure CN120071372A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the cross-technical field of computer vision and natural language processing, and in particular to a multimodal content extraction and structured parsing system and method for complex document images. Background Art
[0002] Most of the current document automatic reading and comprehension systems first use optical character recognition (OCR) technology to extract text information, and then perform semantic analysis on the document based on natural language processing methods. However, for rich visual documents, that is, documents containing rich visual elements such as formulas, tables, and charts, there are two problems with existing document analysis and understanding methods: First, it is difficult to recognize and understand complex non-text objects. Most current methods focus on the recognition of visual text, and there are obvious shortcomings in the semantic understanding of non-text objects, such as the difficulty in capturing the relative position relationship between symbols in the formula and their mathematical semantics, the inability to effectively recognize the vertical and horizontal line structure, cross-cell merging logic, and the correlation between tables and context descriptions, as well as the lack of modeling capabilities for the relative size of chart elements and data mapping relationships, resulting in the inability to fully extract the key semantic information of non-text objects; second, task separation and model fragmentation. The existing technology adopts staged modeling for tasks, first performing document analysis, extracting each sub-region, and then further identifying the sub-region. For non-text objects, existing methods usually model and train in stages, first extracting basic visual components and then combining the components for multimodal modeling and prediction, which leads to parameter redundancy, complex maintenance, and the inability to effectively utilize cross-modal correlation information.
[0003] Therefore, faced with large-scale document data, how to extract semantic information from visual content and establish a unified visual document object analysis and understanding mechanism to achieve rich visual document analysis and understanding is a key issue that needs to be solved urgently.
[0004] Therefore, the present invention is just produced based on the above shortcomings. Summary of the invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a visual document content extraction and analysis system.
[0006] The present invention is achieved through the following technical solutions:
[0007] A visual document content extraction and analysis system, characterized in that: it includes a multi-task unified modeling module and a content extraction and information extraction decoding module. The multi-task unified modeling module includes a multi-modal feature extraction sub-module that can extract multi-modal features in the input document image and perform pixel mask reconstruction prediction and context mining on the input document image through visual language mask prediction technology, a multi-modal representation optimization sub-module that generates a unified multi-modal representation by fusing visual, text, and location information through contrastive learning, and a task prompt generation sub-module that generates task prompt vectors through prompt learning and unifies different tasks into a sequence generation parsing framework. The content extraction and information extraction decoding module includes a multi-modal context fusion sub-module that combines visual features, text semantics, and layout information to generate a unified semantic representation, and a unified decoding sub-module that directly corresponds document regions to semantic labels through a shared decoder and balances the region detection and semantic assignment tasks using a dynamic weight joint loss function.
[0008] The visual document content extraction and analysis system as described above, characterized in that: the multi-modal representation optimization sub-module realizes the adaptation of domain-specific tasks by inserting an adapter module.
[0009] The visual document content extraction and analysis system as described above, characterized in that: the visual language mask prediction technology includes performing pixel-level mask processing on text, table, formula, and illustration regions in the document image and reconstructing the structural and content semantic information of the mask region through cross-modal context relationship mining.
[0010] The visual document content extraction and analysis system as described above, characterized in that the specific steps for the shared decoder to achieve end-to-end decoding are: generating the bounding box coordinates of the document region based on the unified semantic representation; assigning a semantic label to each region, and the label includes at least one of title, text, table, formula, and illustration; synchronously optimizing the region detection and semantic assignment tasks through the dynamic weight joint loss function.
[0011] The visual document content extraction and analysis system as described above, characterized in that: the dynamic weight joint loss function includes a region detection loss and a semantic classification loss, and the weights are dynamically adjusted according to the training process to balance the region detection and semantic assignment tasks.
[0012] A visual document content extraction and analysis method, characterized in that it includes the following steps:
[0013] S1. Extract the multi-modal features of the document image through the multi-task unified modeling module and generate task prompt vectors;
[0014] S2. Based on the multi-modal features and the task prompt vector, use the content extraction and information extraction decoding module to synchronously complete document area parsing, semantic label assignment, and structured information output.
[0015] The visual document content extraction and analysis method as described above is characterized in that in the step S1, it includes:
[0016] S11. Extract multi-modal features in the input document image and perform pixel mask reconstruction prediction and context mining on the input document image through visual language mask prediction technology;
[0017] S12. Insert an adapter module to fuse visual, text, and location information to optimize the multi-modal representation space;
[0018] S13. Generate a task prompt vector through prompt learning to unify the multi-modal recognition task into a sequence generation parsing framework.
[0019] The visual document content extraction and analysis method as described above is characterized in that in the step S2, it includes:
[0020] S21. Use the multi-modal context modeling module to fuse visual features, text semantics, and layout information to generate a unified semantic representation;
[0021] S22. Correlate the document area with the semantic label through a shared decoder, and use a dynamic weight joint loss function to balance region detection and semantic assignment.
[0022] The visual document content extraction and analysis method as described above is characterized in that it further includes step S3: Collect rich visual documents through diverse data sources, generate a multi-modal document test set through a grid candidate layout algorithm and annotate it; S4: Use the multi-modal document test set to verify the multi-task processing ability and robustness of the model.
[0023] The visual document content extraction and analysis method as described above is characterized in that: The grid candidate layout algorithm includes candidate sampling, grid construction, best match search, and iterative layout filling.
[0024] Compared with the prior art, the present invention has the following advantages:
[0025] 1. First, the unified modeling of the present invention realizes the unified processing of multiple tasks, aligns tasks such as table recognition, formula recognition, chart recognition, text recognition, and illustration recognition to the same multi-modal representation framework, and significantly improves the collaborative efficiency between different tasks and the generalization performance of the model by sharing and reusing model parameters. Second, through the content extraction and information extraction synchronous decoding framework, an efficient mapping from document content extraction to structured semantic information allocation is successfully achieved, enabling the present invention to simultaneously complete multi-modal task processing and in-depth parsing of semantic information in complex document scenarios, effectively solving the efficiency loss and consistency problems caused by task separation in traditional methods, and providing a unified, efficient, and robust solution for the parsing of complex documents.
[0026] 2. The present invention also uses a high-quality document test set covering multi-modal content such as text, tables, formulas, illustrations, and emojis to verify the multi-task processing ability and robustness of the model, ensuring that the visual document content extraction and analysis system can not only perform well on the training data but also adapt to actual complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a schematic diagram of the visual document content extraction and analysis system of the present invention;
[0028] Figure 2 is a schematic diagram of the modeling method of the multi-task unified modeling module of the present invention;
[0029] Figure 3 is a schematic diagram of the end-to-end content extraction and information extraction decoding method of the present invention;
[0030] Figure 4 is a schematic diagram of the grid candidate best matching algorithm of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0031] The present invention will be further described below with reference to the accompanying drawings:
[0032] As Figure 1 shown, a visual document content extraction and analysis system includes a multi-task unified modeling module and a content extraction and information extraction decoding module.
[0033] To achieve efficient sharing and parsing of multiple tasks, the multi-task unified modeling module includes a multi-modal feature extraction sub-module, which can extract multi-modal features in the input document image, such as text paragraphs, tables, formulas, illustrations, emojis, etc., and perform pixel mask reconstruction prediction on the input document image through visual language mask prediction technology, and through context mining, mine visual context information to enhance the robustness and semantic representation ability of feature extraction; a multi-modal representation optimization sub-module, which uses contrastive learning to fuse visual, text, and position information to generate a unified multi-modal representation and optimize the multi-modal representation space; a task prompt generation sub-module, which generates task prompt vectors through prompt learning and unifies different tasks including table recognition, formula recognition, chart recognition, text recognition, and illustration recognition into a sequence generation parsing framework, solving the complexity problem caused by task separation in traditional methods.
[0034] To achieve end-to-end content extraction and information extraction decoding, the content extraction and information extraction decoding module includes a multi-modal context fusion sub-module, which uses the multi-modal context modeling module to combine visual features, text semantics, and layout information to generate a unified semantic representation, and on the basis of feature extraction, completes the accurate parsing of the content of the document area, effectively solving the parsing problems brought by complex layouts and cross-modal associations; a unified decoding sub-module, which directly corresponds the document area with semantic labels (such as "title", "text", "table", etc.) through a shared decoder and uses a dynamic weight joint loss function to balance the region detection and semantic assignment tasks, thereby reducing error accumulation and improving parsing efficiency.
[0035] For different task requirements, the multi-modal representation optimization sub-module realizes the adaptation of domain-specific tasks by inserting an adapter module. For different task requirements, a small adapter module is inserted between each Transformer layer in the pre-trained model to deeply fuse visual, text, and position information and optimize the multi-modal representation space. In actual tasks, this method can effectively retain general knowledge while adapting to domain-specific tasks, significantly improving robustness and adaptability.
[0036] The visual language mask prediction technology includes performing pixel-level mask processing on the text, table, formula, and illustration areas in the document image and reconstructing the structural and content semantic information of the mask area through cross-modal context relationship mining.
[0037] The task prompt generation sub-module uses a learnable prompt vector to unify the parsing processes of different tasks into a sequence generation method based on the Transformer architecture.
[0038] The specific steps for the shared decoder to achieve end-to-end decoding are as follows: generating the bounding box coordinates of the document regions based on the unified semantic representation; assigning semantic labels to each region, where the labels include at least one of title, text, table, formula, and illustration, and using a Softmax classifier to output the class probabilities; synchronously optimizing the region detection and semantic assignment tasks through a dynamic weight joint loss function.
[0039] The dynamic weight joint loss function includes a region detection loss and a semantic classification loss, and the weights are dynamically adjusted according to the training process to balance the region detection and semantic assignment tasks, and the weights are automatically adjusted according to the task uncertainty.
[0040] The present invention also discloses a method for visual document content extraction and analysis, including the following steps:
[0041] S1. Extracting the multi-modal features of the document image through a multi-task unified modeling module and generating a task prompt vector;
[0042] S2. Based on the multi-modal features and the task prompt vector, using the content extraction and information extraction decoding module to synchronously complete document region parsing, semantic label assignment, and structured information output.
[0043] Step S1 includes:
[0044] S11. Extracting the multi-modal features in the input document image and performing pixel mask reconstruction prediction and context mining on the input document image through visual language mask prediction technology;
[0045] S12. Inserting an adapter module to fuse visual, text, and location information and optimizing the multi-modal representation space;
[0046] S13. Generating a task prompt vector through prompt learning and unifying the multi-modal recognition task into a sequence generation parsing framework.
[0047] Step S2 includes:
[0048] S21. Using a multi-modal context modeling module to fuse visual features, text semantics, and layout information to generate a unified semantic representation;
[0049] S22. Corresponding the document regions with the semantic labels through a shared decoder and using a dynamic weight joint loss function to balance region detection and semantic assignment.
[0050] The method for visual document content extraction and analysis further includes step S3: collecting rich visual documents through diverse data sources, generating a multi-modal document test set through a grid candidate layout algorithm and annotating it; S4: using the multi-modal document test set to verify the multi-task processing ability and robustness of the model. Step S3 is carried out before step S1 or simultaneously with step S1, and step S4 is carried out after step S2. Step S3 is the data preparation stage, aiming to use a grid candidate layout algorithm as Figure 4 shown to generate a high-quality, multi-modal document test set, and combined with accurate annotation to ensure the diversity and applicability of the training data. S1 and S2 are the model training stages, aiming to input the document image into a multi-task unified modeling module to complete multi-modal feature extraction and task prompt learning, and on this basis, using a decoding module to synchronously optimize content extraction and semantic label assignment, and adopting a dynamically adjusted joint loss function to ensure multi-task collaboration and performance improvement. Step S4 is the model verification stage, aiming to use the test set to verify the multi-task processing ability and robustness of the model, and further verify the model performance through actual application scenarios such as social media content analysis.
[0051] Specifically, the grid candidate layout algorithm includes candidate sampling, grid construction, best match search, and iterative layout filling. Candidate sampling is to diversely select elements such as text, tables, formulas, and illustrations from a document element library. Grid construction is to construct an adaptive grid structure according to the document type and define layout constraints. Best match search is to search for the optimal grid position for each element based on semantic relevance and layout rules. Iterative layout filling is to iteratively adjust the element positions through an optimization algorithm to maximize the layout score and generate a high-quality test set.
[0052] Generally speaking, the unified modeling method and synchronous decoding framework proposed by the present invention can efficiently process multi-modal tasks in complex documents and achieve synchronous optimization of content extraction and semantic information extraction. It mainly consists of a unified modeling module and a content extraction and information extraction decoding module. The trained model can be directly applied to different document parsing tasks without the need for specialized task customization. The model has strong generalization and transferability, is applicable to various document types, such as scanned documents, social media screenshots, scientific papers, etc., and there is no need to design a specialized output head for each task. The trained model can not only share and reuse parameters in different task scenarios, but also the prediction accuracy is positively correlated with the model scale. With the increase of the model scale, the performance is significantly improved, which provides broad space for the further optimization and development of the model. Although training requires a large dataset, these datasets are easy to obtain, and through the automatic pairing of multi-modal data sources, a large-scale and high-quality document test set can be quickly constructed. In addition, the unified modeling method of the present invention solves the complexity problem brought by multi-task separation. Through task sharing and optimization, the collaborative efficiency between tasks is greatly improved, and it can handle non-text objects in complex documents, such as tables, formulas, illustrations, etc. This method verifies the advantages of the present invention in multi-modal tasks and provides strong technical support for future applications in a wider range of fields, such as document analysis, sentiment analysis, etc. In summary, the present invention not only improves the efficiency of multi-task processing, but also enhances the generalization ability and application scope of the model, providing an innovative solution for complex document parsing and multi-modal task implementation.
Claims
1. A visual document content extraction and analysis system, characterized in that: It includes a multi-task unified modeling module and a content extraction and information extraction decoding module. The multi-task unified modeling module includes a multi-modal feature extraction submodule that can extract multi-modal features in an input document image and perform pixel mask reconstruction prediction and context mining on the input document image through visual language mask prediction technology, a multi-modal representation optimization submodule that generates a unified multi-modal representation by fusing vision, text and position information using contrast learning, and a task prompt generation submodule that generates task prompt vectors through prompt learning and unifies different tasks into a sequence generation and parsing framework. The content extraction and information extraction decoding module includes a multi-modal context fusion submodule that generates a unified semantic representation by combining visual features, text semantics and layout information, and a unified decoding submodule that directly corresponds document regions to semantic labels through a shared decoder and uses a dynamic weight joint loss function to balance region detection and semantic allocation tasks.
2. The visual document content extraction and analysis system according to claim 1, characterized in that: The multimodal representation optimization submodule realizes the adaptation of domain-specific tasks by inserting an adapter module.
3. The visual document content extraction and analysis system according to claim 1, characterized in that: The visual language mask prediction technology includes pixel-level mask processing of text, tables, formulas, and illustration areas in document images and reconstructing the structure and content semantic information of the mask area through cross-modal contextual relationship mining.
4. The visual document content extraction and analysis system according to claim 1, characterized in that: The specific steps of the shared decoder to achieve end-to-end decoding are: generating bounding box coordinates of the document area based on the unified semantic representation; assigning a semantic label to each area, wherein the label includes at least one of a title, text, table, formula, and illustration; and synchronously optimizing the area detection and semantic assignment tasks through the dynamic weight joint loss function.
5. The visual document content extraction and analysis system according to claim 1, characterized in that: The dynamic weight joint loss function includes region detection loss and semantic classification loss, and the weight is dynamically adjusted according to the training process to balance the region detection and semantic assignment tasks.
6. A method for extracting and analyzing visual document content, characterized in that: The following steps are involved: S1, extracting multimodal features of document images through a multi-task unified modeling module and generating task prompt vectors; S2. Based on the multimodal features and the task prompt vector, the content extraction and information extraction decoding modules are used to synchronously complete the document region parsing, semantic tag assignment and structured information output.
7. The method for extracting and analyzing visual document content according to claim 6, characterized in that: The step S1 includes: S11, extracting multimodal features from the input document image and performing pixel mask reconstruction prediction and context mining on the input document image through visual language mask prediction technology; S12, insert the adapter module to fuse visual, textual and position information and optimize the multimodal representation space; S13. Generate task prompt vectors through prompt learning and unify multimodal recognition tasks into a sequence generation and parsing framework.
8. The method for extracting and analyzing visual document content according to claim 6, characterized in that: The step S2 includes: S21. Use the multimodal context modeling module to fuse visual features, text semantics and layout information to generate a unified semantic representation; S22. Use a shared decoder to map document regions to semantic labels, and use a dynamic weighted joint loss function to balance region detection and semantic assignment.
9. The method for extracting and analyzing visual document content according to claim 6, characterized in that: It also includes step S3: collecting rich visual documents through diversified data sources, generating and annotating a multimodal document test set through a grid candidate layout algorithm; S4: using the multimodal document test set to verify the multi-task processing capability and robustness of the model.
10. The method for extracting and analyzing visual document content according to claim 9, characterized in that: The grid candidate layout algorithm includes candidate sampling, grid construction, best match search and iterative layout filling.
Citation Information
Patent Citations
Multi-modal data processing method and device, electronic equipment and storage medium
CN117453880A
Generative visual document zero sample information extraction method based on large model
CN119314189A
Multimodal multitask machine learning system for document intelligence tasks
US20230245485A1
Cited By
Key point information generation method and device, equipment and storage medium
CN121524344A