Multi-modal document parsing method, device and computer readable storage medium

CN121170830BActive Publication Date: 2026-09-18ZHONGDIAN DATA IND CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511339054.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-18
Publication Date
2026-09-18
Estimated Expiration
2045-09-18

AI Technical Summary

Technical Problem

然而,这种方式在面对包含表格、图形和文本等多模态元素时,可能在解析转换的过程中丢失元素坐标和各元素之间的关联关系,从而导致文档解析的准确性较低的问题

Benefits of technology

[0046] This application provides a multimodal document parsing method. First, it acquires the optical character recognition (OCR) text and document image of the multimodal document. Then, it determines the spatial layout information of each text word in the OCR text in the document image, and generates a document object tree based on each text word and its corresponding spatial layout information. Based on the document object tree, it determines the bounding box of each text word in the document image. The document image is segmented into multiple candidate regions, and the target pixels located in each candidate region within each bounding box are determined. Based on each target pixel, the parsed document is determined.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170830B_ABST
    Figure CN121170830B_ABST
Patent Text Reader

Abstract

This application discloses a multimodal document parsing method, device, and computer-readable storage medium. The application relates to the field of document processing technology. The method includes: acquiring optical character recognition (OCR) text and a document image of a multimodal document; determining the spatial layout information of each text word in the OCR text within the document image; generating a document object tree based on each text word and its corresponding spatial layout information; determining the bounding boxes of each text word in the document image based on the document object tree; segmenting the document image into multiple candidate regions; determining the target pixels located within each candidate region from each pixel within each bounding box; and determining the parsed document based on each target pixel. This application can improve the accuracy of multimodal document parsing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of document processing technology, and in particular to a multimodal document parsing method, device, and computer-readable storage medium. Background Technology

[0002] Driven by the digital wave, the efficient and accurate parsing and conversion of paper documents to achieve electronic storage and convenient use of information has become a key requirement in many fields.

[0003] Currently, OCR (Optical Character Recognition) is commonly used in document parsing. This involves scanning, photographing, or other methods to acquire information from printed documents and converting it into computer-editable text. However, when dealing with multimodal elements such as tables, graphics, and text, this approach may lose element coordinates and relationships between elements during the parsing process, leading to lower accuracy in document parsing.

[0004] Therefore, improving the accuracy of multimodal document parsing is a problem that urgently needs to be solved. Summary of the Invention

[0005] The main objective of this application is to provide a multimodal document parsing method, device, and computer-readable storage medium, with the aim of improving the accuracy of multimodal document parsing.

[0006] To achieve the above objectives, this application provides a multimodal document parsing method, which includes:

[0007] Obtain the optical character recognition text and document image of the multimodal document;

[0008] Determine the spatial layout information of each text word in the optical character recognition text in the document image, and generate a document object tree based on each text word and the corresponding spatial layout information;

[0009] Determine the bounding box of each text word in the document image based on the document object tree;

[0010] The document image is segmented into multiple candidate regions, and target pixels located in each candidate region are determined within each bounding box. The parsed document is then determined based on each target pixel.

[0011] In one embodiment, the spatial layout information includes bounding box coordinates, and the step of determining the spatial layout information of each text word in the optical character recognition text within the document image, and generating a document object tree based on each text word and the spatial layout information, includes:

[0012] The document image and the optical character recognition text are input into the first preset model;

[0013] The optical character recognition text is segmented using the first preset model to obtain multiple text units.

[0014] The first preset model is used to perform position encoding processing on each of the text words to obtain the bounding box coordinates of each of the text words in the document image;

[0015] The first preset model is used to perform structured representation processing on each text word and the corresponding bounding box coordinates to generate a document object tree.

[0016] In one embodiment, the step of generating a document object tree by performing structured representation processing on each text word and its corresponding bounding box coordinates using the first preset model includes:

[0017] The text features of each text word and the visual features of the image region corresponding to each bounding box coordinate are extracted using the first preset model.

[0018] The text features and the visual features are fused across modalities to obtain fused features;

[0019] A document object tree is generated based on the fusion features.

[0020] In one embodiment, the step of determining the target pixel located in each of the candidate regions among each pixel within each of the bounding boxes includes:

[0021] A pixel-level mask matrix is ​​generated based on each of the candidate regions, wherein the pixel-level mask matrix represents the positional relationship between each of the text words and each of the candidate regions;

[0022] Based on the pixel-level mask matrix, determine whether each pixel within each bounding box is located in each candidate region;

[0023] If so, the pixel located in each of the candidate regions is taken as the target pixel.

[0024] In one embodiment, the step of segmenting the document image into multiple candidate regions includes:

[0025] The document image is input into a second preset model, and a feature map is generated by the feature extraction module of the second preset model.

[0026] The region proposal network in the second preset model generates multiple candidate regions and category labels for each candidate region based on the feature map.

[0027] In one embodiment, the step of determining the parsed document based on each of the target pixels includes:

[0028] Based on the category labels, determine the table-type areas and chart-type areas in each of the candidate areas;

[0029] Based on each of the target pixels, the table-type region and the chart-type region are structurally transformed to obtain a standardized representation;

[0030] The parsed document is determined based on each of the target pixels and the canonical representation.

[0031] In one embodiment, the canonical representation includes tables and charts, and the step of performing a structured transformation on the table-type region and the chart-type region based on each of the target pixels to obtain the canonical representation includes:

[0032] Determine the first pixel belonging to the table-type region among the target pixels, and perform a structured transformation on the table-type region based on each first pixel to obtain the table;

[0033] Determine the second pixel belonging to the chart type region among the target pixels, and perform a structured transformation on the chart type region based on each second pixel to obtain the chart.

[0034] In one embodiment, the step of determining the parsed document based on each of the target pixels and the canonical representation includes:

[0035] For any two adjacent first cells in the same row in the table, the horizontal distance between each first cell is determined based on the boundary coordinates of each first cell. If the horizontal distance is less than a first preset threshold, the first cells are merged.

[0036] For any two adjacent second cells in the same column in the table, the vertical distance between each second cell is determined based on the boundary coordinates of each second cell. If the vertical distance is less than a second preset threshold, the second cells are merged.

[0037] Based on each of the target pixels, the chart, and the table after merging, the parsed document is determined.

[0038] Furthermore, to achieve the above objectives, this application also provides a multimodal document parsing apparatus, the multimodal document parsing apparatus comprising:

[0039] The acquisition module is used to acquire the optical character recognition text and document image of the multimodal document;

[0040] The construction module is used to determine the spatial layout information of each text word in the optical character recognition text in the document image, and generate a document object tree based on each text word and the corresponding spatial layout information;

[0041] The determination module is used to determine the bounding box of each text word in the document image based on the document object tree;

[0042] The parsing module is used to segment the document image into multiple candidate regions, determine the target pixels located in each candidate region among the pixels within each bounding box, and determine the parsed document based on each target pixel.

[0043] In addition, to achieve the above objectives, this application also proposes a multimodal document parsing device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the multimodal document parsing method as described above.

[0044] In addition, to achieve the above objectives, this application also provides a storage medium, which is a computer-readable storage medium, on which a program implementing the multimodal document parsing method is stored. The program implementing the multimodal document parsing method is executed by a processor to implement the steps of the multimodal document parsing method as described above.

[0045] In addition, to achieve the above objectives, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the multimodal document parsing method described above.

[0046] This application provides a multimodal document parsing method. First, it acquires the optical character recognition (OCR) text and document image of the multimodal document. Then, it determines the spatial layout information of each text word in the OCR text in the document image, and generates a document object tree based on each text word and its corresponding spatial layout information. Based on the document object tree, it determines the bounding box of each text word in the document image. The document image is segmented into multiple candidate regions, and the target pixels located in each candidate region within each bounding box are determined. Based on each target pixel, the parsed document is determined.

[0047] In summary, this application first associates the document image of the multimodal document with the optical character recognition (OCR) text, generating a document object tree containing information about each text word and its spatial layout within the document image. This effectively captures the correspondence between text content and image spatial location. Simultaneously, by segmenting the document image to obtain candidate regions and accurately locating target pixels belonging to these candidate regions within the bounding boxes of each text word, it achieves a fine association between text information and image regions. Thus, compared to traditional OCR methods for parsing multimodal documents, this application significantly improves the accuracy of multimodal document parsing by structurally integrating text words, spatial coordinates, and image region information. Attached Figure Description

[0048] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0049] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is a flowchart illustrating the first embodiment of the multimodal document parsing method of this application;

[0051] Figure 2 This is a schematic diagram of the document parsing process involved in an embodiment of the multimodal document parsing method of this application;

[0052] Figure 3 This is a schematic diagram of the overall process involved in one embodiment of the multimodal document parsing method of this application;

[0053] Figure 4 This is a schematic diagram of the module structure of the multimodal document parsing device of this application;

[0054] Figure 5 This is a schematic diagram of the device structure of the hardware operating environment involved in the multimodal document parsing method in this application embodiment.

[0055] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0056] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0057] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0058] The main solution of this application is: to acquire the optical character recognition text and document image of the multimodal document; to determine the spatial layout information of each text word in the optical character recognition text in the document image, and to generate a document object tree based on each text word and the corresponding spatial layout information; to determine the bounding box of each text word in the document image based on the document object tree; to segment the document image into multiple candidate regions, to determine the target pixels located in each candidate region among the pixels within each bounding box, and to determine the parsed document based on each target pixel.

[0059] Currently, OCR (Optical Character Recognition) is commonly used in document parsing. This involves scanning, photographing, or other methods to acquire information from printed documents and converting it into computer-editable text. However, when dealing with multimodal elements such as tables, graphics, and text, this approach may lose element coordinates and relationships between elements during the parsing process, leading to lower accuracy in document parsing.

[0060] Therefore, improving the accuracy of multimodal document parsing is a problem that urgently needs to be solved.

[0061] This application first associates the document image of a multimodal document with optical character recognition (OCR) text, generating a document object tree containing information about each text word and its spatial layout within the document image. This effectively captures the correspondence between text content and image spatial location. Simultaneously, by segmenting the document image to obtain candidate regions and accurately locating target pixels belonging to these candidate regions within the bounding boxes of each text word, it achieves a fine association between text information and image regions. Thus, compared to traditional OCR methods for parsing multimodal documents, this application significantly improves the accuracy of multimodal document parsing by structurally integrating text words, spatial coordinates, and image region information.

[0062] It should be noted that the execution subject of the multimodal document parsing method in various embodiments of this application can be a document parsing system, or a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or a multimodal document parsing device capable of implementing the above functions. This embodiment does not specifically limit this. The following uses a document parsing system as the execution subject as an example to describe this embodiment and the following embodiments.

[0063] Based on this, this application proposes a first embodiment of a multimodal document parsing method, please refer to... Figure 1 The multimodal document parsing method includes steps S10 to S30:

[0064] Step S10: Obtain the optical character recognition text and document image of the multimodal document;

[0065] It should be noted that the text content of the multimodal document image (hereinafter referred to as the document image for distinction) is extracted in advance using OCR to obtain OCR text (i.e., optical character recognition text). A multimodal document refers to a document containing multimodal content to be parsed, and it belongs to paper documents.

[0066] Step S20: Determine the spatial layout information of each text word in the optical character recognition text in the document image, and generate a document object tree based on each text word and the corresponding spatial layout information;

[0067] It should be noted that spatial layout information includes location information and relationships. The location information of text terms in the document image can be the coordinates of the smallest bounding box containing the text term in the document image (i.e., bounding box coordinates). The document object tree includes multiple nodes, each node including node type, node content, location code (i.e., location information), and child nodes. The multimodal document is the root node, and chapter titles, paragraphs, tables, and images are child nodes of the root node. In this way, the relationships between the nodes can be determined.

[0068] In this embodiment, the spatial layout information includes bounding box coordinates, and step S20 may include:

[0069] Step S201: Input the document image and the optical character recognition text into the first preset model;

[0070] Step S202: The optical character recognition text is segmented using the first preset model to obtain multiple text units;

[0071] Step S203: Perform position encoding processing on each of the text words using the first preset model to obtain the bounding box coordinates of each of the text words in the document image;

[0072] Step S204: The first preset model is used to perform structured representation processing on each text word and the corresponding bounding box coordinates to generate a document object tree.

[0073] It should be noted that the first preset model can be the LayoutLMv3 model. LayoutLMv3 is a third-generation multimodal pre-trained model for document understanding proposed by Microsoft. It is used to achieve deep fusion and efficient processing of text, layout, and visual information in documents through a unified text and image mask pre-training mechanism. The first preset model includes a text segmentation module, a positional encoding module, and a structured representation module.

[0074] The document image of the multimodal document and the optical character recognition text are input into the first preset model; the text segmentation module of the first preset model is used to segment the optical character recognition text to obtain multiple text words; the position encoding module of the first preset model is used to encode the position of each text word to obtain the coordinates of the minimum bounding box of each text word in the document image (i.e., bounding box coordinates); finally, the structured representation module of the first preset model is used to perform structured representation processing on each text word and the corresponding bounding box coordinates to generate a document object tree.

[0075] In this embodiment, step S204 may include:

[0076] Step S2041: Extract the text features of each text word and the visual features of the image region corresponding to each bounding box coordinate using the first preset model;

[0077] Step S2042: Perform cross-modal fusion of the text features and the visual features to obtain fused features;

[0078] Step S2043: Generate a document object tree based on the fusion features.

[0079] It should be noted that textual features refer to information extracted from text lexical units that can characterize the semantic and syntactic attributes of lexical units, such as word vectors, part-of-speech tags, and semantic dependency relationships. Visual features refer to information extracted from the document image region corresponding to the bounding box coordinates that can characterize the visual attributes of the region, such as the color distribution, texture features, and edge contour features of the image region. Cross-modal fusion refers to associating and integrating textual and visual features through specific algorithms to eliminate the heterogeneity of the two modalities and form a unified information carrier that can simultaneously reflect the semantic relationship between text and the visual relationship between the image. Furthermore, the fused features are the result of cross-modal fusion and contain the correlation information between text and vision.

[0080] The structured representation module of the first preset model extracts text features for each text word using a text feature extraction component. Simultaneously, for the bounding box coordinates of each text word, the corresponding image region in the document image is determined, and the visual features of this image region are extracted using a visual feature extraction component. Then, the extracted text features and corresponding visual features are input into a cross-modal fusion component. This component eliminates the information differences between the text and visual modalities through feature alignment, weight allocation, and other logic, integrating the two features into a fused feature. Finally, the fused features corresponding to all text words are input into the structured representation component. This component constructs a tree-like data structure (i.e., a document object tree) based on the text semantic and spatial relationships contained in the fused features.

[0081] In one feasible implementation, the document image of the multimodal document and the OCR-recognized text are input into the LayoutLMv3 model; the model performs word segmentation on the OCR-recognized text to generate the input_ids parameter, which represents the token sequence after text segmentation; simultaneously, the model determines the bounding box parameter corresponding to each token, which represents the bounding box coordinates of each text word in the original image, and the coordinates can be represented as [x_min, y_min, x_max, y_max]; subsequently, the model jointly extracts visual features, text features, and layout features through a multimodal pre-training architecture to achieve cross-modal information interaction; finally, the model outputs a structured representation of the document containing spatial location encoding, in which each word node is associated with its visual features, semantic features, and precise coordinates in the image, thereby constructing a hierarchical document object tree, which fully preserves the spatial relationships between text content, original layout, and visual elements.

[0082] Furthermore, it should be noted that the document structured representation includes, but is not limited to: token-level embeddings, where each token includes a vector representation of its semantics (from the text model) and also incorporates visual features of its corresponding position in the document image (from the visual model); positional encoding (bounding box coordinates), where each text term has an explicit, learned positional encoding associated with its bounding box coordinates, used to enable the model to obtain relative positional information between terms; and association relationships, where the model learns the dependencies between all terms through a self-attention mechanism, for example, the model can understand which terms constitute a paragraph and which terms belong to the same table cell.

[0083] Step S30: Determine the bounding box of each text word in the document image based on the document object tree;

[0084] Based on the document object tree, determine the bounding box coordinates of each text word in the document image, and based on the bounding box coordinates, determine the bounding box position of each text word in the document image.

[0085] Step S40: Segment the document image into multiple candidate regions, determine the target pixels located in each candidate region within each bounding box, and determine the parsed document based on each target pixel.

[0086] The document image is segmented into multiple rectangular regions (hereinafter referred to as candidate regions for distinction). The content in each candidate region belongs to a category, such as text, table, or graphic. Each pixel within the bounding box of each text word is determined to belong to a candidate region. Pixels located in the candidate regions are called target pixels for distinction. The target pixels constitute the overlapping area of ​​the word bounding box and the candidate regions. The parsed document is determined based on each target pixel.

[0087] In this embodiment, step S40 may include:

[0088] Step A10: Input the document image into the second preset model, and generate a feature map through the feature extraction module of the second preset model;

[0089] Step A20: The region proposal network in the second preset model generates multiple candidate regions and category labels for each candidate region based on the feature map.

[0090] It should be noted that the second preset model can be a Mask R-CNN (Mask Region-based Convolutional Neural Network) model. The second preset model includes a feature extraction module and a region proposal network.

[0091] The document image is input into a second preset model; a multimodal feature map is generated by the feature extraction module of the second preset model; then, the region proposal network in the second preset model generates multiple candidate regions and category labels for each candidate region based on the feature map. The category label refers to the category of the object within the candidate region, which can be text, tables, line charts, illustrations, bar charts, etc.

[0092] In one feasible implementation, the input document image is first fed into a segmentation network based on Mask R-CNN, and multi-scale feature maps are extracted using the feature extraction module (such as ResNet-FPN) in the model. Then, a large number of candidate regions (ROIs) are generated on the feature maps through RPN (Region Proposal Network). Each ROI (Region of Interest) contains the possible location and preliminary category information of the target image.

[0093] Furthermore, the data structure output by Mask R-CNN includes: bounding boxes, class labels, mask matrix, and confidence scores. The confidence scores and the mask matrix are two distinct variables. Multiple region results are filtered using the NMS algorithm, which utilizes the non-maximum suppression parameter `nms_threshold` to eliminate highly overlapping redundant prediction boxes. The algorithm works by calculating the Intersection over Union (IoU) to predict whether two regions overlap. The `nms_threshold` parameter needs to be manually set; a smaller value results in stricter suppression. For example, when the parameter is set to 0.3, if two boxes have only minor overlap (IoU > 0.3), the one with the lower score will be suppressed. Low-confidence detection boxes are also filtered using the confidence threshold parameter `conf_threshold`, thus obtaining high-precision mixed-image region segmentation results, i.e., the candidate regions.

[0094] In this embodiment, step S40 may include:

[0095] Step B10: Generate a pixel-level mask matrix based on each candidate region, wherein the pixel-level mask matrix represents the positional relationship between each text word and each candidate region;

[0096] A pixel-level mask matrix is ​​generated based on each candidate region. This matrix includes the pixels within the bounding box of each text word and the corresponding pixel value. The pixel value represents the probability or a Boolean value that the pixel belongs to a target candidate region; a value of 1 indicates that the pixel belongs to a candidate region, while a value of 0 indicates that the pixel does not belong to any candidate region. In other words, the pixel-level mask matrix represents the positional relationship between each text word and each candidate region.

[0097] In one feasible implementation, the input document image is fed into a Mask R-CNN-based model to obtain candidate regions. Then, the ROIAlign layer and classification head of the Mask R-CNN model output a pixel-level mask matrix with category labels. The category labels correspond to the candidate regions.

[0098] Step B20: Based on the pixel-level mask matrix, determine whether each pixel within each bounding box is located in each candidate region;

[0099] Step B30: If yes, then the pixel located in each of the candidate regions among the pixels is taken as the target pixel.

[0100] Based on a pixel-level mask matrix, pixels with a value of 1 in the document image are retained as target pixels. That is, the pixels located in the candidate region within the bounding box of each text word are determined, thereby extracting effective information (target pixels) in the document image and improving the accuracy of structural transformation of complex elements.

[0101] In one feasible implementation, the candidate regions are aligned using the ROIAlign layer of the Mask R-CNN model to avoid quantization errors and ensure accurate transmission of spatial information. The aligned features are then input into the classification head and the mask head to output the category label and pixel-level mask matrix corresponding to each candidate region.

[0102] For example, such as Figure 2 The diagram illustrates the document parsing process. For multimodal documents, the document image is first processed by OCR to obtain OCR-recognized text (i.e., optical character recognition text). The document image and OCR-recognized text are then input into the LayoutLMv3 model to obtain the document object tree. The document image is then input into the Mask R-CNN model to obtain multiple candidate regions and a pixel-level mask matrix. The target pixels are determined based on this matrix. Finally, the parsed document is determined based on the target pixels.

[0103] This application's embodiments first associate the document image of a multimodal document with optical character recognition (OCR) text, generating a document object tree containing the location information (boundary box coordinates) of each text word and their associated relationships. This effectively captures the correspondence between text content and image spatial location. Simultaneously, by segmenting the document image to obtain candidate regions and accurately locating target pixels belonging to candidate regions within the bounding boxes of each text word, a fine association between text information and image regions is achieved. Thus, compared to traditional methods of parsing multimodal documents using OCR, this application's embodiments significantly improve the accuracy of multimodal document parsing by structurally integrating text words, spatial coordinates, and image region information.

[0104] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description and will not be repeated hereafter. Based on this, step S40 may include:

[0105] Step C10: Determine the table-type region and the chart-type region among the candidate regions based on each of the category labels;

[0106] Step C20: Perform a structured transformation on the table-type region and the chart-type region based on each of the target pixels to obtain a standardized representation;

[0107] Step C30: Determine the parsed document based on each of the target pixels and the canonical representation.

[0108] Based on the category labels of each candidate region, table-type regions and chart-type regions are determined. These two types of regions undergo structural transformation to obtain a canonical representation, and the parsed document is generated based on each target pixel and the canonical representation.

[0109] In this embodiment, the specification representation includes tables and charts, and step C20 may include:

[0110] Step C201: Determine the first pixel in each of the target pixels that belongs to the table-type region, and perform a structured transformation on the table-type region based on each of the first pixels to obtain the table;

[0111] Step C202: Determine the second pixel in each of the target pixels that belongs to the chart type region, and perform a structured transformation on the chart type region based on each second pixel to obtain the chart.

[0112] For table-type areas, the pixels belonging to the table-type area in each target pixel are identified (hereinafter referred to as the first pixel for distinction). Based on each first pixel, the table-type area is structurally transformed to obtain a table in a preset format, so as to accurately parse the table area in the multimodal document. For chart-type areas, the pixels belonging to the chart-type area in each target pixel are identified (hereinafter referred to as the second pixel for distinction). Based on each second pixel, the chart-type area is structurally transformed to obtain a chart in a preset format, so as to accurately parse the chart area in the multimodal document.

[0113] For example, for table-type elements, regular expressions are used to match data points (such as numbers, percentages, etc.) to generate a structured Markdown table representation; for chart-type elements, the data is converted into the JSON data structure required by the D3.js template.

[0114] In this embodiment, step C30 may include:

[0115] Step C301: For any two adjacent first cells in the same row in the table, determine the horizontal distance between each first cell based on the boundary coordinates of each first cell. If the horizontal distance is less than a first preset threshold, merge each first cell.

[0116] Step C302: For any two adjacent second cells in the same column in the table, determine the vertical distance between each second cell based on the boundary coordinates of each second cell. If the vertical distance is less than a second preset threshold, merge each second cell.

[0117] Step C303: Based on each of the target pixels, the chart, and the table after merging, determine the parsed document.

[0118] It should be noted that a lower limit value for the horizontal distance between two adjacent cells in the same row of the table is preset (hereinafter referred to as the first preset threshold for distinction), and a lower limit value for the vertical distance between two adjacent cells in the same column of the table is preset (hereinafter referred to as the second preset threshold for distinction).

[0119] For any two adjacent cells in the same row of the table (hereinafter referred to as the first cell for distinction), the horizontal distance between the first cells is determined based on their boundary coordinates. Specifically, the variance of the x-coordinates of the left boundaries of the two first cells is used as their horizontal distance. If the horizontal distance is less than a first preset threshold, the two first cells are merged. Similarly, for any two adjacent cells in the same column of the table (hereinafter referred to as the second cell for distinction), the vertical distance between the second cells is determined based on their boundary coordinates. Specifically, the variance of the y-coordinates of the lower boundaries of the two second cells is used as their vertical distance. If the vertical distance is less than a second preset threshold, the two second cells are merged. Finally, the parsed document is determined based on the merged table, each target pixel, and the chart.

[0120] This application embodiment repairs table merging errors by employing the axis alignment algorithm of TextIn. Specifically, in one feasible implementation, when using TextIn's axis alignment algorithm to repair table merging errors, the process first analyzes the boundary coordinate data of rows and columns in the table image to accurately detect abnormal segmentation of cells across rows and columns caused by OCR recognition. The parameter `merge_threshold` is set according to actual needs to determine the pixel threshold for row and column alignment, while the `direction` parameter specifies whether the correction direction is horizontal or vertical. Next, the repair process begins. The first step is merging the bounding boxes: based on the detected abnormally segmented cell coordinate information, a new bounding box (bbox) that can completely enclose all cells to be merged is calculated. The second step is merging the content: according to a predetermined reading order, the text content within all cells to be merged is spliced ​​together to form the complete content of the new cell. The third step is updating the data structure and attributes: in the table's internal data structure, all merged old, atomized cell objects are deleted, and a completely new cell object is created in its original position. Simultaneously, the new cell is assigned the new bounding box (bbox) and new content calculated in the first two steps, thereby eliminating cell breakage and completing the repair of the table merging error.

[0121] For example, such as Figure 3 The diagram shows the overall process. First, multimodal feature extraction is performed based on the document image to obtain the document object tree. Then, element separation and labeling are performed, that is, a pixel-level mask matrix with category labels is determined to retain the target pixels in the document image. The structured elements (table elements and chart elements) in the document image are subjected to structured transformation to obtain a standardized representation. Abnormal cell correction is performed on the table to output the structured document (i.e., the parsed document).

[0122] Thus, after obtaining the valid information in the multimodal document, this embodiment of the application further performs a structured transformation on table and chart elements to obtain a standard representation that is easy to view and understand and matches the information in the multimodal document. It also uses an axis alignment algorithm to handle table merging errors, merging incorrectly segmented cells to restore the correct structure of the table, thereby improving the parsing accuracy of the multimodal document.

[0123] This application also provides a multimodal document parsing device. Please refer to... Figure 4 The multimodal document parsing device includes:

[0124] The acquisition module 10 is used to acquire the optical character recognition text and document image of the multimodal document;

[0125] Construction module 20 is used to determine the spatial layout information of each text word in the optical character recognition text in the document image, and generate a document object tree based on each text word and the corresponding spatial layout information;

[0126] The determination module 30 is used to determine the bounding box of each text word in the document image based on the document object tree;

[0127] The parsing module 40 is used to segment the document image into multiple candidate regions, determine the target pixels located in each candidate region among the pixels within each bounding box, and determine the parsed document based on each target pixel.

[0128] Optionally, the spatial layout information includes bounding box coordinates, and the construction module 20 is further configured to:

[0129] The document image and the optical character recognition text are input into the first preset model;

[0130] The optical character recognition text is segmented using the first preset model to obtain multiple text units.

[0131] The first preset model is used to perform position encoding processing on each of the text words to obtain the bounding box coordinates of each of the text words in the document image;

[0132] The first preset model is used to perform structured representation processing on each text word and the corresponding bounding box coordinates to generate a document object tree.

[0133] Optionally, the building module 20 is further configured to:

[0134] The text features of each text word and the visual features of the image region corresponding to each bounding box coordinate are extracted using the first preset model.

[0135] The text features and the visual features are fused across modalities to obtain fused features;

[0136] A document object tree is generated based on the fusion features.

[0137] Optionally, the parsing module 40 is further configured to:

[0138] A pixel-level mask matrix is ​​generated based on each of the candidate regions, wherein the pixel-level mask matrix represents the positional relationship between each of the text words and each of the candidate regions;

[0139] Based on the pixel-level mask matrix, determine whether each pixel within each bounding box is located in each candidate region;

[0140] If so, the pixel located in each of the candidate regions is taken as the target pixel.

[0141] Optionally, the parsing module 40 is further configured to:

[0142] The document image is input into a second preset model, and a feature map is generated by the feature extraction module of the second preset model.

[0143] The region proposal network in the second preset model generates multiple candidate regions and category labels for each candidate region based on the feature map.

[0144] Optionally, the parsing module 40 is further configured to:

[0145] Based on the category labels, determine the table-type areas and chart-type areas in each of the candidate areas;

[0146] Based on each of the target pixels, the table-type region and the chart-type region are structurally transformed to obtain a standardized representation;

[0147] The parsed document is determined based on each of the target pixels and the canonical representation.

[0148] Optionally, the specification representation includes tables and charts, and the multimodal document parsing device further includes:

[0149] Determine the first pixel belonging to the table-type region among the target pixels, and perform a structured transformation on the table-type region based on each first pixel to obtain the table;

[0150] Determine the second pixel belonging to the chart type region among the target pixels, and perform a structured transformation on the chart type region based on each second pixel to obtain the chart.

[0151] Optionally, the parsing module 40 is further configured to:

[0152] For any two adjacent first cells in the same row in the table, the horizontal distance between each first cell is determined based on the boundary coordinates of each first cell. If the horizontal distance is less than a first preset threshold, the first cells are merged.

[0153] For any two adjacent second cells in the same column in the table, the vertical distance between each second cell is determined based on the boundary coordinates of each second cell. If the vertical distance is less than a second preset threshold, the second cells are merged.

[0154] Based on each of the target pixels, the chart, and the table after merging, the parsed document is determined.

[0155] The multimodal document parsing apparatus provided in this application, employing the multimodal document parsing method described in the above embodiments, can solve the technical problem of how to improve the accuracy of multimodal document parsing. Compared with the prior art, the beneficial effects of the multimodal document parsing apparatus provided in this application are the same as those of the multimodal document parsing method described in the above embodiments, and other technical features in the multimodal document parsing apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0156] This application provides a multimodal document parsing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the multimodal document parsing method in Embodiment 1 above.

[0157] The following is for reference. Figure 5 The diagram illustrates a structural schematic of a multimodal document parsing device suitable for implementing embodiments of this application. The multimodal document parsing device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The multimodal document parsing device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0158] like Figure 5As shown, the multimodal document parsing device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the multimodal document parsing device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the multimodal document parsing device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows multimodal document parsing devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0159] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0160] The multimodal document parsing device provided in this application, employing the multimodal document parsing method described in the above embodiments, can solve the technical problem of how to improve the accuracy of multimodal document parsing. Compared with the prior art, the beneficial effects of the multimodal document parsing device provided in this application are the same as those of the multimodal document parsing method provided in the above embodiments, and other technical features in this multimodal document parsing device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0161] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0162] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0163] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the multimodal document parsing method described in the above embodiments.

[0164] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0165] The aforementioned computer-readable storage medium may be included in a multimodal document parsing device; or it may exist independently and not be assembled into a multimodal document parsing device.

[0166] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a multimodal document parsing device, cause the multimodal document parsing device to: acquire optical character recognition (OCR) text and a document image of the multimodal document; determine the spatial layout information of each text word in the OCR text within the document image, and generate a document object tree based on each text word and the corresponding spatial layout information; determine the bounding box of each text word in the document image based on the document object tree; segment the document image into multiple candidate regions, determine the target pixels located in each candidate region within each bounding box, and determine the parsed document based on each target pixel.

[0167] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0168] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0169] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0170] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described multimodal document parsing method, thereby solving the technical problem of how to improve the accuracy of multimodal document parsing. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the multimodal document parsing method provided in the above embodiments, and will not be repeated here.

[0171] This application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the multimodal document parsing method described above.

[0172] The computer program product provided in this application can improve the accuracy of multimodal document parsing. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiments of this application are the same as the beneficial effects of the multimodal document parsing method provided in the above embodiments, and will not be repeated here.

[0173] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.

Claims

1. A multimodal document parsing method, characterized in that, The multimodal document parsing method includes: Obtain the optical character recognition text and document image of the multimodal document; The spatial layout information of each text word in the optical character recognition text is determined in the document image, and a document object tree is generated based on each text word and the corresponding spatial layout information, wherein the spatial layout information includes bounding box coordinates; Determine the bounding box of each text word in the document image based on the document object tree; The document image is segmented into multiple candidate regions; A pixel-level mask matrix is ​​generated based on each of the candidate regions, wherein the pixel-level mask matrix represents the positional relationship between each of the text words and each of the candidate regions; Based on the pixel-level mask matrix, determine whether each pixel within each bounding box is located in each candidate region; If so, then the pixel located in each of the candidate regions among the pixels is taken as the target pixel; Based on the category labels of each candidate region, determine the table-type region and the chart-type region among the candidate regions; Based on each of the target pixels, the table-type region and the chart-type region are structurally transformed to obtain a standardized representation; The parsed document is determined based on each of the target pixels and the canonical representation.

2. The multimodal document parsing method as described in claim 1, characterized in that, The step of determining the spatial layout information of each text word in the optical character recognition text within the document image, and generating a document object tree based on each text word and the spatial layout information, includes: The document image and the optical character recognition text are input into the first preset model; The optical character recognition text is segmented using the first preset model to obtain multiple text units. The first preset model is used to perform position encoding processing on each of the text words to obtain the bounding box coordinates of each of the text words in the document image; The first preset model is used to perform structured representation processing on each text word and the corresponding bounding box coordinates to generate a document object tree.

3. The multimodal document parsing method as described in claim 2, characterized in that, The step of generating a document object tree by performing structured representation processing on each text word and its corresponding bounding box coordinates using the first preset model includes: The text features of each text word and the visual features of the image region corresponding to each bounding box coordinate are extracted using the first preset model. The text features and the visual features are fused across modalities to obtain fused features; A document object tree is generated based on the fusion features.

4. The multimodal document parsing method as described in claim 1, characterized in that, The step of segmenting the document image into multiple candidate regions includes: The document image is input into a second preset model, and a feature map is generated by the feature extraction module of the second preset model. The region proposal network in the second preset model generates multiple candidate regions and category labels for each candidate region based on the feature map.

5. The multimodal document parsing method as described in claim 1, characterized in that, The standardized representation includes tables and charts. The step of performing a structured transformation on the table-type region and the chart-type region based on each target pixel to obtain the standardized representation includes: Determine the first pixel belonging to the table-type region among the target pixels, and perform a structured transformation on the table-type region based on each first pixel to obtain the table; Determine the second pixel belonging to the chart type region among the target pixels, and perform a structured transformation on the chart type region based on each second pixel to obtain the chart.

6. The multimodal document parsing method as described in claim 5, characterized in that, The step of determining the parsed document based on each of the target pixels and the canonical representation includes: For any two adjacent first cells in the same row in the table, the horizontal distance between each first cell is determined based on the boundary coordinates of each first cell. If the horizontal distance is less than a first preset threshold, the first cells are merged. For any two adjacent second cells in the same column in the table, the vertical distance between each second cell is determined based on the boundary coordinates of each second cell. If the vertical distance is less than a second preset threshold, the second cells are merged. Based on each of the target pixels, the chart, and the table after merging, the parsed document is determined.

7. A multimodal document parsing device, characterized in that, The multimodal document parsing device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the multimodal document parsing method as described in any one of claims 1 to 6.

8. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, it implements the steps of the multimodal document parsing method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Document digital processing method and device and nonvolatile storage medium

    CN119964185A

  • Method and device for processing document image and electronic equipment

    CN120032386A