PDF (Portable Document Format) file table analysis method
By converting PDF files into pictures and using OCR models to extract table information, and combining named entity recognition technology to identify cell types, the problem of poor performance in parsing PDF nested tables in the existing technology is solved, and more efficient and accurate extraction and restoration of table information is achieved.
Patent Information
- Application Number
- CN202510580825.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The prior art performs poorly when parsing nested tables in PDF files, especially when processing PDF files in scanned or image form, which may result in the inability to accurately extract the table contents.
By converting PDF files into image format, the pre-trained OCR model is used to extract table information and convert it into HTML format. Then, the named entity recognition technology is used to identify the key-value type to which the cell belongs, and combine the relative position and the type to which the cell is paired to form a complete key-value pair, and finally saved to Json format.
Improve the accuracy and efficiency of extracting table information from PDF files, especially when dealing with complex nested tables and PDF files in scan or image form, the table information can be effectively restored.
Smart Images

Figure CN120104577A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of PDF file processing, and in particular to a PDF file table parsing method. Background Art
[0002] PDF (Portable Document Format) is a file format developed and widely adopted by Adobe Systems in 1993. PDF was originally designed to achieve consistent display of documents across platforms, that is, to ensure that the display of documents matches the original intentions of creation regardless of the hardware configuration or software environment. This format ensures the integrity and consistency of documents by encapsulating information such as text, fonts, images, and document layout in a single file, making it an ideal choice for electronic document exchange.
[0003] However, this feature of PDF files also brings some challenges, especially when dealing with nested tables. Nested tables refer to complex layouts that contain one or more tables within the cells of a table. Since PDF files are designed to display visual representations of documents rather than store structured data, PDF files do not always retain the original structure information of the document when they are saved. This is in stark contrast to office software such as Microsoft Word or Excel, which can store and retain the original structure information of the document, making it more straightforward and accurate to extract structured data from these files.
[0004] When parsing nested tables in PDF files, the problem becomes even more complicated. Currently, there are a variety of methods and techniques for parsing nested tables in PDF. In the Python programming language, there are several popular libraries that can assist in this task, including PyPDF2, pdfminer.six, and pdfplumber. These libraries provide basic functions for parsing tables in PDF files, but their performance is not ideal when dealing with complex layouts such as nested tables. Especially when the PDF file is an image file scanned from a paper document, these libraries may miss when identifying the table structure, resulting in an inability to accurately extract the table content.
[0005] To solve this problem, OCR (Optical Character Recognition) technology was introduced into PDF table parsing. OCR technology can extract text information and structural information from images, which is particularly important for PDF files in scanned or image form. However, OCR technology faces many challenges in parsing unstructured table data and converting it into structured table data. Although OCR can recognize text information, it is difficult to parse unstructured table data and convert it into structured table data, with low accuracy. At the same time, it is difficult to infer the relationship between cells and restore data based on text information and structural information. Summary of the invention
[0006] The object of the present invention is to provide a PDF file table parsing method, which can improve the accuracy and efficiency of extracting table information from a PDF file.
[0007] To achieve the above object, the present invention provides the following technical solutions: A PDF file table parsing method includes the following steps: PDF document to image format conversion step: convert the PDF document page to be parsed into image format; Table area detection step: Input the converted image into a pre-trained table recognition OCR model, which can identify the table area frame range in the image and output the coordinates of the table area frame in the image ,in is the i-th coordinate point of the table area frame, is the confidence of the i-th coordinate point; Table area cropping step: crop the original image according to the table position coordinates P provided by the model, and only keep the image of the table area frame; Image to HTML format conversion step: Use the pre-trained OCR model to identify the table space structure and the cell text information and structural features contained in the table image, and then convert it from the image format to structured table data based on the extracted table space structure and cell text information and structural features, and further convert it into a table in HTML format; Steps for converting cells across nested tables in HTML format: In HTML format tables, use the table tag to define a table, use the tr tag to define rows in the table, use the td tag to define standard cells, use the rowspan attribute to define the number of rows spanned by a table cell, and use the colspan attribute to define the number of columns spanned by a table cell; Steps for identifying the key value type of a cell: Use the pre-trained named entity recognition model to identify the key value type of the cell: Saving step: Pair adjacent cells in pairs and save them in JSON format.
[0008] Specifically, in the cross-cell conversion step in the nested table in HTML format: Traverse the tr tags in the table tag, read the cells containing td tags one by one, and determine whether the cell contains the rowspan attribute or the colspan attribute: If none of the td cells contain the rowspan attribute or the colspan attribute, directly read the corresponding information of the td cell and mark the cell data as , where c is the number of td cells; If the cell contains the colspan attribute, obtain the colspan attribute value m and the corresponding information of the td cell, and mark the cell data as , where c is the number of td cells; If the cell contains the colspan attribute, obtain the colspan attribute value n and the corresponding information of the td cell, and mark the cell data as , where c is the number of td cells; for i < k < i + n - 1, there is ; Specifically, in the step of identifying the key-value type to which the cell belongs, specifically: Extraction of the cell text series under the tr tag: Traverse the tr tags in the table tag row by row, and splice the text of the td cells contained under the tr tag into a cell text sequence , where c is the number of td cells contained in the i-th tr tag; Generate cell text series annotations using a pre-trained model: Use a pre-trained model to encode the cell text sequence and generate a sequence representation ; Obtain the start vector representation and end vector representation of the entity segment: Input the sequence representation into the linear layer and the linear layer in the key-value type prediction layer to which the cell belongs, and obtain the start vector representation and end vector representation of the information of each entity segment in the sequence representation: Among them, the start vector representation is the vector representation of the start character of the entity segment, and the end vector representation is the vector representation of the end character of the entity segment. is the key value type of the cell. , It is a set of all key value types, which is equivalent to the start vector representation and end vector representation of the key value type identification of different cells. Therefore, for the cell text sequence , the key value type of the cell The confidence sequence of is: Calculate the key value type of the cell Confidence: For cell text sequences middle , which belongs to the key value type of the cell The confidence level is: Specifically, in the saving step, specifically: By identifying the key value type of the cell in the table, traverse each row of cell types row by row: If the cell key sequence If all cell types in the row are key, all cells in this row will be cached; If the cell key sequence middle, and The text is the same, that is, it belongs to a cross-row cell, then it is ignored Corresponding cell type, and converting the format to ; If the cell key sequence middle, and If the text is not the same, As the parent key, As The key of the subordinate node of ; If the cell key sequence Only the first cell in For key, others is val, and are all keys, then As primary key, As The key of the subordinate node of As The key of the subordinate node of is the value corresponding to it.
[0009] Compared with the prior art, the present invention has the following beneficial effects:
[0010] The present invention provides an innovative solution. The present invention first converts the PDF file into an image format, then uses a pre-trained OCR model to extract the table information in the converted image, and converts the extracted table information into HTML format. Next, the named entity recognition technology is used to identify the type of each cell in each row of the table. Finally, by combining the relative position and type of each cell in each row, the corresponding cells are paired to form a complete key-value pair, and then converted into Json format for storage. Please refer to Figure 1 .
[0011] For PDF files that are scanned files or tables in the form of pictures, the present invention can use the OCR model to obtain the structural information and text content in the table. For complex table layouts such as nested tables, the present invention uses named entity recognition technology to distinguish the corresponding types of text in each cell in each row, and combines the relative position and recognition attributes of the cell to effectively restore the table information, thereby improving the accuracy of table parsing. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0013] Figure 1 It is a flowchart of the present invention; Figure 2 For a table example; Figure 3 for Figure 2 The conversion process of the table example shown; Figure 4 is a key-value table; Figure 5 for Figure 2 The complete conversion process for the table example shown. DETAILED DESCRIPTION
[0014] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0015] See Figures 1 to 5 , a PDF file table parsing method, comprising the following steps: PDF document to image format conversion step: Use the pdf2image library in the Python programming language to convert the PDF document page to be parsed into an image format, such as PNG or JPEG. This step converts visual elements such as text, images, and tables in the PDF page into pixel data to provide input for the subsequent table recognition model.
[0016] Table area detection step: Input the converted image into a pre-trained table recognition OCR model, which can identify the table area frame range in the image and output the coordinates of the table area frame in the image ,in is the i-th coordinate point of the table area frame, is the confidence of the i-th coordinate point.
[0017] Table area cropping step: According to the table position coordinates P provided by the model, the original image is cropped to retain only the image of the table area frame. This step is the key to the present invention, which can effectively prevent other elements such as plain text or illustrations in the document from affecting the subsequent table recognition, thereby improving the recognition accuracy.
[0018] Steps for converting images to HTML format: Use a pre-trained OCR model to identify the table space structure and the cell text information and structural features contained in the table image, and then convert it from image format to structured table data based on the extracted table space structure and cell text information and structural features, and further convert it into a table in HTML format.
[0019] Steps for converting cells across nested tables in HTML format: In HTML format tables, use the table tag to define a table, use the tr tag to define rows in the table, use the td tag to define standard cells, use the rowspan attribute to define the number of rows spanned by a table cell, and use the colspan attribute to define the number of columns spanned by a table cell; Before converting HTML tables into structured table information, you need to consider the nested table conversion problem, that is, the cross-cell conversion problem in the table. Figure 3 The specific tag conversion process can be achieved according to the following steps: Traverse the tr tags in the table tag, read the cells containing the td tags one by one, and determine whether the cell contains the rowspan attribute or the colspan attribute: If all cell td do not contain rowspan attributes or colspan attributes, directly read the corresponding information of the td cell and mark the cell data as , where c is the number of td cells; If the cell If it contains the colspan attribute, obtain the value m of the colspan attribute and the corresponding information of the td cell, and mark the cell data as , where c is the number of td cells; If the cell contains the colspan attribute, obtain the value n of the colspan attribute and the corresponding information of the td cell, and mark the cell data as , where c is the number of td cells; for i < k < i + n - 1, there is ; Steps for identifying the key-value type to which the cell belongs: Use a pre-trained named entity recognition model to identify the key-value type to which the cell belongs. Specifically: Extraction of the cell text series under the tr tag: Traverse the tr tags in the table tag line by line, and splice the text of the td cells contained under the tr tag into a cell text sequence , where c is the number of td cells contained in the i-th tr tag; Generate annotations for the cell text series using a pre-trained model: Use a pre-trained model to encode the cell text sequence and generate a sequence representation ; Obtain the start vector representation and end vector representation of the entity segment: Input the sequence representation into the linear layer in the key-value type prediction layer for the cell to which it belongs and the linear layer respectively, to obtain the start vector representation and end vector representation of the information of each entity segment in the sequence representation: Among them, the start vector representation is the vector representation of the start character of the entity segment, and the end vector representation is the vector representation of the end character of the entity segment is the key-value type to which the cell belongs, , is the set of all key-value types, which is equivalent to the start vector representation and end vector representation for identifying different key-value types to which cells belong, Therefore, for the cell text sequence , the confidence sequence of the key-value type to which the cell belongs is: Calculate the key-value type Confidence: For cell text sequences middle , which belongs to the key value type of the cell The confidence level is: Saving steps: Pair adjacent cells and save them in Json format. Specifically, identify the key value type of the cell in the table, such as Figure 5 As shown in step 3, traverse each row of cell types line by line: If the cell key sequence If all cell types in the row are key, all cells in this row will be cached; If the cell key sequence middle, and The text is the same, that is, it belongs to a cross-row cell, then it is ignored Corresponding cell type, and converting the format to ; If the cell key sequence middle, and If the text is not the same, As the parent key, As The key of the subordinate node of ; If the cell key sequence Only the first cell in For key, others is val, and are all keys, then As primary key, As The key of the subordinate node of As The key of the subordinate node of is the value corresponding to it.
[0020] The converted cell Json format is as follows Figure 5 As shown in step 4.
[0021] The beneficial effects of the present invention are:
[0022] The present invention provides an innovative solution. The present invention first converts the PDF file into an image format, then uses a pre-trained OCR model to extract the table information in the converted image, and converts the extracted table information into HTML format. Next, the named entity recognition technology is used to identify the type of each cell in each row of the table. Finally, by combining the relative position and type of each cell in each row, the corresponding cells are paired to form a complete key-value pair, and then converted into Json format for storage. Please refer to Figure 1 .
[0023] For PDF files that are scanned files or tables in the form of pictures, the present invention can use the OCR model to obtain the structural information and text content in the table. For complex table layouts such as nested tables, the present invention uses named entity recognition technology to distinguish the corresponding types of text in each cell in each row, and combines the relative position and recognition attributes of the cell to effectively restore the table information, thereby improving the accuracy of table parsing.
[0024] The above are only preferred embodiments of the present invention, and are not intended to limit the present invention in any form. Although the present invention has been disclosed as above in the form of preferred embodiments, it is not intended to limit the present invention. Any technical personnel in the field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. A PDF file table parsing method, characterized in that: The following steps are involved: PDF document to image format conversion step: convert the PDF document page to be parsed into image format; Table area detection step: Input the converted image into a pre-trained table recognition OCR model, which can identify the table area frame range in the image and output the coordinates of the table area frame in the image ,in is the i-th coordinate point of the table area frame, is the confidence of the i-th coordinate point; Table area cropping step: crop the original image according to the table position coordinates P provided by the model, and only keep the image of the table area frame; Image to HTML format conversion step: Use the pre-trained OCR model to identify the table space structure and the cell text information and structural features contained in the table image, and then convert it from the image format to structured table data based on the extracted table space structure and cell text information and structural features, and further convert it into a table in HTML format; Steps for converting cells across nested tables in HTML format: In HTML format tables, use the table tag to define a table, use the tr tag to define rows in the table, use the td tag to define standard cells, use the rowspan attribute to define the number of rows spanned by a table cell, and use the colspan attribute to define the number of columns spanned by a table cell; The key value type identification step of the cell: using the pre-trained named entity recognition model to identify the key value type of the cell; Saving steps: Pair adjacent cells and save them in Json format.
2. The PDF file table parsing method according to claim 1, characterized in that: In the step of converting cells across nested tables in the HTML format: Traverse the tr tags in the table tag, read the cells containing the td tags one by one, and determine whether the cell contains the rowspan attribute or the colspan attribute: If all cell td do not contain rowspan attributes or colspan attributes, directly read the corresponding information of the td cell and mark the cell data as , where c is the number of td cells; If the cell Contains the colspan attribute, then obtains the colspan attribute value m and the corresponding information of the td cell, and marks the cell data as , where c is the number of td cells; If the cell contains the colspan attribute, obtain the value n of the colspan attribute and the corresponding information of the td cell, and mark the cell data as , where c is the number of td cells; for i < k < i + n - 1, there is .
3. The PDF file table parsing method according to claim 1, characterized in that: In the step of identifying the key value type to which the cell belongs, specifically: Extract the cell text series under the tr tag: traverse the tr tags in the table tag row by row, and splice the td cell text contained in the tr tag into a cell text sequence , where c is the number of td cells contained in the i-th tr tag; Generate cell text series annotations using pre-trained models: Generate cell text series annotations using pre-trained models Encode and generate sequence representation ; Get the start vector representation and end vector representation of the entity fragment: Represent the sequence Input into the linear layer of the key value type prediction layer of the cell and the linear layer , get the start vector representation and end vector representation of each entity fragment information in the sequence representation: The start vector represents the vector representation of the start character of the entity segment, and the end vector represents the vector representation of the end character of the entity segment. is the key value type of the cell. , It is a set of all key value types, which is equivalent to the start vector representation and end vector representation of the key value type identification of different cells. Therefore, for the cell text sequence , the key value type of the cell The confidence sequence of is: Calculate the key value type of the cell Confidence: For cell text sequences middle , which belongs to the key value type of the cell The confidence level is: 。 4. The PDF file table parsing method according to claim 1, characterized in that: In the saving step, specifically: By identifying the key value type of the cell in the table, traverse each row of cell types line by line: If all cells in the cell key value sequence are of type key, all cells in this row are cached; If the cell key sequence middle, and The text is the same, that is, it belongs to a cross-row cell, then it is ignored Corresponding cell type, and convert the format to ; If the cell key sequence middle, and If the text is not the same, As the parent key, As The key of the subordinate node of ; If the cell key sequence Only the first cell in For key, others is val, and are all keys, then As primary key, As The key of the subordinate node of As The key of the subordinate node of is the value corresponding to it.
Citation Information
Patent Citations
Multi-modal intelligent question-answering system based on large model and construction method and device
CN119783819A
System and method for facilitating content display on portable devices
US20090125802A1
Cited By
Table identification method and electronic equipment
CN120764508A