A method for parsing tables in PDF files

By converting PDF files into pictures and using OCR model and named entity recognition technology, the accuracy problem of nested table analysis is solved, and structured table data is efficiently extracted from PDF files.

CN120104577BActive Publication Date: 2025-07-08GUANGZHOU ZIWEIYUN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510580825.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-07-08
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively parse nested tables in PDF files, especially PDF files in scanned or image formats, resulting in inaccurate extraction of table content.

Method used

Convert PDF files into image format, use pre-trained OCR models to detect table areas and crop them, combine named entity recognition technology to identify cell types, and convert them into structured tables through HTML format, and finally save them to Json format.

Benefits of technology

Improve the accuracy and efficiency of nested table parsing, and can effectively restore the table structure information in scan or image PDF files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104577B_ABST
    Figure CN120104577B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for parsing tables in PDF files, comprising the following steps: a step of converting a PDF document into a picture format: converting the pages of the PDF document to be parsed into a picture format; a step of detecting table areas: inputting the converted picture into a pre-trained table recognition OCR model, which can recognize the range of the table area frames in the picture; a step of cropping table areas: cropping the original picture according to the table position coordinates P provided by the model, and only retaining the picture of the table area frame part. A step of converting the picture into an HTML format: using a pre-trained OCR model to recognize the table space structure, the cell text information and the structural features contained in the table picture, and then based on the extracted table space structure, cell text information and structural features, converting it from the picture format to structured table data and further converting it into a table in HTML format. The present invention can improve the accuracy and efficiency of extracting table information from PDF files.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of PDF file processing, and particularly to a method for parsing tables in PDF files. Background Art

[0002] PDF (Portable Document Format) is a file format developed and widely adopted by Adobe Systems in 1993. The original design intention of PDF is to achieve cross-platform consistent display of documents, that is, regardless of the hardware configuration or software environment, it can ensure that the display effect of the document matches the original intention during creation. By encapsulating information such as text, fonts, images, and document layout in a single file, this format ensures the integrity and consistency of the document, making it an ideal choice for electronic document exchange.

[0003] However, this characteristic of PDF files also brings some challenges, especially when dealing with nested tables. A nested table refers to a complex layout where one or more tables are contained within the cells of another table. Since the original design intention of PDF files is to display the visual representation of documents rather than store structured data, PDF files do not always retain the original structure information of the document when saved. This is in sharp contrast to office software (such as Microsoft Word or Excel), which can store and retain the original structure information of the document, making it more direct and accurate to extract structured data from these files.

[0004] When parsing nested tables in PDF files, the problems become more complex. Currently, there are various methods and technologies for parsing nested tables in PDF. In the Python programming language, there are several popular libraries that can assist in this task, including PyPDF2, pdfminer.six, and pdfplumber, etc. These libraries provide basic functions for parsing tables in PDF files, but their performance is not ideal when dealing with complex layouts such as nested tables. Especially when the PDF file is an image file obtained by scanning a paper document, these libraries may miss the recognition of the table structure, resulting in the inability to accurately extract the table content.

[0005] To solve this problem, OCR (Optical Character Recognition) technology is introduced into PDF form parsing. OCR technology can extract text information and structural information from images, which is particularly important for PDF files in scanned or image form. However, OCR technology faces many challenges when parsing and converting unstructured tabular data into structured tabular data. Although OCR can recognize text information, it is difficult to parse and convert unstructured tabular data into structured tabular data with high accuracy. At the same time, it is difficult to deduce the relationship between cells and restore data based on text information and structural information. Summary of the Invention

[0006] The object of the present invention is to provide a method for parsing tables in PDF files, which can improve the accuracy and efficiency of extracting table information from PDF files.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A method for parsing tables in PDF files, comprising the following steps:

[0009] Step of converting PDF document to image format: Convert the page of the PDF document to be parsed into image format;

[0010] Step of detecting table area: Input the converted image into a pre-trained table recognition OCR model, which can recognize the range of the table area box in the image and output the coordinates of the table area box in the image , where is the i-th coordinate point of the table area box, is the confidence of the i-th coordinate point;

[0011] Step of cropping table area: Crop the original image according to the table position coordinates P provided by the model, and only retain the image of the table area box part;

[0012] Step of converting image to HTML format: Use a pre-trained OCR model to recognize the table space structure and the cell text information and structural features contained in the table image, and then based on the extracted table space structure and cell text information and structural features, convert it from image format to structured tabular data and further convert it to an HTML format table;

[0013] Steps for cross-cell conversion in a nested table in HTML format: In a table in HTML format, a table is defined using the table tag, rows in the table are defined using the tr tag, standard cells are defined using the td tag, the number of rows a table cell spans is defined using the rowspan attribute, and the number of columns a table cell spans is defined using the colspan attribute;

[0014] Steps for identifying the key-value type to which a cell belongs: Use a pre-trained named entity recognition model to identify the key-value type to which a cell belongs:

[0015] Saving step: Pair adjacent cells in pairs and save them in Json format.

[0016] Specifically, in the steps for cross-cell conversion in the nested table in the HTML format:

[0017] Traverse the tr tags in the table tag, read one by one the cells containing the td tag, and determine whether the cell contains the rowspan attribute or the colspan attribute:

[0018] If none of the td cells contain the rowspan attribute or the colspan attribute, directly read the information corresponding to the td cell, and mark the cell data as , where c is the number of td cells;

[0019] If the cell contains the colspan attribute, obtain the value m of the colspan attribute and the information corresponding to the td cell, and mark the cell data as , where c is the number of td cells;

[0020] If the cell contains the colspan attribute, obtain the value n of the colspan attribute and the information corresponding to the td cell, and mark the cell data as , where c is the number of td cells; for i < k < i + n - 1, there is ;

[0021] Specifically, in the steps for identifying the key-value type to which a cell belongs, specifically:

[0022] Extraction of the cell text series under the tr tag: Traverse the tr tags in the table tag row by row, and splice the text of the td cells contained under the tr tag into a cell text sequence , where c is the number of td cells contained in the i-th tr tag;

[0023] Generate cell text series annotations using a pre-trained model: Use a pre-trained model for the cell text sequence Encode and generate sequence representations ;

[0024] Obtain the start vector representation and end vector representation of the entity segment: Input the sequence representation into the linear layers in the key-value type prediction layer of the cell respectively and the linear layer to obtain the start vector representation and end vector representation of the information of each entity segment in the sequence representation:

[0025]

[0026]

[0027] Among them, the start vector representation is the vector representation of the start character of the entity segment, and the end vector representation is the vector representation of the end character of the entity segment, is the key-value type of the cell, , is the set of all key-value types, which is equivalent to the start vector representation and end vector representation recognized by different key-value types of cells,

[0028] Therefore, for the cell text sequence , the confidence sequence of the key-value type of the cell is:

[0029]

[0030]

[0031] Calculate the confidence of the key-value type of the cell: For the cell text sequence in , the confidence of belonging to the key-value type of the cell is:

[0032]

[0033] Specifically, in the saving step, specifically:

[0034] By identifying the key-value type of the cell in the table, traverse each row of cell types row by row:

[0035] If all cell types in the cell key-value sequence are keys, cache all cells in this row;

[0036] If in the cell key-value sequence , and If the text is the same, that is, it belongs to a multi-line cell, it is ignored Corresponding cell type and convert the format to ;

[0037] If the cell key value sequence In And The text is different, then As the superior key As The key of the subordinate node of, and convert the format to ;

[0038] If in the cell key value sequence Only the first one in the cell is Is the key, and the others Are vals, and Are all keys, then As the primary key As The key of the subordinate node of, As The key of the subordinate node of, Is the corresponding value for it

[0039] Compared with the prior art, the beneficial effects of the present invention are:

[0040] The present invention provides an innovative solution. The invention first converts a PDF file into a picture format, then uses a pre-trained OCR model to extract the table information in the converted picture, and converts the extracted table information into an HTML format. Next, it uses named entity recognition technology to identify the type of each cell in each row of the table. Finally, by combining the relative positions and types of each cell in each row, the corresponding cells are paired in pairs to form a complete key-value pair and saved in Json format. Please refer to Figure 1 .

[0041] For a table in the form of a scanned file or picture in a PDF file, the present invention can use an OCR model to obtain the structural information and text content in its table. For a complex table layout such as a nested table, the present invention uses named entity recognition technology to distinguish the text corresponding types of each cell in each row, and effectively restores the table information by combining the relative position and recognition attributes of the cell, improving the accuracy of table parsing BRIEF DESCRIPTION OF THE DRAWINGS

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0043] Figure 1 is the flowchart of the present invention;

[0044] Figure 2 is an example of a table;

[0045] Figure 3 is Figure 2 the conversion process of the shown table example;

[0046] Figure 4 is a key-value table;

[0047] Figure 5 is Figure 2 the complete conversion process of the shown table example. Detailed implementation manners

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments.

[0049] See Figures 1 to 5 , a method for parsing a table in a PDF file, including the following steps:

[0050] PDF document to image format conversion step: Use the pdf2image library in the Python programming language to convert the pages of the PDF document to be parsed into an image format, such as PNG or JPEG. This step converts the visual elements such as text, images, and tables in the PDF page into pixel data, providing input for the subsequent table recognition model.

[0051] Table area detection step: Input the converted image into a pre-trained table recognition OCR model, which can identify the range of the table area box in the image and output the coordinates of the table area box in the image , where is the i-th coordinate point of the table area box, is the confidence of the i-th coordinate point.

[0052] Table area cropping step: According to the table position coordinates P provided by the model, crop the original image, and only retain the image of the table area frame part. This step is the key to the present invention, which can effectively prevent the influence of other elements such as pure text or illustrations in the document on the subsequent table recognition, thereby improving the recognition accuracy.

[0053] Image to HTML format conversion step: Use a pre-trained OCR model to recognize the table space structure and the cell text information and structural features contained in the table image, and then based on the extracted table space structure and cell text information and structural features, convert it from the image format to structured table data, and further convert it into an HTML format table.

[0054] Cross-cell conversion step in nested tables in HTML format: In an HTML format table, use the table tag to define a table, use the tr tag to define the rows in the table, use the td tag to define standard cells, use the rowspan attribute to define the number of rows that a table cell spans, and use the colspan attribute to define the number of columns that a table cell spans;

[0055] Before converting the HTML format table into structured table information, it is necessary to consider the nested table conversion problem, that is, the cross-cell conversion problem in the table. Please refer to Figure 3 . The specific conversion process of the tags can be implemented according to the following steps:

[0056] Traverse the tr tags in the table tag, read the cells containing the td tag one by one, and judge whether the cell contains the rowspan attribute or the colspan attribute:

[0057] If none of the td cells contain the rowspan attribute or the colspan attribute, directly read the information corresponding to the td cell, and mark the cell data as , where c is the number of td cells;

[0058] If the cell contains the colspan attribute, obtain the colspan attribute value m and the information corresponding to the td cell, and mark the cell data as , where c is the number of td cells;

[0059] If the cell contains the colspan attribute, obtain the colspan attribute value n and the information corresponding to the td cell, and mark the cell data as , where c is the number of td cells; for i < k < i + n - 1, there is ;

[0060] Cell key-value type identification steps: Use a pre-trained named entity recognition model to identify the key-value type to which the cell belongs. Specifically:

[0061] Extraction of cell text series under tr tag: Traverse the tr tags in the table tag line by line, and splice the text of the td cells contained under the tr tag into a cell text sequence , where c is the number of td cells contained in the i-th tr tag;

[0062] Generate cell text series annotations using a pre-trained model: Use a pre-trained model to encode the cell text sequence and generate a sequence representation ;

[0063] Obtain the start vector representation and end vector representation of the entity segment: Input the sequence representation into the linear layer in the cell key-value type prediction layer and the linear layer respectively, to obtain the start vector representation and end vector representation of the information of each entity segment in the sequence representation:

[0064]

[0065]

[0066] Among them, the start vector representation is the vector representation of the start character of the entity segment, and the end vector representation is the vector representation of the end character of the entity segment is the key-value type to which the cell belongs, , is the set of all key-value types, which is equivalent to the start vector representation and end vector representation of different cell key-value type identifications,

[0067] Therefore, for the cell text sequence , the confidence sequence of the key-value type to which the cell belongs is:

[0068]

[0069]

[0070] Calculate the confidence of the key-value type to which the cell belongs: For the cell text sequence in , the confidence of belonging to the key-value type to which the cell belongs is:

[0071]

[0072] Saving steps: Pair adjacent cells and save them in Json format. Specifically, identify the key value type of the cell in the table, such as Figure 5 As shown in step 3, traverse each row of cell types line by line:

[0073] If the cell key sequence If all cell types in the row are key, all cells in this row will be cached;

[0074] If the cell key sequence middle, and The text is the same, that is, it belongs to a cross-row cell, then it is ignored Corresponding cell type, and convert the format to ;

[0075] If the cell key sequence middle, and If the text is not the same, As the parent key, As The key of the subordinate node of ;

[0076] If the cell key sequence Only the first cell in For key, others is val, and If both are keys, As primary key, As The key of the subordinate node of As The key of the subordinate node of is the value corresponding to it.

[0077] The converted cell Json format is as follows Figure 5 As shown in step 4

[0078] The beneficial effects of the present invention are:

[0079] The present invention provides an innovative solution. First, the PDF file is converted into a picture format, and then the pre-trained OCR model is used to extract the table information in the converted picture, and the extracted table information is converted into HTML format. Next, the named entity recognition technology is used to identify the type of each cell in each row of the table. Finally, by combining the relative positions and types of the cells in each row, the corresponding cells are paired in pairs to form complete key-value pairs, and saved in Json format. Please refer to Figure 1 。

[0080] For the tables in the form of scanned files or pictures of PDF files, the present invention can use the OCR model to obtain the structural information and text content in the tables. For the complex table layout of nested tables, the present invention uses the named entity recognition technology to distinguish the corresponding types of the text in each cell in each row, and effectively restores the table information by combining the relative position and recognition attributes of the cell, improving the accuracy of table parsing.

[0081] The above are only the preferred embodiments of the present invention, and there is no limitation to the present invention in any form. Although the present invention has been disclosed above with the preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to the above-disclosed technical content within the scope of the technical solution of the present invention to obtain equivalent embodiments with equivalent changes. However, as long as it does not depart from the technical solution content of the present invention, any brief modifications, equivalent changes and modifications made to the above embodiments according to the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A method for parsing a table in a PDF file, characterized in that, Including the following steps: PDF document to image format conversion step: converting the pages of the PDF document to be parsed into image format; Table area detection steps: Input the converted image into a pre-trained table recognition OCR model, which can identify the range of the table area box in the image and output the coordinates of the table area box in the image , where is the i-th coordinate point of the table area box, is the confidence of the i-th coordinate point; Table area cropping step: cropping the original image according to the table position coordinates P provided by the model, and only retaining the image of the table area frame part; Image to HTML format conversion step: using a pre-trained OCR model to recognize the table space structure, the cell text information and the structural features contained in the table image, and then based on the extracted table space structure, cell text information and structural features, converting it from image format to structured table data, and further converting it into an HTML format table; Cross-cell conversion in the HTML format nested table step: in the HTML format table, using the table tag to define a table, using the tr tag to define the rows in the table, using the td tag to define standard cells, using the rowspan attribute to define the number of rows spanned by the table cell, and using the colspan attribute to define the number of columns spanned by the table cell; Cell key-value type identification step: using a pre-trained named entity recognition model to identify the key-value type to which the cell belongs, specifically: Extraction of cell text sequence under tr tag: Traverse the tr tags in the table tag line by line, and splice the text of the td cells contained under the tr tag into a cell text sequence where c is the number of td cells contained in the i-th tr tag; Generate cell text sequence annotations using a pre-trained model: Use a pre-trained model to encode the cell text sequence and generate a sequence representation ; Obtain the start vector representation and end vector representation of the entity segment: Input the sequence representation into the linear layers in the key-value type prediction layer of the cell and the linear layer respectively, and obtain the start vector representation and end vector representation of the information of each entity segment in the sequence representation: Among them, the start vector represents the vector representation of the start character of the entity segment, and the end vector represents the vector representation of the end character of the entity segment. is the key-value type to which the cell belongs. , is the set of all key-value types, which is equivalent to the start vector representation and end vector representation for identifying the key-value types to which different cells belong. Therefore, for the cell text sequence , the confidence sequence of the key-value type to which the cell belongs is as follows: Calculate the confidence of the key-value type to which the cell belongs For the cell text sequence in , the confidence of belonging to the key-value type to which the cell belongs is as follows: ; Saving step: pairing adjacent cells in pairs and saving them in Json format.

2. The PDF file table parsing method according to claim 1, wherein: In the cross-cell conversion step in the HTML format nested table: Traverse the tr tags in the table tag, read the cells containing the td tag one by one, and judge whether the cell contains the rowspan attribute or the colspan attribute: If none of the cells contain the rowspan or colspan attribute, directly read the information corresponding to the cell and mark the cell data as , where c is the number of cells; If the cell contains the colspan attribute, obtain the value m of the colspan attribute and the corresponding information of the td cell, and mark the cell data as , where c is the number of td cells; If the cell contains the colspan attribute, obtain the value n of the colspan attribute and the corresponding information of the td cell, and mark the cell data as , where c is the number of td cells; for i < k < i + n - 1, there is .

3. The PDF file table parsing method according to claim 1, wherein In the saving step, specifically: By identifying the key-value type to which the cells in the table belong, traverse each row of cell types row by row: If all cell types in the cell key value sequence are keys, cache all cells in this row; If the cell key value sequences in and the text is the same, that is, it belongs to a multi-line cell, then ignore the corresponding cell type and convert the format to ; If the cell key value sequences in and the text is not the same, then is used as the superior key, is used as the key of the subordinate node of, and the format is converted to ; If the cell key value sequence only the first one in the cell is the key, and the others are the vals, and all are keys, then take as the primary key, as the key of the subordinate node of as the key of the subordinate node of and use it as the corresponding value.

Citation Information

Patent Citations

  • Multi-modal intelligent question-answering system based on large model and construction method and device

    CN119783819A

  • System and method for facilitating content display on portable devices

    US20090125802A1