A method and device for identifying tables in PDF documents
By executing the software module on the terminal device and using the special table recognition module to identify special tables in the PDF document, the problem of low accuracy in table recognition in the prior art is solved, and the accurate identification and extraction of tables in the PDF document is achieved.
Patent Information
- Application Number
- CN202111209477.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-18
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2041-10-18
AI Technical Summary
The prior art is not very accurate when identifying tables in PDF documents, especially the identification of special tables without crossed lines is inaccurate.
By executing the software module on the terminal device, obtaining the original target page of the PDF document, and using the special table identification module to identify whether there are special tables, including wireless box tables and tables with only horizontal lines. This module is identified using camelot's stream module and Hough transform module.
It improves the recognition accuracy of tables in PDF documents, especially when identifying special tables, and can accurately extract the table contents, avoiding the low accuracy problem caused by recognition failure.
Smart Images

Figure CN114067323B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a method and device for extracting a table from a PDF document. Background Art
[0002] As we all know, Portable Document Format (PDF) has been widely used in various industries such as finance, IT, electronics and education. Each industry accumulates a large amount of PDF document data, which records a large amount of text, pictures, tables and other information. In some cases, it is necessary to extract some key information from PDF documents, which may be text or tables. For example, there are a large number of tables in some PDF prospectuses. There are existing methods for identifying tables, but the recognition accuracy is not high.
[0003] Therefore, how to improve the recognition accuracy of tables in PDF documents is an urgent problem that needs to be solved. Summary of the invention
[0004] The present application provides a method and device for identifying tables in PDF documents, which can improve the recognition accuracy of tables in PDF documents.
[0005] In a first aspect, a method for identifying a table in a PDF document is provided, comprising: obtaining an original target page of a PDF document; when it is identified that no regular table exists in the original target page, or when the recognition result of the regular table in the original target page does not meet a preset condition, identifying whether a special table exists in the original target page through a special table recognition module, wherein the regular table refers to a table with crossed lines, the special table refers to a table without crossed lines, and the special table recognition module is a module for identifying a table without crossed lines.
[0006] The above method can be executed by a software module on a terminal device. The software module obtains a PDF document from a server. When a regular table is not successfully identified in the original target page of the PDF document, the software module will again identify whether there is a special table in the original target page through a special table identification module. Compared with the prior art, which inaccurately identifies special tables in PDF documents and causes the special tables to be unable to be accurately extracted, the present application can accurately extract the special tables in the PDF document through the special table identification module, thereby improving the recognition accuracy of the tables in the PDF document.
[0007] Optionally, the special table includes: a wireless frame table and a table with only horizontal lines, the special table identification module includes: a stream module and a Hough transform module of Camelot, and the method further includes: when the stream module determines that the special table exists in the original target page, and when the stream module determines that the special table is the wireless frame table, the content in the wireless frame table is identified by the stream module; when the Hough transform module determines that the special table exists in the original target page, and when the Hough transform module determines that the special table is a table with only horizontal lines, the content of the wireless frame area between adjacent horizontal lines in the table with only horizontal lines is identified by the stream module. It can be seen that the special tables existing in the PDF document can be accurately identified by the stream module and the Hough transform module of Camelot, thereby avoiding the problem of low accuracy of table recognition in the PDF document due to failure in special table recognition.
[0008] Optionally, all text elements in the original target page are extracted by a text parsing module; when the difference between the coordinates of the first text element in all the text elements and the coordinates of the row-level endpoint element is greater than a first distance threshold, the first text element is determined to be a target text element, the row-level endpoint element is any text element in all the text elements that is closest to the left edge of the original target page, and the row where the first text element is located only includes the first text element; when the difference between the coordinates of the second text element in all the text elements and the coordinates of the third text element is greater than a second distance threshold, the second text element and the third text element are determined to be target text elements, and the second text element and the third text element are text elements in the same row. The text parsing module is used to identify all text elements in the original target page and the corresponding coordinate positions, and based on these coordinate positions, the text elements that truly belong to the table can be accurately determined, thereby further improving the recognition accuracy of the table in the PDF document.
[0009] Optionally, when the special table exists in the original target page, the special table recognition module recognizes the recognition result of the special table in the original target page as a first result, and the text parsing module recognizes the recognition result of the special table in the original target page as a second result, and the method further includes: when the first result is the same as the second result, filling the first result or the second result into the converted special table; when the first result is different from the second result, identifying the text elements of the coordinate positions of the first result and the second result through OCR, and filling the text elements identified by OCR into the converted special table. It can be seen that when the first result is different from the second result, using the OCR recognition result to determine the text elements in the converted special table can effectively avoid the situation where the text elements in different cells in the special table appear serially.
[0010] Optionally, a table label and row and column attributes are set for the converted special table, the table label is used to display the converted special table, and the row and column labels are used to indicate the number of rows and columns of the converted special table.
[0011] Optionally, when the special table does not exist in the original target page, a screenshot is taken of the area where the first text element, the second text element and the third text element are located.
[0012] For example, the area where the first text element, the second text element and the third text element are located is a statistical chart or a process structure diagram. By taking a screenshot of the area, some non-tabular graphic information can be identified, thereby avoiding the problem of missing non-tabular graphic information in the PDF.
[0013] Optionally, the recognition result of the regular table in the original target page does not meet the preset condition, including: the recognition accuracy of the regular table in the original target page is lower than the accuracy threshold. Evaluating the recognition result of the regular table using the recognition accuracy can ensure the accuracy of the recognition result.
[0014] Optionally, before identifying whether a special table exists in the original target page by using a special table identification module, the method further includes: identifying whether a regular table exists in the original target page by using a lattic module of Camelot.
[0015] In a second aspect, a device for identifying a table in a PDF document is provided, comprising a module for executing any one of the methods in the first aspect.
[0016] According to a third aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes any one of the methods according to the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0018] Figure 1 A schematic diagram of a method for identifying a table in a PDF document according to an embodiment of the present invention;
[0019] Figure 2 This is a schematic diagram of a table structure in an embodiment of the present invention;
[0020] Figure 3 A schematic diagram of a non-tabular structure in an embodiment of the present invention;
[0021] Figure 4 It is a schematic diagram of the process of parsing a table by the camelot table parsing module in an embodiment of the present invention;
[0022] Figure 5 A schematic diagram of a process of parsing a table by a camelot table parsing module according to an embodiment of the present invention;
[0023] Figure 6 It is a schematic diagram of the process of parsing a special table in stream mode of Camelot in an embodiment of the present invention;
[0024] Figure 7 A schematic diagram of the process of generating Html tag data in an embodiment of the present invention;
[0025] Figure 8 Schematic diagram of a device for identifying a table in a PDF document according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0027] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0028] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0029] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0030] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0031] The present application is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0032] In order to solve the problem of low recognition accuracy of tables in PDF documents, this application proposes a method for recognizing tables in PDF documents, which can improve the recognition accuracy of tables in PDF documents. Figure 1 As shown, the method can be executed by a software module of a terminal device, and the method includes:
[0033] S101, obtaining an original target page of a PDF document.
[0034] Exemplarily, the software module obtains the PDF document to be analyzed from the PDF document storage server, then recognizes the text and table in the original target page of the PDF document, and stores and displays the recognition results in other forms, for example, storing the recognition results in the form of Html tags and displaying them in the form of web pages to facilitate user copying and pasting.
[0035] S102, when it is identified that there is no regular table in the original target page, or when the recognition result of the regular table in the original target page does not meet the preset conditions, a special table recognition module is used to identify whether there is a special table in the original target page, wherein a regular table refers to a table with crossed lines, a special table refers to a table without crossed lines, and the special table recognition module is a module for identifying tables without crossed lines.
[0036] Exemplarily, the lattic module of Camelot is used to identify whether there is a regular table in the original target page. The above Camelot is an algorithm for extracting table data from PDF, wherein the lattic module (i.e., lattic mode) of Camelot is used to identify regular tables of wireframe type. Regular tables refer to tables with crossed lines, such as Figure 2 As shown, 201 shows a representation of a conventional table. For example, the lattic module identifies the horizontal and vertical lines of the conventional table in the original target page by setting the line_scale parameter, and determines the cells in the table by obtaining the intersection of the horizontal and vertical lines, thereby parsing the table structure of the conventional table and identifying the text content in the conventional table. The above-mentioned line_scale parameter is used by the lattic module to determine the length of the horizontal and vertical lines. The larger the line_scale parameter value is set, the shorter the length of the horizontal and vertical lines that the lattic module can identify. For example, the line_scale parameter can be set to 40, or other values can be set. The specific setting of the line_scale parameter value can be adjusted according to actual needs, and this application does not impose any restrictions on this.
[0037] Exemplarily, the recognition result of the regular table in the original target page does not meet the preset condition, including: the recognition accuracy of the regular table in the original target page is lower than the accuracy threshold. The preset condition includes: the recognition accuracy threshold and the table blank rate threshold. When there is returned data after the lattic module identifies that there is a regular table in the PDF document, wherein the returned data includes recognition accuracy and / or blank rate (i.e., recognition accuracy, or blank rate, or recognition accuracy and blank rate), when the recognition accuracy of the lattic module identifying that there is a regular table in the PDF document is lower than the accuracy threshold, it is determined that there is no regular table in the PDF document (i.e., recognition of the presence of a regular table in the PDF document fails); when the blank rate of the lattic module identifying that there is a regular table in the PDF document is lower than the blank rate threshold, it is determined that there is no regular table in the PDF document (i.e., recognition of the presence of a regular table in the PDF document fails); when the recognition accuracy of the lattic module identifying that there is a regular table in the PDF document is lower than the accuracy threshold, and when the blank rate of the lattic module identifying that there is a regular table in the PDF document is lower than the blank rate threshold, it is determined that there is no regular table in the PDF document (i.e., recognition of the presence of a regular table in the PDF document fails). The above recognition accuracy refers to the precision value of the lattic module identifying that there is a regular table in the PDF document; the above blank rate refers to the ratio of the spaces in the regular table identified by the lattic module to the total number of spaces in the original target page.
[0038] Exemplarily, when it is recognized that there is no regular table in the original target page, or when the recognition result of the regular table in the original target page does not meet the preset conditions, the software module can identify whether there is a special table in the original target page through a special table recognition module, wherein the special table refers to a table without cross lines, and the special table recognition module is a module for identifying a table without cross lines. The above software module can accurately extract the special table existing in the PDF document through the special table recognition module, thereby improving the recognition accuracy of the table in the PDF document.
[0039] Exemplarily, the special table includes: a wireless frame table and a table with only horizontal lines, and the special table identification module includes: a stream module and a Hough transform module of Camelot; when the stream module determines that there is a special table in the original target page, and when the stream module determines that the special table is a wireless frame table, the stream module is used to identify the content in the wireless frame table; when the Hough transform module determines that there is a special table in the original target page, and when the Hough transform module determines that the special table is a table with only horizontal lines, the stream module is used to identify the content of the wireless frame area between adjacent horizontal lines in the table with only horizontal lines. For example, when the stream module recognizes that there is a wireless frame table in the original target page, the stream module will identify the table structure of the wireless frame table and the text content in the wireless frame table, such as Figure 2 As shown, 202 shows a form of a wireless frame table; when the Hough transform module detects that there is a table with only horizontal lines on the original target page, the Hough transform module obtains the area where the horizontal lines are located (that is, the area where the table with only horizontal lines is located), and then the stream module identifies the table structure of the table with only horizontal lines and the text content in the table with only horizontal lines, such as Figure 2 As shown, 203 shows a table with only horizontal lines.
[0040] Exemplarily, all text elements in the original target page are extracted through a text parsing module; when the difference between the coordinates of the first text element among all text elements and the coordinates of the row-level endpoint element is greater than a first distance threshold, the first text element is determined to be the target text element, the row-level endpoint element is any text element among all text elements that is closest to the left edge of the original target page, and the row where the first text element is located only includes the first text element; when the difference between the coordinates of the second text element among all text elements and the coordinates of the third text element is greater than a second distance threshold, the second text element and the third text element are determined to be the target text elements, and the second text element and the third text element are text elements in the same row.
[0041] The above-mentioned text parsing module can be an Itext parsing module, which is used to extract all text elements in a PDF document, wherein a text element refers to the text content in the original target page of a PDF document, and the text element can be a single character or a character group; for example, a single character is a Chinese character, and a character group is a four-character idiom or a phrase consisting of two Chinese characters. This application does not impose any restrictions on the length of the character group, and the length of the character group can be set according to actual conditions.
[0042] The difference between the coordinates of the first text element and the coordinates of the line-level endpoint element refers to the distance between the first text element and the line-level endpoint element, which is a positive integer greater than 0; and the difference between the coordinates of the second text element and the coordinates of the third text element refers to the distance between the second text element and the third text element, which is also a positive integer greater than 0.
[0043] The first distance threshold refers to the maximum value of the horizontal coordinates of the text elements in each row; the second distance threshold refers to the maximum difference between the horizontal coordinates of the text elements in each row. The target text elements refer to text elements that do not have a regular text format, and the target text elements include elements in tables, elements in statistical charts, elements in flow charts, and elements in scatter plots.
[0044] For example, the Itext parsing module obtains the target PDF document to be analyzed and the name and unique identifier (ID) of the target PDF document from the server where the PDF source is located. The ID is used to distinguish the target PDF document from other non-target PDF documents. After that, the Itext parsing module identifies all text elements in the original target page of the target PDF document through a replication interface (e.g., a RenderListener interface), and also identifies the coordinates, text family name, text size, and boldness of these text elements. The text family name refers to the font of the text element, such as Songti or Heiti; the text size refers to the font size of the text element, such as size 4 or 5; and boldness refers to whether the font of the text element is bold.
[0045] The Itext parsing module performs row-level summarization on all identified text elements. For example, the Itext parsing module sorts the text elements according to the vertex ordinate of each text element, and groups the ordinates of different text elements that meet the preset rules into one row to complete the row-level summarization of all text elements. Since the vertex ordinates of different text elements in certain rows of a PDF document may not be completely consistent, the above preset rules can be set to group these different text elements into one row when the vertex ordinate spacing of different text elements does not exceed a preset distance, wherein the preset distance refers to the maximum difference in the vertex ordinate spacing of different text elements, for example, the preset distance can be 5 pixels or 3 pixels, and the preset distance can be set according to actual conditions, and this application does not limit this.
[0046] After the Itext parsing module summarizes all text elements at the line level, it uses the line-level text positioning algorithm to perform special line-level text element identification on the line-level text elements after line-level summarization. The line-level text positioning algorithm is as follows:
[0047] When a row-level text element contains only one text element (i.e., the first text element), if the Itext parsing module determines that the difference between the coordinates of the first text element and the coordinates of the row-level endpoint element is greater than the first distance threshold, the first text element is determined to be a special row-level text element (i.e., the target text element); when a row-level text element contains more than one text element, if the Itext parsing module determines that the difference between the coordinates of the second text element in the row and the coordinates of the third text element is greater than the second distance threshold, the second text element and the third text element are determined to be special row-level text elements (i.e., the target text element), and the second text element and the third text element are text elements in the same row.
[0048] For example, when row A contains only one text element C1, the Itext parsing module will traverse all the row-level elements in the original target page to find a text element among all row-level elements that is closest to the left edge of the original target page and the horizontal coordinate corresponding to the text element, for example, the horizontal coordinate corresponding to the text element is X1; based on this X1 coordinate, the setting parameters in the Itext parsing module are set to n (for example, n is 30 pixels), and the first distance threshold is X1+n; the Itext parsing module determines whether the horizontal coordinates of all row-level elements in the original target page are greater than X1+n. If the Itext parsing module determines that the horizontal coordinate of a row-level element is greater than X1+n, the row-level element A is determined to be a special row-level text element. If the Itext parsing module determines that the horizontal coordinate of a row-level element is less than or equal to X1+n, the row-level element is determined to be a non-special row-level text element. The non-special row-level text element refers to a text element in a regular text format in a PDF document.
[0049] For another example, when there is more than one text element in row B, for example, the second distance threshold is 10 pixels, there are two text elements in row B, and the two text elements are the second text element and the third text element, respectively, wherein the horizontal coordinate of the second text element is X2, and the horizontal coordinate of the third text element is X3, when the Itext parsing module determines that the difference between the coordinates of the second text element and the coordinates of the third text element (that is, the spacing between the second text element and the third text element (that is, X2-X3)) is greater than the second distance threshold (that is, 10), the row-level element of B is determined to be a special row-level text element; when the Itext parsing module determines that the difference between the coordinates of the second text element and the coordinates of the third text element (that is, the spacing between the second text element and the third text element (that is, X2-X3)) is less than or equal to the second distance threshold (that is, 10), the row-level element of B is determined to be a non-special row-level text element.
[0050] Exemplarily, when there is a special table in the original target page, the special table recognition module recognizes the recognition result of the special table in the original target page as the first result, and the text parsing module recognizes the recognition result of the special table in the original target page as the second result. When the first result is the same as the second result, the first result or the second result is filled into the converted special table; when the first result is different from the second result, the text elements of the coordinate positions of the first result and the second result are recognized by OCR, and the text elements recognized by OCR are filled into the converted special table.
[0051] For example, when there is a special table in the original target page, the special table recognition module is used to recognize the table structure and text elements of the special table in the original target page as the first result; and the text parsing module recognizes the text elements of the special table in the original target page as the second result; the coordinate integration processing module (i.e., the coordinate integration and processing module) in the software module determines whether the text elements in the first result are the same as the text elements in the second result according to the coordinate position of the table structure (i.e., the converted special table) and the coordinate position of the special row-level text elements; if the coordinate integration processing module determines that the text elements in the first result are the same as the text elements in the second result, the text elements in the first result or the text elements in the second result are filled in into the table structure of the first result (i.e. the converted special table); if the coordinate integration processing module determines that the text elements in the first result are not the same as the text elements in the second result, the software module performs OCR recognition on the text elements at the coordinate position of the first result and the text elements at the coordinate position of the second result through the OCR module, and fills the text elements at the coordinate position of the first result recognized by OCR into the table structure of the first result (i.e. the converted special table), or fills the text elements at the coordinate position of the second result recognized by OCR into the table structure of the second result (i.e. the converted special table), wherein the above-mentioned coordinate integration processing module is used for table structure analysis and filling text elements in the table structure. For example, table structure analysis can be used to analyze the coordinate position of the table structure and the coordinate position of the text elements in the table structure. It can be seen that when the text elements in the first result are not the same as the text elements in the second result, the text elements in the converted special table are determined by using the OCR recognition results, which can effectively avoid the situation where the text elements of different cells in the special table appear in series.
[0052] Exemplarily, when the special table does not exist in the original target page, the area where the first text element, the second text element and the third text element are located is captured. When the special table does not exist in the original target page, it indicates that the special row-level text element is not an element in the special table. For example, the first text element, the second text element and the third text element are not special row-level text elements. The first text element, the second text element and the third text element may be text elements in some statistical charts, flow charts and scatter plots, such as Figure 3 As shown, 301 is a display form of a statistical chart, and 302 is a display form of a flowchart. The above-mentioned special row-level text elements may be text elements such as category 1, category 2, category 3 and category 4 in 301, or may be text elements such as start, get data, yes, no, etc. in 302. When the first text element, the second text element and the third text element are not special row-level text elements, the software module can take a screenshot of the area where the first text element, the second text element and the third text element are located, and process the screenshot result in the manner of image recognition. The area refers to the area formed by the coordinate position corresponding to the first text element, the coordinate position corresponding to the second text element and the coordinate position corresponding to the third text element, for example, the area is the area where the statistical chart displayed by 301 is located, or the area where the flowchart displayed by 302 is located.
[0053] Exemplarily, table tags and row and column attributes are set for the converted special table. The table tag is used to display the converted special table, and the row and column tags are used to indicate the number of rows and columns of the converted special table. The above-mentioned converted special table refers to the special table existing in the PDF document after it is recognized by the special table recognition module in the software module. The software module sets the table tag (for example,
[0054] Figures 4 to 8
[0055] Figure 4 Figure 5
[0056] Figure 6
[0057] Figure 7
[0058] Figure 8 Figure 8
[0059] Figure 1
[0060]
[0061]
[0062]
[0063]
[0064]
[0065]
[0066]
[0067]
[0068]
[0069]
[0070]
[0071]
[0072] Tags) and set row attribute parameters (e.g., rowspan parameters) and column attribute parameters (e.g., colspan parameters) according to the contents of the merged cells, and display the converted special table in Html format to facilitate user copy and paste. For ease of understanding, the following is an exemplary description of the overall process steps of the method for identifying tables in PDF documents provided by this application. As shown, the software module obtains the PDF document to be analyzed from the PDF document storage server (i.e., the PDF source), and performs Camelot table parsing on the PDF. For example, the complete wireframe table (i.e., regular table) is parsed (i.e., recognized) by using Camelot's lattic module (i.e., Camelot complete wireframe table parsing module), and the parsed complete wireframe table data is returned to the software module; the wireless frame table (i.e., special table) is parsed (i.e., recognized) by using Camelot's stream module (i.e., Camelot wireless frame table parsing module), and the parsed wireless frame table data is returned to the software module; the software module performs coordinate correction and text element filling processing on the complete wireframe table data and the wireless frame table data by using the coordinate integration and processing module; the Html generation and storage module adds tags and stores the data processed by the coordinate integration and processing module. As shown, when the Itext parsing module finishes parsing the text elements in the PDF document, the Itext parsing module sends a table parsing request to the camelot table parsing module in the software module, and the camelot table parsing module includes: a camelot lattic module and a camelot stream module. The camelot table parsing module performs camelot table parsing on the PDF document and sets the line_scale parameter in the camelot table parsing module (for example, to 40); if the camelot lattic module parses that there is a regular table in the PDF document (i.e., the parsing is successful) and sends a request to the camelot table parsing module, If the analysis module returns the regular table data, the Camelot table analysis module determines whether the recognition accuracy and the blank rate in the regular table data meet the preset conditions; if they do not meet the preset conditions, the line_scale parameter is reduced and then analyzed; if the preset conditions are met after the line_scale parameter is reduced, the regular table data analyzed by the Camelot lattic module is returned to the above-mentioned coordinate integration and processing module; if the preset conditions are still not met after the line_scale parameter is reduced, it can be determined that there is no regular table in the original target page of the PDF document, and the area where the special row-level text element is located is screenshotted and the screenshot data is returned to the above-mentioned coordinate integration and processing module.As shown, when the Itext parsing module parses the text elements in the PDF document, the Itext parsing module sends a table parsing request to the stream module (i.e., stream mode) of Camelot in the software module. The stream module of Camelot will parse the special table in the PDF document and return the parsed data to the above-mentioned coordinate integration and processing module. As shown, the above-mentioned coordinate integration and processing module generates sorted table data after coordinate correction and text element filling processing on the table parsed by the Camelot table parsing module. The software module generates Html tag data based on the sorted table data and stores the page number of the PDF document for easy display. The present application provides a schematic diagram of the structure of a device for starting an aging test. The dotted line in indicates that the unit or the module is optional. The device 800 can be used to implement the method described in the above-mentioned method embodiment. The device 800 can be a terminal device or a server or a chip. The device 800 includes one or more processors 801, and the one or more processors 801 can support the device 800 to implement the method in the corresponding method embodiment. The processor 801 can be a general-purpose processor or a special-purpose processor. For example, the processor 801 can be a central processing unit (CPU). The CPU can be used to control the device 800, execute the software program, and process the data of the software program. The device 800 may also include a communication unit 805 to implement the input (reception) and output (transmission) of signals. For example, the device 800 may be a chip, and the communication unit 805 may be the input and / or output circuit of the chip, or the communication unit 805 may be the communication interface of the chip, and the chip may be used as a component of a terminal device. For another example, the device 800 may be a terminal device, and the communication unit 805 may be the transceiver of the terminal device, or the communication unit 805 may be the transceiver circuit of the terminal device. The device 800 may include one or more memories 802, on which a program 804 is stored, and the program 804 can be run by the processor 801 to generate instructions 803, so that the processor 801 performs the method described in the above method embodiment according to the instructions 803. Optionally, data may also be stored in the memory 802. Optionally, the processor 801 may also read data stored in the memory 802, which may be stored at the same storage address as the program 804, or may be stored at a different storage address than the program 804. The processor 801 and the memory 802 may be provided separately or integrated together, for example, integrated on a system-on-chip (SOC) of a terminal device. The specific manner in which the processor 801 executes the method for identifying a table in a PDF document may refer to the relevant description in the method embodiment.It should be understood that each step of the above method embodiment can be completed by a hardware-based logic circuit or a software-based instruction in the processor 801. The processor 801 can be a CPU, a digital signal processor (DSP), a field programmable gate array (FPGA) or other programmable logic device, such as a discrete gate, a transistor logic device or a discrete hardware component. The present application also provides a computer program product, which implements the method described in any method embodiment of the present application when the computer program product is executed by the processor 801. The computer program product can be stored in the memory 802, such as a program 804, which is finally converted into an executable target file that can be executed by the processor 801 after preprocessing, compiling, assembling and linking. The present application also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a computer, the method described in any method embodiment of the present application is implemented. The computer program can be a high-level language program or an executable target program. The computer-readable storage medium is, for example, a memory 802. The memory 802 may be a volatile memory or a non-volatile memory, or the memory 802 may include both volatile memory and non-volatile memory. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SynchLink DRAM, SLDRAM) and direct memory bus random access memory (Direct Rambus RAM, DRRAM).Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the above-described devices and equipment and the technical effects produced can refer to the corresponding process and technical effects in the above-mentioned method embodiments, and will not be repeated here. In several embodiments provided in this application, the disclosed system, device and method can be implemented in other ways. For example, some features of the method embodiments described above can be ignored or not executed. The device embodiments described above are only schematic, and the division of units is only a logical function division. There may be other division methods in actual implementation, and multiple units or components can be combined or integrated into another system. In addition, the coupling between the units or the coupling between the components can be direct coupling or indirect coupling, and the above coupling includes electrical, mechanical or other forms of connection. The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, not to limit them. Although the present application is described in detail with reference to the above-mentioned embodiments, those of ordinary skill in the art should understand that it is still possible to modify the technical solutions recorded in the above-mentioned embodiments, or to replace some of the technical features therein by equivalent, and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the scope of protection of the present application.
Claims
1. A method for identifying tables in PDF documents, It is characterized in that The method comprises: Get the original target page of the PDF document; When it is identified that there is no regular table in the original target page, or when the identification result of the regular table in the original target page does not meet the preset condition, a special table identification module is used to identify whether there is a special table in the original target page, wherein the regular table refers to a table with crossed lines, the special table refers to a table without crossed lines, and the special table identification module is a module for identifying a table without crossed lines; Extract all text elements in the original target page through a text parsing module; When the difference between the coordinates of the first text element among all the text elements and the coordinates of the row-level endpoint element is greater than a first distance threshold, the first text element is determined to be a target text element, the row-level endpoint element is any text element among all the text elements that is closest to the left edge of the original target page, the row where the first text element is located only includes the first text element, the target text element refers to a text element that does not have a regular text format, and the target text element includes an element in a table, an element in a statistical chart, an element in a flow chart, and an element in a scatter plot; When the difference between the coordinates of the second text element and the coordinates of the third text element in all the text elements is greater than the second distance threshold, the second text element and the third text element are determined to be target text elements, the second text element and the third text element are text elements in the same row, the first distance threshold refers to the maximum value of the horizontal coordinates of the text elements in each row; the second distance threshold refers to the maximum difference between the horizontal coordinates of the text elements in each row.
2. The method according to claim 1, It is characterized in that The special table includes: a table without a frame and a table with only horizontal lines, the special table identification module includes: a stream module and a Hough transform module of Camelot, and the method further includes: When the stream module determines that the original target page has the special table, and when the stream module determines that the special table is the wireless frame table, identifying the content in the wireless frame table through the stream module; When the Hough transform module determines that the special table exists in the original target page, and when the Hough transform module determines that the special table is a table with only horizontal lines, the stream module is used to identify the content of the wireless frame area between adjacent horizontal lines in the table with only horizontal lines.
3. The method according to claim 1 or 2, It is characterized in that When the special table exists in the original target page, the special table recognition module recognizes the special table in the original target page as a first result, and the text parsing module recognizes the special table in the original target page as a second result. The method further comprises: When the first result is the same as the second result, filling the first result or the second result into the converted special table; When the first result is different from the second result, the text elements at the coordinate positions of the first result and the second result are identified by OCR, and the text elements identified by OCR are filled into the converted special table.
4. The method according to claim 3, It is characterized in that The method further comprises: A table label and row and column attributes are set for the converted special table, wherein the table label is used to display the converted special table, and the row and column labels are used to indicate the number of rows and columns of the converted special table.
5. The method according to claim 1 or 2, It is characterized in that The method comprises: When the special table does not exist in the original target page, a screenshot is taken of the area where the first text element, the second text element and the third text element are located.
6. The method according to claim 1 or 2, It is characterized in that The recognition result of the regular table in the original target page does not meet the preset conditions, including: The recognition accuracy of the regular table in the original target page is lower than the accuracy threshold.
7. The method according to claim 1 or 2, It is characterized in that Before the special table identification module identifies whether there is a special table in the original target page, the method further includes: It is identified through the lattic module of Camelot whether the regular table exists in the original target page.
8. A device for identifying a table in a PDF document, It is characterized in that The device comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that the device executes the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Extraction method and device for table data in pdf document and storage medium
CN110516048A
Table extraction method and device in PDF document, equipment and medium
CN110795919A