A PDF content extraction method, device and equipment
By removing edge format information from PDF files and extracting PDF body information using machine learning models, the problems of low efficiency and insufficient accuracy of PDF content extraction in the prior art are solved, and more efficient and accurate content extraction is achieved.
Patent Information
- Application Number
- CN202011406023.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-04
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-12-04
AI Technical Summary
The prior art is less efficient and has low accuracy when extracting content of interest from massive PDF files, and is easily interfered by page elements of page-independent content.
By receiving the PDF file to be processed, the PDF text information is determined and edge format information such as header, footer and page number are removed, leaving only the PDF text information for subsequent identification. Using machine learning models, including convolutional neural networks and LSTM structures, page information feature maps are extracted, PDF body information is determined, and content extraction is performed through pre-trained page layout model.
It improves the accuracy and efficiency of content extraction of PDF documents, reduces the image size recognized by subsequent programs, eliminates interference from page edge elements, and enhances the accuracy of content recognition and extraction.
Smart Images

Figure CN113807158B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of PDF recognition, and in particular to a PDF content extraction method, device, equipment and computer-readable storage medium. Background Art
[0002] With the development of society, PDF (Portable Document Format), a portable file format, has become popular because it can be migrated between common operating platforms and can reliably restore every character and color of the file when printing. In daily life, we often convert edited documents into PDF format to facilitate the reliable dissemination of information. Especially in recent decades, the amount of information has continued to rise, and a large amount of data has emerged in the form of PDF format. When we want to extract the content of interest from a large number of PDF files, it becomes very difficult. Because it is very easy to convert common readable files such as HTML and word into PDF files, but it is difficult to reversely convert PDF files into readable files. Based on the above situation, many scholars and companies have extracted PDF text and PDF tables. At present, the main work focuses on establishing algorithm models and parsing systems from the perspective of rule engines and deep learning, focusing on image recognition and text recognition of PDF pages, but the efficiency is often poor and it is easy to be disturbed by page elements with irrelevant content on the page, reducing the efficiency of content recognition and content extraction.
[0003] Therefore, how to solve the poor efficiency and low accuracy of PDF document content extraction is a problem that needs to be solved urgently by those skilled in the art. Summary of the invention
[0004] The purpose of the present invention is to provide a PDF content extraction method, device, equipment and computer-readable storage medium to improve the content extraction accuracy and extraction efficiency of PDF documents.
[0005] In order to solve the above technical problems, the present invention provides a PDF content extraction method, comprising:
[0006] Receive PDF files to be processed;
[0007] Determine PDF text information according to the PDF file to be processed;
[0008] The PDF content extraction information is obtained according to the PDF body information.
[0009] Optionally, in the PDF content extraction method, determining PDF text information according to the PDF file to be processed includes:
[0010] Acquire sample page information according to the PDF file to be processed;
[0011] According to the sample page information, a page information feature graph is obtained using a machine learning model;
[0012] The PDF body information is determined by the PDF file to be processed and the page information feature map.
[0013] Optionally, in the PDF content extraction method, obtaining PDF content extraction information according to the PDF body information includes:
[0014] Using the PDF text information, obtaining the block information to be identified and the category information corresponding to the area information to be identified through a pre-trained page layout model;
[0015] The PDF content extraction information is obtained by using the block to be identified through an identification method corresponding to the corresponding category information.
[0016] Optionally, in the PDF content extraction method, obtaining the PDF content extraction information by using the block to be identified through an identification method corresponding to the corresponding category information includes:
[0017] When the category information is text block information, obtaining paragraph start information and paragraph end information of the block to be identified;
[0018] The PDF content extraction information is determined according to the paragraph start information, the paragraph end information and the preset writing order information.
[0019] Optionally, in the PDF content extraction method, determining the PDF content extraction information according to the paragraph start information, the paragraph end information and the preset writing order information includes:
[0020] Get text dividing line information;
[0021] The PDF content extraction information is determined according to the paragraph start information, the paragraph end information, the text dividing line information and the preset writing order information.
[0022] Optionally, in the PDF content extraction method, obtaining the PDF content extraction information by using the block to be identified through an identification method corresponding to the corresponding category information includes:
[0023] When the category information is table information, obtaining table data block coordinate information;
[0024] Determine single-column horizontal coordinate information according to the table data block information;
[0025] Obtaining single-row vertical coordinate information according to the table data block coordinate information and the single-column horizontal coordinate information;
[0026] The PDF content extraction information is determined according to the table data block coordinate information, the single-column horizontal coordinate information, and the single-row vertical coordinate information.
[0027] Optionally, in the PDF content extraction method, determining the single-column horizontal coordinate information according to the table data block information includes:
[0028] The single-column horizontal coordinate information is determined by using the table data block information through a mean shift algorithm with a characteristic number of 1.
[0029] A PDF content extraction device, comprising:
[0030] A receiving module, used for receiving the PDF file to be processed;
[0031] A text determination module, used to determine PDF text information according to the PDF file to be processed;
[0032] The extraction module is used to obtain PDF content extraction information according to the PDF text information.
[0033] A PDF content extraction device, comprising:
[0034] An instruction input device, used for inputting operation instructions;
[0035] Memory for storing computer programs;
[0036] A processor is used to implement the steps of any one of the above-mentioned PDF content extraction methods when executing the computer program.
[0037] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any one of the above-mentioned PDF content extraction methods are implemented.
[0038] The PDF content extraction method provided by the present invention receives a PDF file to be processed; determines PDF text information according to the PDF file to be processed; and obtains PDF content extraction information according to the PDF text information. The present invention pre-processes the PDF file to be processed, removes the format information located at the edge of the PDF file such as the header, footer and page number of the PDF file, and only leaves the PDF text information for subsequent recognition. Compared with the prior art, the image size to be recognized by the subsequent program is reduced, and the page edge elements that assist reading but do not carry content information are excluded, leaving only the PDF text information related to the content, which greatly improves the recognition and extraction efficiency of the content by the subsequent program. At the same time, due to the removal of interference information such as headers and footers, the accuracy of the subsequent content recognition is also improved. The present invention also provides a PDF content extraction device, equipment and computer-readable storage medium with the above-mentioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions of the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0040] Figure 1 A flowchart of a specific implementation of the PDF content extraction method provided by the present invention;
[0041] Figure 2 A flowchart of another specific implementation of the PDF content extraction method provided by the present invention;
[0042] Figure 3 A flowchart of another specific implementation of the PDF content extraction method provided by the present invention;
[0043] Figure 4 A flowchart of another specific implementation of the PDF content extraction method provided by the present invention;
[0044] Figure 5 A flowchart of another specific implementation of the PDF content extraction method provided by the present invention;
[0045] Figure 6 The present invention provides a schematic structural diagram of a specific implementation of the PDF content extraction device. DETAILED DESCRIPTION
[0046] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0047] The core of the present invention is to provide a PDF content extraction method, a flowchart of a specific implementation method thereof is shown as follows: Figure 1 As shown, it is called specific implementation mode 1, including:
[0048] S101: receiving a PDF file to be processed.
[0049] S102: Determine PDF text information according to the PDF file to be processed.
[0050] The above-mentioned determination of the PDF body information through the PDF file to be processed can be achieved through machine learning methods, such as through LSTM structure or CNN neural network training, by learning several sample pages in the same PDF document, finding the commonalities between the sample pages, and then automatically excluding information such as headers, footers and page numbers, leaving only the PDF body information; or, the PDF page can be directly cropped according to preset rules, such as cropping images of preset lengths at the upper and lower ends of the PDF page, and using the remaining images as the PDF body information. Of course, other methods can also be used according to actual conditions.
[0051] S103: Obtain PDF content extraction information according to the PDF text information.
[0052] The PDF body information may include body text, tables, titles, comments, images and other information, which can be located and classified respectively, and structured according to the classification results.
[0053] The purpose of the present invention is to use a deep learning algorithm to identify the page layout of a PDF, that is, to locate and classify information such as body text, title, annotation, table, image, etc., and to extract and structure the text of the title and body text area according to the classification results; according to the position of the table area, the metadata of the table is extracted to structure the table. For the deep learning model of PDF page layout, we use the model architecture of yolov4. In order to make the model suitable for PDF documents, the present invention adds a preprocessing module, which generates a feature map of the target page through a convolutional network, and extracts the page layout style of PDF through an LSTM structure. The layout characteristics here are also represented by a feature map, and then the two feature maps are fused and sent to yolov4 for calculation to obtain the location and category of the page content. Screenshots can be taken based on the location information of the page, and text extraction (including the location information of the text in the screenshot) is performed using a deep learning model related to OCR. Finally, the text is formatted and the table is formatted according to the category information.
[0054] The PDF content extraction method provided by the present invention receives a PDF file to be processed; determines PDF text information according to the PDF file to be processed; and obtains PDF content extraction information according to the PDF text information. The present invention pre-processes the PDF file to be processed, removes the format information located at the edge of the PDF file such as the header, footer and page number of the PDF file, and only leaves the PDF text information for subsequent recognition. Compared with the prior art, the image size to be recognized by the subsequent program is reduced, and the page edge elements that assist reading but do not carry content information are excluded, leaving only the PDF text information related to the content, which greatly improves the recognition and extraction efficiency of the content by the subsequent program. At the same time, due to the removal of interference information such as headers and footers, the accuracy of the subsequent content recognition is also improved.
[0055] On the basis of the specific implementation mode 1, the method for obtaining the PDF text information is further limited to obtain a specific implementation mode 2, and its flow chart is as follows: Figure 2 As shown, including:
[0056] S201: receiving a PDF file to be processed.
[0057] The PDF file to be processed in this specific implementation is a PDF file in which each page is converted into an image, wherein the image files generated by the same PDF can be stored in the same path. The images can be obtained through open source frameworks such as PDFBOX and PYMUPDF.
[0058] S202: Acquire sample page information according to the PDF file to be processed.
[0059] The sample page information is a page image file used to extract page information features. Since the headers, footers and other non-text contents of the same PDF file are roughly the same, the corresponding page information features can be extracted without too many pages. Therefore, the sample page information is usually 3 to 4 pages of PDF page image information. For convenience, it is generally sampled backward from the first page, that is, the first N pages of the PDF file to be processed (N is a positive integer greater than zero). When the subsequent pages are insufficient, forward sampling can be performed; if all are not satisfied, that is, the entire PDF document does not have N pages, as many pages as possible are obtained, and the insufficient positions are used <token>Or page1 instead. <token>It is an image object of the same size as page1, and its read-in tensor (the high-dimensional array corresponding to the image) metadata is 0.
[0060] S203: Obtain a page information feature graph using a machine learning model based on the sample page information.
[0061] The machine learning model includes computer technologies such as computer deep learning models or knowledge engines, wherein the convolutional neural network of the computer deep learning model can be used to obtain the page information feature map through the sample page information. Still using the above example, if the sample page information includes four pages of PDF, named page1, page2, page3, and page4 respectively, let page1 to page4 be processed by CNN (convolutional neural network) and merged to form a layout_weight feature map, and the layout_weight feature map is further processed by a CNN structure to make it consistent with the size of the single-page image file in the PDF file to be processed after CNN processing, so as to obtain the page information feature map.
[0062] S204: Determine the PDF body information through the PDF file to be processed and the page information feature map.
[0063] The image file of each page of the PDF file to be processed is multiplied by the above-mentioned page information feature map in turn to obtain page_attention_feature_map. The main function of page_attention_feature_map is to reduce the dimension of the original feature map of the read page, generate the regional attention feature map of the page by using the page layout correlation of the same PDF, and determine the PDF text information by judging the attention distribution.
[0064] S205: Obtain PDF content extraction information according to the PDF text information.
[0065] The method for extracting the PDF content extraction information from the PDF body information may adopt a pre-trained page layout model, which may be a yolov4 model.
[0066] In this specific implementation, a method for obtaining the PDF body information is specifically provided. A page information feature map of the PDF file to be processed is obtained through a pre-trained machine learning model, so as to determine where the PDF body content is and where the headers, footers and other elements that need to be discarded are, thereby extracting the PDF body information with high accuracy and having great applicability.
[0067] On the basis of the second specific implementation mode, the method for obtaining the PDF content extraction information is further limited to obtain a third specific implementation mode, and its flow chart is as follows: Figure 3 As shown, including:
[0068] S301: receiving a PDF file to be processed.
[0069] S302: Acquire sample page information according to the PDF file to be processed.
[0070] S303: Obtain a page information feature graph using a machine learning model based on the sample page information.
[0071] S304: Determine the PDF body information through the PDF file to be processed and the page information feature map.
[0072] S305: Utilizing the PDF text information, obtaining the block information to be identified and the category information corresponding to the region information to be identified through a pre-trained page layout model.
[0073] The category information can be regarded as a label for the block information to be identified, and the block information to be identified is classified into a body text block, an image block, a table block, etc.
[0074] S306: Using the block to be identified, and using an identification method corresponding to the corresponding category information, the PDF content extraction information is obtained.
[0075] Based on the second specific implementation mode, this specific implementation mode further describes in depth the process of obtaining the main PDF content extraction information through the page layout model, wherein the page layout model can prepare training and verification data sets based on the company's historical annotation data and public data sets related to other fields.
[0076] During the training process, the machine learning model and the page layout model can be integrated and trained uniformly. As a preferred implementation,
[0077] During model training, the data is unbalanced, and the focus of the current task happens to be on those types with small data volumes, such as tables and titles. In order to allow the model to better learn the characteristics of these areas, an influence factor is added to the loss function to increase the learning ability of these categories.
[0078] The loss function of this model is mainly divided into three parts: border loss, classification loss, and confidence loss. The border loss of Yolov4 uses CIoU loss, which does not need any modification; the confidence loss also does not need to be changed, because the higher the confidence, the better, and there is no difference between categories; what the present invention wants to modify is the category loss caused by classification.
[0079] Modified category loss function:
[0080]
[0081] where is the impact factor of the Φ(c) category, It is the cross entropy loss belonging to category c, multiplied by an influence factor to distinguish the importance of different categories.
[0082] In this specific implementation, the PDF body information is extracted through the page layout model, and the PDF body information is divided into one or more block information to be identified, each block information to be identified corresponds to a category information that marks its category, and different identification methods are called for different categories of blocks to be identified to extract the PDF content extraction information. Of course, when the PDF body information of the PDF file to be processed includes multiple categories of blocks to be identified (such as body text, images, tables, etc.), the PDF content extraction information finally outputted can be a collection of the PDF content extraction information obtained through each block to be identified. Using different content extraction methods for different categories greatly improves the accuracy of the PDF content extraction information finally extracted.
[0083] On the basis of the third specific implementation mode, the method for obtaining the specific type of PDF content extraction information is further limited to obtain the fourth specific implementation mode, and its flow chart is as follows: Figure 4 As shown, including:
[0084] S401: receiving a PDF file to be processed.
[0085] S402: Acquire sample page information according to the PDF file to be processed.
[0086] S403: Obtain a page information feature graph using a machine learning model based on the sample page information.
[0087] S404: Determine the PDF body information through the PDF file to be processed and the page information feature map.
[0088] S405: Utilizing the PDF text information, obtaining the block information to be identified and the category information corresponding to the region information to be identified through a pre-trained page layout model.
[0089] S406: When the category information is text block information, obtain paragraph start information and paragraph end information of the block to be identified.
[0090] The pre-trained OCR deep learning model can be used for text recognition and area positioning to obtain the block to be recognized whose category information is the text. Furthermore, the data is organized into json type data for easy storage and call.
[0091] The paragraph start information and the paragraph end information are to match the beginning and end of the text block, and determine whether the current text block is the beginning or end of a paragraph based on the rules. The rules can be based on whether there is a first line indent at the beginning, whether there is a line break at the end, etc.
[0092] S407: Determine the PDF content extraction information according to the paragraph start information, the paragraph end information and the preset writing order information.
[0093] The writing order information is information reflecting the order of text reading, such as the order from top to bottom and from left to right. The program can sort multiple text block information according to their position distribution on the PDF page according to the preset writing order to generate the text content.
[0094] As a preferred implementation, the determining the PDF content extraction information according to the paragraph start information, the paragraph end information and the preset writing order information includes:
[0095] Get text dividing line information;
[0096] The PDF content extraction information is determined according to the paragraph start information, the paragraph end information, the text dividing line information and the preset writing order information.
[0097] If there is a situation where the document area is segmented vertically, then horizontal jumping can be achieved by identifying the segmentation, the vertical distance between the front and back text blocks and other features. The program executes different text block matching work in a loop. In order to allow the program to adaptively change columns, the present invention adds a deadline to the rule. If there is a vertical segmentation (such as a blank area with a width exceeding the preset value), the deadline is set to the vertical center line of the segmentation, and the next text block to be matched must be above the deadline. If the text blocks above the deadline have been taken, the deadline is reset to 0, and the matching is completed until all the target types of text blocks of the current page are connected, that is, all text blocks are matched. After the above steps, each page of the PDF document meets the reading order. Of course, the document whose body text is horizontally separated is also applicable. After detecting the text segmentation line, the position, start information, and end information of each area to be identified are combined to determine the arrangement order of the text in the area to be identified. For example, if there is a horizontal segmentation line that divides the body text in a page into left and right parts, the left text block can be extracted from top to bottom first, and then the right text block can be extracted from top to bottom.
[0098] It should be noted that the task of using OCR to recognize text can also be placed in the head part of the page layout model, allowing the model to perform region positioning and region classification while extracting text. Using a model with multiple tasks is more conducive to improving the performance of the model, and all we need to do is add a text recognition branch and a loss function for text recognition.
[0099] In this specific implementation, the integration is mainly based on information such as the position of the text block, the writing order, whether the text is the beginning and the end, and the block information to be identified is specifically limited to text block information, that is, the content extraction method when the main text is the text. On the basis of recognizing the text, a method for determining the reading and splicing order between text blocks in different positions is also introduced, so that the main text finally extracted is fluent and does not require secondary sorting, which greatly improves the content extraction efficiency.
[0100] On the basis of the specific implementation mode 3, the operation method when the PDF content extraction information is a table is further discussed to obtain a specific implementation mode 5, the flow chart of which is shown in FIG5, including:
[0101] S501: receiving a PDF file to be processed.
[0102] S502: Acquire sample page information according to the PDF file to be processed.
[0103] S503: Obtain a page information feature graph using a machine learning model based on the sample page information.
[0104] S504: Determine the PDF body information through the PDF file to be processed and the page information feature map.
[0105] S505: Utilizing the PDF text information, obtaining the block information to be identified and the category information corresponding to the region information to be identified through a pre-trained page layout model.
[0106] S506: When the category information is table information, obtain table data block coordinate information.
[0107] S507: Determine single-column horizontal coordinate information according to the table data block information.
[0108] As a preferred implementation, the single-column horizontal coordinate information is determined by using the table data block information through a mean shift algorithm with a characteristic number of 1. The specific operation method is as follows:
[0109] Determine the number of columns in the table based on the clustering algorithm. The data involved in clustering are: the starting point of the horizontal coordinate of the text block in the table area; or the midpoint of the horizontal coordinate of the text block in the table area. The clustering algorithm mainly classifies data based on the clustering of data. The selection of the two groups of data can be made according to the following rules: If a large number of text blocks are aligned on the left boundary (the boundary uses a soft boundary), the horizontal coordinate starting point data set is used for clustering; if a large number of text blocks are aligned on the horizontal coordinate midpoint, the corresponding data set is used for clustering. The "large number" in the above text can be determined based on the alignment ratio threshold. The alignment ratio threshold does not need to be set very high. The column segmentation of table data is generally clear. A column of data can be clustered using a clustering algorithm. The task itself is not difficult.
[0110] If you know how many columns there are, using the K-means algorithm is a good choice, but for a non-specified PDF table type, the model must adaptively find the number of columns in the table. The present invention proposes a mean shift algorithm Mean-models-shift1 with a feature number of 1 based on the idea of the mean shift algorithm. The mean shift clustering algorithm is mainly used to cluster samples in multidimensional space. The main parameters include the sliding window radius r of the mean shift. This parameter is used to assist in finding the mean center in the algorithm. The setting of the radius in actual operations does not have a significant impact on the algorithm results. In the case of one-dimensional features, the relevant parameters and rules are changed to obtain the Mean-models-shift1 algorithm, including:
[0111] 1) Determine a one-dimensional window radius r and randomly generate up to len(x) / 2 center points within the sample distribution interval.
[0112] 2) Generate a sliding window with a radius of r for each center point and start sliding; each time you slide to a new area, calculate the mean (or mode) in the sliding window as the new center point and update it as the center of the current sliding window. The number of samples in the sliding window is recorded as the sample density in the sliding window, and the algorithm will always move the center of the sliding window to the point with high density.
[0113] 3) When multiple windows overlap, the sliding window with the highest density is retained.
[0114] 4) The window is updated iteratively until the density of the window no longer changes.
[0115] Among them, x is the clustering object, its feature number is 1, len(x) is the sample size, and the final output object is the category center after clustering.
[0116] In addition, if the majority models are used as the new window, the input x needs to be preprocessed, that is, a threshold is set to unify similar points.
[0117] This specific implementation method mainly uses a clustering algorithm to find the identification points of the table columns, and divides the columns according to the clustering results. The distinction of rows is mainly based on the horizontal alignment of table row data. The data is first sorted by the vertical axis, and then the rows are divided based on the rules.
[0118] According to the center point, the data block of the current column is obtained from the original data to observe whether there is a data block in the same row. If not, the current column division is maintained. If there are multiple data blocks in the same row and the number of occurrences is greater than the set threshold, the column is split. The threshold is related to the number of columns of the table data.
[0119] S508: Obtaining single-row vertical coordinate information according to the table data block coordinate information and the single-column horizontal coordinate information.
[0120] Specifically, according to the upper boundary ordinate and the lower boundary ordinate of the data block, a soft boundary error can be used to determine whether the current data is in the same row, and the soft boundary refers to an error range that allows a certain boundary alignment, and does not require complete alignment. The upper boundary ordinate and the lower boundary ordinate are ordinates determined by extending upward and downward by a preset distance from the uppermost end and the lowermost end of the table data block, respectively.
[0121] Furthermore, it is important to note that there may be cross-row data blocks in the table. If a data block is not aligned with the current row, but its upper boundary is greater than the lower boundary of the current row, and its lower boundary is less than the upper boundary of the next row, that is, the data block between the two rows is processed and marked as a cross-row data block, and the row number information of the associated row is marked.
[0122] It is particularly important to note that there may be cross-column data blocks in the table. After obtaining the single-row vertical coordinate information, the data block of the current column can be obtained from the original data based on the center point to observe whether there are data blocks in the same row. If not, the current column division is maintained. If there are multiple data blocks in the same row and the number of occurrences is greater than a set threshold, the column is split, where the threshold is related to the number of columns of the table data.
[0123] S509: Determine the PDF content extraction information according to the table data block coordinate information, the single-column horizontal coordinate information, and the single-row vertical coordinate information.
[0124] The layout of current PDF documents is relatively complex and there is no fixed format. In particular, three-line tables are used for table objects in many disciplines. For borderless tables, current table extraction methods have great problems, mainly manifested in inaccurate data unit division and poor column distinction. There is a high probability that the recognition result will merge columns. The clustering algorithm based on the horizontal axis in this specific implementation method actually reduces the dimension of the data and eliminates the influence of the vertical axis. Because the column division is mainly the horizontal axis and the vertical axis has no influence, this operation does not lose the amount of information of the column division task. The column information of the table obtained by the horizontal axis clustering algorithm is more accurate and the processing efficiency is improved. In addition, since the specific implementation method no longer searches for the calibrated table border, but directly determines the "soft border" of the cell according to the coordinates of the table data block information, the positional relationship between the obtained table layout and the cell is also more accurate.
[0125] The PDF content extraction device provided by an embodiment of the present invention is introduced below. The PDF content extraction device described below and the PDF content extraction method described above can be referred to each other.
[0126] Figure 6 The structure diagram of the PDF content extraction device provided by the embodiment of the present invention is shown in FIG. Figure 6 The PDF content extraction device may include:
[0127] The receiving module 100 is used to receive the PDF file to be processed;
[0128] A text determination module 200, configured to determine PDF text information according to the PDF file to be processed;
[0129] The extraction module 300 is used to obtain PDF content extraction information according to the PDF text information.
[0130] As a preferred implementation, the text determination module 200 includes:
[0131] A sample acquisition unit, used for acquiring sample page information according to the PDF file to be processed;
[0132] A page feature unit, used to obtain a page information feature graph using a machine learning model according to the sample page information;
[0133] The body unit is used to determine the PDF body information through the PDF file to be processed and the page information feature map.
[0134] As a preferred implementation, the extraction module 300 includes:
[0135] A block category determination unit, used to obtain the block information to be identified and the category information corresponding to the area information to be identified by using the PDF body information and a pre-trained page layout model;
[0136] The extraction unit is used to obtain the PDF content extraction information by using the block to be identified through the identification method corresponding to the corresponding category information.
[0137] As a preferred implementation, the extraction module 300 includes:
[0138] A text block determination unit, used for obtaining paragraph start information and paragraph end information of the block to be identified when the category information is text block information;
[0139] The text content extraction unit is used to determine the PDF content extraction information according to the paragraph start information, the paragraph end information and the preset writing order information.
[0140] As a preferred implementation, the extraction module 300 includes:
[0141] A dividing line determination unit, used to obtain text dividing line information;
[0142] The dividing line text extraction unit is used to determine the PDF content extraction information according to the paragraph start information, the paragraph end information, the text dividing line information and the preset writing order information.
[0143] As a preferred implementation, the extraction module 300 includes:
[0144] A table data block determining unit, used for obtaining table data block coordinate information when the category information is table type information;
[0145] A single-column horizontal coordinate determining unit, used to determine single-column horizontal coordinate information according to the table data block information;
[0146] A row vertical coordinate determining unit is used to obtain single row vertical coordinate information according to the table data block coordinate information and the single column horizontal coordinate information;
[0147] The table content extraction unit is used to determine the PDF content extraction information according to the table data block coordinate information, the single column horizontal coordinate information and the single row vertical coordinate information.
[0148] As a preferred implementation, the extraction module 300 includes:
[0149] The mean shift unit is used to determine the single-column horizontal coordinate information by using the table data block information and a mean shift algorithm with a characteristic number of 1.
[0150] The PDF content extraction device provided by the present invention is used to receive the PDF file to be processed through the receiving module 100; the text determination module 200 is used to determine the PDF text information according to the PDF file to be processed; and the extraction module 300 is used to obtain the PDF content extraction information according to the PDF text information. The present invention pre-processes the PDF file to be processed, removes the format information located at the edge of the PDF file such as the header, footer and page number of the PDF file, and only leaves the PDF text information for subsequent recognition. Compared with the prior art, the image size to be recognized by the subsequent program is reduced, and the page edge elements that assist reading but do not carry content information are excluded, leaving only the PDF text information related to the content, which greatly improves the recognition and extraction efficiency of the content by the subsequent program. At the same time, due to the removal of interference information such as headers and footers, the accuracy of the subsequent content recognition is also improved.
[0151] The PDF content extraction device of this embodiment is used to implement the aforementioned PDF content extraction method. Therefore, the specific implementation of the PDF content extraction device can be seen in the embodiment of the PDF content extraction method in the previous text. For example, the receiving module 100100, the text determination module 200200, and the extraction module 300300 are respectively used to implement steps S101, S102 and S103 in the aforementioned PDF content extraction method. Therefore, its specific implementation can refer to the description of the corresponding embodiments of each part, which will not be repeated here.
[0152] A PDF content extraction device, comprising:
[0153] An instruction input device, used for inputting operation instructions;
[0154] Memory for storing computer programs;
[0155] A processor is used to implement the steps of any of the above-mentioned PDF content extraction methods when executing the computer program. The PDF content extraction method provided by the present invention receives a PDF file to be processed; determines the PDF text information according to the PDF file to be processed; and obtains PDF content extraction information according to the PDF text information. The present invention pre-processes the PDF file to be processed, removes the format information located at the edge of the PDF file such as the header, footer and page number of the PDF file, and only leaves the PDF text information for subsequent recognition. Compared with the prior art, the size of the image to be recognized by the subsequent program is reduced, and the page edge elements that assist reading but do not carry content information are excluded, leaving only the PDF text information related to the content, which greatly improves the efficiency of subsequent program recognition and extraction of content. At the same time, due to the removal of interference information such as headers and footers, the accuracy of subsequent content recognition is also improved.
[0156] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the PDF content extraction method described in any one of the above are implemented. The PDF content extraction method provided by the present invention receives a PDF file to be processed; determines PDF text information according to the PDF file to be processed; and obtains PDF content extraction information according to the PDF text information. The present invention pre-processes the PDF file to be processed, removes the format information located at the edge of the PDF file such as the header, footer and page number of the PDF file, and only leaves the PDF text information for subsequent recognition. Compared with the prior art, the image size to be recognized by the subsequent program is reduced, and the page edge elements that assist reading but do not carry content information are excluded, leaving only the PDF text information related to the content, which greatly improves the efficiency of the subsequent program in identifying and extracting the content. At the same time, due to the removal of interference information such as headers and footers, the accuracy of the subsequent content recognition is also improved.
[0157] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0158] It should be noted that, in this specification, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0159] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0160] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0161] The PDF content extraction method, device, equipment and computer-readable storage medium provided by the present invention are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method and core ideas of the present invention. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.< / token> < / token>
Claims
1. A PDF content extraction method, characterized in that: include: S101: receiving a PDF file to be processed; S102: Determine PDF text information according to the PDF file to be processed; The PDF body information is determined by the PDF file to be processed, which can be achieved by a machine learning method. By learning several sample pages in the same PDF document, the commonality between the sample pages can be found through LSTM structure or CNN neural network training, so as to automatically exclude the header, footer and page number information and leave only the PDF body information. Alternatively, the PDF page can be directly cropped according to a preset rule, and images of a preset length at the upper and lower ends of the PDF page can be cropped, and the remaining images can be used as the PDF body information. S103: Obtain PDF content extraction information according to the PDF text information; The PDF body information includes body text, tables, titles, comments, and image information, which are located and classified respectively, and structured according to the classification results; Use deep learning algorithms to analyze the page layout of PDF, that is, locate and classify the main text, title, annotation, table, and image information, and extract and structure the text in the title and main text area according to the classification results; extract the metadata of the table according to the location of the table area to structure the table; the deep learning model of PDF page layout uses the model architecture of yolov4; in order to make the model suitable for PDF documents, add a preprocessing module, which generates a feature map of the target page through a convolutional network, and extracts the page layout style of PDF through an LSTM structure. The layout characteristics here are also represented by a feature map, and then the two feature maps are fused and sent to yolov4 for calculation to obtain the location and category of the page content; screenshots can be taken based on the location information of the page, and text extraction is performed using OCR-related deep learning models; finally, text formatting and table formatting are performed according to category information; The process also includes: S201: receiving a PDF file to be processed; The PDF document to be processed is a PDF file in which each page is converted into an image, wherein the image files generated by the same PDF can be stored in the same path; the image can be obtained through PDFBOX and PYMUPDF open source architectures; S202: Acquire sample page information according to the PDF file to be processed; The sample page information is a page image file used to extract page information features. Since the header and footer non-text content of the same PDF file are roughly the same, the corresponding page information features can be extracted without too many pages. Therefore, the sample page information is 3 to 4 pages of PDF page image information. For convenience, all samples are sampled backward from the first page, that is, the first N pages of the PDF file to be processed, N is a positive integer greater than zero, and forward sampling can be performed when the subsequent pages are insufficient; if all are not satisfied, that is, the entire PDF document does not have N pages, as many pages as possible are obtained, and the insufficient positions are used <token>Or page1 instead; <token> It is a picture object of the same size as page1, and its read-in tensor metadata is 0;< / token> < / token> S203: Obtaining a page information feature graph using a machine learning model according to the sample page information; The machine learning model includes a computer deep learning model or a knowledge engine computer technology, wherein the convolutional neural network of the computer deep learning model can be used to obtain the page information feature map through the sample page information; S204: Determine the PDF text information according to the PDF file to be processed and the page information feature map; Multiply the image file of each page of the PDF file to be processed by the page information feature map in sequence to obtain page_attention_feature_map; The main function of page_attention_feature_map is to reduce the dimension of the original feature map of the page read in, generate the regional attention feature map of the page by using the page layout correlation of the same PDF, and determine the PDF text information by judging the attention distribution; S205: Obtain PDF content extraction information according to the PDF text information; The method for extracting the PDF content information from the PDF body information can use a pre-trained page layout model, which is a yolov4 model; The process also includes: S301: receiving a PDF file to be processed; S302: Acquire sample page information according to the PDF file to be processed; S303: Obtaining a page information feature graph using a machine learning model according to the sample page information; S304: Determine the PDF text information according to the PDF file to be processed and the page information feature map; S305: using the PDF text information, obtaining the block information to be identified and the category information corresponding to the block information to be identified through a pre-trained page layout model; The category information can be regarded as a label for the block information to be identified, and the block information to be identified is classified into a body text block, an image block or a table block; S306: using the block to be identified, and obtaining the PDF content extraction information through the identification method corresponding to the corresponding category information; During model training, the data is unbalanced, and the current task focuses on those types with small data volumes. In order to enable the model to better learn the characteristics of these areas, an influencing factor is added to the loss function to increase the learning ability of these categories. The loss function of this model is mainly divided into three parts: border loss, classification loss, and confidence loss. The border loss of Yolov4 uses CIoU loss, which does not need any modification. The confidence loss also does not need to be changed, because the higher the confidence, the better, and there is no difference between categories. What needs to be modified is the category loss caused by classification. Modified category loss function: ; Where Φ(c) is the impact factor of the category; ; is the cross entropy loss belonging to category c, multiplied by an impact factor to distinguish the importance of different categories; Also includes: S401: receiving a PDF file to be processed; S402: Acquire sample page information according to the PDF file to be processed; S403: Obtaining a page information feature graph using a machine learning model according to the sample page information; S404: Determine the PDF text information according to the PDF file to be processed and the page information feature map; S405: using the PDF text information, obtaining the block information to be identified and the category information corresponding to the block information to be identified through a pre-trained page layout model; S406: When the category information is text block information, obtain paragraph start information and paragraph end information of the block to be identified; Use the pre-trained OCR deep learning model to perform text recognition and area positioning, and obtain the block to be recognized whose category information is the text of the main text; further, organize the data into json type data for easy storage and call; The paragraph start information and the paragraph end information are to match the beginning and the end of the text block, and determine whether the current text block is the beginning and the end of the paragraph according to the rules. The rules can be determined based on whether there is a first line indentation at the beginning and whether there is a line break at the end; S407: Determine the PDF content extraction information according to the paragraph start information, the paragraph end information and the preset writing order information; The writing order information is information reflecting the order of reading text, in a top-down, left-to-right order. The program sorts the multiple text block information according to their positions on the PDF page according to the preset writing order to generate the text content; The step of determining the PDF content extraction information according to the paragraph start information, the paragraph end information and the preset writing order information includes: Get text dividing line information; Determining the PDF content extraction information according to the paragraph start information, the paragraph end information, the text dividing line information and the preset writing order information; If there is a vertical document area segmentation, then horizontal jumping can be achieved by identifying the segmentation and the vertical distance features of the previous and next text blocks; the program executes different text block matching work in a loop. In order to allow the program to adaptively change columns, a deadline is added to the rule. If there is a vertical segmentation, the deadline is set to the vertical center line of the segmentation, and the next text block to be matched must be above the deadline; if the text blocks above the deadline have been taken, the deadline is reset to 0, and the matching is continued until all the target type of text blocks of the current page are connected, that is, all text blocks are matched; after the above steps, each page of the PDF document meets the reading order. Of course, the same applies to documents whose body text is horizontally divided. After detecting the text segmentation line, the position, start information, and end information of each block information to be identified are combined to determine the arrangement order of the text in the block to be identified. If there is a horizontal segmentation line that divides the body text in a page into left and right parts, the left text block can be extracted from top to bottom first, and then the right text block can be extracted from top to bottom; The task of using OCR to recognize text can also be placed in the head part of the page layout model, allowing the model to perform region positioning and region classification while extracting text; using a model with multiple tasks is more conducive to improving the performance of the model, and all we need to do is add a text recognition branch and a loss function for text recognition; Also includes: S501: receiving a PDF file to be processed; S502: Acquire sample page information according to the PDF file to be processed; S503: Obtaining a page information feature graph using a machine learning model according to the sample page information; S504: Determine the PDF text information according to the PDF file to be processed and the page information feature map; S505: using the PDF text information, obtaining the block information to be identified and the category information corresponding to the block information to be identified through a pre-trained page layout model; S506: When the category information is table information, obtain table data block coordinate information; S507: Determine single-column horizontal coordinate information according to the table data block information; Using the table data block information, the single-column horizontal coordinate information is determined by a mean shift algorithm with a characteristic number of 1. The specific operation method is as follows: The number of columns of the table is determined based on the clustering algorithm; the data involved in clustering is: the starting point of the horizontal coordinate of the text block in the table area; or the midpoint of the horizontal coordinate of the text block in the table area; the clustering algorithm is mainly classified according to the clustering of the data, and the selection of the two groups of data can be made according to the following rules: if a large number of text blocks are aligned on the left edge, the horizontal coordinate starting point data set is used for clustering; if a large number of text blocks are aligned on the horizontal coordinate midpoint, the corresponding data set is used for clustering; the "large number" in the above text can be determined based on the alignment ratio threshold, and the alignment ratio threshold does not need to be set very high. The column segmentation of the table data is clear, and a column of data can be clustered using the clustering algorithm, and the task itself is not difficult; We can know how many columns there are and use the K-means algorithm. However, for a non-specified PDF table type, the model must adaptively find the number of columns in the table. Based on the idea of the mean shift algorithm, a mean shift algorithm Mean-models-shift1 with a feature number of 1 is proposed. The mean shift clustering algorithm is mainly for clustering samples in multidimensional space. The main parameters are the sliding window radius r of the mean shift. This parameter is used to assist in finding the mean center in the algorithm. In actual operations, the setting of the radius will not have a significant impact on the algorithm results. In the case of one-dimensional features, the relevant parameters and rules are changed to obtain the Mean-models-shift1 algorithm, including: 1) Determine a one-dimensional window radius r and randomly generate up to len(x) / 2 center points within the sample distribution interval; 2) Generate a sliding window with a radius of r for each center point and start sliding; each time sliding to a new area, calculate the mean or mode in the sliding window as the new center point and update it as the center of the current sliding window; the number of samples in the sliding window is recorded as the sample density in the sliding window, and the algorithm will always move the center of the sliding window to the point with high density; 3) When multiple windows overlap, the sliding window with high density is retained; 4) The window is updated iteratively until the density of the window no longer changes; Among them, x is the clustering object, its feature number is 1, len(x) is the sample size, and the final output object is the category center after clustering; In addition, if the majority model is used as a new window, the input x needs to be preprocessed, that is, a threshold is set to unify similar points; Use clustering algorithm to find the identification points of table columns, and divide the columns according to the clustering results. The distinction of rows is mainly based on the horizontal alignment of table row data. First, sort the data by the vertical axis, and then divide the rows based on the rules; Get the data block of the current column from the original data according to the center point, and observe whether there is a data block in the same row. If not, keep the current column division. If there are multiple data blocks in the same row and the number of occurrences is greater than the set threshold, split the column. The threshold is related to the number of columns of the table data. S508: Obtaining single-row vertical coordinate information according to the table data block coordinate information and the single-column horizontal coordinate information; Specifically, according to the upper boundary ordinate and the lower boundary ordinate of the data block, a soft boundary error can be used to determine whether the current data is in the same row, and the soft boundary refers to an error range that allows a certain boundary alignment, and does not require complete alignment; wherein the upper boundary ordinate and the lower boundary ordinate are ordinates determined by extending upward and downward by a preset distance from the uppermost end and the lowermost end of the table data block, respectively; There are cross-row data blocks in the table. If a data block is not aligned with the current row, but its upper boundary is greater than the lower boundary of the current row, and its lower boundary is less than the upper boundary of the next row, that is, the data block between the two rows is processed, marked as a cross-row data block, and the row number information of the associated row is marked; after obtaining the single-row ordinate information, the data block of the current column can be obtained from the original data according to the center point to observe whether there are data blocks in the same row. If not, the current column division is maintained. If there are multiple data blocks in the same row and the number of occurrences is greater than a set threshold, the column is split, where the threshold is related to the number of columns of the table data; S509: Determine the PDF content extraction information according to the table data block coordinate information, the single-column horizontal coordinate information, and the single-row vertical coordinate information; The layout of current PDF documents is relatively complex and there is no fixed format. Three-line tables are used as table objects in many disciplines. For borderless tables, current table extraction methods have great problems, mainly manifested in inaccurate data unit division and poor column distinction. There is a high probability that the recognition result will merge columns. The clustering algorithm based on the horizontal axis of this method actually reduces the dimension of the data and eliminates the influence of the vertical axis. Because the column division is mainly the horizontal axis and the vertical axis has no influence, this operation does not lose the amount of information of the column division task. The column information of the table obtained by the horizontal axis clustering algorithm is more accurate and the processing efficiency is improved. In addition, since the method no longer searches for the calibrated table border, but directly determines the "soft border" of the cell according to the coordinates of the table data block information, the positional relationship between the obtained table layout and the cell is also more accurate.
2. A PDF content extraction device, using the PDF content extraction method according to claim 1, characterized in that: include: A receiving module, used for receiving the PDF file to be processed; A text determination module, used to determine PDF text information according to the PDF file to be processed; The extraction module is used to obtain PDF content extraction information according to the PDF text information.
3. A PDF content extraction device, characterized in that: include: An instruction input device, used for inputting operation instructions; Memory for storing computer programs; A processor, configured to implement the steps of the PDF content extraction method as claimed in claim 1 when executing the computer program.
4. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the PDF content extraction method according to claim 1 are implemented.
Citation Information
Patent Citations
Electronic reader and document typesetting method thereof
CN101986290A
Table recognition method and device, computer device and storage medium
CN110334585A
PDF document table extraction method, device and equipment and computer readable storage medium
CN110390269A
Visual deep learning-based document information fragmentation extraction method
CN110991403A
Resume analysis method and system based on deep learning
CN111737969A