A method and system for extracting valid information of PDF document page elements

By constructing the initial PDF document information extraction model and document analysis ruleset, combining the page object coordinates and shape information, the identification errors and text content confusion in the extraction of effective page elements in PDF documents are solved, achieving more efficient and accurate information extraction.

CN114611466BActive Publication Date: 2025-06-17SOUTHERN POWER GRID DIGITAL GRID RESEARCH INSTITUTE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210259864.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-16
Publication Date
2025-06-17
Estimated Expiration
2042-03-16

AI Technical Summary

Technical Problem

When the prior art extracts valid information of page elements in PDF documents, it is easy to have the problem of identification errors and confusing text content. In addition, traditional methods may introduce errors in the information processing process, and the recognition rate is not high.

Method used

By constructing an initial PDF document information extraction model, and combining the document analysis rule set to generate a PDF document information extraction rule model, using page object coordinates and shape information, combining the front and back page and page characteristics, text information is extracted to ensure the accuracy and completeness of the information.

Benefits of technology

It effectively overcomes the problems of identification errors and confusing text content, improves the effectiveness and accuracy of PDF text information, and ensures the refined processing of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114611466B_ABST
    Figure CN114611466B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for extracting valid information of PDF document page elements, including the following: constructing an initial PDF document information extraction model and storing it in a first storage area; obtaining a document parsing rule set; generating a PDF document information extraction rule model according to the initial PDF document information extraction model and the document parsing rule set and storing it in a second storage area; constructing a PDF document information extraction model for extracting valid information of the PDF document according to the initial PDF document information extraction model and the PDF document information extraction rule model; by setting a first interval time, updating the PDF document information extraction model according to the initial PDF document information extraction model and the document parsing rule set; the present invention obtains texts from the front and back pages respectively according to the text information at the top and bottom of the page to complement the missing text information of this page, summarizes the text information in units of pages, and the information is more refined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer information processing, and more particularly, to a method and system for extracting effective information of PDF document page elements. Background Art

[0002] PDF is the most common document format for daily communication, and its display and printing effects are not affected by the operating system and operating devices. A large amount of text and image information is contained in PDF documents. The method for extracting effective information of PDF pages is to extract the text information contained in PDF documents, and through a series of information processing, combined with the context, filter out interference information and useless information, and extract the effective information and store it in a specified manner. The traditional method is to convert PDF into images through OCR recognition technology, and after layout analysis, character segmentation, and text recognition, the recognized results are output. This method needs to perform intelligent analysis in the above steps, which may introduce errors and there are certain problems with the recognition rate. Another method is to perform parsing on all the text in the PDF document through text parsing technology to obtain all the text content. Since the encoding and display characters in the PDF document are not completely corresponding, and the position where the text is displayed does not completely correspond to the actual position in the document, this method cannot extract characters through the internal code of the characters extracted, and even if the characters are successfully extracted, there may be a situation where the text content is disordered. Summary of the Invention

[0003] In order to solve the above problems, the purpose of the present invention is to provide a text extraction method that combines the coordinates and shape information of PDF page objects, combines the context of the previous and next pages and page features, overcomes the situation of incorrect recognition, disordered text content, and interference of invalid information in the prior art, and improves the effectiveness of PDF text information. Especially when the information of each page of the document needs to be processed independently, the effective information of this page can be determined according to the context splicing of the previous and next pages.

[0004] In order to achieve the above technical purpose, the present application provides a method for extracting effective information of PDF document page elements, including the following steps:

[0005] Construct an initial PDF document information extraction model and store it in the first storage area. The initial PDF document information extraction model is used to generate a time-effective PDF document information extraction rule model;

[0006] Obtain a document parsing rule set, which is used to represent a set of rules for parsing a PDF document to obtain the effective information of the PDF document;

[0007] Generate a PDF document information extraction rule model according to the initial PDF document information extraction model and the document parsing rule set, and store it in the second storage area;

[0008] Construct a PDF document information extraction model for extracting the valid information of a PDF document according to the initial PDF document information extraction model and the PDF document information extraction rule model.

[0009] By setting a first interval time, update the PDF document information extraction model according to the initial PDF document information extraction model and the document parsing rule set.

[0010] Preferably, in the process of using the PDF document information extraction model to obtain the valid information of a PDF document, open the PDF document, and parse and arrange the objects of the PDF document in the order from top to bottom and from left to right, where the objects include text boxes, pictures, rectangular boxes, and curves.

[0011] Preferably, after the process of obtaining the objects of the PDF document, search for the marker positions of the key markers in the objects of the PDF document, and divide the page into multiple different information regions according to the marker positions, and extract the information of different regions and store it in variables.

[0012] Preferably, in the process of searching for the positions of the key markers in the objects, the key markers are used to represent lines that exceed the length of the dash;

[0013] The marker positions are used to represent the division of information regions.

[0014] Preferably, in the process of dividing the page into multiple different information regions, the different information regions include the title, chapter number, and page number of the header, the body region, and the footnote, page number, and annotation of the footer.

[0015] Preferably, in the process of obtaining the valid information of the PDF document, judge whether a paragraph starts according to different information regions. If not, splice the last punctuation mark to the end of the previous page's text to the beginning of the current page's text.

[0016] Preferably, in the process of obtaining the valid information of the PDF document, obtain the density information of the text in the body region. If the density information is less than the specified percentage, then use OCR recognition. Otherwise, extract according to the position of the text in the page region, where the process of extracting the text includes: page number extraction verification, title extraction, body extraction, and annotation extraction.

[0017] Preferably, in the process of judging whether a paragraph starts, determine the coordinate range of the valid information according to the marker position, and exclude all objects outside this range;

[0018] Traverse the objects within the coordinate range in sequence, extract the text information, and based on the preference for incomplete information, determine whether the paragraph at the top or bottom is the beginning. The basis for judgment is that the first character at the top is not indented, and / or the end at the bottom is not a punctuation mark.

[0019] The present invention also discloses a system for extracting valid information of PDF document page elements, including:

[0020] A data acquisition module for acquiring PDF documents;

[0021] A data parsing module for constructing an initial PDF document information extraction model and storing it in the first storage area. The initial PDF document information extraction model is used to generate a time-effective PDF document information extraction rule model; obtaining a document parsing rule set, which is used to represent a set of rules for parsing a PDF document to obtain valid information of the PDF document; generating a PDF document information extraction rule model according to the initial PDF document information extraction model and the document parsing rule set, and storing it in the second storage area; constructing a PDF document information extraction model for extracting valid information of the PDF document according to the initial PDF document information extraction model and the PDF document information extraction rule model; updating the PDF document information extraction model according to the initial PDF document information extraction model and the document parsing rule set by setting a first interval time;

[0022] A data correction module for recording the update time point of each valid information in the generation of valid information of the PDF document and the used document parsing rules, generating a PDF document information extraction index model according to the valid information, the update time point, and the document parsing rules; merging the PDF document information extraction index model into the PDF document information extraction model; and judging the accuracy of the corresponding valid information according to the document parsing rules and the update time point according to a second interval time.

[0023] Preferably, the extraction system further includes:

[0024] A first storage unit for storing the initial PDF document information extraction model;

[0025] A second storage unit for storing the PDF document information extraction rule model;

[0026] A third storage unit for storing the PDF document information extraction index model.

[0027] The present invention discloses the following technical effects:

[0028] The present invention uses specific elements (objects) in PDF pages as markers to define the effective range of information extraction. Within the effective range, text objects are arranged in coordinate order and then the text is extracted sequentially to prevent the text content from being disordered. According to the text information at the top and bottom of the page, text is obtained from the front and back pages respectively to complement the missing text information on this page, and the text information is summarized in units of pages, making the information more refined. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0030] Figure 1 is the information extraction flowchart of the present invention;

[0031] Figure 2 is the schematic diagram of the information complement preference process of the present invention;

[0032] Figure 3 is the schematic diagram of the method process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] In order to make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Usually, the components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application to be protected, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0034] As Figures 1-3 shown, the present invention provides a method for extracting effective information of PDF document page elements, including the following steps:

[0035] Construct an initial PDF document information extraction model and store it in the first storage area. The initial PDF document information extraction model is used to generate a time-sensitive PDF document information extraction rule model;

[0036] Obtain a document parsing rule set, which is used to represent a set of rules for parsing a PDF document to obtain valid information of the PDF document;

[0037] Generate a PDF document information extraction rule model according to the initial PDF document information extraction model and the document parsing rule set, and store it in the second storage area;

[0038] Construct a PDF document information extraction model for extracting valid information of the PDF document according to the initial PDF document information extraction model and the PDF document information extraction rule model;

[0039] By setting a first interval time, update the PDF document information extraction model according to the initial PDF document information extraction model and the document parsing rule set.

[0040] Further preferably, in the process of using the PDF document information extraction model to obtain valid information of the PDF document, open the PDF document, and parse and arrange the objects of the PDF document in the order from top to bottom and from left to right, where the objects include text boxes, pictures, rectangular boxes, and curves.

[0041] Further preferably, after the process of obtaining the objects of the PDF document, search for the marked positions of the key marks in the objects according to the objects of the PDF document, and divide the page into multiple different information regions according to the marked positions, and extract the information of different regions and store it in variables.

[0042] Further preferably, in the process of searching for the positions of the key marks in the objects, the key marks are used to represent lines longer than the dash length;

[0043] The marked position is used to represent the division of information regions.

[0044] Further preferably, in the process of dividing the page into multiple different information regions, the different information regions include the title, chapter number, and page number of the header, the body region, and the footnote, page number, and annotation of the footer.

[0045] Further preferably, in the process of obtaining valid information of the PDF document, judge whether a paragraph starts according to different information regions. If not, take the last punctuation mark to the end from the previous page's body text and splice it to the beginning of the body text.

[0046] Further preferably, in the process of obtaining valid information of the PDF document, obtain the density information of the text in the body region. If the density information is less than the specified percentage, then use OCR recognition. Otherwise, extract according to the position of the text in the page region, where the process of extracting the text includes: page number extraction verification, title extraction, body text extraction, and annotation extraction.

[0047] Further preferably, during the process of determining whether a paragraph starts, according to the marked position, determine the coordinate range of the valid information, and exclude all objects outside this range;

[0048] Traverse the objects within the coordinate range in sequence, extract the text information, and based on the preference for incomplete information, judge whether the paragraph is the beginning for the text information at the top or bottom, where the basis for judgment is: the first character at the top is not indented, and / or the end at the bottom is not a punctuation mark.

[0049] The present invention also discloses a system for extracting valid information of PDF document page elements, including:

[0050] A data acquisition module for acquiring PDF documents;

[0051] A data parsing module for constructing an initial PDF document information extraction model and storing it in a first storage area. The initial PDF document information extraction model is used to generate a time-effective PDF document information extraction rule model; obtain a document parsing rule set, which is used to represent a set of rules for parsing a PDF document to obtain valid information of the PDF document; generate a PDF document information extraction rule model according to the initial PDF document information extraction model and the document parsing rule set, and store it in a second storage area; construct a PDF document information extraction model for extracting valid information of the PDF document according to the initial PDF document information extraction model and the PDF document information extraction rule model; update the PDF document information extraction model according to the initial PDF document information extraction model and the document parsing rule set by setting a first interval time;

[0052] A data correction module records the update time point of each valid information in the generated valid information of the PDF document and the used document parsing rules, generates a PDF document information extraction index model according to the valid information, the update time point, and the document parsing rules; merges the PDF document information extraction index model into the PDF document information extraction model; and judges the accuracy of the corresponding valid information according to the document parsing rules and the update time point according to a second interval time.

[0053] Further preferably, the extraction system further includes:

[0054] A first storage unit for storing the initial PDF document information extraction model;

[0055] A second storage unit for storing the PDF document information extraction rule model;

[0056] A third storage unit for storing the PDF document information extraction index model.

[0057] Example 1: The most necessary and original technical solution for PDF document extraction: such asFigure 1 As shown; based on the above technical solution, improvements have been made, that is, in addition to the necessary input documents, it also includes some necessary input information, such as storage methods (local documents or database systems), information completion preferences (completing according to page context or not), as Figure 2 shown.

[0058] Furthermore, only the folder path where the input document is located is specified, and the program automatically traverses all eligible PDF documents, automatically extracts the valid information of all documents and saves it according to the regulations.

[0059] The method for extracting text from PDF documents provided by the present invention has the following specific process: Open the specified PDF document, and if it does not exist, directly report an error and exit;

[0060] Parse all the objects in the page and arrange them from top to bottom and from left to right according to the coordinates, generally including text boxes, pictures, rectangular boxes, curves, etc.

[0061] Search for the positions of key markers in the obtained objects (the markers generally have lines longer than the length of the dash (only those within a certain range from the top and bottom of the page are valid), the positions where the first or last special-sized characters appear), and divide the page into multiple different information regions according to the marker positions (these regions are generally the titles / section numbers / page numbers in the header, the main text region, the footnotes / page numbers / annotations in the footer, etc.), extract the information of different regions and store it in variables;

[0062] Store information entries in units of pages. To improve efficiency, it is also possible to store information uniformly after the entire document is parsed, that is, place D1 as shown Figure 1 after the document closing operation.

[0063] When necessary, the valid information of this page can be supplemented to a certain extent through the information of the previous and subsequent pages. For example, if a page does not have a page number, the page number of this page can be calculated and finally determined through the page number information of the previous and subsequent pages and the electronic page number information of the electronic document.

[0064] It should be noted that: Similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0065] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. All should be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for extracting valid information of PDF document page elements, characterized in that, Including the following steps: Construct an initial PDF document information extraction model and store it in the first storage area. The initial PDF document information extraction model is used to generate a time-effective PDF document information extraction rule model; Obtain a document parsing rule set, which is used to represent a set of rules for parsing a PDF document to obtain valid information of the PDF document; Generate a PDF document information extraction rule model according to the initial PDF document information extraction model and the document parsing rule set, and store it in the second storage area; Construct a PDF document information extraction model for extracting valid information of the PDF document according to the initial PDF document information extraction model and the PDF document information extraction rule model; By setting a first interval time, update the PDF document information extraction model according to the initial PDF document information extraction model and the document parsing rule set; During the process of using the PDF document information extraction model to obtain valid information of the PDF document, open the PDF document, and parse and arrange the objects of the PDF document in the order from top to bottom and from left to right. Among them, the objects include text boxes, pictures, rectangular boxes, and curves.

2. The method for extracting valid information of PDF document page elements according to claim 1, characterized in that: After the process of obtaining the objects of the PDF document, search for the marked positions of the key marks in the objects according to the objects of the PDF document, and divide the page into multiple different information areas according to the marked positions, and extract the information of different areas and store it in variables.

3. The method for extracting valid information of PDF document page elements according to claim 2, characterized in that: During the process of searching for the positions of the key marks in the objects, the key marks are used to represent lines longer than the dash length; The marked position is used to represent the division of information areas.

4. The method for extracting valid information of PDF document page elements according to claim 3, characterized in that: During the process of dividing the page into multiple different information areas, the different information areas include the title, chapter number, and page number of the header, the body area, and the footnote, page number, and note of the footer.

5. The method for extracting valid information of PDF document page elements according to claim 4, characterized in that: During the process of obtaining valid information of the PDF document, judge whether a paragraph starts according to the different information areas. If not, take the last punctuation mark from the previous page's text to the end and splice it to the beginning of the text.

6. The method for extracting valid information of PDF document page elements according to claim 5, characterized in that: During the process of obtaining valid information of the PDF document, obtain the density information of the text in the body area. If the density information is less than the specified percentage, use OCR recognition. Otherwise, extract according to the position of the text in the page area. Among them, the process of extracting the text includes: page number extraction verification, title extraction, body extraction, and note extraction.

7. The method for extracting valid information of PDF document page elements according to claim 5, characterized in that: During the process of judging whether a paragraph starts, determine the coordinate range of the valid information according to the marked position, and exclude all objects outside this range; Traverse the objects within the coordinate range in sequence, extract text information, and judge whether the paragraph is the beginning according to the text information at the top or bottom according to the preference for incomplete information. The basis for judgment is: the first character at the top is not indented, and / or the end at the bottom is not a punctuation mark.

8. A system for extracting valid information of PDF document page elements, characterized in that, Content for implementing the method for extracting valid information of PDF document page elements according to any one of claims 1-7; The system includes: A data collection module for collecting PDF documents; A data parsing module, which is used to build an initial PDF document information extraction model and store it in the first storage area. The initial PDF document information extraction model is used to generate a time-sensitive PDF document information extraction rule model; obtain a document parsing rule set, which is used to represent a set of rules for parsing a PDF document to obtain valid PDF document information; generate a PDF document information extraction rule model according to the initial PDF document information extraction model and the document parsing rule set, and store it in the second storage area; build a PDF document information extraction model for extracting the valid information of the PDF document according to the initial PDF document information extraction model and the PDF document information extraction rule model; update the PDF document information extraction model according to the initial PDF document information extraction model and the document parsing rule set by setting a first interval time. A data correction module, which records the update time point and the used document parsing rules of each valid information in the generated valid information of the PDF document, and generates a PDF document information extraction index model according to the valid information, the update time point, and the document parsing rules; merge the PDF document information extraction index model into the PDF document information extraction model; and judge the accuracy of the corresponding valid information according to the document parsing rules and the update time point at a second interval time.

9. The system for extracting valid information of PDF document page elements according to claim 8, characterized in that: The extraction system further includes: A first storage unit, which is used to store the initial PDF document information extraction model; A second storage unit, which is used to store the PDF document information extraction rule model; A third storage unit, which is used to store the PDF document information extraction index model.

Citation Information

Patent Citations

  • Information extraction method and apparatus for PDF file

    CN106951400A

  • PDF document paragraph automatic extraction system and device based on deep learning

    CN111259623A