Method and system for automatically extracting PDF (Portable Document Format) document information

Through the combination of preprocessing and layout-aware deep learning model, the identification error and poor applicability problems when automatically extracting key information from massive PDF documents are solved, and high-precision and automated information extraction effect is achieved.

CN120014661APending Publication Date: 2025-05-16SHANGHAI MARITIME UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510055952.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When the prior art automatically extracts key information from massive PDF documents, it faces the problems of identification error, poor applicability and lack of contextual semantics of the information.

Method used

Preprocessing steps are used to form a data set, and the pre-trained layout-aware deep learning model is used for fine-tuning and training of the model, combining multimodal information of text, layout and image, classifying text and extracting key information paragraphs.

Benefits of technology

It improves the accuracy and efficiency of information extraction, avoids the limitations of OCR technology, achieves high automation and wide adaptability, and is suitable for the extraction of key information in professional documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014661A_ABST
    Figure CN120014661A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic PDF (Portable Document Format) document information extraction method, which comprises the following steps of: performing preprocessing operation on a PDF document to form a data set required by a fine tuning training model; then, a pre-trained layout perception deep learning model is utilized, and model fine tuning and training are carried out in combination with data set information; and finally, processing multi-modal information including texts, layouts and images by utilizing the trained model, classifying the texts and extracting key information paragraphs. According to the method, the text and the layout information are combined, so that the precision and efficiency of information extraction are remarkably improved, the limitation of the OCR technology is avoided, the method has high automation and wide adaptability and expansibility, is particularly suitable for key information extraction of professional documents, and provides a more efficient and accurate solution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of information processing and document analysis, and in particular relates to a method and system for automatically extracting PDF document information. Background Art

[0002] With the popularity of electronic documents, especially in the fields of academia, industry, law and business, PDF format has become the main carrier for information publishing and sharing. PDF documents have a fixed structure and consistent display across platforms, making them an important way to transmit high-precision information. However, with the increase in document size and complexity, it becomes increasingly important to automatically extract key information from PDF documents. Efficient automated extraction technology can not only help enterprises and academic institutions quickly obtain key information and reduce manual workload, but also play an important role in large-scale document processing and data mining. Therefore, research and development of PDF information automated extraction technology has important practical significance.

[0003] At present, the mainstream methods of PDF information extraction mainly rely on optical character recognition (OCR) and template-based text parsing technology. The latest research shows that the technology based on the fusion of visual and textual information has obvious advantages in improving the accuracy of information extraction. For example, research work in recent years, such as the application of machine learning and deep learning technology in PDF information extraction, combined with natural language processing (NLP) and layout information solutions. These studies combine text, layout and image information together, greatly improving the parsing ability of complex documents, especially in scenes containing tables and images.

[0004] Although existing technologies have made some progress in PDF information extraction, they still face many challenges. First, traditional OCR methods rely on image recognition and often perform poorly when dealing with complex structures or multilingual documents, especially when dealing with documents with complex layouts such as embedded charts, symbols or annotations. Recognition errors or data omissions are prone to occur. Secondly, template matching-based solutions have poor applicability and cannot flexibly cope with documents of different styles and formats. Templates often need to be customized for each document type, which is time-consuming and difficult to expand. In addition, most current technologies fail to make full use of the layout and visual information in the document, resulting in the extracted information lacking contextual semantics and structural accuracy. With the continuous growth in the number of electronic documents, how to quickly and accurately automatically extract key information from massive PDF documents remains an urgent problem to be solved. Summary of the invention

[0005] The technical problem to be solved by the present invention is to provide a method and system for automatically extracting information from PDF documents, which solves the problem of how to automatically extract key information from massive PDF documents accurately and quickly in the prior art.

[0006] The present invention adopts the following technical solutions to solve the above technical problems: An automatic method for extracting information from PDF documents preprocesses the PDF documents to form a data set required for fine-tuning a training model; then, a pre-trained layout-aware deep learning model is used to fine-tune and train the model in combination with the data set information; finally, the trained model is used to process multimodal information including text, layout, and images, to classify text and extract key information paragraphs.

[0007] The preprocessing operation on the PDF document includes: Step 1: Extract the text information of the PDF document page. Use the text extraction tool to directly extract the text content and its layout information from the PDF document line by line, and obtain the position information of the line at the same time. Step 2. Extract the PDF page image, use the document image conversion tool to convert the PDF page into an image format, adjust the image size to match the actual size of the page, and save the processed image; Step 3: Data preparation and annotation: Use annotation tools to annotate PDF page images and generate the data set required for fine-tuning the model.

[0008] In step 1, the PDF documents in the specified folder are traversed to extract the content of each page, obtain the page size, each line of text and its corresponding line position information.

[0009] During the annotation process, import the preprocessed PDF page image, set the annotation category and the target paragraph, and annotate in paragraphs so that each selected area represents a complete paragraph and is attached with a corresponding label.

[0010] After the annotation is completed, the results are exported and matched with the extracted text and layout information, and the text lines with high matching degree are filtered out and a dataset is generated.

[0011] The specific process of model fine-tuning and training is as follows: First, load the pre-trained model and prepare the dataset. Then, set the training parameters. Then, fine-tune the model by combining the text, layout, and image information extracted from the PDF document. Finally, save the best model checkpoint by regularly evaluating the model performance.

[0012] Use the fine-tuned and trained model to extract key information from the newly input PDF document, split the text and location information of the page into several parts, process each part independently, and perform reasoning on the processed PDF data, parse the model output, and obtain the category and content of each line of text in the PDF page.

[0013] The system for automatically extracting information from PDF documents comprises an information input module, a data processing module and a result output module, wherein the information input module is used to input PDF documents, the data processing module applies the method to extract and mark key information of PDF document pages, and the result output module is used to visually display the processed PDF documents.

[0014] The data processing module aggregates the extracted text lines by category, annotates the extracted text paragraphs with bounding boxes of different colors, and saves the results.

[0015] A computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, all or part of the steps of the method are called.

[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. Avoid recognition errors and improve extraction accuracy The present invention uses a PDF document extraction tool to directly extract text content from PDF by line, rather than relying on OCR technology. OCR often has recognition errors when processing complex layouts, special characters (such as formulas, symbols) and multilingual texts, affecting accuracy. By directly extracting the original text in the document and its precise location information, the recognition errors in OCR technology are avoided, especially when processing documents with complex symbols or professional documents, ensuring high-precision text extraction.

[0017] 2. Combine layout awareness and use row information to improve information extraction accuracy The present invention processes the document content by lines by obtaining the position information of the text and the page layout. The method combined with the layout information can better capture the paragraphs and tables in the document, not only based on the content, but also understand the relative position of each line of text in the page. This method is particularly effective for processing documents with complex layouts, greatly improving the accuracy and efficiency of paragraph extraction.

[0018] 3. Process in segments to ensure complete information In the process of information extraction, the present invention processes the text content and layout information of the page by segmentation while retaining the complete image information. This segmentation processing method avoids information loss caused by model input limitations and ensures that all text paragraphs can be fully extracted even if the page content is large, thereby ensuring the continuity and integrity of information extraction.

[0019] 4. Efficient labeling, reducing labor costs The present invention adopts paragraph-level annotation in the data preparation and annotation process, and selects the target paragraph through the annotation tool and attaches the category label. Compared with line-by-line annotation, paragraph-level annotation is not only more in line with actual application needs, but also greatly reduces the annotation time and labor cost, while improving the efficiency and accuracy of annotation.

[0020] 5. Automated process to improve extraction efficiency Traditional PDF information extraction methods usually require a lot of manual intervention and manual configuration. The present invention combines a pre-trained multimodal model based on document layout perception to achieve a high degree of automation from preprocessing to information extraction, reduce manual operations, and have the ability to process PDF files in batches, which is particularly suitable for large-scale document processing needs.

[0021] 6. Visualize results and enhance interpretability The extracted key information is annotated on the PDF page image, with different colors to distinguish categories, providing an intuitive visual display. This display method not only makes it easier for users to verify the accuracy of the extraction results, but also enhances the interpretability of information extraction, making key information easier to understand and apply.

[0022] The present invention significantly improves the accuracy and efficiency of information extraction by combining text and layout information, avoids the limitations of OCR technology, has high automation, wide adaptability and extensibility, is particularly suitable for extracting key information from professional documents, and provides a more efficient and accurate solution. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a picture of the recognition and extraction effect of common OCR tools.

[0024] Figure 2 This is a schematic diagram of the entire process of extracting key information provided by the present invention.

[0025] Figure 3 This is a flow chart of the model training provided by the present invention.

[0026] Figure 4 This is a table showing the changing trends of the training effect during the model training process of an embodiment of the present invention.

[0027] Figure 5 It is a schematic diagram showing the model extraction effect in an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The structure and working process of the present invention will be further described below in conjunction with the accompanying drawings.

[0029] In view of some problems in the prior art, the present invention proposes a method for extracting key information from PDF documents by combining text extraction tools and layout-aware models. The purpose is to solve the following technical problems: 1. How to efficiently and accurately extract key information from a large number of PDF files in batches, especially extract key paragraphs from PDF documents with complex layouts.

[0030] 2. How to combine the document’s layout information to improve the accuracy of information extraction, especially when it comes to tables, charts, and complex text structures.

[0031] 3. How to handle the diversity and complexity of information in different types of documents to ensure strong adaptability and high versatility.

[0032] 4. How to reduce manual intervention, automate information extraction, improve processing efficiency and reduce labor costs.

[0033] The solution adopted is, first, to directly obtain the text content and its corresponding position information (bounding box) line by line from the PDF document through the text extraction tool, avoiding the recognition errors that may be caused by OCR and retaining the layout structure of the text, such as tables and paragraphs. This method is faster and more efficient than OCR, greatly saving processing time. Then, the pre-trained layout-aware deep learning model is used to fine-tune and train the model in combination with the extracted text and layout information. The model is capable of processing multimodal information of text, layout, and images, so as to better understand the complex structure and typesetting layout of the document. When processing complex PDF documents containing tables, charts, symbols, and multiple languages, the model can accurately classify text and extract key information paragraphs. It greatly improves the efficiency and accuracy of PDF information extraction, reduces manual intervention, and realizes automated processing. It is especially suitable for large-scale information extraction applications of professional documents such as electronic component manuals.

[0034] An automatic method for extracting information from PDF documents preprocesses the PDF documents to form a data set required for fine-tuning a training model; then, a pre-trained layout-aware deep learning model is used to fine-tune and train the model in combination with the data set information; finally, the trained model is used to process multimodal information including text, layout, and images, to classify the text and extract key information paragraphs.

[0035] Specific embodiments, such as Figures 1 to 5 As shown, this solution is introduced in detail by taking the method of extracting key information from the data sheet of electronic components as an example.

[0036] A method for automatic extraction of PDF document information is mainly divided into the following six steps: Step 1: PDF document preprocessing. Use the PDF document extraction tool to directly extract text content and its layout information from PDF documents line by line, rather than relying on OCR technology for image text recognition. When extracting line text, the tool will also obtain the position information of the line (such as the line bounding box coordinates), so that when processing paragraphs or tables later, text positioning can be performed more efficiently, thereby improving efficiency and saving time. The specific process includes: traversing the PDF documents in the specified folder, extracting the content of each page, and obtaining the page size, each line of text and its corresponding line position information.

[0037] In this embodiment, a certain number of electronic component PDF data manuals are collected as a set of documents to be processed. A PDF document extraction tool is used to directly extract text, page size and layout information from the PDF file, including the position information of each line of text. This ensures that the position of the text is accurately identified when the text is extracted, avoiding errors caused by OCR recognition.

[0038] Step 2: Page image extraction. In order to further utilize the visual information in the PDF, use the document image conversion tool to convert the PDF page to image format. By adjusting the image size to be consistent with the actual size of the page, the information can be visualized later. The specific process includes: converting the PDF page to image format, reading the previously saved page size information, adjusting the image size, and saving the processed image.

[0039] This step ensures the precise positioning of the text information on the image and saves the adjusted image file for subsequent annotation work and visual display of information extraction.

[0040] Step 3: Data preparation and annotation. Use annotation tools to annotate PDF page images and generate the dataset required for fine-tuning the model. During the annotation process, you need to annotate in paragraphs, ensuring that each boxed area represents a complete paragraph and is labeled accordingly. The specific process includes: importing the preprocessed PDF page image, setting the annotation category, annotating the target paragraph and attaching the corresponding label. After the annotation is completed, export the results and match them with the extracted text and layout information, filter out the text lines with high matching degree and generate the dataset.

[0041] In this embodiment, an annotation tool is used to annotate some PDF data manuals, and the categories include "features", "descriptions", and "applications". The annotation is done in paragraphs to ensure that the target paragraphs are completely selected and attached with corresponding category labels. Paragraphs without corresponding labels are marked as "others" when constructing the dataset; the annotation file is parsed, the annotation information is extracted and matched with the text content and layout information extracted in the preprocessing stage, the matching degree is calculated to ensure that the annotated paragraphs are accurately aligned with the text lines, and a standardized dataset is generated; the dataset is divided into a training set and a validation set in a certain ratio to ensure that the model can effectively evaluate its generalization ability during the training process.

[0042] Step 4: Model training, using a pre-trained document layout-aware multimodal model, combined with the processing capabilities of text and layout information to cope with the text structure of complex documents. Set training parameters to ensure that the model achieves optimal performance without overfitting. Use mixed precision training to improve efficiency, especially when processing large-scale documents, which helps reduce the use of computing resources. Regularly evaluate model performance, calculate the accuracy, recall, and F1 score of the validation set, dynamically adjust model parameters, and save the best model checkpoint. The final evaluation results show that the model can efficiently and accurately extract key information paragraphs in the electronic component manual. Figure 4 It shows the changing trends of various performance indicators during the model training process, including changes in F1, training loss, validation loss, accuracy, recall, and precision.

[0043] Step 5: Information extraction. Use the fine-tuned model to extract key information from new PDF documents. To avoid missing information due to model input limitations, the text and location information of the page is divided into several parts, each of which is processed independently to ensure that the page layout is fully preserved. The extraction process includes: reasoning on the processed PDF data, parsing the model output, and obtaining the category and content of each line of text in the PDF page.

[0044] This embodiment extracts key information from the remaining unlabeled PDF data manuals, uses the trained model to process the text of each page of the PDF document, and automatically identifies and extracts key information paragraphs in categories such as "features", "descriptions", and "applications" by combining text and layout information. The extracted information is summarized by category to generate structured text data for further analysis and application.

[0045] Step 6: Output and visualization of results. The extracted key information will be summarized by category and displayed on the PDF page image through visualization tools for easy verification and understanding by users. The specific process includes: summarizing the extracted text lines by category, annotating the extracted text paragraphs with bounding boxes of different colors, and saving the results.

[0046] Generate visualization results for each data sheet, annotating the extracted text paragraphs on the original PDF page image. Different colors are used to distinguish different categories of text paragraphs, providing an intuitive display method to help users quickly verify and understand the distribution of key information. Figure 5 A schematic diagram showing the extraction effect of the present invention is shown, which intuitively shows key information sections such as "characteristics" and "descriptions" automatically extracted from the electronic component data manual, and visually annotates them on the PDF page image.

[0047] It can be seen from this embodiment that the method of the present invention successfully extracts key paragraphs automatically from the electronic component data manual and intuitively displays them through visual display, which significantly improves the efficiency and accuracy of information extraction, reduces manual intervention, and promotes the automation process of document parsing.

[0048] In addition, in the financial field, financial reports and audit reports in PDF format often contain complex tables and paragraphs. This embodiment demonstrates the application of the method of the present invention in financial documents, with the goal of quickly extracting key information from annual financial statements, and mainly includes the following steps: Step 1: Extract text from the PDF document of the financial report, obtain the data in the table and paragraph row by row, and record its layout position information. This process can accurately obtain the column information and row values ​​in the table, avoiding errors caused by OCR processing.

[0049] Step 2: Convert each page of the PDF document into image format, adjust the image size to match the page layout, and retain the visual information of tables and other elements to support subsequent annotation and model training.

[0050] Step 3: Use the annotation tool to select important paragraphs and table data in the financial statements, such as "total assets", "current assets", "net assets" and other key information in the "balance sheet", and add corresponding category labels to each selected area. After the annotation is completed, match the annotation results with the text information to generate a high-quality data set.

[0051] Step 4: Load the pre-trained model and fine-tune the model based on the text and layout information in the financial statements. The model is optimized for tables, symbols, and complex paragraphs to ensure that important data, including those across columns and rows, can be accurately extracted.

[0052] Step 5: Apply the trained model to new financial reports, automatically extract key information such as "assets", "liabilities", and "net profit", and summarize them to generate structured data files.

[0053] Step 6: Visually annotate the extracted information on the original PDF image, such as marking the "Total Assets" field with different color borders, to intuitively display the classification results for user verification and analysis.

[0054] This embodiment verifies that the method of the present invention can quickly and accurately extract key indicators from financial reports with complex layouts, greatly reduce the time for manual data sorting, and improve analysis efficiency.

[0055] Another specific example is the extraction of key information from academic papers. Academic papers usually contain multiple types of information modules, such as abstracts, introductions, methods, results, and conclusions, and their document structures are complex and their layouts are diverse. This example shows the specific application of the method of the present invention in extracting information from academic papers: Step 1: Extract the text content and location information of the paper line by line, especially the information of the title, paragraph, figure description and other areas, and record the layout.

[0056] Step 2: Convert the paper pages into image format to ensure that the layout information of charts and formulas can be preserved in subsequent processing.

[0057] Step 3: Annotate the paragraphs of the paper page images, and classify and annotate modules such as "Abstract", "Research Methods", and "Experimental Results". In particular, separate annotation and labeling of chart titles and important data paragraphs. The annotation results are matched with the extracted text to form a complete data set.

[0058] Step 4: Use the pre-trained layout-aware deep learning model to fine-tune the model on the multimodal information of the paper, especially optimizing the model's ability to handle charts and multilingual text in academic scenarios.

[0059] Step 5: Process the new academic paper PDF document, automatically classify and extract key information paragraphs such as "research purpose", "experimental data", and "conclusion", and generate structured output.

[0060] Step 6: Display the extracted results in a visual way, such as adding different color highlighting in the summary section and experimental results section, and provide structured information export files for further analysis.

[0061] This example verifies that the method of the present invention can efficiently extract important information modules from academic papers, significantly improve the efficiency of information retrieval and processing, and provide a convenient tool for scientific researchers.

[0062] The system for automatically extracting information from PDF documents comprises an information input module, a data processing module and a result output module, wherein the information input module is used to input PDF documents, the data processing module applies the method to extract and mark key information of PDF document pages, and the result output module is used to visually display the processed PDF documents.

[0063] The data processing module aggregates the extracted text lines by category, annotates the extracted text paragraphs with bounding boxes of different colors, and saves the results.

[0064] A computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, all or part of the steps of the method are called.

[0065] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), disk or optical disk, etc., which can store program code.

[0066] Those skilled in the art should understand that they can implement variations by combining the prior art and the above embodiments. Such variations do not affect the essence of the present solution and are not described in detail here.

[0067] It should be understood that the present solution is not limited to the above-mentioned specific implementation methods, and the devices and structures not described in detail should be understood to be implemented in a common manner in the art; any technician familiar with the art can use the above-disclosed methods and technical contents to make many possible changes and modifications to the technical solution of the present solution without departing from the scope of the technical solution of the present solution, or modify it into an equivalent embodiment with equivalent changes, which does not affect the essential content of the present solution. Therefore, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present solution without departing from the content of the technical solution of the present solution still falls within the scope of protection of the technical solution of the present solution.

Claims

1. A method for automatically extracting information from a PDF document, characterized in that: The PDF document is preprocessed to form the data set required for fine-tuning the training model. Then, the pre-trained layout-aware deep learning model is used to fine-tune and train the model in combination with the data set information. Finally, the trained model is used to process multimodal information including text, layout, and images to classify the text and extract key information paragraphs.

2. The method for automatically extracting information from a PDF document according to claim 1, characterized in that: The preprocessing operation on the PDF document includes: Step 1: Extract the text information of the PDF document page. Use the text extraction tool to directly extract the text content and its layout information from the PDF document line by line, and obtain the position information of the line at the same time. Step 2. Extract the PDF page image, use the document image conversion tool to convert the PDF page into an image format, adjust the image size to match the actual size of the page, and save the processed image; Step 3: Data preparation and annotation: Use annotation tools to annotate PDF page images and generate the data set required for fine-tuning the model.

3. The method for automatically extracting information from a PDF document according to claim 2, characterized in that: In step 1, the PDF documents in the specified folder are traversed to extract the content of each page, obtain the page size, each line of text and its corresponding line position information.

4. The method for automatically extracting information from a PDF document according to claim 2, wherein: During the annotation process, import the preprocessed PDF page image, set the annotation category and the target paragraph, and annotate in paragraphs so that each selected area represents a complete paragraph and is attached with a corresponding label.

5. The method for automatically extracting information from a PDF document according to claim 4, characterized in that: After the annotation is completed, the results are exported and matched with the extracted text and layout information, and the text lines with high matching degree are filtered out and a dataset is generated.

6. The method for automatically extracting information from a PDF document according to claim 1, characterized in that: The specific process of model fine-tuning and training is as follows: First, load the pre-trained model and prepare the dataset. Then, set the training parameters. Then, fine-tune the model by combining the text, layout, and image information extracted from the PDF document. Finally, save the best model checkpoint by regularly evaluating the model performance.

7. The method for automatically extracting information from a PDF document according to claim 1, characterized in that: Use the fine-tuned and trained model to extract key information from the newly input PDF document, split the text and location information of the page into several parts, process each part independently, and perform reasoning on the processed PDF data, parse the model output, and obtain the category and content of each line of text in the PDF page.

8. PDF document information automatic extraction system, characterized by: The invention comprises an information input module, a data processing module and a result output module, wherein the information input module is used to input a PDF document, the data processing module applies the method described in any one of claims 1 to 7 to extract and mark key information of a PDF document page, and the result output module is used to visually display the processed PDF document.

9. The system for automatically extracting information from PDF documents according to claim 8, characterized in that: The data processing module aggregates the extracted text lines by category, annotates the extracted text paragraphs with bounding boxes of different colors, and saves the results.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, all or part of the steps of the method described in any one of claims 1 to 8 are called.

Citation Information

Cited By

  • Document structure extraction and model training method and device, equipment and medium

    CN121600519A