Document information extraction method, device and system and storage medium

By combining OCR, layout analysis and large-scale models, the problem of traditional methods being difficult to extract complex document logic hierarchical information is solved, and document information extraction with high accuracy and automation is achieved, which is suitable for a variety of complex document layouts.

CN119942576APending Publication Date: 2025-05-06BEIJING YIDAO BOSHI TECH
View PDF 0 Cites 18 Cited by

Patent Information

Application Number
CN202411872348.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Traditional document information extraction methods are difficult to effectively extract logical hierarchical information hidden in layout in complex documents, especially when document typesetting is complex and content structure is diversified.

Method used

Combining the technical means of OCR, layout analysis and large model, text information is identified through OCR, layout analysis is divided into document layout, and information extraction is used to form a Prompt template for fine-tuning training to achieve accurate extraction of document information.

Benefits of technology

It significantly improves the accuracy of information extraction, adapts to a variety of complex document layouts, reduces manual intervention, generates structured JSON format output, and improves the automation and scalability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942576A_ABST
    Figure CN119942576A_ABST
Patent Text Reader

Abstract

The invention discloses a document information extraction method, device and system and a storage medium, and relates to the field of deep learning, and the method comprises the steps: inputting an original document for preprocessing, and obtaining a document image; the method comprises the following steps: carrying out OCR (Optical Character Recognition) on a document image by adopting a Vision Transform (ViT) to obtain text information and textbox coordinate information corresponding to the text information; determining the category and coordinate information of each layout element frame based on a deep learning yolov8-seg instance segmentation algorithm, and then performing layout region matching on the coordinate information of the layout element frames and the coordinate information of the textbox to obtain text information corresponding to each layout element frame; and taking the category of each layout element box and the corresponding text information as a layout area matching result, forming a Prompt template by combining to-be-extracted document information, taking the Prompt template as the input of a large model and performing fine tuning training, and after the fine tuning training is completed, correctly extracting the document information by the model according to the input. According to the method, character recognition of OCR, layout analysis of layout analysis and language understanding ability of a large model are combined, and key information can be accurately extracted from complex and diversified documents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of deep learning, and in particular to a document information extraction method, device, system and storage medium. Background Art

[0002] In the field of modern information processing, document digitization and automated information extraction have become key technologies and are widely used in many industries such as archive management, legal document processing, and financial statement analysis. How to quickly and accurately extract useful information from massive documents has become one of the main challenges in the field of information processing. Among traditional methods, OCR technology is widely used and can convert text information in paper documents into editable text. In actual applications, the content of a document is not just a simple collection of text. The layout, format, and typesetting structure of the document often contain rich semantic information. Figure 1 As shown in the figure, the document image is a sample of a VAT electronic invoice, which contains a lot of structured information, including titles, QR codes, seals, tables, etc. These structured information are crucial to understanding the content of the document. Simply relying on OCR can only obtain text information, but cannot effectively extract these valuable logical hierarchical information hidden in the layout.

[0003] In order to solve this problem, layout analysis technology is proposed. The core purpose of layout analysis technology is to analyze the visual structure of the document and divide the document image into several layout elements, such as titles, text paragraphs, tables, pictures, etc. By dividing the document into various logical units, it is not only possible to achieve hierarchical information extraction, but also to retain the structure and layout of the document, thereby improving the depth and accuracy of information understanding, which is crucial for subsequent content extraction, summary generation, and information retrieval tasks.

[0004] In recent years, with the rapid rise of large models, such as GPT, LLAMA and other pre-trained language models, information extraction technology has entered a new stage of development. Large models are pre-trained on massive corpora, and have powerful language understanding, reasoning and generation capabilities, and can handle highly complex natural language tasks. Combining OCR and layout analysis technology, the structural information in the document is input into the large model together with the text content, which can fully utilize the semantic understanding ability of the large model to achieve more intelligent document information extraction.

[0005] In summary, the complex document information extraction method that combines OCR, layout analysis and large models overcomes the limitations of traditional single technical solutions, can effectively deal with situations where document layout is complex and content structure is diverse, and provides a more comprehensive and intelligent solution for the automated processing and intelligent information extraction of various types of documents. Summary of the invention

[0006] To this end, the technical solution of the present invention proposes a document information extraction method, device, system and storage medium based on optical character recognition (OCR), layout analysis and large models. The method combines the text recognition of OCR, the layout analysis of layout analysis and the language understanding ability of the large model, and can accurately extract key information from complex and diverse documents.

[0007] According to a first aspect of the technical solution of the present invention, a document information extraction method is provided, comprising:

[0008] S1, document preprocessing step: input the original document for preprocessing to obtain a document image;

[0009] S2, OCR recognition step: using Vision Transformer (ViT) to perform OCR recognition on the document image to obtain text information and text box coordinate information corresponding to the text information;

[0010] S3, layout analysis step: based on the yolov8-seg instance segmentation algorithm of deep learning, determine the category and coordinate information of each layout element frame, then match the layout element frame coordinate information with the text frame coordinate information for layout area, and obtain the text information corresponding to each layout element frame;

[0011] S4, information extraction step: take the category and corresponding text information of each layout element box as the layout area matching result, combine it with the document information to be extracted to form a Prompt template, use the Prompt template as the input of the large model and perform fine-tuning training. After the fine-tuning training is completed, the large model can correctly extract the document information according to the input.

[0012] Further, in S1, the preprocessing operation includes:

[0013] If the original document is a PDF file, the document type is first converted into an image file, and then the image feature is processed to obtain a document image;

[0014] If the original document is in image format, image feature processing is performed directly to obtain the document image.

[0015] Furthermore, the graphic feature processing includes:

[0016] Denoising: Use Gaussian blur algorithm to remove image noise and improve image clarity;

[0017] Grayscale: Convert color images into grayscale images to reduce data redundancy and simplify subsequent processing steps;

[0018] Image scaling: Proper scaling is performed based on the resolution and size of the document image;

[0019] Rotation Correction: Detect tilted text lines through Hough transform, calculate the best rotation angle, and perform rotation correction to ensure horizontal or vertical denoising, grayscale, image scaling, and rotation correction of document content.

[0020] Furthermore, in S3, the categories of the layout elements include document title, table of contents, paragraphs, information blocks, tables, pictures, headers, footers, page numbers, signatures, seals, chart annotations, chart titles, formulas, and columns.

[0021] Furthermore, in S3, the coordinate information of the layout element frame is specifically the coordinates of the four vertices of each layout element frame, and the text frame coordinate information is the coordinates of the four vertices of each text frame.

[0022] Furthermore, the S3 specifically includes:

[0023] Based on the deep learning yolov8-seg instance segmentation algorithm, each layout element frame in the document image is classified and segmented to determine the category and coordinate information of each layout element frame;

[0024] Calculate the coordinate overlap between the text box and the layout element box, match the layout element box coordinate information with the text box coordinate information in the layout area, and when the overlap exceeds a set threshold, determine that the text box matches the layout element box, thereby obtaining the text information corresponding to each of the layout element boxes.

[0025] Furthermore, in S4, in the Prompt template, the document information to be extracted is defined to include two field types: "entity field" and "table field", wherein the extraction result of the "entity field" is output in the form of a key-value pair, and the extraction result of the "table field" is output in the form of a two-dimensional list; the extraction range is defined as the entire document image; and it is required to be output in JSON format.

[0026] Furthermore, in S4, after the two field types are input, they are automatically filled into the "entity field" and "form field" in the Prompt template, and the layout area matching result is filled into the label to form the Prompt template.

[0027] Furthermore, in S4, the fine-tuning training specifically refers to: the large model learns the corresponding relationship between the Prompt template and the label value, and optimizes its own parameters to gradually improve the ability of document information extraction.

[0028] Furthermore, the large model is a large language model (Large Language Model).

[0029] Here, large language models specifically refer to large neural network models trained using deep learning technology, focusing on natural language processing tasks. These models have powerful language understanding and generation capabilities and can be used for tasks such as text generation, machine translation, and information extraction. Prompts are instructions or questions given to the model when interacting with the large language model, which are used to guide the model to generate outputs that meet expectations.

[0030] Additionally, the "Tag Value" is defined as follows:

[0031] When fine-tuning the large model, a prompt is constructed for each sample as the input of the large model; each sample has a corresponding manually annotated answer, that is, a label value label, which is in JSON format and is 100% correct. It is expected that the prediction result of the large model is exactly the same as the label value, that is, the accuracy rate reaches 100%. As fine-tuning proceeds, the prediction result of the large model will get closer and closer to the label value, and the accuracy rate will also increase.

[0032] According to a second aspect of the technical solution of the present invention, there is provided a document information extraction device, the document information extraction device operates based on the document information extraction method according to any one of the above aspects, comprising:

[0033] A document preprocessing unit, used for inputting an original document for preprocessing to obtain a document image;

[0034] The OCR recognition unit is used to perform OCR recognition on the document image using Vision Transformer (ViT) to obtain text information and text box coordinate information corresponding to the text information.

[0035] A layout analysis unit, which is used to determine the category and coordinate information of each layout element frame based on the yolov8-seg instance segmentation algorithm of deep learning, and then perform layout area matching between the layout element frame coordinate information and the text frame coordinate information to obtain text information corresponding to each layout element frame;

[0036] The information extraction unit is used to take the category and corresponding text information of each layout element box as the layout area matching result, combine it with the document information to be extracted to form a Prompt template, use the Prompt template as the input of the large model and perform fine-tuning training, and output the corresponding document information after the fine-tuning training is completed.

[0037] According to the third aspect of the technical solution of the present invention, a document information extraction system is provided, the system comprising: a processor and a memory for storing executable instructions; wherein the processor is configured to execute the executable instructions to execute the document information extraction method as described in any of the above aspects.

[0038] According to a fourth aspect of the technical solution of the present invention, there is provided a computer-readable storage medium, wherein a computer program is stored thereon, and when the computer program is executed by a processor, the document information extraction method as described in any of the above aspects is implemented.

[0039] Beneficial effects of the present invention:

[0040] 1. Improve the accuracy of information extraction: The present invention introduces a layout analysis module that can identify the document layout and finely divide the document into blocks to ensure that the OCR recognition results are correctly attributed, which significantly improves the accuracy of information extraction.

[0041] 2. Adapt to various complex document layouts: The layout analysis module introduced in the present invention can also classify and segment different types of documents, such as newspapers, books, papers and financial reports, and handle multiple layouts, languages ​​and formats. This diverse adaptability enables the system to play a role in various document types, expands the scope of application of information extraction, and solves the problem that traditional technologies cannot effectively deal with complex layouts.

[0042] 3. Automatically generate prompts using prompt templates: The output content of the layout analysis module is directly filled into the prompt template as a prompt for the big model, reducing manual intervention and improving the automation of the system. The prompt template is dynamically generated based on the document content and user needs, ensuring that the big model understands the contextual semantics of the document and provides targeted answers or information extraction results that meet user needs.

[0043] 4. Structured JSON format output: The present invention fine-tunes the large model, transforming the model from a general natural language processing model into an information extraction domain model, which can generate structured output with high accuracy and unified format standards.

[0044] 5. Strong system scalability and compatibility: Each module designed by the present invention has good scalability and can directly integrate the OCR system and layout analysis technology. Such scalability and compatibility enable the system to adapt to the needs of different industries and have a long technical life cycle. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the structures shown in these drawings without paying creative work.

[0046] Figure 1 An example document image is shown.

[0047] Figure 2A system flow chart of an embodiment of the technical solution of the present invention is shown.

[0048] Figure 3 A schematic diagram of a Prompt template according to an embodiment of the technical solution of the present invention is shown.

[0049] Figure 4 A Prompt example diagram of an embodiment of the technical solution of the present invention is shown.

[0050] Figure 5 An example diagram of label values ​​of an embodiment of the technical solution of the present invention is shown.

[0051] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0052] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0053] The terms "first", "second", etc. in the specification and claims of the present disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein, for example.

[0054] In addition, the terms "comprises," "comprising," and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or elements is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to such process, method, product, or apparatus.

[0055] Multiple includes two or more.

[0056] And / or, it should be understood that the term "and / or" used in this disclosure is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. It reduces manpower input and helps business automation, with the characteristics of universality, efficiency, and high precision.

[0057] The present invention relates to a method, device, system and storage medium for extracting document structured information using OCR, layout analysis and large models. By combining OCR technology, layout analysis and large models, text information in documents can be effectively identified, and specific information can be extracted according to user needs. The current mainstream document information extraction methods do not do well in extracting information from complex documents. The main reasons are as follows: layout analysis technology is not used. Plain text corpus does not contain the layout information of the original document. The document is processed into plain text data for information extraction, and structural elements such as tables, seals, and page numbers cannot be identified, and important information is missing. Layout analysis technology can identify document layout, finely divide documents into blocks, and ensure that the OCR recognition results are correctly attributed; large model technology is not used. Traditional NLP methods have few model parameters, weak processing capabilities for long texts and complex documents, poor generalization capabilities, large model parameter scale, and have undergone large-scale corpus pre-training. It contains rich world knowledge, can understand and process long-distance context dependencies, and has stronger generalization capabilities.

[0058] The present invention innovatively combines OCR, layout analysis and large models to achieve information extraction from documents. The main innovations are as follows:

[0059] 1. Introducing layout analysis technology, the OCR results are rearranged according to the layout of layout elements, retaining the original layout information of the document.

[0060] 2. Use big model technology to replace traditional NLP methods. The present invention fine-tunes the big model to enhance its information extraction capability, fills the information that the user wants to extract into the prompt template of the big model, and accurately extracts the information that the user is interested in.

[0061] Specifically, the technical solution of the present invention first provides a document information extraction method, comprising:

[0062] S1. Document preprocessing step: input the original document for preprocessing to obtain a document image.

[0063] In a preferred embodiment, in S1, the preprocessing operation includes:

[0064] If the original document is a PDF file, the document type is first converted into an image file, and then the image feature is processed to obtain a document image;

[0065] If the original document is in image format, image feature processing is performed directly to obtain the document image.

[0066] In a preferred embodiment, the graphic feature processing includes:

[0067] Denoising: Use Gaussian blur algorithm to remove image noise and improve image clarity;

[0068] Grayscale: Convert color images into grayscale images to reduce data redundancy and simplify subsequent processing steps;

[0069] Image scaling: Proper scaling is performed based on the resolution and size of the document image;

[0070] Rotation Correction: Detect tilted text lines through Hough transform, calculate the best rotation angle, and perform rotation correction to ensure horizontal or vertical denoising, grayscale, image scaling, and rotation correction of document content.

[0071] S2, OCR recognition step: using ViT to perform OCR recognition on the document image to obtain text information and text box coordinate information corresponding to the text information.

[0072] S3, layout analysis step: based on the deep learning yolov8-seg instance segmentation algorithm, determine the category and coordinate information of each layout element box, and then match the layout element box coordinate information with the text box coordinate information to obtain the text information corresponding to each layout element box.

[0073] In a preferred embodiment, in S3, the categories of layout elements include document title, table of contents, paragraphs, information blocks, tables, pictures, headers, footers, page numbers, signatures, stamps, chart annotations, chart titles, formulas, and columns.

[0074] In a preferred embodiment, in S3, the coordinate information of the layout element frame is specifically the coordinates of the four vertices of each layout element frame, and the text frame coordinate information is the coordinates of the four vertices of each text frame.

[0075] In a preferred embodiment, S3 specifically includes:

[0076] Based on the deep learning yolov8-seg instance segmentation algorithm, each layout element frame in the document image is classified and segmented to determine the category and coordinate information of each layout element frame;

[0077] Calculate the coordinate overlap between the text box and the layout element box, match the layout element box coordinate information with the text box coordinate information in the layout area, and when the overlap exceeds a set threshold, determine that the text box matches the layout element box, thereby obtaining the text information corresponding to each of the layout element boxes.

[0078] S4, information extraction step: take the category and corresponding text information of each layout element box as the layout area matching result, combine it with the document information to be extracted to form a Prompt template, use the Prompt template as the input of the large model and perform fine-tuning training, and output the corresponding document information after the fine-tuning training is completed.

[0079] In a preferred embodiment, in S4, in the Prompt template, the document information to be extracted is defined to include two field types: "entity field" and "table field", wherein the extraction result of the "entity field" is output in the form of a key-value pair, and the extraction result of the "table field" is output in the form of a two-dimensional list; the extraction range is defined as the layout area matching result; and it is required to be output in JSON format.

[0080] In a preferred embodiment, in S4, the fine-tuning training specifically means: the label value of the sample is in JSON format, the large model learns the label values ​​corresponding to different prompt templates, and optimizes its own parameters to gradually improve the ability to extract document information.

[0081] In a preferred embodiment, the large model is a Large Language Model.

[0082] The technical solution of the present invention also provides a document information extraction device, which operates based on the document information extraction method described above, including:

[0083] A document preprocessing unit, used for inputting an original document for preprocessing to obtain a document image;

[0084] The OCR recognition unit is used to perform OCR recognition on the document image using ViT to obtain text information and text box coordinate information corresponding to the text information.

[0085] A layout analysis unit, which is used to determine the category and coordinate information of each layout element frame based on the yolov8-seg instance segmentation algorithm of deep learning, and then perform layout area matching between the layout element frame coordinate information and the text frame coordinate information to obtain text information corresponding to each layout element frame;

[0086] The information extraction unit is used to take the category and corresponding text information of each layout element box as the layout area matching result, combine it with the document information to be extracted to form a Prompt template, use the Prompt template as the input of the large model and perform fine-tuning training, and output the corresponding document information after the fine-tuning training is completed.

[0087] The technical solution of the present invention further provides a document information extraction system, which includes: a processor and a memory for storing executable instructions; wherein the processor is configured to execute the executable instructions to execute the document information extraction method as described above.

[0088] The technical solution of the present invention further provides a computer-readable storage medium, wherein a computer program is stored thereon, and when the computer program is executed by a processor, the document information extraction method as described above is implemented.

[0089] Example

[0090] A document information extraction method, device, system and storage medium, the system flow chart of which is as follows Figure 2 As shown, including:

[0091] 1. Document preprocessing module

[0092] The main function of the document preprocessing module is to convert the input document type and process the image features to ensure that the subsequent OCR recognition and layout analysis can be performed on high-quality images. This module consists of the following two parts:

[0093] 1) Document type conversion

[0094] In order to unify the document format, this part converts the input PDF file into an image file. If it is a multi-page PDF, it will be split into multiple single-page documents according to the page number and converted into corresponding image files one by one. For documents in other formats such as JPEG, PNG and other image formats, they will directly enter the subsequent image processing process.

[0095] 2) Image feature processing

[0096] In order to improve the effects of OCR recognition and layout analysis, this part performs a series of optimization processes on document images, including: Denoising: using Gaussian blur algorithm to remove image noise and improve image clarity; Grayscale: converting color images into grayscale images to reduce data redundancy and simplify subsequent processing steps; Image scaling: scaling appropriately according to the resolution and size of the document image; Rotation correction: detecting tilted text lines through Hough transform, calculating the optimal rotation angle, and performing rotation correction to ensure that the document content is horizontal or vertical.

[0097] 2.OCR module

[0098] The OCR model is based on deep learning and can accurately recognize text in multiple fonts and languages. It extracts text areas in documents and converts them into editable text, providing basic data for layout analysis and information extraction. The model uses Visual Transformer (ViT), the core idea of ​​which is to treat images as sequences rather than extracting features layer by layer. The specific steps include: 1) Image encoding: divide the input image into small blocks of 16×16 pixels and convert them into vectors; 2) Position encoding: add position vectors to distinguish the positions of different image blocks in the original image; 3) Self-attention mechanism: capture the relationship between characters and contextual information to help understand the semantics and layout between characters. This module effectively captures character edges and details, and is suitable for handwritten signatures, complex fonts, and low-resolution text recognition. The output results include text information, coordinate information, and confidence. Coordinate information is an important basis for layout analysis.

[0099] 3.Layout Analysis Module

[0100] This module is used for accurate document layout segmentation, identifying and classifying different layout elements, such as titles, directories, paragraphs, tables, etc., and generating structured coordinate information and text content. The present invention uses the yolov8-seg instance segmentation algorithm based on deep learning to segment the document layout. In specific operation, the OCR module outputs the coordinates of the four vertices of each text box, and the layout segmentation algorithm outputs the coordinates of the four vertices of each layout element box. Next, the overlap between the text box and the layout element box is calculated to determine their overlap. When the overlap exceeds the set threshold, the matching of the text box and the layout element box is determined.

[0101] 4. Information extraction module

[0102] This module is divided into two parts: Prompt template and large model fine-tuning.

[0103] 1) Prompt template

[0104] The present invention designs two types of fields to be extracted: entity fields and table fields. The extraction results of entity fields are output in the form of key-value pairs, and the extraction results of table fields are output in the form of two-dimensional lists. After the user enters the two fields to be extracted, they are automatically filled into the "entity fields" and "table fields" in the template, and the layout area matching results are filled into <text>In the tag, a targeted prompt template is formed. The prompt template is as follows Figure 3 shown.

[0105] by Figure 1 For example, assuming that the entity fields that the user wants to extract are invoice code, taxpayer identification number, buyer name, seller name, and invoice issuer, and the table fields are service name, specification model, unit price, amount, and tax amount, after combining the layout area matching results, the final prompt example is as follows Figure 4 shown.

[0106] 2) Fine-tuning large models

[0107] The large model has strong general language capabilities and can become an expert in this task type through specialized fine-tuning training. During the fine-tuning training process, the label value uses the JSON format to represent the expected extraction results. The large model learns the label values ​​corresponding to different prompts, optimizes its own parameters, and gradually improves its ability to extract information from documents.

[0108] by Figure 4 For example, the label value of the model training is Figure 5 As shown, if the "Specification Model" field is empty, "null" is returned.

[0109] It should be noted that, in this article, the term "includes", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.

[0110] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0111] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0112] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation modes, which are merely illustrative rather than restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.< / text>

Claims

1. A document information extraction method, characterized in that: include: S1, document preprocessing step: input the original document for preprocessing to obtain a document image; S2, OCR recognition step: using Vision Transformer to perform OCR recognition on the document image to obtain text information and text box coordinate information corresponding to the text information; S3, layout analysis step: based on the yolov8-seg instance segmentation algorithm of deep learning, determine the category and coordinate information of each layout element frame, then match the layout element frame coordinate information with the text frame coordinate information for layout area, and obtain the text information corresponding to each layout element frame; S4, information extraction step: take the category and corresponding text information of each layout element box as the layout area matching result, combine it with the document information to be extracted to form a Prompt template, use the Prompt template as the input of the large model and perform fine-tuning training. After the fine-tuning training is completed, the large model correctly extracts the document information according to the input.

2. The document information extraction method according to claim 1, characterized in that: In S1, the preprocessing operation includes: If the original document is a PDF file, the document type is first converted into an image file, and then the image feature is processed to obtain a document image; If the original document is in image format, image feature processing is performed directly to obtain the document image.

3. The document information extraction method according to claim 2, characterized in that: The graphic feature processing includes: Denoising: Use Gaussian blur algorithm to remove image noise and improve image clarity; Grayscale: Convert color images into grayscale images to reduce data redundancy and simplify subsequent processing steps; Image scaling: Proper scaling is performed based on the resolution and size of the document image; Rotation Correction: Detect tilted text lines through Hough transform, calculate the best rotation angle, and perform rotation correction to ensure horizontal or vertical denoising, grayscale, image scaling, and rotation correction of document content.

4. The document information extraction method according to claim 1, characterized in that: In S3, the categories of the layout elements include document title, table of contents, paragraph, information block, table, picture, header, footer, page number, signature, seal, chart annotation, chart title, formula, and column.

5. The document information extraction method according to claim 1, characterized in that: In S3, the coordinate information of the layout element frame is specifically the coordinates of the four vertices of each layout element frame, and the text frame coordinate information is the coordinates of the four vertices of each text frame.

6. The document information extraction method according to claim 1, characterized in that: The S3 specifically includes: Based on the deep learning yolov8-seg instance segmentation algorithm, each layout element frame in the document image is classified and segmented to determine the category and coordinate information of each layout element frame; Calculate the coordinate overlap between the text box and the layout element box, match the layout element box coordinate information with the text box coordinate information in the layout area, and when the overlap exceeds a set threshold, determine that the text box matches the layout element box, thereby obtaining the text information corresponding to each of the layout element boxes.

7. The document information extraction method according to claim 1, characterized in that: In S4, in the Prompt template, the document information to be extracted is defined to include two field types: "entity field" and "table field", wherein the extraction result of the "entity field" is output in the form of a key-value pair, and the extraction result of the "table field" is output in the form of a two-dimensional list; the extraction range is defined as the entire document image; and it is required to be output in JSON format.

8. The document information extraction method according to claim 7, characterized in that: In S4, after the two field types are input, they are automatically filled into the "entity field" and "form field" in the Prompt template, and the layout area matching result is filled into the label to form the Prompt template.

9. The document information extraction method according to claim 1, characterized in that: In S4, the fine-tuning training specifically refers to: the large model learns the corresponding relationship between the Prompt template and the label value, and optimizes its own parameters.

10. The document information extraction method according to claim 1, characterized in that: The large model is a large language model.

11. A document information extraction device, characterized in that: The document information extraction device operates based on the document information extraction method according to any one of claims 1 to 10, including: A document preprocessing unit, used for inputting an original document for preprocessing to obtain a document image; An OCR recognition unit, used to perform OCR recognition on the document image using Vision Transformer (ViT) to obtain text information and text box coordinate information corresponding to the text information; A layout analysis unit, which is used to determine the category and coordinate information of each layout element frame based on the yolov8-seg instance segmentation algorithm of deep learning, and then perform layout area matching between the layout element frame coordinate information and the text frame coordinate information to obtain text information corresponding to each layout element frame; The information extraction unit is used to take the category and corresponding text information of each layout element box as the layout area matching result, combine it with the document information to be extracted to form a Prompt template, use the Prompt template as the input of the large model and perform fine-tuning training, and output the corresponding document information after the fine-tuning training is completed.

12. A document information extraction system, the system comprising: A processor and a memory for storing executable instructions; characterized in that the processor is configured to execute the executable instructions to perform the document information extraction method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the document information extraction method according to any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Document content matching method and system based on multiple modes

    CN120182990A

  • Method and device for shielding personal sensitive information of clinical test participants

    CN120257368A

  • Visual webpage data crawling method and system based on large model

    CN120632181A

  • Image processing method and device

    CN120783353A

  • Invoice identification method, device and equipment based on multi-modal large model

    CN120808377A