A method for extracting document information

Through the document structure analysis model and seal removal algorithm, combined with printed and handwriting recognition, and using a generative language model, the problems of element recognition and information extraction in complex documents are solved, and efficient and accurate structured data generation is achieved.

CN119964170BActive Publication Date: 2025-07-22TUGUAN (TIANJIN) DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510430398.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-22
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

Existing document analysis methods are difficult to accurately analyze multiple elements in complex documents, especially the interference of headers, footers, and seals on text recognition, and the handwritten and printed text recognition effect is poor, resulting in low information extraction efficiency and insufficient accuracy.

Method used

The document structure analysis model is used to identify document elements, the header and footer are treated as blanks, the overlapping parts of the seal are removed, and the printed and handwritten text recognition models are combined, and the generated language model is used to output structured data.

Benefits of technology

It realizes accurate identification and distinction of multiple document elements, improves the accuracy and efficiency of text recognition, and generates structured data that conforms to standardized formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964170B_ABST
    Figure CN119964170B_ABST
Patent Text Reader

Abstract

The present invention provides a method for extracting document information, including: obtaining a document to be parsed; using a document structure parsing model to parse different elements in the document and give identification bounding boxes for the elements; for the parsed header, footer, QR code, illustration, and trademark parts, processing the images within their bounding boxes into blank images; for the parsed seal part, if the seal and printed text overlap, using an algorithm to remove the seal part and retain the text part covered by the seal, and replacing the text part after removing the seal with the original image at the seal position; extracting the printed text and handwritten text from the processed document image, and recognizing the printed text and handwritten text in the document image; combining the positions of the original table, printed text, and handwritten text in the document image to assemble the recognized text together; based on a generative language large model, designing prompt words to generate the required structured data to be extracted and outputting it in a fixed format.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of natural language processing, and in particular relates to a method for extracting document information. Background Art

[0002] With the development of informatization, automated document processing has been widely applied in many industries such as enterprises, law, and finance. A large number of paper documents or scanned image documents need to be structured for more efficient information storage, retrieval, and analysis. However, current document information processing technologies still have certain limitations when facing complex document structures. For example, in documents with headers and footers, the context will be interfered by the headers and footers when crossing pages, and the seal covering part of the original text may cause errors in character recognition, affecting information extraction.

[0003] Existing document parsing methods mostly rely on templated rules and are difficult to handle different types of document changes. Especially when involving multiple information elements, it is very difficult for traditional methods to accurately extract each element. In addition, existing optical character recognition (OCR) technologies have poor recognition effects on handwritten characters and seals in complex scenarios and have low processing accuracy. These technical defects lead to problems such as low efficiency and insufficient accuracy in document information extraction.

[0004] To solve these problems, there is an urgent need for a structured information extraction method that can accurately parse different elements in a document to improve the automation and intelligence levels of document information processing. Summary of the Invention

[0005] In view of this, the present invention aims to overcome the above-mentioned deficiencies in the prior art and proposes a method for extracting document information.

[0006] To achieve the above object, the technical solution of the present invention is realized as follows:

[0007] The first aspect of the present invention provides a method for extracting document information, including the following steps:

[0008] Step 1: Obtain the document to be parsed and construct a document structure parsing model;

[0009] Step 2: Use the document structure parsing model to parse different elements in the document, including headers, footers, two-dimensional codes, printed characters, handwritten characters, tables, illustrations, trademarks, and seals, and give the recognition bounding boxes for each element;

[0010] Step 3: For the parsed headers, footers, two-dimensional codes, illustrations, and trademark parts, process the images within their bounding boxes into blank images;

[0011] Step 4: For the parsed seal part, if there is an overlap between the seal and the printed text, use the seal removal algorithm to remove the seal part, retain the text part covered by the seal, and replace the text part after removing the seal to the seal position of the original image;

[0012] Step 5: Extract the printed text in the processed document image and recognize the printed text in the document image;

[0013] Step 6: Extract the handwritten text in the processed document image and recognize the handwritten text in the document image;

[0014] Step 7: Assemble the recognized text from the original table, printed text, and handwritten text together according to the position of their original content;

[0015] Step 8: Based on the generative language large model, design prompt words, generate the required structured data to be extracted, and output it in a fixed format.

[0016] Further, in the above step 1, constructing the document structure parsing model includes:

[0017] Collect a real document image dataset, remove the noise on the image, remove the image background, and scale the image to the same size;

[0018] Annotate the document image, use the bounding box method to annotate the document, annotate the header, footer, QR code, printed text, handwritten text, table, illustration, trademark, and seal, obtain the element coordinates, and form the training data;

[0019] Use the object detection algorithm, substitute the annotated document image training data into the object detection algorithm, and perform model training;

[0020] Use the mean average precision and the overlap degree index between the predicted box and the ground truth box to evaluate the model effect.

[0021] Further, in the above step 4, the seal removal algorithm includes:

[0022] Obtain the seal image containing background text;

[0023] Represent the seal image in the Lab color space, and each pixel point is represented by three values where represents the brightness of the seal image color, represents the component of the seal image color from green to red, represents the component of the seal image color from blue to yellow;

[0024] Given the target color, and represent the target color in the Lab color space where represents the target color brightness, represents the component of the target color from green to red, represents the component of the target color from blue to yellow;

[0025] Calculate the Euclidean distance between each pixel point in the seal image and the target color ;

[0026]

[0027] Set a threshold , and calculate the transparency of each pixel point in the seal image through the following formula ;

[0028] ;

[0029] where, is the deviation value, the larger it is, when the larger it is; is the scaling value, the smaller it is, the larger the blur interval of

[0030] Traverse each transparency point and update the transparency of each point using the following formula;

[0031] ;

[0032] where, is the coordinate of the pixel point, represents the square neighborhood of the pixel point with the coordinate is the transparency of the pixel point with the coordinate , represents the updated transparency of the pixel point with the coordinate

[0033] Overlay the original seal image with the transparency using the following formula to obtain the image after removing the seal, where represents the RGB value of the seal image, represents the RGB value of white

[0034] .

[0035] Furthermore, the specific steps of step 5 include:

[0036] Collect the images in the printed text recognition bounding boxes given by the document structure parsing model;

[0037] Label the text within the image of the printed text recognition bounding box;

[0038] Substitute the training data consisting of the image of the printed text recognition bounding box and its labeled text into the text recognition algorithm, set hyperparameters, and encode for model training;

[0039] Evaluate the model performance using the word error rate and accuracy metrics.

[0040] Further, step 6 specifically includes:

[0041] Collect the images within the handwritten text recognition bounding boxes given by the document structure parsing model after parsing;

[0042] Label the text within the image of the handwritten text recognition bounding box;

[0043] Substitute the training data consisting of the image of the handwritten text recognition bounding box and its labeled text into the text recognition algorithm, set hyperparameters, and encode for model training;

[0044] Evaluate the model performance using the word error rate and accuracy metrics.

[0045] Further, step 8 specifically includes:

[0046] Collect the original document text data;

[0047] Design text information extraction prompts, label the fields to be extracted and their values according to the original document text to be parsed, and represent the values of the fields in a set format;

[0048] Assemble the prompts, the original document text to be parsed, the fields to be extracted, and the values of the fields in the set format to form a dataset;

[0049] Based on the open-source generative language large model, substitute the dataset and the data obtained in step 7 into the large model, and the large model parses out the information to be extracted and directly outputs it in the set format.

[0050] The second aspect of the present invention provides a document information extraction device, including:

[0051] The first processing unit is used to obtain the document to be parsed and construct a document structure parsing model;

[0052] The second processing unit is used to parse different elements in the document using the document structure parsing model, including the header, footer, QR code, printed text, handwritten text, table, illustration, trademark, and seal, and give the recognition bounding boxes for each element;

[0053] The third processing unit is used to process the images within the bounding boxes of the parsed header, footer, QR code, illustration, and trademark into blank images;

[0054] A fourth processing unit, which, for the parsed seal part, if there is an overlap between the seal and the printed text, uses a seal removal algorithm to remove the seal part, retains the text part covered by the seal, and replaces the text part after removing the seal with the original position of the seal in the original image;

[0055] A fifth processing unit, which extracts the printed text in the processed document image and recognizes the printed text in the document image;

[0056] A sixth processing unit, which extracts the handwritten text in the processed document image and recognizes the handwritten text in the document image;

[0057] A seventh processing unit, which assembles the recognized texts from the original table, printed text, and handwritten text together according to their original positions;

[0058] An eighth processing unit, which designs prompt words based on a generative language large model, generates the required structured data to be extracted, and outputs it in a fixed format.

[0059] The third aspect of the present invention provides an electronic device, including a processor and a memory communicatively connected to the processor and used to store executable instructions of the processor, and the processor is used to execute the above-mentioned document information extraction method.

[0060] The fourth aspect of the present invention provides a computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, the above-mentioned document information extraction method is realized.

[0061] Compared with the prior art, the document information extraction method of the present invention has the following advantages:

[0062] The present invention uses multiple models such as object detection, handwritten text recognition, printed text recognition, seal recognition, and generative language large model to complete document information extraction.

[0063] The present invention can parse various elements in the document (such as headers, footers, QR codes, handwritten characters, tables, etc.), demonstrating a comprehensive understanding of the document structure and being applicable to various document types.

[0064] By performing blank image processing on the header and footer areas, the present invention can effectively eliminate interference information and improve the accuracy of subsequent text recognition.

[0065] The present invention uses an algorithm to remove the seal overlapping with the printed text and retains the text part covered by the seal, reflecting the fine processing and restoration ability of the document content and solving the problem of overlapping information that is difficult to handle by traditional methods.

[0066] The present invention uses printed text and handwritten text recognition models respectively, and conducts specialized processing for different text forms, improving the accuracy and efficiency of recognition.

[0067] The present invention assembles the recognized text information in combination with its position in the document, and can generate complete structured data, meeting the standardized requirements of data processing and storage.

[0068] The present invention uses a generative language large model for prompt design, which not only improves the flexibility of data extraction, but also can generate diverse output formats according to user needs. Brief Description of the Drawings

[0069] The drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0070] Figure 1 is a schematic flow diagram of a method for extracting document information of the present invention;

[0071] Figure 2 is a schematic flow diagram of constructing a document structure analysis model of the present invention;

[0072] Figure 3 is a schematic flow diagram of a seal removal algorithm of the present invention;

[0073] Figure 4 is a schematic sub - flow diagram of printed text recognition of the present invention;

[0074] Figure 5 is a schematic sub - flow diagram of handwritten text recognition of the present invention;

[0075] Figure 6 is a schematic flow diagram of information extraction of the present invention;

[0076] Figure 7 is a schematic diagram of a document to be analyzed of the present invention;

[0077] Figure 8 is a schematic diagram of the result of analyzing document elements of the present invention;

[0078] Figure 9 is a schematic diagram of the result obtained by processing the header and footer of the present invention;

[0079] Figure 10 is a schematic diagram of the result obtained by removing the seal of the present invention;

[0080] Figure 11 is a schematic diagram of the result obtained by recognizing printed text of the present invention;

[0081] Figure 12Schematic diagram of the result obtained by the present invention for recognizing handwritten characters. Detailed implementation manners

[0082] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.

[0083] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "up", "down", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a set orientation, be constructed and operated in a set orientation, and thus should not be construed as a limitation to the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0084] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installation", "connection", and "connection" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood through specific situations.

[0085] The present invention will be described in detail below with reference to the drawings and in combination with embodiments.

[0086] Embodiment 1:

[0087] The present invention aims at the recognition and processing of multiple elements in a complex document structure, overcomes the deficiencies in the prior art, and provides a method for extracting document information, mainly solving the following technical problems:

[0088] 1. Automatic recognition and parsing of multiple document elements: Provide a method that can automatically recognize different elements in a document (such as headers, footers, printed characters, handwritten characters, tables, seals, two-dimensional codes, etc.), ensure that each element can be accurately calibrated and distinguished, and solve the limitations of traditional templatized methods.

[0089] 2. Overlap processing of seal and text: Solve the problem of the overlap between the seal and printed text. Through image processing algorithms, remove the seal and retain the covered text information to improve the integrity and accuracy of document information.

[0090] 3. Distinguishing and recognizing handwritten and printed texts: Provide an algorithm that can accurately distinguish and separately recognize handwritten and printed texts to ensure the correctness of the extracted document information, especially in documents with mixed different writing styles and text structures.

[0091] 4. Information extraction and formatted output: Combine with the generative language large model. By designing prompt words, extract the key information in the document and structure it into a fixed format (such as JSON), solving the problems of accuracy and formatted output in the information extraction process.

[0092] The flowchart of a document information extraction method of the present invention is as Figure 1 shown, and specifically includes the following steps:

[0093] Step 1: Obtain the document to be parsed and construct a document structure parsing model;

[0094] Step 2: Use the document structure parsing model to parse different elements in the document, including the header, footer, QR code, printed text, handwritten text, table, illustration, trademark, and seal, and give the recognition bounding boxes for each element;

[0095] Step 3: For the parsed header, footer, QR code, illustration, and trademark parts, process the images within their bounding boxes into blank images;

[0096] Step 4: For the parsed seal part, if there is an overlap between the seal and printed text, use the seal removal algorithm to remove the seal part, retain the text part covered by the seal, and replace the text part after removing the seal to the original seal position in the image;

[0097] Step 5: Extract the printed text in the processed document image and recognize the printed text in the document image;

[0098] Step 6: Extract the handwritten text in the processed document image and recognize the handwritten text in the document image;

[0099] Step 7: Assemble the recognized texts from the original table, printed text, and handwritten text together according to their original positions in the content;

[0100] Step 8: Based on the generative language large model, design prompt words to generate the required structured data to be extracted and output in a fixed format (such as json).

[0101] In the present invention, the process of constructing the document structure parsing model is as follows Figure 2 shown and includes:

[0102] 1. Collect a real document image data set, remove the noise on the image, remove the image background, and scale the image to the same size;

[0103] 2. Annotate the document image, use bounding boxes to annotate the document, and annotate the header, footer, QR code, printed text, handwritten text, table, illustration, trademark, and seal to obtain the element coordinates and form the training data;

[0104] 3. Use the object detection algorithm, substitute the labeled document image training data into the algorithm, and perform model training;

[0105] 4. Use the mean average precision (mAP) and the overlap degree (IoU) of the predicted box and the ground truth box to evaluate the model effect.

[0106] In the present invention, the process of the seal removal algorithm is as follows Figure 3 shown and includes:

[0107] 1. Obtain a seal image containing background text;

[0108] 2. Represent the seal image in the Lab color space, and each pixel is represented by three values where represents the brightness of the seal image color, represents the component of the seal image color from green to red, represents the component of the seal image color from blue to yellow;

[0109] 3. Given the target color and represent the target color in the Lab color space , where represents the target color brightness, represents the component of the target color from green to red, represents the component of the target color from blue to yellow;

[0110] 4. Calculate the Euclidean distance between each pixel in the seal image and the target color ;

[0111] ;

[0112] 5. Set the threshold , and calculate the transparency of each pixel in the seal image through the following formula ;

[0113] ;

[0114] Among them, is the deviation value, the larger when the larger; is the scaling value, the smaller the larger the fuzzy interval of

[0115] 6. Traverse each transparency point and use the following formula to update the transparency of each point to make the image smoother;

[0116] ;

[0117] Among them, is the coordinate of the pixel point, represents the square neighborhood of the pixel point with the coordinate of , is the transparency of the pixel point with the coordinate of , represents the updated transparency of the pixel point with the coordinate of ;

[0118] 7. Use the following formula to superimpose the original seal image with the transparency. Obtain the image after removing the seal. Among them represents the RGB value of the seal image, represents the RGB value of white

[0119] ;

[0120] In the present invention, the printed text recognition sub-process is as Figure 4 shown, including:

[0121] 1. Collect the image in the printed text recognition bounding box given by the document structure parsing model after parsing;

[0122] 2. Mark the text in the image of the printed text recognition bounding box;

[0123] 3. Use the text recognition algorithm, substitute the training data composed of the image of the printed text recognition bounding box and its marked text into the text recognition algorithm, set the hyperparameters, and encode for model training;

[0124] 4. Use the character error rate (CER) and accuracy rate indicators to evaluate the model effect.

[0125] In the present invention, the handwritten text recognition sub-process is as Figure 5 shown, including:

[0126] 1. Collect the image in the handwritten text recognition bounding box given by the document structure parsing model after parsing;

[0127] 2. Label the text within the image with the bounding box for handwritten text recognition.

[0128] 3. Use an optical character recognition (OCR) algorithm. Substitute the training data consisting of the image with the bounding box for handwritten text recognition and its labeled text into the OCR algorithm, set the hyperparameters, and perform model training through encoding.

[0129] 4. Evaluate the model performance using metrics such as character error rate (CER) and accuracy.

[0130] In the present invention, the document information extraction process is as Figure 6 shown and includes:

[0131] 1. Collect the original document text data.

[0132] 2. Design text information extraction prompts. Based on the original document text, label the fields to be extracted and their values, and represent the values of the fields in a set format (such as JSON).

[0133] 3. Assemble the prompts, the original text, the fields to be extracted, and the values of the fields in the set format to form a dataset.

[0134] 4. Based on an open-source generative language large model, substitute the dataset and the data obtained in step 7 into the large model. The large model analyzes and obtains the information to be extracted and directly outputs it in the set format.

[0135] The solution of the present invention will be illustrated below through specific examples.

[0136] As Figure 7 shown, for the original document to be parsed, use a document structure parsing model to parse different elements in the document, including the header, footer, QR code, printed text, handwritten text, table, illustration, trademark, seal, etc., and give the recognition bounding boxes for the elements, as Figure 8 shown. For the parsed header and footer parts, regardless of whether the boxes include handwritten text, printed text, or other elements, process the image within their bounding boxes into a blank image, as Figure 9 shown. For the parsed seal part, if the seal overlaps with the printed text, use an algorithm to remove the seal part and retain the text part covered by the seal. Replace the text part after removing the seal with the seal position in the original image, as Figure 10 shown. Extract the printed text from the processed document image and use a printed text recognition model to recognize the printed text in the document image, as Figure 11 shown. Extract the handwritten text from the processed document image and use a handwritten text recognition model to recognize the handwritten text in the document image, as Figure 12As shown in the figure. Combine the positions of the original table, printed text, and handwritten text in the document image, and assemble the recognized text together. This text data contains all the information that needs to be extracted.

[0137] The following is an example of recognition. The extracted text data is:

[0138] XX Advertising Insurance Co., Ltd.

[0139] Report Letter on the Filing of XXX Project Performance Bond Insurance

[0140] I. Company Information

[0141] XX Advertising Insurance Co., Ltd. (hereinafter referred to as "XX Advertising") is the only insurance company operating property insurance business under the financial company - China XXXX Insurance Group. Its headquarters is located in City A, with a registered capital of XX billion yuan. The company has a long operating history, advanced and standardized management, professional and honest services, and a good brand image. It has long received extensive praise and approval from the majority of insurance consumers and all sectors of society. Among them, the table data is shown in Table 1.

[0142] Table 1

[0143]

[0144] Design prompt words according to the information that needs to be extracted, substitute the above-assembled text data and prompt words into the generative language large model. The generative language large model generates the required structured data to be extracted and outputs it in a fixed format (such as json). Users can flexibly set prompt words according to actual needs for information extraction, such as extracting any information that needs to be extracted, such as the tender document number, project name, project contract end number, etc.

[0145] Taking this embodiment as an example, the prompt words can be set as

[0146] key: Buyer, type: String

[0147] Output:

[0148] { "key": "Buyer", "value": "A certain company"}

[0149] The fields that need to be extracted can be set as follows:

[0150] Extracted field: Name of the applicant (bidder), field type: text

[0151] Extracted field: Contact phone number of the applicant (bidder), field type: text

[0152] Extracted field: Name of the insured (tenderer), field type: text

[0153] Extracted field: Contact phone number of the insured (tenderer), field type: text

[0154] Finally, the output result can be obtained

[0155] { "key": "Name of the applicant (bidder)", "value": "XX Construction Co., Ltd."},

[0156] { "key": "Contact phone number of the applicant (bidder)", "value": "010 - 12345XXX"},

[0157] { "key": "Name of the insured (tenderer)", "value": "XX Real Estate Development Co., Ltd."},

[0158] { "key": "Contact phone number of the insured (tenderer)", "value": "021 - 12345XXX"}

[0159] Example 2:

[0160] A document information extraction device, comprising:

[0161] A first processing unit, configured to obtain a document to be parsed and construct a document structure parsing model;

[0162] A second processing unit, configured to use the document structure parsing model to parse different elements in the document, including headers, footers, QR codes, printed text, handwritten text, tables, illustrations, trademarks, and seals, and give recognition bounding boxes for each element;

[0163] A third processing unit, configured to process the images within the bounding boxes of the parsed headers, footers, QR codes, illustrations, and trademarks into blank images;

[0164] A fourth processing unit, configured to, for the parsed seal part, if the seal overlaps with printed text, use a seal removal algorithm to remove the seal part, retain the text part covered by the seal, and replace the text part after removing the seal with the seal position in the original image;

[0165] A fifth processing unit, configured to extract the printed text in the processed document image and recognize the printed text in the document image;

[0166] A sixth processing unit, configured to extract the handwritten text in the processed document image and recognize the handwritten text in the document image;

[0167] A seventh processing unit, configured to assemble the recognized text together in combination with the positions of the original table, printed text, and handwritten text in the document image;

[0168] An eighth processing unit is configured to design prompt words based on a generative language large model, generate the structured data to be extracted, and output it in a fixed format.

[0169] Embodiment III:

[0170] An electronic device includes a processor and a memory communicatively connected to the processor and configured to store executable instructions of the processor. The processor is configured to execute the above-mentioned document information extraction method.

[0171] Embodiment IV:

[0172] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned document information extraction method is implemented.

[0173] The foregoing are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for extracting document information, characterized in that: It includes the following steps: Step 1: Obtain the document to be parsed and construct a document structure parsing model; Step 2: Use the document structure parsing model to parse different elements in the document, including headers, footers, QR codes, printed text, handwritten text, tables, illustrations, trademarks, and seals, and give the recognition bounding boxes for each element; Step 3: For the parsed headers, footers, QR codes, illustrations, and trademark parts, process the images within their bounding boxes into blank images; Step 4: For the parsed seal part, if the seal overlaps with printed text, use a seal removal algorithm to remove the seal part, retain the text part covered by the seal, and replace the text part after removing the seal with the original image at the seal position; Step 5: Extract the printed text from the processed document image and recognize the printed text in the document image; Step 6: Extract the handwritten text from the processed document image and recognize the handwritten text in the document image; Step 7: Assemble the text recognized from the original table, printed text, and handwritten text according to the positions of their original contents; Step 8: Collect the text data of the original document to be parsed, design text information extraction prompts, mark the fields to be extracted and the values of the fields according to the text of the original document to be parsed, represent the values of the fields in a set format, assemble the prompts, the text of the original document to be parsed, the fields to be extracted, and the values of the fields in the set format to form a dataset. Based on the open-source generative language large model, substitute the dataset and the data obtained in Step 7 into the large model, and the large model parses and obtains the information to be extracted and directly outputs it in the set format.

2. The method for extracting document information according to claim 1, wherein: In the above Step 1, constructing the document structure parsing model includes: Collect a real document image dataset, remove the noise on the images, remove the image background, and scale the images to the same size; Annotate the document images, use bounding boxes to annotate the files, annotate the headers, footers, QR codes, printed text, handwritten text, tables, illustrations, trademarks, and seals to obtain the element coordinates and form training data; Use the object detection algorithm, substitute the annotated document image training data into the object detection algorithm for model training; Use the mean average precision and the overlap degree index between the predicted box and the ground truth box to evaluate the model effect.

3. A method for extracting document information according to claim 1, characterized in that: In the above Step 4, the seal removal algorithm includes: Obtain a seal image containing background text; The seal image is represented in the Lab color space, and each pixel is composed of three values which are used to represent the brightness of the seal image color, used to represent the component of the seal image color from green to red, and used to represent the component of the seal image color from blue to yellow; Given a target color and representing the target color in the Lab color space , where represents the brightness of the target color, represents the component of the target color from green to red, represents the component of the target color from blue to yellow; Calculate the Euclidean distance between each pixel point in the seal image and the target color ; ; Set a threshold , and calculate the transparency of each pixel of the seal image through the following formula ; ; Among them, is the deviation value, the larger it is, when the larger it is; is the scaling value, the smaller it is, the larger the fuzzy interval of Traverse each pixel point and update the transparency of each pixel point using the following formula; ; Among them, is the coordinate of a pixel point, represents the square neighborhood of the pixel point with the coordinate of , is the transparency of the pixel point with the coordinate of , represents the updated transparency of the pixel point with the coordinate of . Overlay the original seal image with the transparency using the following formula to obtain the image after removing the seal, where represents the RGB values of the seal image, represents the RGB values of white; 。 4. A method for extracting document information according to claim 1, characterized in that: The above Step 5 specifically includes: Collect the images in the printed text recognition bounding boxes given by the document structure parsing model after parsing; Annotate the text in the images of the printed text recognition bounding boxes; Substitute the training data composed of the images of the printed text recognition bounding boxes and their annotated text into the text recognition algorithm, set the hyperparameters, and encode for model training; Use the word error rate and accuracy rate indicators to evaluate the model effect.

5. The method for extracting document information according to claim 1, characterized in that: The above Step 6 specifically includes: Collect the images in the handwritten text recognition bounding boxes given by the document structure parsing model after parsing; Annotate the text in the images of the handwritten text recognition bounding boxes; The training data consisting of the image of the handwritten text recognition bounding box and its labeled text is input into the text recognition algorithm, hyperparameters are set, and encoding is performed for model training; The model effect is evaluated using the word error rate and accuracy metrics.

6. A document information extraction device, characterized in that: Including: The first processing unit is used to obtain the document to be parsed and construct a document structure parsing model; The second processing unit is used to parse different elements in the document using the document structure parsing model, including headers, footers, QR codes, printed text paragraphs, handwritten characters, tables, illustrations, trademarks, dates, seals, and give the recognition bounding boxes of the elements; The third processing unit is used to process the image within the bounding box of the parsed header and footer parts into a blank image; The fourth processing unit is used to, for the parsed seal part, if the seal and printed text overlap, use the seal removal algorithm to remove the seal part, retain the text part covered by the seal, and replace the text part after removing the seal to the seal position in the original image; The fifth processing unit is used to extract the printed text in the processed document image and recognize the printed text in the document image; The sixth processing unit is used to extract the handwritten text in the processed document image and recognize the handwritten text in the document image; The seventh processing unit is used to assemble the recognized text from the original table, printed text, and handwritten text together according to their original positions in the content; The eighth processing unit is used to collect the text data of the original document to be parsed, design text information extraction prompt words, label the fields to be extracted and the values of the fields according to the text of the original document to be parsed, represent the values of the fields in a set format, assemble the prompt words, the text of the original document to be parsed, the fields to be extracted, and the values of the fields in the set format to form a data set, and based on the open-source generative language large model, input the data set and the data obtained by the seventh processing unit into the large model, and the large model parses and obtains the information to be extracted and outputs it directly in the set format.

7. An electronic device, comprising a processor and a memory communicatively connected to the processor and configured to store executable instructions of the processor, wherein: The processor is used to execute a document information extraction method according to any one of claims 1-5 above.

8. A computer-readable storage medium storing a computer program, characterized in that: The computer program, when executed by the processor, implements a document information extraction method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Instruction-based document image processing method and system

    CN116580411A

  • Company seal identification method and related equipment thereof

    CN116824600A