Document information extraction method and device, equipment and medium

By combining OCR and VLM methods, the problem of information extraction from various document formats and complex content is solved. By obtaining structured data of non-text elements through OCR and aligning it with VLM results, a knowledge tree is constructed, achieving efficient and accurate information extraction from complex documents.

CN120635913APending Publication Date: 2025-09-12BEIJING ZHITONG YUNLIAN TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510751972.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies have difficulty in efficiently extracting information from documents with multiple formats, complex content, and mixed languages, especially when they contain complex images, tables, formulas, and other elements, the accuracy and efficiency of information extraction are insufficient.

Method used

A method combining optical character recognition technology (OCR) and visual language model (VLM) is used to first convert the document into image format. The structured data of non-text elements is obtained through OCR and converted into image format. The results of the visual language model are combined for element-level alignment and supplementation, and the style features and directory rules of the original document elements are used to construct a knowledge tree.

Benefits of technology

It improves the accuracy and efficiency of information extraction from documents with complex layouts, tables, images and other information sources, and achieves comprehensive and accurate extraction of multi-language mixed documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635913A_ABST
    Figure CN120635913A_ABST
Patent Text Reader

Abstract

The invention provides a document information extraction method and device, equipment and a medium. The method comprises the steps that a document is converted into an image format; processing the image through an optical character recognition technology to obtain a first element list; the recognized non-text elements are converted into picture elements to be stored; sequentially processing the image based on the visual language model to obtain a second element list; the first element list and the second element list are subjected to element level alignment, and text element matching is achieved through text elements in the context and the second element list on the basis of position information and context information of all non-text elements in the first element list. The non-text elements and the corresponding picture elements are correspondingly inserted into a second element list according to the matched relative positions; automatically extracting a document directory list based on style features of original document elements; and based on the document directory list and the corrected second element list, constructing a directory-based knowledge tree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information data processing, and in particular to a document information extraction method, device, equipment and medium. Background Art

[0002] With the development of artificial intelligence (AI), particularly computer vision and natural language processing (NLP), document information extraction has become a crucial research area in the information society. Traditional document information extraction methods rely primarily on text analysis and extraction, and are suitable for documents with relatively structured and standardized formats. However, traditional methods often struggle with graphical representations, complex layouts, and mixed-language documents. This is particularly true when dealing with scanned documents or documents containing complex graphics, tables, formulas, and other elements, where the accuracy and efficiency of information extraction often struggle to meet requirements.

[0003] In recent years, Visual Language Model (VLM) technology has made significant progress. By combining visual and linguistic information, it can establish connections between images and text, providing more accurate information extraction capabilities. In particular, in document image processing, VLMs can combine image content and text information to extract key information from documents, and can also understand document structure, improving information extraction.

[0004] Although the existing technology has made certain progress, the information extraction of documents with various document formats, complex content and mixed languages ​​is still not very accurate and comprehensive. In view of this, there is an urgent need to provide an information extraction method that can improve the accuracy of document information, especially for documents containing complex images, tables, formulas, directories and other information. Summary of the Invention

[0005] In order to overcome the problems existing in the related art, the present disclosure provides a document information extraction method, device, equipment and medium to solve the technical problems in the related art.

[0006] One or more embodiments of this specification provide a document information extraction method, including the steps of: Convert documents into image formats; After sequentially processing each image using optical character recognition technology, a first element list of structured data is obtained. The first element list includes the text content of each element, the element's position coordinates, the element's category label, and reading order data. The recognized non-text elements are then converted into image elements for storage. After sequentially processing each image based on the visual language model, a second element list of structured data is obtained, wherein the second element list includes a description text of each element and reading order data; Performing element-level alignment on the first element list and the second element list, and then matching text elements by using text elements in the context and the second element list based on the position information and context information of each non-text element in the first element list, and inserting the non-text elements and corresponding image elements into the second element list according to the matching positions, to obtain a revised second element list; Automatically extract the document directory list based on the style features of the original document elements, combined with the title numbering rules and common directory dictionaries; Based on the document directory list and the revised second element list, a directory-based knowledge tree is constructed.

[0007] One or more embodiments of this specification provide a document information extraction device, including: Format conversion module, used to convert documents into image formats; A first element list acquisition module is configured to sequentially process each image using optical character recognition technology to obtain a first element list of structured data, the first element list including the text content of each element, element position coordinates, element category label, and reading order data; and then convert the recognized non-text elements into image elements and store them; a second element list acquisition module, configured to sequentially process each image based on a visual language model to obtain a second element list of structured data, the second element list including a description text of each element and reading order data; A second element list completion module is used to align the first element list and the second element list at the element level, and then, based on the position information and context information of each non-text element in the first element list, match the text elements in the context with the second element list, and insert the non-text elements and the corresponding image elements into the second element list according to the matching positions to obtain a revised second element list; Document directory list extraction module, which is used to automatically extract the document directory list based on the style features of the original document elements, combined with the title numbering rules and common directory dictionaries; The knowledge tree generation module is used to construct a directory-based knowledge tree based on the document directory list and the revised second element list.

[0008] One or more embodiments of this specification provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the document information extraction method described above when executing the computer program.

[0009] One or more embodiments of this specification provide a computer-readable storage medium storing a computer program, wherein the computer program implements the document information extraction method described above when executed by a processor.

[0010] The present disclosure provides a document information extraction method, device, equipment and medium, which have the advantage that, in order to solve the problem that there are still certain limitations in the processing of multiple document formats, complex content and multi-language mixed typesetting, the document is first converted into an image, and then the visual language model is used to sequentially obtain the semantic information of the text, picture, table and other elements and their context in each image. Since OCR technology is relatively good at obtaining table content recognition, non-text element information such as pictures, tables, formulas, etc. is obtained through OCR technology and converted into image format, and the two structured data context information are combined to determine the corresponding elements and relative positions of non-text elements such as pictures, tables, formulas in the second element list, and replace or supplement them, so as to use the OCR extraction results to verify the visual language model extraction results, supplement the deficiencies of the visual language model in visual element extraction, and effectively improve the accuracy and efficiency of document information extraction. For documents with complex typesetting, tables, images and other information sources, the required information can be comprehensively and accurately extracted. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] Figure 1 A flowchart of a document information extraction method provided in one or more embodiments of this specification; Figure 2 An exemplary document image provided for one or more embodiments of this specification; Figure 3 A first element list data diagram for recognizing a document image using an exemplary OCR technology provided in one or more embodiments of this specification; Figure 4 A data list of a second element list of a document image recognized by an exemplary visual language model provided in one or more embodiments of this specification; Figure 5 A data list of a second element sublist after supplementation of exemplary elements provided in one or more embodiments of this specification; Figure 6A data list of exemplary title lists provided for one or more embodiments of this specification; Figure 7 The final obtained exemplary knowledge tree provided for one or more embodiments of this specification; Figure 8 A block diagram of a document information extraction device provided for one or more embodiments of this specification; Figure 9 A schematic diagram of the structure of a computer device provided in one or more embodiments of this specification. DETAILED DESCRIPTION

[0013] In order to help those skilled in the art better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0014] The present invention will be described in detail below with reference to specific implementation methods and the accompanying drawings.

[0015] Method Example According to an embodiment of the present invention, a document information extraction method is provided, such as Figure 1 FIG. 1 is a flowchart of a document information extraction method provided in this embodiment. The document information extraction method according to an embodiment of the present invention includes the following steps: Step S1, converting the document into image format; Step S2: sequentially processing each image using optical character recognition (OCR) technology to obtain a first element list of structured data, which includes data such as the text content of each element, element position coordinates, element category label, and reading order; then, the recognized non-text elements are converted into image elements and stored; Among them, the elements may include text blocks, pictures, tables, formulas and other elements. Based on the reading order, the layout type corresponding to each semantic paragraph (element) in the document is determined, wherein the text reading order includes left-right reading order and top-bottom reading order, and the layout type includes paragraph layout, table layout and image layout, etc.

[0016] Step S3, processing each image sequentially based on the visual language model to obtain a second element list of structured data, the second element list including data such as description text and reading order of each element; In step S4, the first element list and the second element list are aligned at the element level, and then based on the position information and context information of each non-text element in the first element list, the text elements are matched with the text elements in the context and the second element list, and the non-text elements and the corresponding image elements are inserted into the second element list according to the matching positions to obtain a corrected second element list.

[0017] Step S5: Automatically extract the document directory list based on the style features of the original document elements, combined with the title numbering rules and the commonly used directory dictionary.

[0018] Step S6: constructing a directory-based knowledge tree based on the document directory list and the revised second element list.

[0019] The document information extraction method provided in this embodiment, in order to solve the problem that there are still certain limitations in the processing of multiple document formats, complex content and multi-language mixed typesetting, first converts the document into an image, and then uses the visual language model to sequentially obtain the semantic information of the text, picture, table and other elements and their context in each image. Since OCR technology is relatively good at obtaining table content recognition, non-text element information such as pictures, tables, formulas, etc. will be obtained through OCR technology and converted into image format, and the two structured data context information will be combined to determine the corresponding elements and relative positions of non-text elements such as pictures, tables, formulas in the second element list, and replace or supplement them, so as to use the OCR extraction results to verify the visual language model extraction results, supplement the deficiencies of the visual language model in visual element extraction, and effectively improve the accuracy and efficiency of document information extraction. For documents with complex typesetting, tables, images and other information sources, the required information can be extracted comprehensively and accurately.

[0020] In this embodiment, the document may be converted into an image format in step S1 by taking screenshots of the document page by page or converting the document into multiple images, wherein the document includes a word document, PDF, PPT, etc.

[0021] In this embodiment, the results of OCR image recognition are not limited to simple text content, but also include rich metadata such as location information, confidence score, language information, formatting information, and table data. This additional information enables OCR technology to play a greater role in various application scenarios, such as automated document processing, data entry, and image search.

[0022] OCR tools usually provide the location information (such as bounding boxes) of each recognized text block so that users can understand the location of these texts in the original image. Figure 3 For example, to be recognized by OCR tools Figure 2The results for the document image shown in the figure are as follows. The third element in the list is a figure element within the recognized text, and the eighth element is a text element within the recognized text. The figure is a picture containing the letters SZIC. For documents containing tables, OCR tools can extract the table content and convert it into structured data (such as a two-dimensional array or CSV format).

[0023] In this embodiment, for documents containing tables, the visual language model can extract the table content and convert it into structured data (such as a two-dimensional array or CSV format). Furthermore, it can understand the context and structure of the table. The visual language model can not only recognize text, but also identify and describe non-text elements such as images and formulas, and provide corresponding contextual information. Furthermore, the visual language model can automatically extract table of contents elements from the document and generate a complete table of contents structure based on heading numbering rules and style features.

[0024] In order to supplement the lack of non-text elements such as tables, images, and formulas in the visual language model results and ensure the integrity of the information extraction results, this embodiment also includes the following steps: In step 21, the images, tables, and formulas identified by OCR technology are independently captured and stored in the cloud to preserve their original visual information.

[0025] Step 22: Perform semantic analysis on the table elements, convert them into corresponding table semantic texts, and retain them as text elements in the first element list.

[0026] In this embodiment, the specific implementation process of the semantic conversion of table elements includes: first, by analyzing the page elements of the document image, identifying the basic lines and text blocks and other elements that constitute the table, and by analyzing the geometric positions and relative relationships of these elements, determining the boundaries, row and column structure of the table and the position of the cells in the page; then, based on the determined cell position, extracting the text content, identifying the text alignment, font size and style information, so as to accurately reconstruct the data in the table cells and obtain the table semantic text.

[0027] The core idea of ​​this embodiment is to use the excellent characteristics of OCR technology in recognizing pictures, tables, and formulas, and use the OCR recognition results to supplement the visual language model with element information such as pictures, tables, and formulas based on location information and context, thereby improving the recognition results of the visual language model.

[0028] In this embodiment, when the visual language model is used to identify the content of each image, the content of the corresponding text page is incomplete and non-text elements such as pictures and tables are lost. For example, Figure 4 , which is recognized by the visual language model Figure 2The obtained element list does not include the third image element (the picture element of the letters GB (Figure)) and the eighth element (the picture of the letters SZIC) in the above OCR recognition list. Therefore, it is necessary to perform element and element content integrity verification to ensure the integrity of the recognition results of the visual language model. It also includes element-level alignment and information verification of the recognition results of the visual language model through the OCR recognition results to determine whether there are any non-text elements missing in the recognition results of the visual language model. The text elements in the first element sublist and the second element sublist can be compared one by one according to the recognition results. The position of the non-text element in the second element sublist is determined according to the position relationship of the non-text element in the first element sublist relative to the context element. The specific steps include: Step 101: Based on the first element sublist and the second element sublist of the same image, using the information of each element in the first element sublist as the comparison standard and the element as the comparison unit, the second element in the second element sublist is traversed and the text content is matched with the first element in the first element sublist in the order or reverse order. If the similarity is greater than a preset value, the process proceeds to step 102; if no corresponding element is matched, the process continues to traverse the second element sublist; Step 102: Determine whether the element category tags of the upper and lower elements of the first element are non-text elements. If so, the corresponding non-text elements are inserted above or below the second element according to their positions relative to the first element. If not, the upper and lower elements of the first element are text elements, and the second element sublist is traversed. Figure 5 As shown, it is a second element sublist supplemented based on the elements of the first element list.

[0029] In addition to determining whether non-text elements are missing and supplemented by traversing the two element lists in sequence as described above, in another embodiment, in step S101, at least two elements in the first element sublist are used as comparison criteria, or element-level matching and comparison are implemented through steps S101 and S102 to determine the loss of non-text elements, and the lost non-text elements are supplemented to the second element sublist accordingly. This execution step can be obtained by reasoning based on steps S101 and S102 and will not be repeated here.

[0030] Preferably, in this embodiment, it is determined whether the text content of each text element in the second element sublist is missing, which can make up for the problem that the visual language model does not fully recognize the image content. For step S101, the second element is matched with the first element in the first element sublist for element text content comparison. When the similarity is greater than a preset value, the following steps are further included: Step 1011, using the element text content of the first element as the comparison standard, based on the preset size window sliding and matching distance, the second element text content is matched with the first element text content in reverse order to determine whether the second element text content is missing text content. If the text content is missing, all text characters after the last matching text character position of the first element are sequentially filled to the end of the second element text content according to the last matching text character position.

[0031] The execution process of step S101 and step S102 is described below through a specific example.

[0032] Suppose that the text content of a text element obtained by OCR recognition is the first string abcdefghijk, and the text content of the same text element with the highest similarity obtained through visual language model recognition is the second string abcdefg. The matching distance is 3 and the sliding window length is 1. The last 3 characters efg of the second string are flipped to obtain the comparison substring gfe, and then the first string is matched in reverse order. Of course, the first string can also be flipped to obtain the string kjihgfedcba and then matched in forward order. The idea is the same. Therefore, based on the window setting, kji, jih, ihg, hgf, gfe, fed... are taken in turn to compare the similarity with the comparison substring gfe. When the similarity is greater than 90, the comparison substring is determined to be at the end of the first string. Based on the first string, it can be seen that there is a character hijk after efg. Therefore, the character hijk is the text content lost by the visual language model recognition, and it can be supplemented to the end.

[0033] In a specific embodiment, we take the last 20-30 characters of each element string in the visual language model recognition result as the comparison substring and perform flip matching to avoid mismatching caused by the comparison substring being too short.

[0034] In this embodiment, step S4 implements position matching based on the position information and context information of each non-text element in the first element list and the second element list, and inserts the image element corresponding to the non-text element into the corresponding matched position in the second element list. This is to add non-text elements such as images and tables to the second element list according to the matching position based on the document page order, thereby enhancing the comprehensiveness of information extracted by the visual language model. The specific steps are as follows: Step S41, based on the first element list, determining the non-text element and its position coordinate information according to the element category label; Step S42: Obtain the context of the non-text element, perform similarity matching and alignment with the text elements in the second element list, and determine the position mapping of the non-text element in the second element list based on the non-text element relative to the context information, and assign and obtain the insertion position of the non-text element in the second element list.

[0035] Step S43: inserting the image element corresponding to the non-text element into the second element list based on the insertion position.

[0036] In this embodiment, step S5 automatically extracts the document directory list based on the style features of the original document elements, combined with the title numbering rules and the commonly used directory dictionary. The specific implementation steps are as follows: Step S51, identifying titles: using style features and regular expressions to identify titles in the document, and identifying all titles from the original document.

[0037] In this embodiment, title recognition is achieved through visual language modeling (VLM) or optical character recognition (OCR) technology, and style features (such as font size, color, and bold) and style rules are combined to determine which text blocks are titles. Style feature extraction uses image processing technology to extract style features (such as font size and color) of text areas. Style rules define a set of style rules to identify document titles. For example, text with a font size greater than a certain threshold and bold text is likely a title.

[0038] Step S52, Title Numbering and Sorting: Determine the title hierarchy according to the numbering rules. Once the titles are identified, they need to be categorized and sorted according to the title numbering rules. Common title numbering rules include numerical numbers (e.g., 1.1, 1.2, 2.1, etc.), alphabetical numbers (e.g., A, B, C, etc.), and unnumbered titles.

[0039] Step S53, title level correction: using a predefined directory dictionary to correct the title identified in step S52 and its corresponding level information; In order to improve the accuracy of title recognition, this embodiment uses a predefined common directory dictionary (including common titles and their corresponding levels). This dictionary can help correct the recognition results, especially when the style of some titles is not obvious.

[0040] Step S54, generate a directory list: organize the titles according to the hierarchical relationship and generate a structured directory list. Finally, organize the identified titles according to the hierarchical relationship into a structured directory list. Figure 6 This step usually involves hierarchical analysis and sorting to ensure the correctness of the directory structure.

[0041] In this embodiment, in order to ensure that the directory list generated in step S5 or step S54 is correct, the directory list is further marked and filtered using a hierarchical analysis method to correct the hierarchical relationship. Finally, the complete directory structure is generated, which specifically includes the following steps: Step S501, title structure extraction; each title in the original document is analyzed, the title text is segmented according to the preset delimiters, and its format information is extracted; through the mapping rules, the components of each title are converted into structured information; for example, the title "I. Work Progress" will be extracted as a structure list ['cnumber', 'other'], and the title "1.1 Work Objectives" will be extracted as a structure list ['number', 'number', 'other']. This process helps to construct a format structure list of the title as a basis for subsequent hierarchical allocation.

[0042] The format information in this embodiment may include chapter number, section number, letter number, other content and other information.

[0043] Step S502: Historical dictionary construction: During the title structure extraction process, the title structures that have been assigned levels and their corresponding level information are recorded to form a historical dictionary. This dictionary can help determine whether the newly identified title already exists in the previous level allocation and avoid duplicate level allocation.

[0044] Step S503: Special headings are assigned a hierarchy. Starting with the first heading in the original document, it is initially assigned a level 1 heading (level 1). During processing, any headings that conform to a specific format are directly assigned a level 1 heading. These special headings are not recorded in the historical dictionary because they belong to a fixed format or predefined category and have a predetermined hierarchy.

[0045] In this embodiment, the specific format titles include known titles such as "Chapter 1", "Preface", and "Table of Contents", which are conventionally used as first-level titles; Step S504, hierarchical allocation rules; for other titles, if they do not conform to the specific format and have not appeared in the historical dictionary, the hierarchy will be incremented by one level based on the hierarchical title determined above, that is, the default hierarchy is increased by 1. If the new title contains a cnumber (i.e., chapter number, such as 123123) and its hierarchy is higher than the title that does not contain a cnumber, that is, the priority of the cnumber is higher than that of Arabic numerals (such as 123) or letters (such as abc), then the title will be discarded to ensure the rationality and consistency of the hierarchical allocation. If the title conforms to the hierarchical allocation rules, the structure and hierarchy of the title will be recorded in the historical dictionary to verify that there will not be a situation where the Chinese character number is higher than the alphanumeric number (subtitle), so as to facilitate the hierarchy judgment of subsequent titles.

[0046] Step S505: Update the title hierarchy. After all title hierarchy information is automatically assigned and verified, it is updated into the corresponding title attributes, forming a complete document structure hierarchy and obtaining a revised table of contents. This result provides accurate hierarchical information support for subsequent document processing, formatting, and content analysis.

[0047] This embodiment implements a hierarchical automatic assignment method for title format analysis through the above-mentioned hierarchical analysis method, which is suitable for hierarchical assignment of non-uniform title formats. Especially when processing titles containing multiple formats such as PPT documents, when the hierarchical relationship is difficult to judge, it can achieve efficient and accurate title hierarchy identification and classification.

[0048] In this embodiment, step S6 constructs a directory-based knowledge tree based on the document directory list and the revised second element list. The specific steps are as follows: Step S61, parsing the document content: extracting all elements in the revised second element list and their content data; wherein the content data may include the page number, page coordinates, original text information, semantic information, etc.

[0049] Step S62, constructing a hierarchical structure: constructing a tree structure according to the hierarchical relationship in the revised directory list.

[0050] Step S63, filling in non-title elements: adding links and extended attributes of non-text elements (such as paragraphs, pictures, tables, etc.) to the corresponding chapters.

[0051] Step S64, generating structured output: converting the constructed knowledge tree into a structured output format (such as JSON).

[0052] This method can not only clearly display the hierarchical structure of the document, but also facilitate subsequent data processing and analysis. For details, please refer to Figure 7 The following is the final knowledge tree.

[0053] In this embodiment, the extended attribute information for non-text elements includes extensions for table and image elements, supplementing the table or image element screenshot file upload address, table element table file upload address, table element HTML text, and table element semantic content text attributes. This allows this useful information to be returned to the system or user along with the element for further information analysis and display.

[0054] Device embodiment According to an embodiment of the present invention, a document information extraction device is provided. Figure 8 FIG. 1 is a block diagram of a document information extraction device provided in this embodiment. The document information extraction device according to an embodiment of the present invention includes: A format conversion module 10, for converting a document into an image format; The first element list acquisition module 20 is used to sequentially process each image using optical character recognition technology to obtain a first element list of structured data. The first element list includes data such as the text content of each element, the element position coordinates, the element category label and the reading order; and then convert the identified non-text elements into image elements and store them.

[0055] Among them, the elements may include text blocks, pictures, tables, formulas and other elements. Based on the reading order, the layout type corresponding to each element in the document is determined, wherein the text reading order includes left-right reading order and top-down reading order, and the layout type includes paragraph layout, table layout and image layout, etc.

[0056] The second element list acquisition module 30 is used to process each image sequentially based on the visual language model to obtain a second element list of structured data. The second element list includes data such as the description text of each element and the reading order.

[0057] The second element list completion module 40 is used to align the first element list and the second element list at the element level, and then match the text elements in the context with the second element list based on the position information and context information of each non-text element in the first element list, and insert the non-text elements and the corresponding image elements into the second element list according to the matching positions to obtain a corrected second element list.

[0058] The document directory list extraction module 50 is used to automatically extract the document directory list based on the style features of the original document elements, combined with the title numbering rules and the commonly used directory dictionary.

[0059] The knowledge tree generation module 60 is configured to construct a directory-based knowledge tree based on the document directory list and the modified second element list.

[0060] The document information extraction device provided in this embodiment, in order to solve the problem that there are still certain limitations in the processing of various document formats, complex content and multi-language mixed typesetting, firstly, the format conversion module 10 converts the document into an image, and then uses the visual language model to sequentially obtain the semantic information of the text, picture, table and other elements and their context in each image. Since OCR technology is relatively good at obtaining table content recognition, non-text element information such as pictures, tables, formulas, etc. will be obtained through OCR technology and converted into image format, and the context information of the two element lists will be combined to determine the corresponding elements and relative positions of non-text elements such as pictures, tables, formulas in the second element list, and replace or supplement them, so as to use the OCR extraction results to verify the visual language model extraction results, supplement the deficiencies of the visual language model in visual element extraction, and effectively improve the accuracy and efficiency of document information extraction. For documents with complex typesetting, tables, images and other information sources, the required information can be extracted comprehensively and accurately.

[0061] In this embodiment, the second element list completion module 40 includes a traversal submodule and a non-text element confirmation submodule; the traversal submodule is used to traverse the second element of the second element sublist based on the first element sublist and the second element sublist of the same image, using the information of each element in the first element sublist as a comparison standard and the element as a comparison unit, and perform element text content matching and comparison with the first element in the first element sublist in order / reverse order. When the similarity is greater than a preset value, feedback is given to the non-text element confirmation submodule; if no corresponding element is matched, the second element sublist is continued to be traversed; The non-text element confirmation submodule is used to determine whether the element category labels of the upper and lower elements of the first element are non-text elements. If so, the corresponding non-text elements are inserted above or below the second element according to their positions relative to the first element; if not, the upper and lower elements of the first element are text elements, and the second element sublist continues to be traversed.

[0062] In this embodiment, a text content completion submodule is also included, which is used to match and compare the element text content of the second element with the first element in the first element sublist when the traversal submodule determines that the second element is matched with the first element in the first element sublist. When the similarity is greater than a preset value, the element text content of the first element is used as the comparison standard, and based on the preset size window sliding and matching distance, the second element text content is matched with the first element text content in reverse order to determine whether the second element text content is missing. If the text content is missing, all text characters after the last matching text character position of the first element are sequentially filled to the end of the second element text content according to the last matching text character position.

[0063] The embodiment of the present invention is an apparatus embodiment corresponding to the above-mentioned method embodiment. The specific operations of the processing steps of each module can be understood by referring to the description of the method embodiment, and will not be repeated here.

[0064] like Figure 9 As shown, the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the document information extraction method in the above embodiment is implemented.

[0065] The present invention also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the document information extraction method in the above embodiment is implemented.

[0066] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0067] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments. The device and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. A person of ordinary skill in the art can understand and implement it without making any creative efforts.

[0068] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and the contents not described in detail in the specification of the present invention are common knowledge to those skilled in the art.

Claims

1. A document information extraction method, characterized in that: Including steps: Convert documents into image formats; After sequentially processing each image using optical character recognition technology, a first element list of structured data is obtained. The first element list includes the text content of each element, the element's position coordinates, the element's category label, and reading order data. The recognized non-text elements are then converted into image elements for storage. After sequentially processing each image based on the visual language model, a second element list of structured data is obtained, wherein the second element list includes a description text and reading order data of each element; Aligning the first element list and the second element list at the element level, and then matching the text elements in the context with the second element list based on the position information and context information of each non-text element in the first element list, and inserting the non-text elements and the corresponding image elements into the second element list according to the matched relative positions, to obtain a revised second element list; Automatically extract the document directory list based on the style features of the original document elements, combined with the title numbering rules and common directory dictionaries; Based on the document directory list and the revised second element list, a directory-based knowledge tree is constructed.

2. The document information extraction method according to claim 1, wherein: Also includes the steps: Take screenshots and store the images, tables, and formulas recognized by OCR technology; Perform semantic analysis on the table elements, convert them into corresponding table semantic texts, and retain them as text elements in the first element list.

3. The document information extraction method according to claim 1, wherein: Based on the position information and context information of each non-text element in the first element list, text elements are matched by the text elements in the context with the second element list, and the non-text elements and corresponding image elements are inserted into the second element list according to the matching positions. The specific steps are as follows: 1) Based on the first element sublist and the second element sublist of the same image, using the information of each element in the first element sublist as the comparison standard and the element as the comparison unit, the second element in the second element sublist is traversed and the text content is matched with the first element in the first element sublist in the order or reverse order. If the similarity is greater than a preset value, the process proceeds to the next step 2); if no corresponding element is matched, the process continues to traverse the second element sublist; 2) Determine whether the element category tags of the upper and lower elements of the first element are non-text elements. If so, the corresponding non-text elements are inserted above or below the second element according to their positions relative to the first element; if not, the upper and lower elements of the first element are text elements, and continue traversing the second element sublist.

4. The document information extraction method according to claim 3, wherein: The process of matching and comparing the element text content with the first element in the first element sublist, when the similarity is greater than a preset value, further includes the following steps: Taking the element text content of the first element as the comparison standard, based on the preset window sliding size and matching distance, the second element text content is matched with the first element text content in reverse order to determine whether the second element text content is missing. If the text content is missing, all text characters after the last matching text character position of the first element are sequentially filled to the end of the second element text content according to the last matching text character position.

5. The document information extraction method according to claim 1, wherein: The method of matching elements based on the position information and context information of each non-text element in the first element list and the second element list, and inserting the non-text element and the corresponding image element into the second element list according to the matched relative position, includes the following steps: Based on the first element list, determining the non-text element and its position coordinate information according to the element category label; Obtaining context information of the non-text element, performing similarity matching and alignment with the text elements in the second element list, determining a position mapping of the non-text element in the second element list based on the non-text element relative to the context information, and assigning and obtaining an insertion position of the non-text element in the second element list; The image element corresponding to the non-text element is inserted into the second element list based on the insertion position.

6. A document information extraction device, characterized in that: include: Format conversion module, used to convert documents into image formats; A first element list acquisition module is configured to sequentially process each image using optical character recognition technology to obtain a first element list of structured data, the first element list including the text content of each element, element position coordinates, element category label, and reading order data; and then convert the recognized non-text elements into image elements and store them; a second element list acquisition module, configured to sequentially process each image based on a visual language model to obtain a second element list of structured data, the second element list including a description text of each element and reading order data; A second element list completion module is used to align the first element list and the second element list at the element level, and then match the text elements in the context with the second element list based on the position information and context information of each non-text element in the first element list, and insert the non-text elements and the corresponding image elements into the second element list according to the matching positions to obtain a corrected second element list; Document directory list extraction module, which is used to automatically extract the document directory list based on the style features of the original document elements, combined with the title numbering rules and common directory dictionaries; The knowledge tree generation module is used to construct a directory-based knowledge tree based on the document directory list and the revised second element list.

7. The document information extraction device according to claim 6, wherein: The second element list completion module includes a traversal submodule and a non-text element confirmation submodule; A traversal submodule is configured to traverse the second element of the second element sublist based on the first element sublist and the second element sublist of the same image, using the information of each element in the first element sublist as a comparison criterion and the element as a comparison unit, and to perform a text content match comparison with the first element in the first element sublist in order or in reverse order. When the similarity is greater than a preset value, the result is fed back to the non-text element confirmation submodule; if no corresponding element is matched, the second element sublist is continued to be traversed; The non-text element confirmation submodule is used to determine whether the element category labels of the upper and lower elements of the first element are non-text elements. If so, the corresponding non-text elements are inserted above or below the second element according to their positions relative to the first element; if not, the upper and lower elements of the first element are text elements, and the second element sublist continues to be traversed.

8. The document information extraction device according to claim 6, wherein: It also includes a text content completion submodule, which is used to match and compare the element text content of the second element with the first element in the first element sublist when the traversal submodule determines that the second element is matched with the first element in the first element sublist. When the similarity is greater than a preset value, the element text content of the first element is used as the comparison standard, and based on the preset size window sliding and matching distance, the second element text content is matched with the first element text content in reverse order to determine whether the second element text content is missing. If the text content is missing, all text characters after the last matching text character position of the first element are sequentially filled to the end of the second element text content according to the last matching text character position.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the document information extraction method according to any one of claims 1 to 5 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the document information extraction method according to any one of claims 1 to 5 is implemented.

Citation Information

Cited By

  • Multi-format text analysis method and system based on large language model

    CN122113896A