Method and System for Configuring Extraction Based on Multimodal Document Information
Through the multimodal document information configuration extraction method, combined with PDF text extraction, OCR image text recognition and table extraction technology, the problem of accuracy and inefficiency of document information extraction in the existing technology is solved, and efficient and accurate information extraction and user-friendly operation experience are achieved.
Patent Information
- Application Number
- CN202411056081.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-02
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-08-02
AI Technical Summary
The prior art has problems of accuracy and inefficiency in document information extraction, especially when processing complex documents and meeting user personalized needs, it is difficult to achieve efficient and accurate information extraction.
The multimodal document information configuration extraction method is adopted, and PDF text extraction, OCR image text recognition and table extraction technologies are integrated to build a comprehensive processing strategy, and efficient and accurate information extraction is achieved through preliminary analysis, recognition pattern selection, customized extraction strategies and cache mechanisms.
It improves the accuracy and efficiency of information extraction, supports users to define personalized extraction rules without coding, reduces user work burden, and improves the flexibility and user experience of document information extraction.
Smart Images

Figure CN118865419B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document information extraction; specifically, it relates to a method and system for extracting multi-modal document information configuration. Background Art
[0002] With the popularization of digital office work, the management and information extraction of a large number of paper documents and electronic documents have become an urgent need.
[0003] Although existing OCR technologies can achieve automatic recognition of text to a certain extent, they still face challenges in terms of accuracy in precise information extraction. For example, existing technologies often have difficulty accurately distinguishing and extracting information within a specific range, or are inefficient when processing documents with mixed table and image layouts, resulting in incomplete or frequent errors in information extraction.
[0004] Especially for tasks that require highly customized processing, such as only extracting specific paragraphs in a document, specific rows and columns of a table, or text that meets a specific regular expression, existing technical solutions are limited in terms of flexibility and accuracy. In addition, for user-specific requirements, such as specific format conversion and data structuring, complex post-processing procedures are often required in existing technologies.
[0005] In addition, after existing solutions using OCR technology recognize the content of a document, the end user often has to participate in the verification and correction of the recognition results. This requires the user to check the text recognized by OCR one by one and manually add or correct the incorrect or missing parts. In addition, for further extraction of text information, the user may also need to write custom scripts or rules to define the extraction logic and scope, which has a relatively high technical threshold, greatly increasing the user's workload and reducing work efficiency.
[0006] Therefore, there is an urgent need to develop a more accurate, efficient, and user-friendly document information extraction solution, making the information extraction process smoother and more efficient, reducing the burden on the user in result verification and rule formulation, and eliminating the need for the user to perform cumbersome manual adjustments or coding work to solve the above-mentioned difficulties and pain points of the existing technology. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to propose a method and system for configuration extraction based on multi-modal document information. Aiming at the processing limitation problem caused by the existing technology relying only on a single recognition scheme, the present invention integrates three key technologies: PDF text extraction, OCR image text recognition, and table extraction, constructs a comprehensive and efficient document information extraction system, and forms a comprehensive processing strategy. It not only overcomes the shortcomings of individual technologies in complex document processing, solves the limitations of the existing technology in complex document processing, but also ensures the efficient and comprehensive extraction of information from various documents through multi-dimensional recognition and parsing, realizes the automatic and accurate extraction of information from various documents, and improves the breadth and depth of information extraction. Aiming at the problem that the existing technology is not flexible enough in information extraction customization and requires users to have programming capabilities, the present invention provides a simple configuration interface and preset templates, allowing users to easily define personalized extraction rules without coding, reducing the user's workload, and improving the work efficiency of document information extraction.
[0008] The present invention provides a method for configuration extraction based on multi-modal document information, including the following steps:
[0009] S1. Conduct a preliminary analysis of the input document, identify the file type and format, and preprocess the document. The preprocessing includes: format recognition, conversion (such as PDF format, image format, etc.), resolution detection and optimization adjustment, to prepare for subsequent recognition;
[0010] S2. According to the document type and content distribution, select the most suitable recognition mode and configure the corresponding recognition parameters. Among them, the recognition modes include: direct PDF text extraction, OCR recognition, or table recognition. The recognition process adopts a caching mechanism, using the file hash value as the key and the recognition result as the value, to avoid excessive waiting time for multiple configuration recognitions;
[0011] Specifically, analyze whether the document is a native PDF, a scanned PDF, or a PDF containing a table. Among them, a native PDF document is directly generated by typesetting software such as Word or WPS, and the text information in it is a directly retrievable text stream rather than an image, and direct text extraction is used, with a fast extraction speed; the text in a scanned PDF exists in the form of an image rather than a retrievable text, and OCR recognition is used, with a slow extraction speed; for a document containing a table, whether it is composed of text or an image, it may have a complex row and column structure and merged cells, and table recognition is used to maintain the table structure.
[0012] The selection of the most suitable recognition mode and the configuration of the corresponding recognition parameters include:
[0013] If the direct PDF text extraction recognition mode is selected, directly extract the text content from the PDF document;
[0014] If the recognition mode of OCR recognition is selected, the original structure and size are maintained;
[0015] If the recognition mode of table recognition is selected, through advanced layout analysis, accurately recognize the table structure, including row and column division, merged cell processing, and convert the table data into structured data;
[0016] S3. For the recognition results of different types of recognition modes, adopt customized extraction strategies and recognition extraction rules, and set extraction elements by category: for the recognition results directly extracted from PDF text, match and extract by constructing regular expressions; for the recognition results of OCR recognition, combine the position information and accurately obtain the text according to the coordinate range; for the recognition results of table recognition, complete information extraction by selecting cells;
[0017] S4. Determine the document page number position corresponding to each recognition extraction rule, and accurately extract the required information; respond with the extraction results in JSON format or call back through a custom callback address;
[0018] S5. Record and save the recognition extraction rules and recognition extraction configurations (to facilitate the automated processing of future similar documents or the reuse of rules), support the use of the completed recognition extraction configurations through multiple interfaces, integrate multiple file upload methods (supporting multi-channel file input, which can automatically download and uniformly process by the user uploading local files or entering network links), flexible configuration management, and an efficient caching mechanism to achieve intelligent extraction and processing of document information. This not only improves the flexibility and efficiency of information extraction, but also optimizes resource utilization, ensuring the convenience and response speed of user operations.
[0019] Further, the method of matching and extracting the recognition results of directly extracting PDF text in step S3 by constructing regular expressions includes:
[0020] Adopt regular expressions as a powerful text matching tool, and accurately locate the target content by defining keywords and their prefixes and suffixes; if there are multiple matching items in the extraction results, use a pre-set specific delimiter (such as a comma, semicolon, or line break) to separate the extracted content to ensure the structured and easy-to-process nature of the output results.
[0021] Further, the method of accurately obtaining text according to the coordinate range by combining the position information for the recognition results of OCR recognition in step S3 includes:
[0022] Realize the range extraction function by analyzing the text block coordinate information in the recognition results, and traverse the position data (X-axis, Y-axis coordinates and their width and height) of each text block; according to the extraction range configured by the user (such as the upper left corner coordinates and the lower right corner coordinates), judge whether the text block falls into the specified area to achieve accurate content capture;
[0023] Further, the method for extracting information by selecting cells from the recognition result of table recognition in step S3 includes:
[0024] For table data, extraction rules are configured according to the table structure, specifying the number of rows and columns to be extracted; by parsing the table structure, the content of cells within the specified range is accurately extracted according to the configured information, supporting flexible extraction of complex table structures.
[0025] Further, the process of regular expression matching extraction in step S3 is clearly presented in the form of a graphical user interface. The method for keyword overlap processing and highlighting display in regular expression matching extraction includes the following steps:
[0026] S31. Traverse all regular expression configurations, find all text matching items according to the regular expression, and obtain the index of the matching text in the original text, denoted as the matching item array;
[0027] S32. Sort the matching item array by the starting index, and record the first matching item as variable current;
[0028] S33. Traverse the remaining elements of the array. For each element next, check whether the starting index of this element next is less than the ending index of current to determine whether there is an overlap; if there is no overlap, set next as current and continue the next loop; if there is an overlap, perform the following operations respectively according to different overlap situations:
[0029] If next overlaps with the tail of current, then cut off the part of next in current, add the overlap weight of next to the weight of current, and update current to next;
[0030] If next is completely within current, then divide current into three parts by next. Each part in the three parts updates its own starting index. At the same time, add the overlap weight of next to the weight of current, and insert the last part into the array, and then re - execute step S32;
[0031] If the front segment of next and the rear segment of current overlap, then split out the overlapping part as newItem, the overlapping weight of newItem is the superposition of next and current, update the starting index of each part, and then execute step S32;
[0032] S34. Return the array of the final matching results without overlap;
[0033] S35. Remove all the original highlights in the text. According to the processed matching information, replace the text one by one, add the corresponding-level highlight HTML tags according to the overlapping weights, and update the innerHTML of the DOM element to achieve highlight display.
[0034] The algorithm for overlapping processing of regular expression matching extraction in the present invention is based on the principles of array sorting and iterative processing. First, sort all the matching results according to the starting positions, and then traverse the array to compare whether the current item overlaps with the next item; if an overlap is found, split out the overlapping part, set the overlapping weight, and then re-sort the split part according to the starting position. Repeat the above steps until there is no overlapping part. This enables users to intuitively adjust the highlight level dynamically according to the keyword overlapping situation.
[0035] Further, the method of selecting the most suitable recognition mode and configuring the corresponding recognition parameters in the S2 step includes: querying the database to obtain detailed configuration details, including the recognition page range and specific extraction rules; analyzing the configuration details to set specific parameters for subsequent recognition and extraction, ensuring that the processing process meets the user's requirements.
[0036] Further, the method of adopting a customized extraction strategy and recognition extraction rules and setting extraction elements by category in the S3 step further includes:
[0037] For complex information structures encountered in document processing other than direct PDF text extraction, OCR recognition, and table recognition, through an integrated sandboxed JavaScript engine, highly customized data processing capabilities are provided while ensuring that the system security is not threatened.
[0038] The present invention also provides a system for configuration extraction based on multi-modal document information, which executes the method for configuration extraction based on multi-modal document information as described above, including:
[0039] A preprocessing module: used for preliminary analysis of the input document, identifying the file type and format, and preprocessing the document. The preprocessing includes: format recognition, conversion, resolution detection and optimization adjustment;
[0040] Configuration Recognition Mode Module: It is used to select the most suitable recognition mode according to the document type and content distribution, and configure the corresponding recognition parameters. Among them, the recognition modes include: direct PDF text extraction, OCR recognition, or table recognition. The recognition process adopts a caching mechanism, using the file hash value as the key and the recognition result as the value, to avoid excessive waiting time for multiple configuration recognitions. The selection of the most suitable recognition mode and the configuration of the corresponding recognition parameters include: if the direct PDF text extraction recognition mode is selected, the text content is directly extracted from the PDF document; if the OCR recognition mode is selected, the original structure and size are maintained; if the table recognition mode is selected, through advanced layout analysis, the table structure is accurately recognized, including row and column division and merged cell processing, and the table data is converted into structured data.
[0041] Recognition and Extraction Rule Module: It is used to adopt customized extraction strategies and recognition and extraction rules for the recognition results of different types of recognition modes, and set extraction elements by category. For the recognition results of direct PDF text extraction, they are extracted by constructing a regular expression match; for the recognition results of OCR recognition, combined with the position information, the text is accurately obtained according to the coordinate range; for the recognition results of table recognition, the information is extracted by selecting cells.
[0042] Accurate Information Extraction Module: It is used to determine the document page number position corresponding to each recognition and extraction rule, and accurately extract the required information. The extraction result is responded in JSON format or called back through a custom callback address.
[0043] Save Extraction Configuration Module: It is used to record and save the recognition and extraction rules and recognition and extraction configurations, support using the completed recognition and extraction configurations through multiple interfaces, integrate multiple file upload methods, flexible configuration management, and an efficient caching mechanism, and realize the intelligent extraction and processing of document information.
[0044] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method for configuring and extracting based on multi-modal document information as described above are implemented.
[0045] The present invention also provides a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for configuring and extracting based on multi-modal document information as described above are implemented.
[0046] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0047] The method and system for multi-modal document information configuration extraction provided by the present invention can precisely define the extraction range in multiple dimensions (such as keyword context, coordinate range, or table item position), improving the accuracy and adaptability of the extraction process; design an algorithm for processing the overlapping problem of regular expression matching results, dynamically adjust the highlighting level by analyzing the positional relationship of the matching text, avoid the visual chaos caused by highlighting superposition, and improve the accuracy of highlighting processing and the readability of the document; provide a graphical user interface that allows users to intuitively preview the recognition results and highlighting effects, support instant feedback and adjustment, and enhance the intuitiveness and convenience of user operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention.
[0049] In the drawings:
[0050] Figure 1 is a schematic flowchart of the recognition and extraction configuration of an embodiment of the present invention;
[0051] Figure 2 is a schematic flowchart of the basic process of recognition and extraction of an embodiment of the present invention;
[0052] Figure 3 is a flowchart of the method for multi-modal document information configuration extraction of the present invention;
[0053] Figure 4 is a flowchart of the method for keyword overlap processing and highlighting display in regular expression matching extraction of the present invention;
[0054] Figure 5 is a schematic diagram of the composition of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are only examples of devices and products consistent with some aspects of the present disclosure as detailed in the appended claims.
[0056] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. The singular forms "a", "the", and "said" used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0057] It should be understood that although the terms first, second, third, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0058] The embodiments of the present invention will be further described in detail below.
[0059] The embodiments of the present invention provide a method for extracting based on multi-modal document information configuration. Refer to Figure 3 as shown, including the following steps:
[0060] S1. Perform a preliminary analysis on the input document, identify the file type and format, and preprocess the document. The preprocessing includes: format recognition, conversion (such as PDF format, picture format, etc.), resolution detection and optimization adjustment;
[0061] S2. Select the most suitable recognition mode according to the document type and content distribution, and configure the corresponding recognition parameters; wherein, the recognition mode includes: direct PDF text extraction, OCR recognition, or table recognition; a caching mechanism is adopted during the recognition process, with the file hash value as the key and the recognition result as the value, to avoid excessive waiting time for multiple configuration recognitions;
[0062] Preferably, query the database to obtain detailed configuration details, including the range of recognized pages and specific extraction rules; analyze the configuration details to set specific parameters for subsequent recognition and extraction.
[0063] Analyze whether the document is a native PDF, a scanned PDF, or a PDF containing tables. Among them, a native PDF document is directly generated by typesetting software such as Word or WPS, and the text information in it is a directly retrievable text stream rather than an image. The text can be directly extracted, and the extraction speed is fast. The text in a scanned PDF exists in the form of an image rather than a retrievable text. OCR recognition is used, and the extraction speed is slow. Whether a document containing tables is composed of text or images, it may have complex row and column structures and merged cells. Using table recognition can maintain the table structure.
[0064] Select the most suitable recognition mode and configure the corresponding recognition parameters, including:
[0065] If the recognition mode of direct PDF text extraction is selected, directly extract the text content from the PDF document;
[0066] If the recognition mode of OCR recognition is selected, maintain the original structure and size;
[0067] If the recognition mode of table recognition is selected, through advanced layout analysis, accurately identify the table structure, including row and column division and merged cell processing, and convert the table data into structured data;
[0068] S3. For the recognition results of different types of recognition modes, adopt customized extraction strategies and recognition extraction rules, and set extraction elements by category: for the recognition results of direct PDF text extraction, construct a regular expression to match and extract; for the recognition results of OCR recognition, combine the position information and accurately obtain the text according to the coordinate range; for the recognition results of table recognition, complete the information extraction by selecting cells;
[0069] Specifically, for the recognition results of direct PDF text extraction, the method of constructing a regular expression to match and extract includes:
[0070] Use a regular expression as a powerful text matching tool to accurately locate the target content by defining keywords and their prefixes and suffixes; if there are multiple matching items in the extraction result, use a pre-set specific delimiter (such as a comma, semicolon, or line break) to separate the extracted content to ensure the structured and easy-to-process nature of the output result.
[0071] For the recognition results of OCR recognition, the method of combining the position information and accurately obtaining the text according to the coordinate range includes:
[0072] The range extraction function is realized by analyzing the text block coordinate information in the recognition result, and the position data (X-axis, Y-axis coordinates and their width and height) of each text block is traversed; according to the extraction range configured by the user (such as the upper left corner coordinates and the lower right corner coordinates), it is judged whether the text block falls into the specified area to achieve accurate content capture;
[0073] For the recognition results of table recognition, the methods for extracting information by selecting cells include:
[0074] For table data, extraction rules are configured according to the table structure, and the number of rows and columns to be extracted are specified; by parsing the table structure, the content of the cells within the specified range is accurately extracted according to the configuration information, supporting flexible extraction of complex table structures.
[0075] For complex information structures encountered in document processing other than direct extraction of PDF text, OCR recognition, and table recognition, through an integrated sandboxed JavaScript engine, highly customized data processing capabilities are provided while ensuring the security of the system is not threatened.
[0076] In this embodiment, the process of regular expression matching extraction is clearly presented in the form of a graphical user interface. For the method of keyword overlap processing and highlighting display in the regular expression matching extraction, see Figure 4 shown, including the following steps:
[0077] S31. Traverse all regular expression configurations, find all text matching items according to the regular expression, and obtain the indexes of the matching text in the original text, which are recorded as the matching item array;
[0078] S32. Sort the matching item array by the starting index, and record the first matching item as the variable current;
[0079] S33. Traverse the remaining elements of the array. For each element next, check whether the starting index of this element next is less than the ending index of current to determine whether there is an overlap; if there is no overlap, set next as current and continue the next loop; if there is an overlap, perform the following operations respectively according to different situations of the overlap:
[0080] If next overlaps with the tail of current, then cut off the part of next in current, add the overlap weight of next to the weight of current, and update current to next;
[0081] If next is completely within current, then current is divided into three parts by next. Each part updates its own starting index. At the same time, the overlapping weight of next is added to the weight of current, and the last part is inserted into the array. Then, step S32 is executed again;
[0082] If the front segment of next and the rear segment of current overlap, then the overlapping part is split out as newItem. The overlapping weight of newItem is the superposition of next and current. Update the starting index of each part, and then execute step S32;
[0083] S34. Return the array of the final matching results without overlap;
[0084] S35. Remove all the original highlights in the text. According to the processed matching information, replace the text one by one. Add the corresponding level of highlight HTML tags according to the overlapping weight, and update the innerHTML of the DOM element to achieve highlight display.
[0085] The algorithm for overlapping processing in regular expression matching extraction in this embodiment is based on the principles of array sorting and iterative processing. First, all matching results are sorted by the starting position, and then the array is traversed to compare whether the current item and the next item overlap; if overlap is found, the overlapping part is split out, and the overlapping weight is set, and then the split part is sorted again by the starting position. Repeat the above steps until there is no overlapping part. This enables users to intuitively adjust the highlight level dynamically according to the keyword overlapping situation.
[0086] S4. Determine the page number position of the document corresponding to each item recognition and extraction rule, and accurately extract the required information; respond with the extraction result in JSON format or call back through a custom callback address;
[0087] S5. Record and save the recognition and extraction rules and recognition and extraction configurations (for the convenience of future automated processing of similar documents or reuse of rules, see Figure 1 as shown), support the use of the completed recognition and extraction configurations through multiple interfaces, integrate multiple file upload methods (supporting multi-channel file input, which can automatically download and uniformly process by user uploading local files or inputting network links), flexible configuration management, and an efficient caching mechanism to achieve intelligent extraction and processing of document information. This not only improves the flexibility and efficiency of information extraction but also optimizes resource utilization, ensuring the convenience and response speed of user operations. Figure 2 The basic process of recognition and extraction in this embodiment is shown.
[0088] The embodiment of the present invention also provides a system for multi-modal document information configuration extraction, which executes the method for multi-modal document information configuration extraction as described above, including:
[0089] Preprocessing module: used to perform preliminary analysis on the input document, identify the file type and format, and preprocess the document. The preprocessing includes: format recognition, conversion, resolution detection, and optimization adjustment;
[0090] Configuration recognition mode module: used to select the most suitable recognition mode according to the document type and content distribution, and configure the corresponding recognition parameters; among them, the recognition modes include: direct PDF text extraction, OCR recognition, or table recognition; the recognition process adopts a caching mechanism, using the file hash value as the key and the recognition result as the value, to avoid the long waiting time for multiple configuration recognitions; the selection of the most suitable recognition mode and the configuration of the corresponding recognition parameters include: if the direct PDF text extraction recognition mode is selected, directly extract the text content from the PDF document; if the OCR recognition mode is selected, keep the original structure and size; if the table recognition mode is selected, accurately identify the table structure through advanced layout analysis, including row and column division, merged cell processing, and convert the table data into structured data;
[0091] Recognition and extraction rule module: used to adopt customized extraction strategies and recognition and extraction rules for the recognition results of different types of recognition modes, and set extraction elements by category: for the recognition results of direct PDF text extraction, extract by constructing a regular expression match; for the recognition results of OCR recognition, accurately obtain the text based on the position information and coordinate range; for the recognition results of table recognition, complete information extraction by selecting cells;
[0092] Accurate information extraction module: used to determine the document page number position corresponding to each recognition and extraction rule, and accurately extract the required information; respond with the extraction result in JSON format or call back through a custom callback address;
[0093] Save extraction configuration module: used to record and save the recognition and extraction rules and recognition and extraction configurations, support using the completed recognition and extraction configurations through multiple interfaces, integrate multiple file upload methods, flexible configuration management, and an efficient caching mechanism, and realize the intelligent extraction and processing of document information.
[0094] The embodiment of the present invention also provides a computer device, Figure 5 which is a schematic structural diagram of a computer device provided by the embodiment of the present invention; see the attached drawing Figure 5As shown in the figure, the computer device includes: an input device 23, an output device 24, a memory 22, and a processor 21; the memory 22 is used to store one or more programs; when the one or more programs are executed by the one or more processors 21, the one or more processors 21 implement the method for multi-modal document information configuration extraction provided in the above embodiments; wherein the input device 23, the output device 24, the memory 22, and the processor 21 can be connected through a bus or other means. Figure 5 Taking connection through the bus as an example.
[0095] The memory 22, as a readable and writable storage medium of a computing device, can be used to store software programs and computer-executable programs, such as the program instructions corresponding to the method for multi-modal document information configuration extraction described in the embodiments of the present invention; the memory 22 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the device, etc.; in addition, the memory 22 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices; in some instances, the memory 22 can further include a memory remotely set relative to the processor 21, and these remote memories can be connected to the device through a network. Examples of the above network include but are not limited to the Internet, an enterprise internal network, a local area network, a mobile communication network, and combinations thereof.
[0096] The input device 23 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function control of the device; the output device 24 can include a display device such as a display screen.
[0097] The processor 21 executes various functional applications and data processing of the device by running software programs, instructions, and modules stored in the memory 22, that is, implements the above method for multi-modal document information configuration extraction.
[0098] The computer device provided above can be used to execute the method for multi-modal document information configuration extraction provided in the above embodiments, and has corresponding functions and beneficial effects.
[0099] An embodiment of the present invention further provides a storage medium containing computer-executable instructions, and the computer-executable instructions are used to execute the method for extracting based on multi-modal document information configuration provided in the above embodiment when executed by a computer processor. The storage medium is any of various types of memory devices or storage devices, including: installation media such as CD-ROMs, floppy disks or magnetic tape devices; computer system memories or random access memories such as DRAM, DDRRAM, SRAM, EDO RAM, Rambus RAM, etc.; non-volatile memories such as flash memories, magnetic media (such as hard disks or optical storage); registers or other similar types of memory elements, etc.; the storage medium may also include other types of memories or combinations thereof; additionally, the storage medium may be located in a first computer system in which the program is executed, or may be located in a different second computer system, and the second computer system is connected to the first computer system through a network (such as the Internet); the second computer system may provide program instructions to the first computer for execution. The storage medium includes two or more storage media that may reside in different locations (such as in different computer systems connected through a network). The storage medium may store program instructions (such as specifically implemented as a computer program) executable by one or more processors.
[0100] Of course, for a storage medium containing computer-executable instructions provided by an embodiment of the present invention, the computer-executable instructions are not limited to the method for extracting based on multi-modal document information configuration described in the above embodiment, and may also execute related operations in the method for extracting based on multi-modal document information configuration provided by any embodiment of the present invention.
[0101] So far, the technical solution of the present invention has been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the protection scope of the present invention.
[0102] The above are only the preferred embodiments of the present invention and are not used to limit the present invention; for those skilled in the art, the present invention may have various changes and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for extracting multimodal document information configuration, characterized in that: The following steps are involved: S1. Perform preliminary analysis on the input document, identify the file type and format, and pre-process the document. The pre-processing includes: format recognition, conversion, resolution detection and optimization adjustment; S2. Select the most suitable recognition mode according to the document type and content distribution, and configure the corresponding recognition parameters; wherein the recognition mode includes: direct extraction of PDF text, OCR recognition or table recognition; the recognition process adopts a cache mechanism, with the file hash value as the key and the recognition result as the value, to avoid long waiting time for multiple configurations of recognition; The selecting the most suitable recognition mode and configuring corresponding recognition parameters includes: If you select the recognition mode of PDF text direct extraction, the text content will be directly extracted from the PDF document; If you select the OCR recognition mode, the original structure and size will be kept; If you select the table recognition mode, the table structure will be accurately recognized through advanced layout analysis, including row and column division and cell merging, to convert the table data into structured data. S3. Customized extraction strategies and recognition extraction rules are used for the recognition results of different types of recognition modes, and extraction elements are set by category: for the recognition results of direct PDF text extraction, regular expression matching is constructed for extraction; for the recognition results of OCR recognition, the text is accurately obtained according to the coordinate range in combination with the location information; for the recognition results of table recognition, information extraction is completed by selecting cells; S4. Determine the document page position corresponding to each recognition and extraction rule, and accurately extract the required information; respond to the extraction result in JSON format or call back through a custom callback address; S5. Record and save the recognition and extraction rules and recognition and extraction configurations, support the use of completed recognition and extraction configurations through multiple interfaces, integrate multiple file upload methods, flexible configuration management and efficient caching mechanisms, and realize intelligent extraction and processing of document information; The process of regular expression matching extraction in step S3 is presented in the form of a graphical user interface. The method for keyword overlap processing and highlighting in regular expression matching extraction includes the following steps: S31. Traverse all regular expression configurations, search for all text matching items according to the regular expression, and obtain the index of the matching text in the original text, which is recorded as a matching item array; S32, sort the matching item array according to the starting index, and record the first matching item as the variable current; S33, traverse the remaining elements of the array, for each element next, check whether the starting index of the element next is less than the ending index of current, and determine whether there is overlap; if there is no overlap, set next to current and continue the next loop; if there is overlap, perform the following operations according to different overlap situations: If the tail of next overlaps with the tail of current, the part of next in current is cut off, the overlapping weight of next is added to the weight of current, and current is updated to next; If next is completely in current, current is divided into three parts by next, each of the three parts updates its own starting index, and the overlapping weight of next is added to the weight of current, and the last part is inserted into the array, and then step S32 is re-executed; If the first segment of next and the second segment of current overlap, the overlapping portion is split into newItem, the overlapping weight of newItem is the superposition of next and current, the starting index of each portion is updated, and then step S32 is executed; S34, returning an array of final matching results without overlap; S35. Remove all existing highlights in the text, replace the text one by one according to the processed matching information, add highlight HTML tags of corresponding levels according to the overlapping weights, and update the innerHTML of the DOM element to achieve highlight display.
2. The method for extracting multimodal document information configuration according to claim 1, characterized in that: The recognition result of the direct extraction of PDF text in step S3 is extracted by constructing a regular expression matching method, which includes: Regular expressions are used as text matching tools to accurately locate target content by defining keywords and their prefixes and suffixes. If there are multiple matches in the extraction results, pre-set specific delimiters are used to separate the extracted content to ensure the structured and easy-to-handle output results.
3. The method for extracting multimodal document information configuration according to claim 1, characterized in that: The method of combining the OCR recognition result in step S3 with the position information to accurately obtain the text according to the coordinate range includes: Range extraction is achieved by analyzing the coordinate information of the text blocks in the recognition results, and the location data of each text block is traversed; according to the extraction range configured by the user, it is determined whether the text block falls into the specified area to achieve accurate content capture.
4. The method for extracting multimodal document information configuration according to claim 1, characterized in that: The method for extracting information from the recognition result of the table recognition in step S3 by selecting cells includes: For table data, configure extraction rules according to the table structure and specify the number of rows and columns to be extracted; by parsing the table structure, accurately extract the cell content in the specified range based on the configuration information, and support flexible extraction of complex table structures.
5. The method for extracting multimodal document information configuration according to claim 1, characterized in that: The method of selecting the most suitable recognition mode and configuring corresponding recognition parameters in step S2 includes: querying a database to obtain detailed configuration details, including the recognition page range and specific extraction rules; analyzing the configuration details to set specific parameters for subsequent recognition and extraction.
6. The method for extracting multimodal document information configuration according to claim 1, characterized in that: The method of adopting a customized extraction strategy and identifying extraction rules in step S3 and setting extraction elements by category also includes: For complex information structures encountered in document processing other than direct PDF text extraction, OCR recognition, and table recognition, the integrated sandboxed JavaScript engine provides highly customized data processing capabilities while ensuring that system security is not threatened.
7. A system for extracting multimodal document information configuration, executing a method for extracting multimodal document information configuration as claimed in any one of claims 1 to 6, characterized in that: include: Preprocessing module: used to perform preliminary analysis on input documents, identify file types and formats, and preprocess documents. The preprocessing includes: format recognition, conversion, resolution detection and optimization adjustment; Configuration recognition mode module: used to select the most suitable recognition mode according to the document type and content distribution, and configure the corresponding recognition parameters; wherein, the recognition modes include: direct extraction of PDF text, OCR recognition or table recognition; the recognition process adopts a cache mechanism, with the file hash value as the key and the recognition result as the value, to avoid long waiting time for multiple configurations of recognition; the selection of the most suitable recognition mode and configuration of the corresponding recognition parameters include: if the recognition mode of direct extraction of PDF text is selected, the text content is directly extracted from the PDF document; if the recognition mode of OCR recognition is selected, the original structure and size are maintained; if the recognition mode of table recognition is selected, the table structure is accurately identified through advanced layout analysis, including row and column division and cell merging processing, and the table data is converted into structured data; Recognition and extraction rule module: It is used for the recognition results of different types of recognition modes, adopts customized extraction strategies and recognition and extraction rules, and sets extraction elements by category: for the recognition results of direct extraction of PDF text, it is extracted by building regular expression matching; for the recognition results of OCR recognition, it combines the location information and accurately obtains the text according to the coordinate range; for the recognition results of table recognition, it completes information extraction by selecting cells; Accurate information extraction module: used to determine the document page position corresponding to each recognition and extraction rule, and accurately extract the required information; respond to the extraction results in JSON format or call back through a custom callback address; Saving and extraction configuration module: used to record and save the recognition and extraction rules and recognition and extraction configurations, support the use of completed recognition and extraction configurations through multiple interfaces, integrate multiple file upload methods, flexible configuration management and efficient caching mechanism, and realize intelligent extraction and processing of document information.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method for extracting multimodal document information configuration based on any one of claims 1 to 6 are implemented.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method for extracting multimodal document information configuration based on any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method for integrating and extracting multi-mode contents of multi-type documents
CN116627912A
PDF document data processing and information extraction device and method
CN117095419A