Document content matching method and system based on multiple modes
By performing picture direction correction, OCR text recognition and text direction correction for multimodal documents, combined with multimodal document information extraction and modal complementarity enhancement processing, the problem of difficulty in extracting multimodal document information in the prior art is solved, and higher information extraction accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510654917.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing document content matching methods have limitations when dealing with multimodal documents, which are difficult to adapt to different document styles, fonts, typesettings and structures, resulting in limited generalization capabilities, especially when dealing with complex nested lists, multi-column layouts, irregular tables and charts, which increases the difficulty of information extraction.
A multimodal document content matching method is adopted, and the images to be identified are acquired and pre-processed, including image direction correction, OCR text recognition and text direction correction, followed by multimodal document information extraction and modal complementary enhancement processing, and finally information matching is performed to return the matching result.
It significantly improves the accuracy and practicality of information extraction, enhances the understanding of multimodal documents, improves generalization ability and adaptability, and can effectively match and extract information in different application scenarios.
Smart Images

Figure CN120182990A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of character recognition, and particularly relates to a multimodal-based document content matching method and system. Background Art
[0002] Document content matching refers to automatically extracting key data from unstructured documents and converting it into a structured format for output. Its core processes include: document preprocessing, i.e., document sample normalization, OCR, and image optimization; layout analysis, i.e., identifying text regions, tables, images, and lists; information recognition, i.e., understanding the document content through entity recognition, relationship extraction, table parsing, and key-value pair extraction. Deep learning methods such as CNNs (Convolutional Neural Networks), RNNs (Recurrent Neural Networks), Transformers (deep learning models), and Graph Neural Networks are used to improve the extraction accuracy. Document layout analysis is the first step in document understanding and extraction. It involves identifying the various components in a document, such as text blocks, headings, lists, images, and tables, and determining their positional relationships on the page. This process is crucial for subsequent information extraction because different layout elements often carry different types of information, and their positional relationships can provide additional contextual information. Table parsing focuses on extracting the table structure and the data therein from the document. Tables are widely used in detailed lists and statistical data due to their highly structured nature and can quickly convey a large amount of information. Information matching refers to finding and pairing keywords (Keys) and corresponding values (Values) in a document. This method is widely used in information extraction, such as extracting specific terms or data from contracts, invoices, or application forms. For multimodal documents, existing document content matching methods have certain limitations. First, due to the huge differences in document styles, fonts, layouts, and structures, it is difficult for a single model to adapt to all situations, resulting in limited generalization ability. Second, complex nested lists, multi-column layouts, irregular tables, and charts, especially cross-page elements, increase the difficulty of information extraction, and the diversity and complexity of tables also pose challenges to automatic parsing. Although the deep learning models adopted are good at pattern recognition, they still have deficiencies in understanding the deep semantics and context of documents, especially when dealing with multiple languages and technical terms. Therefore, there is an urgent need to provide a matching method suitable for multimodal document content to adapt to different application scenarios. Summary of the Invention
[0003] To solve the above problems, the present invention provides a multimodal-based document content matching method and system to solve the problem that existing document content matching methods have certain limitations for multimodal documents.
[0004] A multimodal-based document content matching method includes: Obtain the image to be recognized and preprocess the image to be recognized; Perform picture orientation correction on the preprocessed image to be recognized to obtain a corrected image; Perform OCR character recognition on the corrected image and mark the text boxes; Perform text orientation correction on the corrected image based on the text boxes to obtain corrected text; Extract multi-modal document information based on the corrected text, and perform modal complementary enhancement processing on the extracted multi-modal document information; Match the processed multi-modal document information and return the matching result.
[0005] According to a specific embodiment of the present invention, obtaining the image to be recognized and preprocessing the image to be recognized includes: Obtain the image to be recognized and analyze the resolution of the image to be recognized, determine whether the resolution of the image to be recognized is less than the preset resolution, and if it is less than the preset resolution, adjust the image to be recognized to the preset resolution.
[0006] According to a specific embodiment of the present invention, the preset resolution is 1920.
[0007] According to a specific embodiment of the present invention, performing picture orientation correction on the preprocessed image to be recognized to obtain a corrected image includes: Use a global key point direction detection model to detect the global key points of the table in the image to be recognized, and the global key points include the four corner points of each cell; Based on the global key points, determine whether there is distortion and deformation between different text lines of the document. If there is distortion and deformation, correct the text lines of the image to be recognized to obtain a corrected image.
[0008] According to a specific embodiment of the present invention, performing OCR character recognition on the corrected image and marking the text boxes includes: Perform binarization and smoothing denoising processing on the corrected image to obtain an enhanced processing image of the corrected image; Use an OCR character detection model to perform OCR character recognition on the enhanced processing image, and mark text boxes at the recognized characters.
[0009] According to a specific embodiment of the present invention, performing text orientation correction on the corrected image based on the text boxes to obtain corrected text includes: Calculate the inclination angle of the text box, and determine whether the text box is shifted according to the inclination angle. If it is shifted, correct the text box to obtain corrected text.
[0010] According to a specific embodiment of the present invention, multi-modal document information extraction is performed based on the corrected text, and modal complementary enhancement processing is performed on the extracted multi-modal document information, including: Extract multi-modal document information and corresponding text boxes from the corrected text; Detect the blank spaces between the text boxes, and mark blank boxes at the positions of the blank spaces for filling; Identify line break document information based on the multi-modal document information and combine and splice it with the relevant text; Filter out misaligned text according to the slope relationship between the text boxes.
[0011] According to a specific embodiment of the present invention, information matching is performed on the processed multi-modal document information and the matching result is returned, including: Extract keywords from the processed multi-modal document information, and determine the corresponding values according to the context of the paragraphs where the keywords are located; Establish a correspondence between the keywords and the values, and form key-value pairs based on the correspondence; Perform information matching according to the key-value pairs and return the matching result.
[0012] According to a specific embodiment of the present invention, after the preprocessed image to be recognized is corrected in picture direction, it further includes: Use the methods of guided learning and contrast learning to identify and classify different seals in the corrected image.
[0013] A multi-modal based document content matching system, including: An image acquisition and processing module, configured to acquire an image to be recognized and perform preprocessing on the image to be recognized; An image correction module, configured to correct the picture direction of the preprocessed image to be recognized to obtain a corrected image; A character recognition module, configured to perform OCR character recognition on the corrected image and mark text boxes; A text correction module, configured to correct the text direction of the corrected image based on the text boxes to obtain a corrected text; An information extraction and processing module, configured to perform multi-modal document information extraction based on the corrected text, and perform modal complementary enhancement processing on the extracted multi-modal document information; An information matching module, configured to perform information matching on the processed multi-modal document information and return the matching result.
[0014] According to a specific embodiment of the present invention, the image acquisition and processing module further includes: An image acquisition module, configured to acquire an image to be recognized; An image processing module for analyzing the resolution of the image to be recognized, determining whether the resolution of the image to be recognized is less than a preset resolution, and if it is less than the preset resolution, adjusting the image to be recognized to the preset resolution, where the preset resolution is 1920.
[0015] According to a specific embodiment of the present invention, the image correction module further includes: A key point detection module for detecting the global key points of the table in the image to be recognized, where the global key points include the four corner points of each cell; A judgment and correction module for judging whether there is distortion and deformation between different text lines of the document based on the global key points, and if there is distortion and deformation, correcting the text lines of the image to be recognized to obtain a corrected image.
[0016] According to a specific embodiment of the present invention, the text recognition module further includes: An enhancement processing module for performing binarization and smoothing denoising processing on the corrected image to obtain an enhanced processing image of the corrected image; A recognition module for performing OCR text recognition on the enhanced processing image and marking a text box at the recognized text.
[0017] According to a specific embodiment of the present invention, the text correction module further includes: An angle calculation module for calculating the inclination angle of the text box; A correction module for judging whether the text box is shifted according to the inclination angle, and if it is shifted, correcting the text box to obtain corrected text.
[0018] According to a specific embodiment of the present invention, the information extraction and processing module further includes: An information extraction module for extracting multi-modal document information and the corresponding text boxes from the corrected text; A filling module for detecting the blank spaces between the text boxes and marking and filling blank boxes at the positions of the blank spaces; A splicing module for identifying line break document information based on the multi-modal document information and combining and splicing it with relevant texts; A filtering module for filtering misaligned texts according to the slope relationship between the text boxes.
[0019] According to a specific embodiment of the present invention, the information matching module further includes: A keyword extraction module for extracting keywords from the processed multi-modal document information and determining the corresponding values according to the context of the paragraphs where the keywords are located; A key-value pair generation module for establishing a correspondence between the keywords and the values and forming key-value pairs based on the correspondence; A matching module for matching information according to key-value pairs and returning a matching result.
[0020] According to a specific embodiment of the present invention, the system further includes: A seal recognition and classification module for recognizing and classifying different seals in the corrected image by using the methods of guided learning and contrast learning.
[0021] An electronic device includes: a processor and a memory, and a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned multimodal-based document content matching method.
[0022] A computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the above-mentioned multimodal-based document content matching method.
[0023] Compared with the prior art, a multimodal-based document content matching method and system provided by the present invention have the following advantages: 1. Deep feature fusion: Compared with traditional information extraction models that can usually only process single-modal data, the present invention can effectively fuse the features of multiple data sources such as multimodal text content, text location, and images through a deep learning network, forming a richer representation, so as to obtain higher accuracy in information extraction tasks.
[0024] 2. Cross-modal association mining: The present invention utilizes the internal connection between different modalities to identify and eliminate noise during information extraction, improving the reliability of the recognition result.
[0025] 3. Enhanced modal complementarity: When the data of a certain modality is missing or of poor quality, the present invention can obtain additional information from other modalities to make up for this deficiency. If the image information is unclear, the model can use the image title or description text to assist in understanding the image content, thereby improving the accuracy of retrieval.
[0026] 4. Improved context awareness ability: The present invention can utilize context information to improve the effect of information extraction, better understand the context, and extract more accurate information fragments.
[0027] 5. Enhanced adaptability and generalization ability: Since multimodal models are trained on diverse datasets, they often have stronger generalization ability and can perform well on unseen data. They can not only handle information extraction within a specific domain but also flexibly adapt to cross-domain or cross-modal tasks, showing broader practicality.
[0028] 6. Real-time Performance and Efficiency Improvement: With the development of algorithm optimization and hardware acceleration technologies, while maintaining high accuracy, multi-modal models have gradually increased their processing speed, enabling real-time or near-real-time information extraction to meet the needs of real-time analysis and rapid response. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0030] Figure 1 It is a flowchart of a multi-modal based document content matching method provided according to an embodiment of the present invention.
[0031] Figure 2 It is a flowchart of a method for obtaining an image to be recognized and preprocessing the image to be recognized provided according to an embodiment of the present invention.
[0032] Figure 3 It is a flowchart of a method for correcting the orientation of a picture provided according to an embodiment of the present invention.
[0033] Figure 4 It is a flowchart of a character recognition method provided according to an embodiment of the present invention.
[0034] Figure 5 It is a flowchart of a method for correcting the text orientation provided according to an embodiment of the present invention.
[0035] Figure 6 It is a flowchart of a method for multi-modal document information extraction and modal complementary enhancement provided according to an embodiment of the present invention.
[0036] Figure 7 It is a flowchart of a method for information matching provided according to an embodiment of the present invention.
[0037] Figure 8 It is a schematic diagram of picture orientation correction provided according to an embodiment of the present invention.
[0038] Figure 9 It is a schematic diagram of text orientation correction provided according to an embodiment of the present invention.
[0039] Figure 10 It is a flowchart of ocr detection and recognition provided according to an embodiment of the present invention.
[0040] Figure 11 It is a schematic diagram of text box correction provided according to an embodiment of the present invention.
[0041] Figure 12 It is a schematic diagram of blank space recognition and filling according to an embodiment of the present invention.
[0042] Figure 13 It is a schematic diagram of cross-row information merging and matching provided according to an embodiment of the present invention.
[0043] Figure 14 It is a flow chart of an overall method for multimodal document content matching provided according to an embodiment of the present invention.
[0044] Figure 15 It is a structural diagram of a document content matching system based on multimodality provided according to an embodiment of the present invention.
[0045] Figure 16 is a structural diagram of an image acquisition and processing module provided according to an embodiment of the present invention.
[0046] Figure 17 is a structural diagram of an image correction module provided according to an embodiment of the present invention.
[0047] Figure 18 is a structural diagram of a text recognition module provided according to an embodiment of the present invention.
[0048] Figure 19 is a structural diagram of a text correction module provided according to an embodiment of the present invention.
[0049] Figure 20 is a structural diagram of an information extraction and processing module provided according to an embodiment of the present invention.
[0050] Figure 21 is a structural diagram of an information matching module provided according to an embodiment of the present invention.
[0051] Figure 22 It is a schematic diagram of the structure of a computer device provided according to an embodiment of the present invention.
[0052] Reference numerals: 01- Image acquisition and processing module; 02- Image correction module; 03- Text recognition module; 04- Text correction module; 05- Information extraction and processing module; 06- Information matching module; 07- Seal recognition and classification module; 011-image acquisition module; 012-image processing module; 021-key point detection module; 022-judgment and correction module; 031-enhanced processing module; 032-identification module; 041-angle calculation module; 042-correction module; 051 - Information extraction module; 052 - Filling module; 053 - Stitching module; 054 - Filtering module; 061 - Keyword extraction module; 062 - Key - value pair generation module; 063 - Matching module. Detailed implementation manners
[0053] To enable those skilled in the art to more clearly understand the concepts and ideas of the present invention, the present invention will be described in detail below in conjunction with specific embodiments. It should be understood that the embodiments given herein are only a part of all possible embodiments of the present invention. After reading the specification of this application, those skilled in the art are capable of making improvements, modifications, or substitutions to part or all of the following embodiments, and these improvements, modifications, or substitutions are also included within the scope of protection required by the present invention.
[0054] In this article, terms such as "advance notice", "entry into the station", and other similar words are not intended to imply any order, quantity, or importance, but are only used to distinguish different elements. In this article, terms such as "a", "an", and other similar words are not intended to mean that there is only one thing, but rather that the relevant description only refers to one of the things, and the thing may have one or more. In this article, terms such as "comprise", "include", and other similar words are intended to represent a logical relationship and should not be regarded as representing a spatial structure relationship. For example, "A includes B" is intended to mean that logically B belongs to A, rather than meaning that B is located inside A in terms of space. Additionally, the meanings of terms such as "comprise", "include", and other similar words should be regarded as open - ended rather than closed - ended. For example, "A includes B" is intended to mean that B belongs to A, but B does not necessarily constitute all of A, and A may also include other elements such as C, D, E, etc.
[0055] In this article, terms such as "embodiment", "the present embodiment", "one embodiment", "a single embodiment" do not mean that the relevant description only applies to a specific embodiment, but rather that these descriptions may also apply to one or more other embodiments. Those skilled in the art should understand that in this article, any description made for a certain embodiment can be substituted, combined, or otherwise combined with the relevant descriptions in one or more other embodiments, and the new embodiments generated by substitution, combination, or other means of combination are easily conceivable by those skilled in the art and fall within the scope of protection of the present invention.
[0056] Embodiment 1 Additional aspects and advantages of the embodiments of the present invention will be given in part in the following description, become apparent in part from the following description, or be understood through the practice of the embodiments of the present invention. In combination with Figures 1 - 14 , the embodiments of the present invention provide a multimodal - based document content matching method, including: S1: Obtain the image to be recognized and preprocess the image to be recognized.
[0057] S2: Correct the orientation of the preprocessed image to be recognized to obtain a corrected image.
[0058] S3: Perform OCR character recognition on the corrected image and mark the text boxes.
[0059] S4: Correct the text orientation of the corrected image based on the text boxes to obtain corrected text.
[0060] S5: Extract multi-modal document information based on the corrected text, and perform modal complementary enhancement processing on the extracted multi-modal document information.
[0061] S6: Match the processed multi-modal document information and return the matching result.
[0062] Traditional information extraction can usually only process single-modal data. For multi-modal documents, there are huge differences in document styles, fonts, layouts, and structures. Therefore, traditional information extraction methods for single models are difficult to adapt to all situations, resulting in limited generalization ability. Multi-modal documents have specific complex nested lists, multi-column layouts, irregular tables, and charts, especially cross-page elements, which increase the difficulty of information extraction. Existing deep learning models still have deficiencies in understanding the deep semantics and context of documents, especially when dealing with multiple languages and technical terms. For information extraction in documents, not only semantics but also document layout need to be considered to adapt to the content of different scenarios. Based on the above problems, the multi-modal-based document content matching method proposed in the present invention not only depends on the detection of cells in table content parsing, but also considers the influence of document content and row-column relationships, improving the generalization ability of table parsing. Through improvements in aspects such as deep feature fusion, cross-modal association, and modal complementary enhancement of various data source features such as multi-modal text content and text position, the accuracy and practicality of information extraction are significantly improved.
[0063] Specifically, step S1 of obtaining the image to be recognized and preprocessing the image to be recognized includes: S11: Obtain the image to be recognized and analyze the resolution of the image to be recognized.
[0064] S12: Determine whether the resolution of the image to be recognized is less than the preset resolution. If it is less than the preset resolution, adjust the image to be recognized to the preset resolution.
[0065] In the embodiments of the present invention, the image to be recognized is a chart - type document image. Since there are a large number of cells in chart - type documents and the number of rows in documents of the same type of table varies greatly, to ensure the accuracy of the ocr detection results, the embodiments of the present invention adjust the sizes of samples with different resolutions to ensure within a reasonable resolution range. In a specific embodiment of the present invention, the preset resolution is 1920. If it is determined that the resolution of the image to be recognized is less than 1920, then the resolution of the image to be recognized is adjusted to 1920.
[0066] Specifically, step S2 corrects the orientation of the pre - processed image to be recognized, and the obtained corrected image includes: S21: Use a global key - point direction detection model to detect the global key points of the table in the image to be recognized. The global key points include the four corner points of each cell.
[0067] S22: Based on the global key points, determine whether there is distortion and deformation between different text lines of the document. If there is distortion and deformation, then correct the text lines of the image to be recognized to obtain a corrected image.
[0068] When dealing with complex tables, merging cells, spanning rows, and spanning columns are one of the most common problems. Merging cells means that some cells occupy the space of multiple rows or columns, which is difficult to accurately parse in traditional table recognition algorithms. By detecting the key points of cells, that is, the four corner points of cells, the present invention can accurately identify the merged cells and the space spanning rows and columns.
[0069] The working principle of cell key - point detection is to use a deep - learning model, such as a region - based convolutional neural network, to locate the boundary of each cell in the table. These models are trained to recognize the shape and position of cells. Even for large merged cells that span multiple cells, by identifying the cell boundaries, it is possible to accurately distinguish independent cells and merged cells, thus solving the problem of complex table layouts.
[0070] The embodiments of the present invention use the GPDD (Global Document Key Point Direction Detection) model to detect the global key points of the table in the image to be recognized, such as Figure 8As shown, the global key points include the four corner points of each cell. This model can not only detect the global key points of the document, but also judge whether there is distortion and deformation between different text lines according to the connection between the key points. If there is distortion and deformation, the document will be corrected at the line level. The global key point direction detection model (GPDD) used in the present invention can restore the document with distorted text lines to a positive standard document, effectively solve problems such as loss of layout information due to row and column distortion and content misalignment of the text, and provide effective layout information for later multi-modal information matching. By correcting the direction of the picture, the text lines can be adjusted to the horizontal or vertical direction to reduce the impact of picture distortion on the accuracy of text recognition, and thus improve the clarity and recognizability of the text.
[0071] Specifically, step S3 for performing OCR text recognition on the corrected image and marking text boxes includes: S31: Perform binarization and smoothing denoising processing on the corrected image to obtain an enhanced processed image of the corrected image.
[0072] S32: Use an OCR text detection model to perform OCR text recognition on the enhanced processed image, and mark text boxes at the recognized text.
[0073] Inside the cell, the typesetting of the text may be very complex, including multi-line text, different font sizes, italics, bold, and text alignment. Traditional table recognition methods may deviate in these situations, resulting in misreading or loss of text content. The present invention can effectively solve these problems by using text boxes. A text box is a rectangular boundary that encloses a text area and can determine the exact position of the text. These text boxes are generated by specialized text detection models that can identify the presence of text and locate it at specific positions within the cell. Through precise text positioning, the present invention can ensure that all text content is correctly read.
[0074] In a specific embodiment of the present invention, by performing binarization and smoothing denoising processing on the corrected image, the detection and recognition ability of the document is enhanced and fine-tuned, thereby improving the clarity and recognizability of the text. In a specific embodiment of the present invention, the OCR detection and recognition model open-sourced by Baidu is used to perform OCR text recognition on table-type documents, and text boxes are marked at the recognized text.
[0075] By combining the text box with the key points of the cell, the present invention can accurately determine the position of the text within the cell. Using the combination of cell key point detection and text detection box can more comprehensively understand the structure and content of the table. When dealing with merged cells, cell key point detection helps to identify the boundaries of the cells, while the text box ensures that the text content is accurately positioned and read. This not only improves the accuracy of table recognition but also enhances the robustness of the system, enabling it to maintain high performance in various complex table layouts.
[0076] Specifically, step S4 performs text direction correction on the corrected image based on the text box, and the obtained corrected text includes: S41: Calculate the tilt angle of the text box.
[0077] S42: Determine whether the text box has shifted according to the tilt angle. If it has shifted, correct the text box to obtain the corrected text.
[0078] In the embodiment of the present invention, it is judged whether there is a shift by calculating the tilt angle of the text box. As Figure 9 shown, the left side is the text box before correction. By calculating its tilt angle, it is judged that there is a shift. To optimize the position of the text box, the globally detected result of the text box is used to correct the shifted text box to obtain the corrected text box. By correcting the direction of the text box, the text direction is further corrected to improve the clarity and recognizability of the text. Through the correction of the chart cells, the present invention can effectively alleviate the situation where the overall layout of the cell cannot be truly restored when only using the text box coordinates to obtain the table cell information. The present invention maps and restores the real table layout information to the positive document according to the local and global features of the cell position, greatly improving the recognition rate of the model for small samples with missed detection and misdetection in information extraction.
[0079] Specifically, step S5 performs multi-modal document information extraction based on the corrected text and performs modal complementary enhancement processing on the extracted multi-modal document information, including: S51: Extract multi-modal document information and the corresponding text box from the corrected text.
[0080] S52: Detect the blank spaces between the text boxes and mark blank boxes at the positions of the blank spaces for filling.
[0081] S53: Identify the line break document information based on the multi-modal document information and combine and splice it with the relevant text.
[0082] S54: Filter the misaligned text according to the slope relationship between the text boxes.
[0083] Traditional information extraction methods are usually divided into three steps. After inputting the picture samples, OCR detection and recognition, document layout structure detection, and visual coding are performed respectively. However, traditional information extraction methods only consider the text semantic features and ignore the position features of the text boxes. Due to resolution problems, some samples have text undetected. To ensure the integrity of text content extraction, the present invention fills the undetected text blank lines. First, text information and corresponding text boxes are extracted from the corrected document, and whether there are blank spaces between the text boxes is detected. If there are, blank boxes are marked at the positions of the blank spaces for filling. Then, line break texts are identified from the document and combined and spliced with their related texts. As Figure 13 shown, texts with multiple lines are grouped into one bounding box. For line break texts, due to document layout problems, the relationship between adjacent lines is damaged. To ensure the integrity of the extracted content, the present invention labels and identifies line break text boxes and splices the corresponding line break texts of the texts. For example, the key value value1 corresponding to the keyword key1 and the text line1 are texts with multiple lines, so the key value value1 and the text line1 are grouped into one bounding box. Due to text skew and distortion, the same-line relationship is damaged. To ensure that the same-line text relationship can be effectively recognized, the present invention performs secondary filtering and re-matching on the misaligned texts according to the slope relationship between the text boxes, thereby improving the accuracy of text recognition. The present invention can effectively process the cross-line information matching problem by performing modality complementary enhancement processing on the extracted multi-modal document information. Since the multi-modal sample annotation cost is relatively high, only using a small number of samples to fine-tune the model cannot accurately handle the business requirements of specific scenarios. Therefore, the present invention can achieve cross-line matching of multi-scenario texts through modality complementary enhancement processing.
[0084] Specifically, step S6 performs information matching on the processed multi-modal document information and returns the matching result, including: S61: Extract keywords from the processed multi-modal document information and determine the corresponding values according to the context of the paragraphs where the keywords are located.
[0085] S62: Establish the corresponding relationship between the keywords and the values, and form key-value pairs based on the corresponding relationship.
[0086] S63: Perform information matching according to the key-value pairs and return the matching result.
[0087] Information matching refers to finding and pairing keywords (Keys) and corresponding values (Values) in a document. In the embodiments of the present invention, a multimodal document utilizes context information to improve the information matching effect. By understanding the context of the paragraph where the keyword is located, the correct range of the value is determined. According to the context and semantics of the keyword, the corresponding value is extracted from the document, and then the association between the keyword and the value is established to form a key-value pair data structure. Finally, information matching is performed based on the key-value pair and the matching result is returned. The present invention extracts keywords from the multimodal document and combines context semantic understanding, improving the context awareness ability and significantly enhancing the accuracy of information matching.
[0088] Specifically, after step S2 corrects the orientation of the preprocessed image to be recognized, the following steps are further included: Use guided learning and contrastive learning methods to recognize and classify different seals in the corrected image.
[0089] Since the differences between the seals of different operators are small and the number of samples is limited, in order to enable the model to effectively learn the differences between different seals, the present invention uses guided learning and contrastive learning methods to recognize and classify different seals in the document, and can identify the subtle differences between different categories of samples.
[0090] Embodiment 2 Based on the above method, the embodiments of the present invention further provide a multimodal-based document content matching system, as Figures 15 - 21 shown, including: An image acquisition and processing module 01, configured to obtain an image to be recognized and preprocess the image to be recognized.
[0091] An image correction module 02, configured to correct the orientation of the preprocessed image to be recognized to obtain a corrected image.
[0092] A text recognition module 03, configured to perform OCR text recognition on the corrected image and mark text boxes.
[0093] A text correction module 04, configured to correct the text orientation of the corrected image based on the text boxes to obtain corrected text.
[0094] An information extraction and processing module 05, configured to extract multimodal document information based on the corrected text and perform modal complementary enhancement processing on the extracted multimodal document information.
[0095] An information matching module 06, configured to perform information matching on the processed multimodal document information and return a matching result.
[0096] Compared with traditional information extraction which can usually only process single-modal data, the present invention improves in aspects such as deep feature fusion, cross-modal association, and enhancement of modal complementarity for various data source features of multi-modal text content, text location, etc., significantly improving the accuracy and practicality of information extraction.
[0097] Specifically, the image acquisition and processing module 01 further includes: The image acquisition module 011 is used to acquire the image to be recognized.
[0098] The image processing module 012 is used to analyze the resolution of the image to be recognized, determine whether the resolution of the image to be recognized is less than the preset resolution. If it is less than the preset resolution, the image to be recognized is adjusted to the preset resolution, and the preset resolution is 1920.
[0099] In the embodiment of the present invention, the image to be recognized acquired by the image acquisition module 011 is a chart-type document image. Since there are a large number of cells in the chart-type document and the number of rows in documents of the same type of table varies greatly, to ensure the accuracy rate of ocr detection results, the embodiment of the present invention adjusts the sizes of samples with different resolutions through the image processing module 012 to ensure within a reasonable resolution range. In a specific embodiment of the present invention, the preset resolution is 1920. If it is determined that the resolution of the image to be recognized is less than 1920, the resolution of the image to be recognized is adjusted to 1920.
[0100] Specifically, the image correction module 02 further includes: The key point detection module 021 is used to detect the global key points of the table in the image to be recognized. The global key points include the four corner points of each cell.
[0101] The judgment and correction module 022 is used to judge whether there is distortion and deformation between different text lines of the document based on the global key points. If there is distortion and deformation, the text lines of the image to be recognized are corrected to obtain a corrected image.
[0102] The embodiment of the present invention uses the key point detection module 021 to detect the global key points of the table in the image to be recognized. As Figure 8 shown, the global key points include the four corner points of each cell. This model can not only detect the global key points of the document, but also judge whether there is distortion and deformation between different text lines according to the connection between the key points. If there is distortion and deformation, the document is corrected at the line level. Through the judgment and correction module 022 to correct the picture direction, the text lines can be adjusted to the horizontal or vertical direction to reduce the impact of picture distortion on the character recognition accuracy rate, thereby improving the clarity and recognizability of the text.
[0103] Specifically, the text recognition module 03 further includes: An enhancement processing module 031 for performing binarization and smoothing denoising processing on the corrected image to obtain an enhanced processed image of the corrected image.
[0104] An identification module 032 for performing OCR character recognition on the enhanced processed image and marking text boxes at the recognized characters.
[0105] In an embodiment of the present invention, the enhancement processing module 031 performs binarization and smoothing denoising processing on the corrected image to perform enhanced fine-tuning on the document detection and recognition capabilities, thereby improving the clarity and recognizability of the characters. The recognition module 032 performs OCR character recognition on table-type documents and marks text boxes at the recognized characters.
[0106] Specifically, the text correction module 04 further includes: An angle calculation module 041 for calculating the tilt angle of the text box.
[0107] A correction module 042 for determining whether the text box has shifted according to the tilt angle. If a shift occurs, the text box is corrected to obtain corrected text.
[0108] In an embodiment of the present invention, the angle calculation module 041 calculates the tilt angle of the text box for determining whether a shift occurs, and the correction module 042 corrects the shifted text box to obtain corrected text. By correcting the direction of the text box, the correction of the text direction is realized, so as to improve the clarity and recognizability of the characters.
[0109] Specifically, the information extraction and processing module 05 further includes: An information extraction module 051 for extracting multi-modal document information and corresponding text boxes from the corrected text.
[0110] A filling module 052 for detecting the blank spaces between the text boxes and marking blank boxes at the positions of the blank spaces for filling.
[0111] A splicing module 053 for identifying line break document information based on the multi-modal document information and combining and splicing it with the relevant text.
[0112] A filtering module 054 for filtering out misaligned text according to the slope relationship between the text boxes.
[0113] First, the present invention extracts text information and corresponding text boxes from the corrected document through the information extraction module 051. The filling module 052 detects whether there are blank spaces between the text boxes. If so, blank boxes are marked at the positions of the blank spaces for filling. Then, the splicing module 053 identifies line-breaking text from the document and combines and splices it with the relevant text. Finally, the filtering module 054 performs secondary filtering on the misaligned text and rematches it, thereby improving the accuracy of text recognition.
[0114] Specifically, the information matching module 06 further includes: A keyword extraction module 061, which is used to extract keywords from the processed multi-modal document information and determine the corresponding values according to the context of the paragraphs where the keywords are located.
[0115] A key-value pair generation module 062, which is used to establish the correspondence between keywords and values and form key-value pairs based on the correspondence.
[0116] A matching module 063, which is used to perform information matching according to the key-value pairs and return the matching results.
[0117] First, the present invention extracts keywords from the multi-modal document information through the keyword extraction module 061, determines the correct range of values by understanding the context of the paragraphs where the keywords are located, extracts the corresponding values from the document according to the context and semantics of the keywords, then establishes the association between keywords and values through the key-value pair generation module 062 to form a key-value pair data structure, and finally performs information matching through the matching module 063 and returns the matching results. By extracting keywords from the multi-modal document and combining context semantic understanding, the present invention improves the context awareness ability and significantly improves the accuracy of information matching.
[0118] Specifically, the system further includes: A seal recognition and classification module 07, which is used to recognize and classify different seals in the corrected image by using the methods of guided learning and contrast learning.
[0119] Since the differences between the seals of different operators are small and the number of samples is limited, in order to enable the model to effectively learn the differences between different seals, the present invention uses the seal recognition and classification module 07 to recognize and classify different seals in the document, and can recognize the subtle differences between different categories of samples.
[0120] Embodiment 3 As Figure 22 shown, the embodiment of the present invention further provides an electronic device, including: a processor and a memory. A computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the above-mentioned multi-modal-based document content matching method. The device in the present invention can be a server, a PC, a PAD, a mobile phone, etc.
[0121] Furthermore, an embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the above-mentioned multi-modal based document content matching method.
[0122] In summary, a multi-modal based document content matching method and system provided by the present invention have the following advantages: 1. Deep feature fusion: Compared with traditional information extraction models that can usually only process data of a single modality, the present invention can effectively fuse features of multiple data sources such as multi-modal text content, text location, images, etc. through a deep learning network to form a richer representation, thereby obtaining higher accuracy in information extraction tasks.
[0123] 2. Cross-modal correlation mining: The present invention utilizes the internal connection between different modalities to identify and eliminate noise during the information extraction process, improving the reliability of the recognition results.
[0124] 3. Enhanced modal complementarity: When data of a certain modality is missing or of poor quality, the present invention can obtain additional information from other modalities to make up for this deficiency. If the image information is unclear, the model can use the image title or description text to assist in understanding the image content, thereby improving the accuracy of retrieval.
[0125] 4. Improved context awareness: The present invention can utilize context information to improve the effect of information extraction, better understand the context, and extract more accurate information fragments.
[0126] 5. Enhanced adaptability and generalization ability: Since multi-modal models are trained on diverse datasets, they often have stronger generalization ability and can perform well on unseen data. They can not only handle information extraction within a specific domain but also flexibly adapt to cross-domain or cross-modal tasks, demonstrating broader practicality.
[0127] 6. Improved real-time performance and efficiency: With the development of algorithm optimization and hardware acceleration technologies, multi-modal models have gradually increased their processing speed while maintaining high accuracy, enabling real-time or near-real-time information extraction to meet the requirements of real-time analysis and rapid response.
[0128] The concepts, principles, and ideas of the present invention have been described in detail above in connection with specific embodiments (including examples and instances). Those skilled in the art should understand that the embodiments of the present invention are not limited to the forms given above. After reading this application document, those skilled in the art may make any possible improvements, substitutions, and equivalent forms to the steps, methods, systems, and components in the above embodiments. These improvements, substitutions, and equivalent forms should be regarded as falling within the scope of the present invention. The protection scope of the present invention is only subject to the claims.
Claims
1. A document content matching method based on multimodality, characterized in that: include: Acquire an image to be identified and preprocess the image to be identified; Performing image direction correction on the preprocessed image to be recognized to obtain a corrected image; Performing OCR text recognition on the corrected image and marking a text box; Performing text direction correction on the corrected image based on the text frame to obtain a corrected text; Extracting multimodal document information based on the corrected text, and performing modality complementary enhancement processing on the extracted multimodal document information; The processed multimodal document information is matched and the matching results are returned.
2. The multimodal document content matching method according to claim 1, characterized in that: The acquiring the image to be identified and preprocessing the image to be identified comprises: An image to be identified is acquired and a resolution of the image to be identified is analyzed to determine whether the resolution of the image to be identified is less than a preset resolution. If the resolution is less than the preset resolution, the image to be identified is adjusted to the preset resolution.
3. The multimodal document content matching method according to claim 2, characterized in that: The preset resolution is 1920.
4. The multimodal document content matching method according to claim 1, characterized in that: The performing image direction correction on the preprocessed image to be recognized to obtain the corrected image comprises: Using a global key point direction detection model to detect global key points of the table in the image to be identified, the global key points include four corner points of each cell; Based on the global key points, it is determined whether distortion and deformation occur between different text lines of the document. If distortion and deformation occur, the text lines of the image to be recognized are corrected to obtain a corrected image.
5. The multimodal document content matching method according to claim 1, characterized in that: The performing OCR text recognition on the corrected image and marking the text box comprises: Performing binarization and smoothing denoising processing on the corrected image to obtain an enhanced processed image of the corrected image; An OCR text detection model is used to perform OCR text recognition on the enhanced processed image, and a text box is marked at the recognized text.
6. The document content matching method based on multimodality according to claim 1, characterized in that: The performing text direction correction on the corrected image based on the text frame to obtain the corrected text comprises: The tilt angle of the text box is calculated, and it is determined whether the text box sends an offset according to the tilt angle. If an offset occurs, the text box is corrected to obtain a corrected text.
7. The document content matching method based on multimodality according to claim 1, characterized in that: The extracting of multimodal document information based on the corrected text and performing modality complementary enhancement processing on the extracted multimodal document information includes: Extracting multimodal document information and corresponding text boxes from the corrected text; Detecting blank spaces between the text boxes and marking blank boxes at the positions of the blank spaces for filling; Identifying line break document information based on the multimodal document information and combining and splicing it with related text; The staggered texts are filtered according to the slope relationship between the text boxes.
8. The document content matching method based on multimodality according to claim 1, characterized in that: The step of matching the processed multimodal document information and returning the matching result includes: Extracting keywords from the processed multimodal document information, and determining corresponding values according to the paragraph context where the keywords are located; Establishing a correspondence between the keyword and the value, and forming a key-value pair based on the correspondence; Information matching is performed according to the key-value pairs and matching results are returned.
9. The document content matching method based on multimodality according to claim 1, characterized in that: After the pre-processed image to be recognized is corrected in picture direction, the method further includes: Guided learning and contrastive learning methods are used to identify and classify different seals in the rectified images.
10. A document content matching system based on multimodality, characterized in that: include: An image acquisition and processing module, used for acquiring an image to be identified and preprocessing the image to be identified; An image correction module is used to correct the image direction of the pre-processed image to be recognized to obtain a corrected image; A text recognition module, used for performing OCR text recognition on the corrected image and marking a text box; A text correction module, used for performing text direction correction on the correction image based on the text frame to obtain a corrected text; An information extraction and processing module, used to extract multimodal document information based on the corrected text, and perform modality complementary enhancement processing on the extracted multimodal document information; The information matching module is used to match the processed multimodal document information and return the matching result.
11. The multimodal document content matching system according to claim 10, characterized in that: The image acquisition and processing module also includes: An image acquisition module, used to acquire an image to be identified; The image processing module is used to analyze the resolution of the image to be identified and determine whether the resolution of the image to be identified is less than a preset resolution. If it is less than the preset resolution, the image to be identified is adjusted to the preset resolution, which is 1920.
12. The multimodal document content matching system according to claim 10, characterized in that: The image correction module also includes: A key point detection module, used to detect global key points of the table in the image to be identified, wherein the global key points include four corner points of each cell; The judgment and correction module is used to judge whether distortion and deformation occur between different text lines of the document based on the global key points. If distortion and deformation occur, the text lines of the image to be recognized are corrected to obtain a corrected image.
13. The multimodal document content matching system according to claim 10, characterized in that: The text recognition module also includes: An enhancement processing module, used for performing binarization and smoothing denoising processing on the corrected image to obtain an enhanced processed image of the corrected image; The recognition module is used to perform OCR text recognition on the enhanced processed image and mark a text box at the recognized text.
14. The multimodal document content matching system according to claim 10, characterized in that: The text correction module also includes: An angle calculation module, used to calculate the tilt angle of the text box; The correction module is used to determine whether the text box sends an offset according to the tilt angle, and if an offset occurs, correct the text box to obtain a corrected text.
15. The multimodal document content matching system according to claim 10, characterized in that: The information extraction and processing module also includes: An information extraction module, used to extract multimodal document information and corresponding text boxes from the corrected text; A filling module, used to detect blank spaces between the text boxes and mark blank spaces at the positions of the blank spaces for filling; A splicing module, used for identifying line break document information based on the multimodal document information and combining and splicing it with related text; The filtering module is used to filter the staggered texts according to the slope relationship between the text boxes.
16. The multimodal document content matching system according to claim 10, characterized in that: The information matching module also includes: A keyword extraction module, used to extract keywords from the processed multimodal document information and determine corresponding values according to the paragraph context where the keywords are located; A key-value pair generation module, used to establish a corresponding relationship between the keyword and the value, and form a key-value pair based on the corresponding relationship; The matching module is used to match the information according to the key-value pair and return the matching result.
17. The multimodal document content matching system according to claim 10, characterized in that: The system further comprises: The seal recognition and classification module is used to recognize and classify different seals in the rectified image using guided learning and contrastive learning methods.
18. An electronic device, characterized in that: include: A processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the multimodal document content matching method described in any one of claims 1 to 9.
19. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the multimodal document content matching method described in any one of claims 1 to 9.
Citation Information
Patent Citations
OCR character position correction method and device, storage medium and electronic equipment
CN112085014A
Bank account statement identification method
CN117037198A
Document information extraction method, device and system and storage medium
CN119942576A
Engineering drawing label identification method and system based on multi-modal information extraction
CN119964171A
RPA and ai-based table information extraction method and apparatus, device and medium
WO2022062798A1
Cited By
Document table structure identification method oriented to water conservancy large model retrieval enhancement
CN120633613A
A document table structure recognition method for enhanced retrieval of large water conservancy models
CN120633613B
Method and framework for optimizing scanning copy content recognition quality by using large model
CN120689894A