Document identification method, system and equipment based on multi-modal large model and medium

Through the document recognition method of multimodal large model, the generalization and robustness of document recognition solutions are solved, efficient and accurate identification of various document types are achieved, customized development costs are reduced, and the recognition accuracy of complex structures and long numbers is improved.

CN120340054APending Publication Date: 2025-07-18BEIJING ZHIPU PILOT TECHNOLOGY CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510444829.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing document recognition solutions have problems such as weak generalization, poor robustness and inaccurate long-digit recognition, resulting in high cost of customized development, sensitive to noise interference and low recognition accuracy.

Method used

Document recognition methods based on multimodal large models are adopted, including image preprocessing, multimodal large model inference, OCR recognition and result verification, and the cross-modal understanding ability of the VLM model and JSON template configuration are used, and the recognition accuracy is improved by combining the SVTR V2 model.

Benefits of technology

It realizes efficient and accurate identification of multiple document types, reduces customized development costs, improves the accuracy of identification of complex structures and long numbers, and adapts to the personalized needs of different customers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340054A_ABST
    Figure CN120340054A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to a multi-modal large model-based document identification method, system and device and a medium, and the method comprises the following steps: 1) image preprocessing: preprocessing a document image input by a user; 2) multi-modal large model reasoning: based on the pre-processed document image, the configured JSON template and the cue word template, performing reasoning by a multi-modal large model to obtain a JSON result; 3) OCR identification: identifying the document image input by the user by using an OCR identification technology to obtain an OCR identification result; and 4) verification: performing similarity comparison on the OCR recognition result and the JSON result, and determining a document recognition result based on a similarity comparison result. The method is high in generalization, can adapt to various types of receipts, and can provide an efficient and accurate recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and relates to a document recognition method, system, device and medium, especially a document recognition method, system, device and medium based on a multimodal large model. Background Art

[0002] Existing document recognition solutions usually adopt the tandem method of an OCR (Optical Character Recognition) model and an LLM (Large Language Model) to achieve information extraction and structured processing. Specifically, this process generally consists of two main steps:

[0003] 1. Use the OCR model to extract all text.

[0004] In this stage, the OCR model is responsible for recognizing and extracting all text in the document. This process includes: (1) Text localization and recognition: The OCR model scans the document image to recognize and extract the text information therein, including text content, numbers, punctuation marks, etc., and usually needs to handle problems such as different fonts, font sizes, text spacing, and printing quality. (2) Image-to-text conversion: The OCR model can convert the scanned picture or the image obtained by taking a photo into editable text data. The OCR model can not only recognize standard machine-printed text, but also handle complex factors such as handwritten text, seals, annotations, signatures, etc. (3) Basic layout understanding: Some advanced OCR models can also partially understand the layout information of the document and recognize structures such as titles, paragraphs, tables, etc. However, the main goal of this stage is to ensure that all text can be extracted to provide basic data for the analysis in the next stage.

[0005] 2. The LLM (Large Language Model) summarizes and structurally analyzes the text content.

[0006] After the OCR model extracts the original text, the next task is to use the LLM (Large Language Model) to further process this text. The functions of the large language model at this stage include: (1) Text understanding and semantic analysis: The LLM (Large Language Model) deeply understands and analyzes the extracted text through natural language processing technology, and can understand the context relationship, grammatical structure and semantic meaning in the text, which enables the LLM (Large Language Model) to identify key information in the text, such as invoice number, date, amount, customer name, product information, etc. (2) Information induction and summary: The LLM (Large Language Model) can induct and summarize the extracted text according to preset rules or user requirements. For example, from a contract containing multiple fields, extract key data such as "signing date", "contract amount" and "signing party", and form a structured output.

[0007] However, the existing document recognition solutions generally have the following problems:

[0008] 1. Existing document recognition solutions have the problem of weak generalization (there are significant costs for customization development).

[0009] Currently, many document recognition solutions (especially traditional OCR or deep learning - based models) are usually developed specifically for certain domains and specific types of documents. For example, one system may be specifically for invoice recognition, while another may be for medical bills. Although this approach can achieve high accuracy in specific tasks, it also brings some problems: (1) High costs for customization development: Since each type of document has different formats, field layouts, semantic contents, etc., traditional models often need to be customized and trained according to the type of document. This means that for each new type of document, the model needs to be re - developed or re - trained, increasing the maintenance and update costs of the system. (2) Weak generalization: Existing document recognition models usually cannot be well migrated to new, unseen document types. For example, when dealing with a new supply chain document or invoices in different formats, existing models may not be able to fully adapt to these changes, resulting in a decrease in recognition accuracy.

[0010] 2. Existing document recognition solutions have the problem of weak robustness (for noises such as watermarks and seals, the existing solutions have poor handling capabilities).

[0011] In many practical applications, documents (such as invoices, contracts, certificates, etc.) often face various interference factors. For example, watermarks: Anti - counterfeiting watermarks, company logos and other elements may affect the accurate recognition of the OCR system; seals and handwritten contents: Seals, signatures or handwritten words may be different from printed words, affecting the effect of image recognition.

[0012] Existing document recognition solutions are often vulnerable to these interference factors (i.e., noises), and it is easy to confuse watermarks, seals and normal text after generating OCR results. Even simple noises may significantly affect the recognition accuracy.

[0013] 3. The VLM model has the problem of being unable to accurately extract long numbers (this problem leads to low accuracy in document recognition).

[0014] The VLM model has made significant progress in recent years. It can handle joint tasks of images and texts and is suitable for tasks such as image recognition and natural language understanding. However, in terms of OCR, the VLM model still faces some challenges, especially for the recognition of long numbers. In many documents (such as bank statements, invoices, contracts, etc.), long numbers (such as ID numbers, bank account numbers, ticket numbers, etc.) are often included. Currently, the VLM model is often not accurate enough in recognizing such long numbers, and it is prone to errors during the process of number recognition.

[0015] Therefore, in view of the defects existing in the above-mentioned prior art, it is necessary to develop a new document recognition method, system, device and medium. Summary of the Invention

[0016] In order to overcome the defects of the prior art, the present invention proposes a document recognition method, system, device and medium based on a multi-modal large model, which has strong generalization ability, can adapt to various types of documents, and can provide efficient and accurate recognition results.

[0017] In order to achieve the above object, the present invention provides the following technical solutions:

[0018] A document recognition method based on a multi-modal large model, characterized by comprising the following steps:

[0019] 1) Image preprocessing: Preprocess the document image input by the user;

[0020] 2) Multi-modal large model inference: Based on the preprocessed document image, the configured JSON template and the prompt word template, perform inference by the multi-modal large model to obtain a JSON result;

[0021] 3) OCR recognition: Use OCR recognition technology to recognize the document image input by the user to obtain an OCR recognition result;

[0022] 4) Verification: Compare the similarity between the OCR recognition result and the JSON result, and determine the document recognition result based on the similarity comparison result.

[0023] Preferably, the step 1) specifically includes:

[0024] 11) Use the PaddleOCR model to perform inference on the document image input by the user to generate an inference result, where the inference result includes all the text on the document image, the confidence value of the text, and the coordinates of the text anchor box;

[0025] 12) Determine the rotation angle and cropping size that the document image input by the user should be based on the inference result, and rotate and crop the document image input by the user according to the rotation angle and cropping size to obtain a cropped document image;

[0026] 13) Use a straight line detection algorithm to determine whether the cropped document image is tilted and correct it when it is tilted.

[0027] Preferably, the step 12) specifically includes:

[0028] 121) Rotate the document image input by the user by four right angles to obtain four rotated document images, and perform inference on the four rotated document images respectively to obtain four inference results. Determine the direction of the text based on the shape of the text anchor boxes in the four inference results, and determine two alternative rotated document images based on the direction of the text;

[0029] 122) Judge through the sum of the confidence values of the text in the inference results corresponding to the two alternative rotated document images, and take the one with the largest sum as the document image with the correct direction;

[0030] 123) Determine the cropping size through the outermost edges of all text anchor boxes in the document image with the correct direction, and crop the document image with the correct direction according to the cropping size to obtain the cropped document image.

[0031] Preferably, step 2) specifically includes:

[0032] 21) Configure the JSON template, that is, configure the fields in the document where text needs to be extracted;

[0033] 22) Concatenate the configured JSON template with the prompt template and input them together with the preprocessed document image into the multi-modal large model for inference to generate an inference result;

[0034] 23) Use the JSON verification module to verify and correct the inference result to obtain the JSON result.

[0035] Preferably, the OCR recognition result in step 3) includes all the text on the single image input by the user, the confidence value of the text, and the coordinates of the text anchor box.

[0036] Preferably, step 4) specifically includes:

[0037] 41) Perform similarity matching on the text in the OCR recognition result and the fields in the JSON result respectively. If the similarity of the matching result is 1, the recognition is correct; if the similarity of the matching result is not 1 and greater than the threshold, list the field as the verification object;

[0038] 42) Use the PaddleOCR model to recognize the coordinates of the verification object and perform image cropping to obtain the field image;

[0039] 43) Use the SVTR V2 model to recognize the field image, and regard the recognition result as the final result of the field and update it to the JSON result to obtain the document recognition result.

[0040] Preferably, the threshold is 0.7.

[0041] In addition, the present invention also provides a document recognition system based on a multi-modal large model, which is characterized by including:

[0042] An image preprocessing module, which is used to preprocess the document image input by the user;

[0043] A multi-modal large model inference module, which is used to perform inference by the multi-modal large model based on the preprocessed document image, the configured JSON template, and the prompt word template to obtain a JSON result;

[0044] An OCR recognition module, which is used to recognize the document image input by the user using OCR recognition technology to obtain an OCR recognition result;

[0045] A verification module, which is used to compare the similarity between the OCR recognition result and the JSON result, and determine the document recognition result based on the similarity comparison result.

[0046] Moreover, the present invention also provides a document recognition device based on a multi-modal large model, which is characterized by including:

[0047] One or more processors;

[0048] A memory for storing one or more programs;

[0049] When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the document recognition method based on the multi-modal large model as described above.

[0050] Finally, the present invention also provides a computer-readable storage medium, on which a computer program is stored, which is characterized in that when the program is executed by a processor, the steps of the document recognition method based on the multi-modal large model as described above are implemented.

[0051] Compared with the prior art, the document recognition method, system, device, and medium based on the multi-modal large model of the present invention have one or more of the following beneficial technical effects:

[0052] 1. The present invention uses the general capabilities of the VLM model, has strong generalization ability, can adapt to different document types, and ensures the accuracy of cold start.

[0053] Traditional document recognition systems usually rely on customized development for specific types of documents. Therefore, when encountering new types of documents, the model must be retrained or adjusted, which increases the development cost and cannot efficiently process various types of documents in practical applications. The present invention utilizes a general prompt template and applies the cross-modal (image-text) understanding ability of the VLM model to be able to process various types of documents. This templatized design has the following characteristics: (1) Strong generality: There is no need to train the model separately for each document type and it can automatically adapt to different formats of documents, such as invoices, contracts, customs declarations, transfer vouchers, etc.; (2) High accuracy: When processing these various types of documents, it can maintain a high accuracy rate (above 80%), showing good performance in practical applications; (3) This generality greatly reduces the dependence on customized development and can still maintain high processing efficiency and accuracy when facing unknown or new types of documents.

[0054] 2. In response to customer requirements, the present invention can improve the accuracy of field extraction at low cost through field configuration.

[0055] In traditional document recognition technologies, usually only the preset fields in the document can be extracted, which cannot meet the customized requirements of customers for certain specific fields or information. At this time, customers often need additional development or manual intervention to complete these tasks. The present invention supports customized field extraction according to the specific requirements of customers through field configuration. Customers can specify the field names to be extracted, and by identifying these fields and matching them with the actual content in the document, the target information can be accurately extracted. The specific advantages include: (1) High customization efficiency: Customers can flexibly select the fields to be recognized according to actual needs. For example, extract the "invoice number" and "amount" fields from an invoice, or extract the "signing date" and "terms" fields from a contract; (2) Improved accuracy: By clearly specifying the fields, the model can focus more on extracting specific content, reducing the probability of misrecognition, thereby improving the recognition accuracy. This function not only enhances the flexibility of recognition but also enables it to adapt to the personalized needs of different customers and meet diverse business scenarios.

[0056] 3. The present invention combines the VLM model and the cutting-edge text recognition model to improve the recognition accuracy.

[0057] Most existing document recognition technologies rely solely on OCR technology or are trained based on vision feature models, lacking the joint understanding of images and text. Simply relying on OCR has poor recognition effects for complex structured documents, especially for table information and long numbers. The present invention combines the VLM model and advanced text recognition technology to improve the overall understanding ability of documents. The VLM model can not only process image information but also perform joint learning in combination with the text in the document, thus achieving the following effects: (1) Multimodal collaboration: The VLM model can simultaneously understand the visual content and text content in the image, thereby improving the recognition effect. For example, by combining the table structure in the image and the field information in the text, structured data in the document can be automatically extracted; (2) Precise processing of complex text: In combination with modern OCR technology, complex text and long numbers can be processed more precisely. For example, long numbers such as the amount in an invoice and the account number in a transfer voucher can be extracted, avoiding the misrecognition problem of traditional VLM models. Brief Description of the Drawings

[0058] Figure 1 It is a flowchart of the document recognition method based on the multimodal large model of the present invention.

[0059] Figure 2 It is a schematic diagram of the composition of the document recognition system based on the multimodal large model of the present invention. Detailed Embodiments

[0060] Before detailing any embodiment of the present invention, it should be understood that in its application, the present invention is not limited to the construction and arrangement details of the components described in the following description or illustrated in the following drawings. The present invention can have other embodiments and can be practiced or carried out in various ways. Additionally, it should be understood that the wording and terms used herein are for the purpose of description and should not be considered restrictive. As used herein, "including" or "having" and their variants are intended to cover the items listed hereinafter and their equivalents as well as additional items. Unless otherwise specified or limited, the terms "mounted", "connected", "supported", and "coupled" and their variants are used broadly and cover direct mounting and indirect mounting, connection, support, and coupling. In addition, "connection" and "coupling" are not limited to physical or mechanical connection or coupling.

[0061] Moreover, on the one hand, in the disclosure of the present invention, the orientation or positional relationship indicated by terms such as "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the above terms should not be construed as limiting the present invention; on the other hand, the term "a" should be understood as "at least one" or "one or more". That is, in one embodiment, the number of an element can be one, while in other embodiments, the number of the element can be multiple. The term "a" should not be construed as limiting the quantity.

[0062] Before introducing the specific content of the present invention, some technical terms used in the present invention will be briefly introduced to facilitate those skilled in the art to better understand the present invention.

[0063] 1. LLM (Large Language Model): LLM refers to the Large Language Model, which is a model trained through deep learning technology and can understand and generate natural language. It can process text data, answer questions, write articles, translate languages, etc. Examples of LLM include GPT, GLM, etc.

[0064] 2. VLM (Vision-Language Model): VLM refers to the Vision-Language Model, which can simultaneously understand and process different types of data, such as text and images. For example, VLM can not only analyze the content of pictures but also understand the related text information, and is widely used in fields such as image-text matching and image description. Common models include CogVLM, etc.

[0065] 3. OCR (Optical Character Recognition): OCR refers to Optical Character Recognition. It is a technology used to convert handwritten or printed text in scanned or photographed documents into editable digital text. For example, after scanning a book, OCR can extract the text in the scanned image and convert it into an editable document.

[0066] 4. Document Recognition: Document recognition refers to the use of automated technology to identify and extract information on documents. These documents can be contracts, invoices, certificates, etc. Document recognition usually involves OCR and other AI technologies to identify key information on documents and then process it.

[0067] 5. Generalization: Generalization refers to the performance ability of a large model when encountering new data. A model with good generalization not only performs well on training data but can also effectively process unseen data. In other words, the stronger the generalization, the wider the scope of application of the model, and it can better handle different situations in practical applications.

[0068] The content of the present invention will be described in detail below. Among them, Figure 1 shows a flowchart of the document recognition method based on a multimodal large model of the present invention. As Figure 1 shown, the document recognition method based on a multimodal large model of the present invention includes the following steps:

[0069] I. Image preprocessing.

[0070] To recognize a document, it is first necessary to preprocess the document image input by the user to facilitate subsequent recognition.

[0071] In the present invention, the preprocessing of the document image input by the user specifically includes:

[0072] 1. Use an existing OCR model, such as the PaddleOCR model, to perform inference on the document image input by the user to generate an inference result. The inference result includes all the text on the document image, the confidence value of the text (the confidence value is related to the direction of the text), and the coordinates of the text anchor box.

[0073] 2. Determine the rotation angle and cropping size that the document image input by the user should be subjected to according to the inference result, and rotate and crop the document image input by the user according to the rotation angle and cropping size to obtain a cropped document image.

[0074] Specifically, first, rotate the document image input by the user by four right angles (that is, rotate four times, each time rotating 90°) to obtain four rotated document images, and perform inference on the four rotated document images respectively to obtain four inference results.

[0075] Next, determine the direction of the text by the shape of the text anchor box in the four inference results and determine two alternative rotated document images based on the direction of the text. Specifically, the shape of the text anchor box can be determined based on the coordinates of the text anchor box, and for a normal text, its text anchor box should be a horizontal rectangle. Therefore, the two document images with horizontal rectangular text anchor boxes are determined as alternative rotated document images.

[0076] Then, make a judgment based on the sum of the confidence values of the text in the inference results corresponding to the two alternative rotated document images (the direction with the largest sum should point to the correct image direction), and use the rotated document image with the largest sum of the confidence values of the text as the document image with the correct direction.

[0077] Finally, determine the cropping size by the outermost edges of all text bounding boxes in the document image with the correct direction, and crop the document image with the correct direction according to the cropping size to remove useless information, so as to obtain the cropped document image.

[0078] 3. Use a line detection algorithm to determine whether the cropped document image is tilted and correct it when it is tilted.

[0079] Finally, use an existing line detection algorithm, such as the Hough line detection, to determine whether the cropped document image is tilted and correct it when it is tilted, so as to obtain the final preprocessed document image.

[0080] II. Multimodal large model inference.

[0081] After image preprocessing, the present invention uses a multimodal large model (the Zhipu closed-source CogVLM-plus model is used in the present invention) for inference to generate a JSON file including the recognition result by the multimodal large model, that is, the JSON result. That is, when performing multimodal large model inference, based on the preprocessed document image, the configured JSON template and the prompt word template, the multimodal large model performs inference to obtain the JSON result, which specifically includes:

[0082] 1. Configure the JSON template.

[0083] In the present invention, when using the multimodal large model for inference, first, a configured JSON template is set, that is, the fields for which text needs to be extracted in the document are configured. The configured JSON template marks the fields for which text needs to be extracted in the document. For example, for a medical image document, the following JSON can be configured: {"Name":"","Physical examination time (year-month-day)":"","Report type":"","Report findings / observations":"","Diagnosis / Opinion":""}.

[0084] By setting the JSON template, field information that needs to be recognized can be provided for the multimodal large model, so that the multimodal large model can output more complete values, thereby improving the overall document recognition accuracy.

[0085] 2. Concatenate the configured JSON template with the prompt word template and input them together with the preprocessed document image into the multimodal large model for inference to generate an inference result.

[0086] Among them, the prompt template is a general template used in the prior art for reasoning and text recognition of documents by a multi-modal large model, aiming to guide and control the multi-modal large model to recognize document images to generate recognition results. This part of the content belongs to the prior art and will not be introduced in detail here for the sake of simplicity.

[0087] 3. Use the JSON verification module to verify and correct the inference result to obtain a JSON result.

[0088] After the multi-modal large model generates an inference result, the JSON verification module is also used to verify and correct the inference result to facilitate obtaining a JSON result that can be directly parsed.

[0089] In this way, through the configured JSON template, the present invention supports customized field extraction according to the specific needs of customers in a field configuration manner. Customers can specify the field names to be extracted, and by identifying these fields and matching them with the actual content in the document, the target information can be accurately extracted. Such specific advantages include: (1) High customization efficiency: Customers can flexibly select the fields to be recognized according to actual needs. For example, extract the "invoice number" and "amount" fields from an invoice, or extract the "signing date" and "terms" fields from a contract. (2) Improved accuracy: By clearly specifying the fields, the multi-modal large model can focus more on extracting specific content, reducing the probability of misrecognition, thereby improving the recognition accuracy. This function not only enhances the flexibility of recognition but also enables it to adapt to the personalized needs of different customers and meet diverse business scenarios.

[0090] III. OCR Recognition.

[0091] Use OCR recognition technology to recognize the document image input by the user to obtain an OCR recognition result.

[0092] In the present invention, an existing OCR model, such as the PaddleOCR model, can be used to recognize the document image input by the user to obtain an OCR result, including all the text on the single image input by the user, the confidence value of the text, and the coordinates of the text anchor box.

[0093] IV. Verification.

[0094] Compare the similarity between the OCR recognition result and the JSON result, and determine the document recognition result based on the similarity comparison result. Specifically, it includes:

[0095] 1. Perform similarity matching on the text in the OCR recognition result and the content of the fields in the JSON result respectively. If the similarity of the matching result is 1, it indicates correct recognition. If the similarity of the matching result is not 1 and is greater than the threshold, for example, the threshold can be taken as 0.7, then list the content of the fields in the JSON result as the verification object. If the similarity is less than the threshold, it is considered that the text in the OCR recognition result has nothing to do with the content of the fields in the JSON result, and no processing is required at this time.

[0096] In the present invention, when calculating the similarity, all the content containing digital strings can be traversed and extracted from the content of the fields in the JSON result, and then all the content containing digital strings can be traversed and extracted from the OCR recognition result, and the ratio of the lengths of the two digital strings is used as the similarity value.

[0097] 2. Use the PaddleOCR model to recognize the coordinates of the verification object and perform image cutting to obtain the field image.

[0098] 3. Use an existing high-precision Chinese scene text recognition model, such as the SVTR V2 model, to recognize the field image. It is found that the SVTR V2 model can well achieve the recognition of long digital strings, and the recognition result is regarded as the final result of the field and updated to the JSON result to obtain the document recognition result.

[0099] The present invention improves the overall understanding ability of documents by combining the VLM model and advanced text recognition technology. The VLM model can not only process image information, but also perform joint learning by combining the text in the document, so as to achieve the following effects: (1) Multimodal collaboration: The VLM model can simultaneously understand the visual content and text content in the image, thereby improving the recognition effect. For example, by combining the table structure in the image and the field information in the text, the structured data in the document can be automatically extracted. (2) Precise processing of complex text: Combining modern OCR technology, such as the SVTR V2 model, can more accurately process complex text and long numbers, such as extracting long numbers such as the amount in the invoice and the account number in the transfer voucher, avoiding the misrecognition problem of the traditional VLM model.

[0100] Figure 2 Shows the schematic diagram of the composition of the document recognition system based on the multimodal large model of the present invention. As Figure 2 shown, the document recognition system based on the multimodal large model of the present invention includes:

[0101] 1. Image preprocessing module.

[0102] The image preprocessing module is used to preprocess the document image input by the user.

[0103] 2. Multimodal large model inference module.

[0104] The multimodal large model inference module is used to perform inference by the multimodal large model based on the preprocessed document images, configured JSON templates, and prompt templates to obtain JSON results.

[0105] 3. OCR recognition module.

[0106] The OCR recognition module is used to recognize the document images input by the user using OCR recognition technology to obtain OCR recognition results.

[0107] 4. Verification module.

[0108] The verification module is used to compare the similarity between the OCR recognition results and the JSON results, and determine the document recognition results based on the similarity comparison results.

[0109] In addition, the present invention also relates to a document recognition device based on a multimodal large model, which includes: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the document recognition method based on the multimodal large model as described above.

[0110] Finally, the present invention also relates to a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, the steps of the document recognition method based on the multimodal large model as described above are implemented.

[0111] The document recognition method, system, device, and storage medium based on the multimodal large model of the present invention have good generalization and can be used for the structured information extraction and recognition of various general documents (which may include complex table information). The specific scenarios are as follows:

[0112] 1. Structured information extraction and recognition of customs declarations.

[0113] Customs declarations usually involve complex data structures, such as various information such as cargo names, categories, quantities, amounts, import and export locations, waybill numbers, invoice numbers, and tariffs. Due to the non-uniformity of customs declaration formats and the existence of a large number of data fields, traditional manual entry methods are prone to errors and have low efficiency.

[0114] 2. Structured information extraction and recognition of transfer vouchers.

[0115] As a record of financial transactions, transfer vouchers contain a large amount of important information, such as transaction date, account number, transaction amount, bank name, payment method, remitter and payee information, etc. These vouchers often appear in printed or handwritten formats, and sometimes there are interference elements such as bank watermarks and signatures. Traditional methods of processing transfer vouchers rely on manual input and are prone to errors.

[0116] 3. Structured information extraction and recognition of medical examination documents.

[0117] Medical examination documents include items such as physical examination reports, laboratory test sheets, diagnosis reports, etc. These documents usually contain information such as patient personal information, examination items, examination results, doctor's suggestions, etc. Many medical documents have different formats, some containing complex tables and some containing handwritten content.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the present invention. Those skilled in the art can modify or equivalently replace the technical solutions of the present invention according to the idea of the present invention, without departing from the essence and scope of the technical solutions of the present invention.

Claims

1. A document recognition method based on a multimodal large model, characterized in that, It includes the following steps: 1) Image preprocessing: Preprocess the document image input by the user; 2) Multimodal large model inference: Based on the preprocessed document image, the configured JSON template, and the prompt template, perform inference by the multimodal large model to obtain a JSON result; 3) OCR recognition: Use OCR recognition technology to recognize the document image input by the user to obtain an OCR recognition result; 4) Verification: Compare the similarity between the OCR recognition result and the JSON result, and determine the document recognition result based on the similarity comparison result.

2. The document recognition method based on a multi-modal large model according to claim 1, wherein The specific content of step 1) includes: 11) Use the PaddleOCR model to perform inference on the document image input by the user to generate an inference result, where the inference result includes all the text on the document image, the confidence value of the text, and the coordinates of the text anchor box; 12) Determine the rotation angle and cropping size that the document image input by the user should be subjected to according to the inference result, and rotate and crop the document image input by the user according to the rotation angle and cropping size to obtain a cropped document image; 13) Use a straight line detection algorithm to determine whether the cropped document image is tilted and correct it when it is tilted.

3. The document recognition method based on a multimodal large model according to claim 2, wherein The specific content of step 12) includes: 121) Rotate the document image input by the user by four right angles to obtain four rotated document images, perform inference on the four rotated document images respectively to obtain four inference results, judge the direction of the text by the shape of the text anchor box in the four inference results, and determine two alternative rotated document images based on the direction of the text; 122) Judge by the sum of the confidence values of the text in the inference results corresponding to the two alternative rotated document images, and the one with the largest sum is used as the document image with the correct direction; 123) Judge the cropping size by the outermost edges of all the text anchor boxes in the document image with the correct direction, and crop the document image with the correct direction according to the cropping size to obtain a cropped document image.

4. The document recognition method based on a multi-modal large model according to claim 1, characterized in that, The specific content of step 2) includes: 21) Configure the JSON template, that is, configure the fields where text needs to be extracted in the document; 22) Concatenate the configured JSON template with the prompt template and input them into the multimodal large model together with the preprocessed document image for inference to generate an inference result; 23) Use the JSON verification module to verify and correct the inference result to obtain a JSON result.

5. The document recognition method based on a multimodal large model according to claim 1, wherein The OCR recognition result in step 3) includes all the text on the document image input by the user, the confidence value of the text, and the coordinates of the text anchor box.

6. The document recognition method based on a multi-modal large model according to claim 1, wherein The specific content of step 4) includes: 41) Perform similarity matching on the text in the OCR recognition result and the fields in the JSON result respectively. If the similarity of the matching result is 1, the recognition is correct; if the similarity of the matching result is not 1 and is greater than the threshold, list the field as an object to be verified; 42) Use the PaddleOCR model to recognize the coordinates of the object to be verified and perform image cropping to obtain a field image; 43) Identify the field image using the SVTR V2 model, and recognize the recognition result as the final result of the field and update it to the JSON result to obtain the document recognition result.

7. The document recognition method based on a multi-modal large model according to claim 6, characterized in that, The threshold is 0.

7.

8. A document recognition system based on a multimodal large model, characterized in that, It includes: An image preprocessing module for preprocessing the document image input by the user; A multi-modal large model inference module for inferring by the multi-modal large model based on the preprocessed document image, the configured JSON template and the prompt word template to obtain the JSON result; An OCR recognition module for using OCR recognition technology to recognize the document image input by the user to obtain the OCR recognition result; A verification module for comparing the similarity between the OCR recognition result and the JSON result, and determining the document recognition result based on the similarity comparison result.

9. A document recognition device based on a multimodal large model, characterized in that, It includes: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the document recognition method based on the multi-modal large model according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the document recognition method based on the multi-modal large model according to any one of claims 1-7.

Citation Information

Cited By

  • Acquired data processing method and device and medium

    CN120612708A

  • A data collection processing method, device and medium

    CN120612708B

  • Invoice information positioning and reading method based on image recognition

    CN121074903A

  • Multi-model signature matching method and device, storage medium and program product

    CN121096031A

  • Container number identification method and device, and computer equipment

    CN121121721A