Graphic and text information extraction method and system based on multi-modal large model and storage medium

Through the end-to-end graphic information extraction method based on multimodal large models, combined with OCR basic ability training and multi-task hybrid training, error propagation, information loss and other problems of graphic information extraction in the existing technology are solved, and high-precision and stable information extraction effect is achieved.

CN120047956APending Publication Date: 2025-05-27BEIJING YIDAO BOSHI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411971194.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art has problems such as error propagation, information loss, high error rate of Chinese character recognition, poor instruction compliance ability and serious hallucinations in the extraction of graphic and text information, which is difficult to meet the requirements of industrial production.

Method used

The end-to-end graphic information extraction method based on multimodal large models is adopted. Through OCR basic ability training and multi-task hybrid training, the character recognition rate and instruction compliance ability of the model are improved, hallucinations are reduced, and the relative position information of the text in the picture is preserved.

Benefits of technology

It significantly improves the accuracy and stability of graphic and text information extraction, enhances the generalization ability of the model and the ability to understand user questions, and solves the problems of high character recognition error rate, weak instruction compliance ability and serious hallucinations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047956A_ABST
    Figure CN120047956A_ABST
Patent Text Reader

Abstract

The invention discloses an image-text information extraction method and system based on a multi-modal large model and a storage medium. The extraction method comprises the following steps: S1, selecting a multi-modal large model; s2, aiming at the multi-modal large model, carrying out an OCR basic capability training task; s3, for the multi-modal large model trained in the S2, multi-task mixed and simultaneous graphics and text information extraction training tasks are carried out, and training data are randomly extracted from multiple tasks; and S4, inputting a to-be-processed image into the multi-modal large model trained in the S3, and outputting graphic and text information in the original image. Through special training task design in the field of image-text information extraction, the character recognition rate and the instruction following capability of a multi-modal large model are greatly improved, illusion is inhibited, a good end-to-end information extraction effect is achieved, and the image-text information extraction precision in industrial production is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of natural language processing and computer vision, and in particular to a method, system, and storage medium for extracting text and image information based on a multimodal large model. Background Art

[0002] Information extraction is an important topic and application scenario in the field of natural language processing, and extracting text information from images requires the technologies of both computer vision and natural language processing.

[0003] Currently, the mainstream method for extracting text and image information in the industry is a cascaded scheme, which mainly consists of two cascaded modules: an OCR (Optical Character Recognition) module and a pure text information extraction module. Specifically, the OCR module is used to extract the text in the picture, and after splicing it into a string, it is sent to the pure text information extraction module for key information extraction. This scheme has three main defects:

[0004] 1. The cascaded scheme will cause error propagation, and the errors in the OCR module will lead to errors and instability in the subsequent pure text information extraction module;

[0005] 2. The cascaded scheme will cause information loss. When the OCR module extracts text from the image, information such as the intuitive relative position, layout, font, and color of the text in the picture will be lost. Especially in the image scenario with complex layout, when the text extracted by the OCR module is spliced into a string, it is difficult to determine the reading order, resulting in semantic confusion. In this case, it is extremely difficult for the subsequent pure text information extraction module to extract the required information;

[0006] 3. Both modules of the cascaded scheme use traditional small models, and their generalization ability for different samples and understanding ability for different user questions are weak, and the versatility is poor;

[0007] With the development of multimodal large models, multimodal large models have also shown certain text and image information extraction capabilities. However, the native multimodal large models still have the following three main defects in the text and image information extraction task, resulting in this scheme not meeting the industrial production requirements:

[0008] 1. High error rate in Chinese character recognition: The current multimodal large models in the industry are positioned for general understanding ability, and only a small number of Chinese and English character recognition tasks are trained in the pre-training stage. Since the number of Chinese characters is too large and the recognition difficulty is high, large-scale data and sufficient training are required to obtain good results. Therefore, the error rate of Chinese text character recognition of the current multimodal large models in the industry is very high, especially for handwritten scenarios and scenarios with slightly blurred handwriting.

[0009] 2. Poor instruction following ability leads to omission and redundancy of output content: For multiple extraction contents specified by the user, the native multimodal large models are very likely to omit. For exampleFigure 1 As shown in , the MiniCPM-V large model missed the content of the "password area" specified by the user to be extracted, missed the content of the second and third rows of the table, and redundantly output the content of "total amount" not specified by the user.

[0010] 3. Serious hallucinations: The original multi-modal large model is very prone to tampering with the original text and outputting fabricated and incorrect content. For example Figure 1 As shown in , the MiniCPM-V large model output an incomplete purchaser's taxpayer identification number, incorrect seller's address and phone number, incorrect tax-inclusive total amount (in capital letters), and only the information in the first row of the table was output. Among the information in the first row, only the unit and tax rate were correctly extracted, and the rest were all incorrect. Summary of the Invention

[0011] To solve the above problems, the technical solution of the present invention proposes a method, system and storage medium for extracting graphic and text information based on a multi-modal large model, which belongs to an end-to-end graphic and text information extraction solution based on a multi-modal large model. By designing specialized training tasks for the field of graphic and text information extraction, the character recognition rate and instruction following ability of the multi-modal large model are greatly improved, hallucinations are suppressed, good end-to-end information extraction effects are achieved, and the accuracy of graphic and text information extraction in industrial production is significantly improved.

[0012] According to the first aspect of the technical solution of the present invention, a method for extracting graphic and text information based on a multi-modal large model is provided, wherein the method for extracting graphic and text information includes:

[0013] S1: Select a multi-modal large model;

[0014] S2: Perform OCR basic ability training tasks for the multi-modal large model;

[0015] S3: Perform graphic and text information extraction training tasks that are multi-task mixed and simultaneous for the multi-modal large model trained in S2, wherein the training data is randomly extracted from multiple tasks;

[0016] S4: Input the image to be processed into the multi-modal large model trained in S3, and output the graphic and text information in the original image.

[0017] Further, in S1, the multi-modal large model is the MiniCPM-V model.

[0018] Further, in S2, the OCR basic ability training is specifically:

[0019] The input is an image containing text content and corresponding instructions, and the instructions require the multi-modal large model to output all the text in the picture;

[0020] The output is all the text on the image, output line by line at the text line level.

[0021] Further, in the S3, the multi-task hybrid graphic and text information extraction training includes:

[0022] A text recognition task, a specified entity information extraction task, a specified table information extraction task, a specified hybrid information extraction task, and a table question-answering task.

[0023] Further, the text recognition task is specifically:

[0024] The input is an image containing text content and a corresponding instruction, and the instruction is to require the multi-modal large model to output all the text in the picture;

[0025] The output is all the text on the image, output line by line at the text line level.

[0026] Further, the specified entity information extraction task is specifically:

[0027] The input is an image containing text content and a corresponding instruction, and the instruction is to require the multi-modal large model to output the specified entity information in the picture;

[0028] The output is the specified entity information on the image, output in json format.

[0029] Further, in the specified entity information extraction task, during the training process, each sample is randomly assigned the fields to be extracted.

[0030] Further, the entity information refers to the information represented in the form of key-value pairs.

[0031] Further, the specified table information extraction task is specifically:

[0032] The input is an image containing text content and a corresponding instruction, and the instruction is to require the multi-modal large model to output the specified table information in the picture;

[0033] The output is the specified table information on the image, output in json format. The outermost layer of the json is a list representing all the table rows, and each element in the list is a row of the table, and all the information in each row is represented by a dictionary.

[0034] Further, in the specified table information extraction task, during the training process, each sample is randomly assigned the fields to be extracted.

[0035] Further, the specified hybrid information extraction task is specifically:

[0036] The input is an image containing text content and corresponding instructions, where the instructions require the multi-modal large model to output all specified entity information and specified table information in the picture;

[0037] The output includes:

[0038] The specified entity information on the image is output in json format;

[0039] The specified table information on the image is output in json format. The outermost layer of json is a list representing all table rows. Each element in the list is a row of the table, and all information in each row is represented by a dictionary.

[0040] Further, in the specified mixed information extraction task, during the training process, each sample is randomly assigned the fields to be extracted.

[0041] Further, the specific table question-answering task is as follows:

[0042] The input is an image containing text content and corresponding instructions, where the instructions require the multi-modal large model to answer the specified information according to the user's question;

[0043] The output is the corresponding specified information in the form of key-value pairs in json format.

[0044] According to the second aspect of the technical solution of the present invention, a graphic and text information extraction system based on a multi-modal large model is provided. The system includes: a processor and a memory for storing executable instructions; wherein, the processor is configured to execute the executable instructions to perform the graphic and text information extraction method based on the multi-modal large model as described in any of the above aspects.

[0045] According to the third aspect of the technical solution of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the graphic and text information extraction method based on the multi-modal large model as described in any of the above aspects.

[0046] The beneficial effects of the present invention:

[0047] 1. Solve the error propagation problem of the current industry mainstream cascade scheme. OCR and information extraction are completed end-to-end through a multi-modal large model, which can significantly improve the system stability and accuracy;

[0048] 2. Solve the information loss problem of the cascade scheme, retain the original visual information such as the relative position and layout of the text in the picture, and support information extraction in image scenarios with complex layouts;

[0049] 3. Based on the pre-trained multi-modal large model, the generalization ability for different samples and the understanding ability for different user questions are significantly enhanced compared with small models;

[0050] 4. Compared with the native pre-trained multi-modal large model, through specialized task design training for the field of graphic and text information extraction, the problems of low character recognition rate, weak instruction following ability, serious hallucination, and poor credibility are solved, and the ability of graphic and text information extraction is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0052] Figure 1 Shows a schematic diagram of the true output of the native multi-modal large model MiniCPM-V in the prior art.

[0053] Figure 2 Shows a flowchart of the method according to an embodiment of the technical solution of the present invention.

[0054] Figure 3 Shows a schematic diagram of the OCR basic ability task training paradigm according to an embodiment of the technical solution of the present invention.

[0055] Figure 4 Shows a schematic diagram of the designated entity information extraction task training paradigm according to an embodiment of the technical solution of the present invention.

[0056] Figure 5 Shows a schematic diagram of the designated table information extraction task training paradigm according to an embodiment of the technical solution of the present invention.

[0057] Figure 6 Shows a schematic diagram of the designated mixed information extraction task training paradigm according to an embodiment of the technical solution of the present invention.

[0058] Figure 7 Shows a schematic diagram of the table question answering task training paradigm according to an embodiment of the technical solution of the present invention.

[0059] Figure 8 Shows a schematic diagram of the user input according to an embodiment of the technical solution of the present invention.

[0060] Figure 9 Shows a schematic diagram of the output of the cascade scheme as a comparative example.

[0061] Figure 10 Shows a schematic diagram of the original output of the base model as a comparative example.

[0062] Figure 11 Shows a schematic diagram of the output according to an embodiment of the technical solution of the present invention.

[0063] The implementation, functional features, and advantages of the present invention will be further described in conjunction with embodiments and with reference to the accompanying drawings. Detailed implementation manners

[0064] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0065] The terms "first", "second", etc. in the specification and claims of the present disclosure are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented, for example, in an order other than those illustrated or described herein.

[0066] In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0067] Multiple, including two or more.

[0068] And / or, it should be understood that for the term "and / or" used in the present disclosure, it is merely an association relationship describing associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone.

[0069] The present invention relates to an end-to-end graphic and text information extraction method. For any formatted image, it can extract any text information specified by the user existing on the image, and has the characteristics of high generality, high robustness, high accuracy, and simple and efficient training and deployment.

[0070] The technical solution of the present invention first provides a graphic and text information extraction method based on a multimodal large model, including:

[0071] S1: Select a multimodal large model.

[0072] In a preferred embodiment, in the S1, the multimodal large model is the MiniCPM-V model.

[0073] S2: Conduct the OCR basic ability training task for the multi-modal large model.

[0074] In a preferred embodiment, in S2, the OCR basic ability training is specifically as follows:

[0075] The input is an image containing text content and a corresponding instruction, and the instruction is to require the multi-modal large model to output all the text in the picture;

[0076] The output is all the text on the image, and it is output line by line at the text line level.

[0077] S3: Conduct the multi-task hybrid and simultaneous graphic and text information extraction training task for the multi-modal large model trained in S2, where the training data is randomly selected from multiple tasks.

[0078] In a preferred embodiment, in S3, the multi-task hybrid graphic and text information extraction training includes:

[0079] Text recognition task, specified entity information extraction task, specified table information extraction task, specified hybrid information extraction task, and table question answering task.

[0080] In a preferred embodiment, the text recognition task is specifically as follows:

[0081] The input is an image containing text content and a corresponding instruction, and the instruction is to require the multi-modal large model to output all the text in the picture;

[0082] The output is all the text on the image, and it is output line by line at the text line level.

[0083] In a preferred embodiment, the specified entity information extraction task is specifically as follows:

[0084] The input is an image containing text content and a corresponding instruction, and the instruction is to require the multi-modal large model to output the specified entity information in the picture;

[0085] The output is the specified entity information on the image, and it is output in json format.

[0086] In a preferred embodiment, in the specified entity information extraction task, the fields to be extracted are randomly specified for each sample during the training process.

[0087] In a preferred embodiment, the entity information refers to the information represented in the form of key-value pairs.

[0088] In a preferred embodiment, the specified table information extraction task is specifically as follows:

[0089] The input is an image containing text content and corresponding instructions, where the instructions require the multimodal large model to output the specified table information in the picture;

[0090] The output is the specified table information on the image, which is output in JSON format. The outermost layer of the JSON is a list representing all table rows, and each element in the list is a row of the table. All information in each row is represented by a dictionary.

[0091] In a preferred embodiment, in the specified table information extraction task, during the training process, each sample is randomly assigned the fields to be extracted.

[0092] In a preferred embodiment, the specified mixed information extraction task is specifically as follows:

[0093] The input is an image containing text content and corresponding instructions, where the instructions require the multimodal large model to output all specified entity information and specified table information in the picture;

[0094] The output includes:

[0095] The specified entity information on the image, which is output in JSON format;

[0096] The specified table information on the image, which is output in JSON format. The outermost layer of the JSON is a list representing all table rows, and each element in the list is a row of the table. All information in each row is represented by a dictionary.

[0097] In a preferred embodiment, in the specified mixed information extraction task, during the training process, each sample is randomly assigned the fields to be extracted.

[0098] In a preferred embodiment, the table question - answering task is specifically as follows:

[0099] The input is an image containing text content and corresponding instructions, where the instructions require the multimodal large model to answer the specified information according to the user's question;

[0100] The output is the corresponding specified information in the form of key - value pairs in JSON format.

[0101] S4: Input the image to be processed into the multimodal large model trained in S3, and output the text - image information in the original image.

[0102] The technical solution of the present invention also provides a text - image information extraction system based on a multimodal large model. The system includes: a processor and a memory for storing executable instructions; wherein, the processor is configured to execute the executable instructions to perform the text - image information extraction method based on a multimodal large model as described in the above aspects.

[0103] The technical solution of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method for extracting graphic and text information based on a multi-modal large model as described in the above aspects.

[0104] Embodiment

[0105] In this embodiment, the open-source generative multi-modal large model MiniCPM-V is selected as the native base large model, and an OCR basic ability training task and five graphic and text information extraction training tasks are designed for multi-task instruction fine-tuning for the graphic and text information extraction task. The schematic flow diagram of the embodiment is as Figure 2 shown.

[0106] OCR Basic Ability Training Task

[0107] Since the current open-source multi-modal large models are positioned as general models and only involve a small number of text recognition tasks in the pre-training stage, and the number of Chinese characters is too large, the text character recognition error rate of the base model is high. The present invention specifically designs an OCR basic ability training task to improve the character recognition ability of the multi-modal large model for the characters on the picture.

[0108] Specifically, the training task is for the user to input a picture containing text content and an instruction requiring the large model to output all the text in the picture. The instruction only needs to clearly express the requirement, such as: "Extract all the text in this picture.", "Recognize and output all the text content on this picture.", etc. The answer given by the model is all the text on this image, output in line breaks at the text line level. Figure 3 This is the paradigm of this task.

[0109] After being specially trained by this task, the character recognition ability of the multi-modal large model for picture characters will be significantly improved.

[0110] Graphic and Text Information Extraction Training Task

[0111] The OCR basic ability training task in the previous stage enables the large model to have strong character recognition ability. In this stage, the graphic and text information extraction ability of the large model is trained on this basis. This training task consists of five sub-tasks: text recognition task, specified entity information extraction task, specified table information extraction task, specified mixed information extraction task, and table question answering task. This stage is multi-task mixed training, and during the training process, these five sub-tasks are trained simultaneously, and the training data for each step is randomly selected from these five sub-tasks.

[0112] The purpose of the mixed training of the five sub-tasks designed by the present invention is as follows:

[0113] 1. Improve the generalization ability of the model: By training multiple types of tasks simultaneously, the model can learn the common features among different tasks, while enhancing its adaptability to specific tasks, thus performing more stably and excellently in different scenarios.

[0114] 2. Accelerate model convergence: Since there may be specific common features among different tasks, the model can learn in a shared feature space, thereby reducing the time required for each task to be trained separately and accelerating the overall training convergence speed of the model.

[0115] 3. Strengthen the context understanding ability of the model: In the text and image information extraction task, the correlation among subtasks is relatively strong. For example, text recognition can provide a basis for the extraction of entity and table information, and table question answering can further utilize the extracted information. Multi-task mixed training can help the model understand the task context and its dependency relationships more comprehensively.

[0116] 4. Achieve collaborative optimization among tasks: Although the goals of different subtasks have different focuses, they all essentially need to extract meaningful information from text and images. Multi-task mixed training can promote collaborative optimization among tasks by sharing network structures and parameters, thereby achieving better comprehensive performance than independent training.

[0117] 5. Enhance the robustness of the model: Since data from different tasks are randomly sampled during training, the model can adapt to the distribution changes of different types of inputs, enhancing its ability to handle abnormal data or complex scenarios.

[0118] 6. Optimize resource utilization: Multi-task mixed training can reduce the overhead of single-task training by sharing computing resources. Especially when dealing with large models and large-scale data, it can utilize hardware resources more efficiently and reduce the overall cost.

[0119] Therefore, the technical effects of the five subtasks are not a simple superposition of the effects of the five subtasks respectively, and none of them can be missing. Through the training in this stage, the model will not only continue to maintain its advantage in character recognition, but also possess a powerful information extraction ability.

[0120] Text recognition task

[0121] The purpose of this training task is to prevent the forgetting of the OCR basic ability in the text and image information extraction training task. The training paradigm of this task is the same as that of the OCR basic ability training task.

[0122] The text recognition task is used to provide a text recognition basis for the entity and table information extraction tasks. If this task is not added in the text and image information extraction training task stage, it will lead to the forgetting of the large model's text recognition ability and the enhancement of hallucinations. The large model is prone to not following the original text content in the picture and tends to fabricate or tamper with the content.

[0123] Specified Entity Information Extraction Task

[0124] The instruction requirement of this training task is to let the large model extract the specified entity information in the picture. The specified entity information extraction task is one of the main application scenarios in the field of graphic and text information extraction. This task is used to enable the model to learn this ability. And the special training and learning of this task for entity information extraction will be beneficial to the "hybrid information extraction task".

[0125] The so-called entity information refers to the information that can be clearly expressed in the form of key-value pairs, such as name, account number, address, phone number, etc. The constructed model answer is in json format for flexible use. During the training process, each sample of this task is randomly assigned the fields to be extracted to avoid overfitting and increase generalization. Figure 4 This is the paradigm of this task.

[0126] Specified Table Information Extraction Task

[0127] The instruction requirement of this training task is to let the large model extract the specified table information in the picture. The specified table information extraction task is one of the main application scenarios in the field of graphic and text information extraction. This task is used to enable the model to learn this ability. And the special training and learning of this task for table information extraction will be beneficial to the "hybrid information extraction task".

[0128] The constructed model answer is in json format. The outermost layer of json is a list representing all table rows. Each element in the list is a row of the table, and all information in each row is represented by a dictionary. During the training process, each sample of this task is randomly assigned the table fields to be extracted to avoid overfitting and increase generalization. Figure 5 This is the paradigm of this task.

[0129] Specified Hybrid Information Extraction Task

[0130] The instruction requirement of this training task is to let the large model extract the specified entity and specified table information in the picture at the same time. The specified hybrid information extraction task is the most commonly used scenario in industrial applications in the field of graphic and text information extraction. This task is used to enable the model to learn this ability.

[0131] The model constructed for this task answers in json format, including entity and table row information, and the format is the same as that of Task 2 and Task 3. During the training process, each sample of this task is randomly assigned the entity and table fields to be extracted to avoid overfitting and increase generalization. Figure 6 This is the paradigm of this task.

[0132] Table Question Answering Task

[0133] The instruction requirement of this training task is to enable the large model to flexibly answer specified information according to the user's question. This task is used to enhance the model's understanding of user instructions and table structures, improve the model's understanding ability of instructions and picture tables, and can significantly improve the accuracy and robustness of entity, table, and mixed information extraction tasks.

[0134] The model answer constructed for this task is in JSON format, and the answers are directly given in key-value pairs. During the training process, the content to be extracted is randomly specified for each sample of this task to avoid overfitting and increase generalization. Figure 7 This is the paradigm of this task.

[0135] Effect demonstration

[0136] User input, such as Figure 8 shown.

[0137] The output result of the cascade scheme, such as Figure 9 shown.

[0138] Due to the loss of relative position information, the seller's address and phone number are output as the buyer's address and phone number, and only the first line is output in the password area.

[0139] The original output result of the base model, such as Figure 10 shown.

[0140] The output result of the minicpm-v base model before training and fine-tuning using the present invention: the buyer's taxpayer identification number result is incomplete, the seller's address and phone number are incorrect, the total amount of price and tax (in capital letters) is incorrect, the password area is not output, and only the first line of table information is output. Only the unit and tax rate in the first line are extracted correctly, and the rest are all incorrect.

[0141] The output result of the present invention, such as Figure 11 shown, and all the output results are correct.

[0142] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including that element.

[0143] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.

[0144] Through the description of the above embodiments, those skilled in the art can clearly understand that the above implementation methods can be realized by means of software plus a necessary general hardware platform. Of course, they can also be realized by hardware. However, in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0145] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these are within the protection scope of the present invention.

Claims

1. A method for extracting graphic information based on a multimodal large model, characterized in that: The graphic information extraction method comprises: S1: Select a large multimodal model; S2: Performing OCR basic capability training tasks on the multimodal large model; S3: For the multimodal large model trained in S2, a multi-task mixed and simultaneous image and text information extraction training task is performed, where the training data is randomly extracted from multiple tasks; S4: Input the image to be processed into the multimodal large model trained in S3, and output the graphic and text information in the original image.

2. The method for extracting graphic information according to claim 1, characterized in that: In S1, the multimodal large model is a MiniCPM-V model.

3. The method for extracting graphic information according to claim 1, characterized in that: In S2, the OCR basic ability training is specifically: The input is an image containing text content and a corresponding instruction, wherein the instruction requires the multimodal large model to output all the text in the image; The output is all the text on the image, wrapped at the text line level.

4. The method for extracting graphic information according to claim 1, characterized in that: In S3, the multi-task mixed image and text information extraction training includes: Text recognition task, specified entity information extraction task, specified table information extraction task, specified mixed information extraction task and table question answering task.

5. The method for extracting graphic information according to claim 4, characterized in that: The text recognition task is specifically: The input is an image containing text content and a corresponding instruction, wherein the instruction requires the multimodal large model to output all the text in the image; The output is all the text on the image, wrapped at the text line level.

6. The method for extracting graphic information according to claim 4, characterized in that: The specific task of extracting the specified entity information is as follows: The input is an image containing text content and a corresponding instruction, wherein the instruction requires the multimodal large model to output the specified entity information in the image; The output is the specified entity information on the image, output in json format.

7. The method for extracting graphic information according to claim 6, characterized in that: In the specified entity information extraction task, the field to be extracted is randomly specified for each sample during the training process.

8. The method for extracting graphic information according to claim 6, characterized in that: The entity information refers to information represented in a key-value pair format.

9. The method for extracting graphic information according to claim 1, characterized in that: The specific task of extracting information from the specified table is: The input is an image containing text content and a corresponding instruction, wherein the instruction requires the multimodal large model to output specified table information in the image; The output is the specified table information on the image, which is output in json format. The outermost layer of json is a list, representing all table rows. Each element in the list is a row of the table, and all information in each row is represented by a dictionary.

10. The method for extracting graphic information according to claim 9, characterized in that: In the specified table information extraction task, the field to be extracted is randomly specified for each sample during the training process.

11. The method for extracting graphic information according to claim 1, characterized in that: The specific mixed information extraction task is: The input is an image containing text content and a corresponding instruction, wherein the instruction requires the multimodal large model to output all specified entity information and specified table information in the image; The output includes: The specified entity information on the image is output in json format; The specified table information on the image is output in JSON format. The outermost layer of JSON is a list, representing all table rows. Each element in the list is a row of the table, and all information in each row is represented by a dictionary.

12. The method for extracting graphic information according to claim 11, characterized in that: In the specified mixed information extraction task, the field to be extracted is randomly specified for each sample during the training process.

13. The method for extracting graphic information according to claim 1, characterized in that: The table question-answering task is specifically as follows: The input is an image containing text content and a corresponding instruction, wherein the instruction requires the multimodal large model to answer specified information according to the user's question; The output is the corresponding specified information in the form of key-value pairs in JSON format.

14. A system for extracting text and image information based on a multimodal large model, the system comprising: A processor and a memory for storing executable instructions; characterized in that the processor is configured to execute the executable instructions to execute the method for extracting graphic and text information based on a multimodal large model according to any one of claims 1 to 13.

15. A computer-readable storage medium, wherein: A computer program is stored thereon, and when the computer program is executed by a processor, the method for extracting graphic and text information based on a multimodal large model according to any one of claims 1 to 13 is implemented.