Document processing method and system based on multi-modal large model

Through multimodal large model and optical character recognition technology PaddleOCR, the process of complex layout documents is generated, and Markdown documents are combined with RAG enhanced retrieval, which solves the problem of graphics and text separation and improves the accuracy and readability of document processing.

CN120340055APending Publication Date: 2025-07-18INSPUR ENTERPRISE CLOUD TECHNOLOGY (SHANDONG) CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510489653.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to achieve the complete separation of graphics and text when processing complex layout documents, resulting in low accuracy in document data processing.

Method used

The multimodal big model is used to convert the document into a target picture, and the structured text information is extracted using PaddleOCR, and the multimodal big model is used to generate a Markdown document with picture links and text summary. It combines RAG enhanced search and Elasticsearch for information retrieval to generate candidate answers that match user queries.

Benefits of technology

The separation of graphics and texts is realized, which significantly reduces the omission rate of key information, improves the accuracy and readability of document processing, and supports diverse query needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340055A_ABST
    Figure CN120340055A_ABST
Patent Text Reader

Abstract

The invention provides a document processing method and system based on a multi-modal large model. According to the method, firstly, a target type document is converted into a target picture based on a multi-modal large model according to a preset document segmentation rule, then structured text information is extracted from the target picture based on a preset cue word and by utilizing an advanced optical character recognition technology PaddleOCR technology, the multi-modal large model is guided to generate a text abstract for each target picture, and the text abstract is extracted from the target picture. And the integrated graphic and text information and the condensed picture content are converted into a Markdown document comprising a picture link of the target picture and text summary information. Through the steps of format recognition, image-text separation, content condensation and the like, a target type document is converted into a format which is easy to manage and retrieve, rapid retrieval of information is achieved by means of RAG retrieval enhancement, and documents and information related to user query can be rapidly found to serve as candidate answers. Compared with a traditional OCR technology, the method has the advantages that the missing rate of key information is remarkably reduced, and the accuracy of document processing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to a document processing method and system based on a multimodal large model. Background Art

[0002] In the current era of rapid development of digital information, documents, as a key carrier for information dissemination and storage, the innovation of their processing technology is crucial. With the continuous acceleration of the digital transformation of enterprises, the demand for processing complex layout documents is showing an increasing trend.

[0003] Currently, traditional document processing is to process documents through optical character recognition technology OCR, and convert the text in the documents into an editable text format for output.

[0004] However, when the prior art processes documents with complex layouts, it is difficult to completely separate the text and images, resulting in relatively low accuracy in document data processing. Summary of the Invention

[0005] An embodiment of the present invention provides a document processing method based on a multimodal large model, which can improve the accuracy of document data processing. The method includes:

[0006] A1: Receive a target type document from the outside, and based on the multimodal large model and a preset document segmentation rule, convert the target type document into a target picture, where the target type document is a document with text and image content;

[0007] A2: Based on a preset prompt, use the optical character recognition technology PaddleOCR to extract the structured text information of each target picture;

[0008] A3: Based on the multimodal large model, parse the target picture and the structured text information to generate a Markdown document including the picture link corresponding to the target picture and text summary information, where the text summary information is a general description of the structured text information;

[0009] A4: Based on RAG enhanced retrieval and the search engine Elasticsearch, perform a retrieval operation on an external knowledge base to generate candidate answers that match the user's query content, where the candidate answers are generated by performing a weighted algorithm operation on the Markdown document and the external knowledge base.

[0010] Preferably,

[0011] The A1 includes:

[0012] Receive a target type document from the outside, and determine the document format type of the target type document;

[0013] When it is determined that the document format type is PPTX, the target type document is parsed into single-page graphics and texts based on the multimodal large model;

[0014] When it is determined that the document format type is PDF, the target type document is parsed according to the configuration parameters of the preset document splitting rules to generate target pictures, where the configuration parameters include: chunk_size, chunk_overlap parameter, and special custom parameter.

[0015] Preferably,

[0016] After A3 and before A4, it further includes:

[0017] Using regular expressions, perform data cleaning operations on the Markdown document to generate a target Markdown document;

[0018] Based on the bge-m3 vector model of RAG enhanced retrieval, convert the content of the target Markdown document into high-dimensional vectors;

[0019] Store the high-dimensional vectors in the ChromaDB vector database.

[0020] Preferably,

[0021] A4 includes:

[0022] Determine at least one weighting factor, where the weighting factors include: relevance, authority, and timeliness of information;

[0023] Set corresponding weight values for each of the weighting factors;

[0024] According to the set weight values, the first formula and the second formula, based on the RAG enhanced retrieval and the search engine Elasticsearch, determine the comprehensive scores of the information in the ChromaDB vector database and the external knowledge base. The first formula is:

[0025] Score M =αX M +βY M +γZ M ;

[0026] Where M is the information in the ChromaDB vector database, Score M is the comprehensive score of the information in the ChromaDB vector database, α is the weight value of the relevance of information, β is the weight value of the authority of information, γ is the weight value of the timeliness of information, XM is the numerical value of the relevance of the ChromaDB vector database information, Y M is the numerical value of the authority of the ChromaDB vector database information, Z M is the numerical value of the timeliness of the ChromaDB vector database information;

[0027] The second formula is:

[0028] Score N = αX N + βY N + γZ N ;

[0029] where N is the information of the external knowledge base, and Score N is the comprehensive score of the information of the external knowledge base, X N is the numerical value of the relevance of the external knowledge base information, Y N is the numerical value of the authority of the external knowledge base information, Z N is the numerical value of the timeliness of the external knowledge base information.

[0030] Based on the calculated comprehensive scores of the information of the ChromaDB vector database and the external knowledge base, determine the information with the higher comprehensive score as the main content of the candidate answer.

[0031] In a second aspect, an embodiment of the present invention provides a document processing system based on a multimodal large model. The system includes:

[0032] A processing module: configured to receive a target type document from the outside, and convert the target type document into a target picture based on a multimodal large model and a preset document segmentation rule, where the target type document is a document with graphic and text content;

[0033] An extraction module: configured to extract structured text information of each of the target pictures of the processing module by using the optical character recognition technology PaddleOCR based on a preset prompt;

[0034] An analysis module: configured to analyze the target pictures of the processing module and the structured text information of the extraction module based on the multimodal large model, and generate a Markdown document including the picture link corresponding to the target picture and text summary information, where the text summary information is a general description of the structured text information;

[0035] Retrieval module: It is used to perform retrieval operations on the external knowledge base based on RAG enhanced retrieval and the search engine Elasticsearch, and generate candidate answers that match the user's query content. Among them, the candidate answers are generated by performing a weighted algorithm operation on the Markdown document of the parsing module and the external knowledge base.

[0036] Preferably,

[0037] The processing module is also used to execute:

[0038] Receive the target type document from the outside, and determine the document format type of the target type document;

[0039] When it is determined that the document format type is PPTX, parse the target type document into single-page pictures and texts based on the multi-modal large model;

[0040] When it is determined that the document format type is PDF, parse the target type document according to the configuration parameters of the preset document splitting rules to generate target pictures, where the configuration parameters include: chunk_size, chunk_overlap parameter, and special custom parameters.

[0041] Preferably,

[0042] After the parsing module and before the retrieval module, it further includes: a vector conversion module;

[0043] The vector conversion module is used to execute:

[0044] Use regular expressions to perform data cleaning operations on the Markdown document to generate a target Markdown document;

[0045] Based on the bge-m3 vector model of the RAG enhanced retrieval, convert the content of the target Markdown document into high-dimensional vectors;

[0046] Store the high-dimensional vectors in the ChromaDB vector database.

[0047] Preferably,

[0048] The retrieval module is also used to execute:

[0049] Determine at least one weighting factor, where the weighting factors include: relevance, authority, and timeliness of information;

[0050] Set corresponding weight values for each of the weighting factors;

[0051] Based on the set weight values, the first formula, and the second formula, and based on the RAG enhanced retrieval and the search engine Elasticsearch, determine the comprehensive score of the information in the ChromaDB vector database and the external knowledge base. The first formula is:

[0052] Score M = αX M + βY M + γZ M ;

[0053] where M is the information in the ChromaDB vector database, Score M is the comprehensive score of the information in the ChromaDB vector database, α is the weight value of the relevance of the information, β is the weight value of the authority of the information, γ is the weight value of the timeliness of the information, X M is the value of the relevance of the information in the ChromaDB vector database, Y M is the value of the authority of the information in the ChromaDB vector database, Z M is the value of the timeliness of the information in the ChromaDB vector database;

[0054] The second formula is:

[0055] Score N = αX N + βY N + γZ N ;

[0056] where N is the information in the external knowledge base, Score N is the comprehensive score of the information in the external knowledge base, X N is the value of the relevance of the information in the external knowledge base, Y N is the value of the authority of the information in the external knowledge base, Z N is the value of the timeliness of the information in the external knowledge base.

[0057] Based on the calculated comprehensive scores of the information in the ChromaDB vector database and the external knowledge base, determine the information with the higher comprehensive score as the main content of the candidate answer.

[0058] In a third aspect, an embodiment of the present invention provides a document processing system based on a multimodal large model, including: at least one memory and at least one processor;

[0059] The at least one memory is used to store machine-readable programs;

[0060] The at least one processor is configured to call the machine-readable program and execute any of the methods described in the first aspect.

[0061] In a fourth aspect, an embodiment of the present invention provides a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor is caused to execute any of the methods described in the first aspect.

[0062] An embodiment of the present invention provides a document processing method and system based on a multi-modal large model. The method first converts a target type document into target pictures based on a preset document segmentation rule and the multi-modal large model to achieve in-depth parsing of the text and picture mixed document. Then, guided by a preset prompt and using the advanced optical character recognition technology PaddleOCR technology, structured text information is extracted from the target pictures, and the multi-modal large model is guided to generate text summaries for each target picture. These text summaries will be effectively integrated with the structured text information extracted from the target pictures to provide a more refined content overview for the user. The integrated text and picture information and the condensed picture content will be converted into a Markdown document including the picture link of the target picture and the text summary information. During the parsing process, the picture and text information will be sorted and associated to clarify the corresponding relationship between the pictures and the text. The existence of the picture link enables the Markdown document to directly reference the corresponding pictures during display, facilitating the user to intuitively view the document content. At the same time, the structured text information is clearly and orderly presented in the form of text summaries, improving the readability and manageability of the document. Through the above steps of format recognition, text and picture separation, and content condensation, the target type document is converted into a format that is easy to manage and retrieve. With the help of RAG retrieval enhancement, information can be quickly retrieved, and relevant documents and information can be quickly found as candidate answers for the user's query. When the user submits a query content, the system will generate candidate answers that match the user's query content according to the Markdown document and the information retrieved from the external knowledge base using a weighted algorithm. Compared with traditional OCR technology, the key information omission rate is significantly reduced, and text and picture separation can be achieved, and the structured text information in the document can be completely extracted, thereby improving the accuracy of document processing. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0064] Figure 1It is a flowchart of a document processing method based on a multimodal large model provided by an embodiment of the present invention;

[0065] Figure 2 It is a flowchart of another document processing method based on a multimodal large model provided by an embodiment of the present invention;

[0066] Figure 3 It is a schematic diagram of a document processing system based on a multimodal large model provided by an embodiment of the present invention;

[0067] Figure 4 It is a schematic diagram of another document processing system based on a multimodal large model provided by an embodiment of the present invention. Detailed implementation manners

[0068] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0069] As Figure 1 shown, an embodiment of the present invention provides a document processing method based on a multimodal large model, and the method may include the following steps:

[0070] Step 101: Receive a target type document from the outside, and convert the target type document into a target picture based on the multimodal large model and a preset document segmentation rule, where the target type document is a document with graphic and text content;

[0071] Step 102: Extract the structured text information of each target picture using the optical character recognition technology PaddleOCR based on a preset prompt;

[0072] Step 103: Parse the target picture and the structured text information based on the multimodal large model to generate a Markdown document including the picture link corresponding to the target picture and the text summary information, where the text summary information is a general description of the structured text information;

[0073] Step 104: Perform a retrieval operation on an external knowledge base based on RAG enhanced retrieval and the search engine Elasticsearch to generate candidate answers matching the user's query content, where the candidate answers are generated by performing a weighted algorithm operation on the Markdown document and the external knowledge base.

[0074] In an embodiment of the present invention, a document processing method based on a multimodal large model is provided. First, according to the preset document segmentation rules, the target type document is converted into target pictures based on the multimodal large model to achieve in-depth parsing of the text-image mixed document. Then, guided by the preset prompt words and using the advanced optical character recognition technology PaddleOCR technology, structured text information is extracted from the target pictures, and the multimodal large model is guided to generate text summaries for each target picture. These text summaries will be effectively integrated with the structured text information extracted from the target pictures to provide a more refined content overview for the user. The integrated text-image information and the condensed picture content will be converted into a Markdown document including the picture link of the target picture and the text summary information. During the parsing process, the pictures and text information will be sorted out and associated to clarify the corresponding relationship between the pictures and the text. The existence of the picture link enables the Markdown document to directly reference the corresponding pictures during display, facilitating the user to intuitively view the document content. At the same time, the structured text information is presented clearly and orderly in the form of text summaries, improving the readability and manageability of the document. Through the above steps of format recognition, text-image separation, and content condensation, the target type document is converted into a format that is easy to manage and retrieve. With the help of RAG retrieval enhancement, information can be quickly retrieved, and relevant documents and information can be quickly found as candidate answers for the user's query. When the user submits a query content, the system will generate candidate answers that match the user's query content based on the Markdown document and the information retrieved from the external knowledge base using a weighted algorithm. Compared with the traditional OCR technology, the omission rate of key information is significantly reduced, and text-image separation can be achieved and the structured text information in the document can be completely extracted, thereby improving the accuracy of document processing.

[0075] In order to efficiently parse the text-image mixed document, in an embodiment of the present invention, step 101 in the above embodiment includes:

[0076] Receive the target type document from the outside and determine the document format type of the target type document;

[0077] When it is determined that the document format type is PPTX, parse the target type document into single-page text-images based on the multimodal large model;

[0078] When it is determined that the document format type is PDF, parse the target type document according to the configuration parameters of the preset document segmentation rules to generate target pictures, where the configuration parameters include: chunk_size, chunk_overlap parameter, and special custom parameters.

[0079] In the embodiments of the present invention, the multi-modal large model can automatically determine the document format type uploaded by the user. For PPTX files, they will be directly parsed into single-page graphics and texts, while for PDF files, a parameter setting interface is provided, allowing users to set parameters such as chunk_size, chunk_overlap, and special custom parameters, and select single-page splitting or default splitting methods according to these parameters to meet the processing requirements of different documents. This flexible and intelligent document splitting rule makes the parsing process more in line with user needs, ensures that the parsing result is closely matched with the original document format, completely solves the problem of format disconnection, maximally restores the original appearance of the document, improves the user's recognition and usability of the parsing result, and at the same time enables users with different professional levels to customize the parsing process according to their own needs.

[0080] To ensure the accuracy of information and improve the retrieval effect of information, in an embodiment of the present invention, after step 103 and before step 104 in the above embodiment, it further includes:

[0081] Using regular expressions, clean the data of the Markdown document to generate a target Markdown document;

[0082] Based on the bge-m3 vector model of RAG enhanced retrieval, convert the content of the target Markdown document into high-dimensional vectors;

[0083] Store the high-dimensional vectors in the ChromaDB vector database.

[0084] In the embodiments of the present invention, to ensure the accuracy and consistency of information, the system will use regular expressions to remove redundant information and unify the data format. At the same time, to improve the retrieval effect, the powerful bge-m3 vector model of RAG enhanced retrieval can be used to convert the document covering image links and text summaries into high-dimensional vectors, and then the ChromaDB vector database is used to store the parsed high-dimensional vectors. The ChromaDB vector database can achieve distributed storage. Using the sharding technology, the system can achieve the efficient storage of large-scale data, ensure the stability and scalability of the data, and at the same time support the dynamic addition and modification of document content, providing great convenience and flexibility for users. In addition to the traditional keyword retrieval, a composite retrieval method based on semantics, tags, and time can also be provided, so as to meet the diverse query needs of users.

[0085] To improve the accuracy of the retrieval result, in an embodiment of the present invention, step 104 in the above embodiment includes:

[0086] Determine at least one weighting factor, where the weighting factor includes: relevance, authority, and timeliness of information;

[0087] Set corresponding weight values for each of the said weighting factors;

[0088] Based on the set weight values, the first formula, and the second formula, and based on the RAG enhanced retrieval and the search engine Elasticsearch, determine the comprehensive scores of the information in the ChromaDB vector database and the external knowledge base. The first formula is:

[0089] Score M = αX M + βY M + γZ M ;

[0090] Where M is the information in the ChromaDB vector database, Score M is the comprehensive score of the information in the ChromaDB vector database, α is the weight value of the relevance of the information, β is the weight value of the authority of the information, γ is the weight value of the timeliness of the information, X M is the value of the relevance of the information in the ChromaDB vector database, Y M is the value of the authority of the information in the ChromaDB vector database, Z M is the value of the timeliness of the information in the ChromaDB vector database;

[0091] The second formula is:

[0092] Score N = αX N + βY N + γZ N ;

[0093] Where N is the information in the external knowledge base, Score N is the comprehensive score of the information in the external knowledge base, X N is the value of the relevance of the information in the external knowledge base, Y N is the value of the authority of the information in the external knowledge base, Z N is the value of the timeliness of the information in the external knowledge base.

[0094] Based on the calculated comprehensive scores of the information in the ChromaDB vector database and the external knowledge base, determine the information with the higher comprehensive score as the main content of the candidate answer.

[0095] In the embodiments of the present invention, when a user submits a query, the Retrieval-Augmented Generation (RAG) technology can effectively integrate the retrieved external knowledge with the content in the target Markdown documents stored in the ChromaDB vector database, and use a weighted algorithm to generate candidate answers that match the user's query, improving the comprehensiveness and accuracy of the answers. Elasticsearch is an efficient search engine that can quickly and accurately find relevant information from external knowledge bases. The weighted algorithm weights the information from different sources based on factors such as relevance, authority, and timeliness to ensure that the generated candidate answers can best meet the user's needs and output more comprehensive and accurate final results for the user. In the embodiments of the present invention, a user interface developed based on the React framework can provide the following functions for the user:

[0096] File management: Users can easily achieve batch uploading, previewing, and deleting of files, greatly improving work efficiency.

[0097] Result display: The system can display the parsed target Markdown documents and retrieval results in real time, including pictures and their condensed content. This intuitive display method greatly improves the user experience.

[0098] As Figure 2 shown, in order to more clearly illustrate the technical solutions and advantages of the present invention, the embodiments of the present invention provide a document processing method based on a multimodal large model, which will be described in detail below. Specifically, it may include the following steps:

[0099] Step 201: Receive a target type document from the outside and determine the document format type of the target type document;

[0100] Step 202: When it is determined that the document format type is PPTX, parse the target type document into single-page pictures and texts based on the multimodal large model;

[0101] Step 203: When it is determined that the document format type is PDF, parse the target type document according to the configuration parameters of the preset document splitting rules to generate target pictures, where the configuration parameters include: chunk_size, chunk_overlap parameter, and special custom parameters.

[0102] Step 204: Based on the preset prompt words, use the optical character recognition technology PaddleOCR to extract the structured text information of each target picture;

[0103] Specifically, using preset prompt words to guide the multimodal large model, two types of prompt words can be set. One type of prompt word is used to guide the multimodal large model to briefly analyze the content of the picture and generate a concise summary, and the other type of prompt word is used to guide it to deeply explore the details of the picture and generate a detailed explanation. Users can select the corresponding type of prompt word according to actual needs and integrate the structured text information extracted from the picture with the text summary generated by the multimodal large model.

[0104] Step 205: Based on the multimodal large model, parse the target picture and the structured text information to generate a Markdown document including the picture link corresponding to the target picture and the text summary information, where the text summary information is a general description of the structured text information;

[0105] Specifically, the system can be developed based on the Python language, use the FastAPI framework to build the backend service, and integrate the Inspur Cloud Sailing Multimodal Large Model and the ChromaDB vector database.

[0106] Step 206: Use regular expressions to perform data cleaning operations on the Markdown document to generate the target Markdown document;

[0107] Step 207: Based on the bge-m3 vector model of RAG enhanced retrieval, convert the content of the target Markdown document into high-dimensional vectors;

[0108] Step 208: Store the high-dimensional vectors in the ChromaDB vector database;

[0109] Step 209: Determine at least one weighting factor, where the weighting factors include: relevance, authority, and timeliness of the information;

[0110] Step 210: Set the corresponding weight values for each weighting factor;

[0111] Step 211: Based on the set weight values, the first formula and the second formula, determine the comprehensive scores of the information in the ChromaDB vector database and the external knowledge base based on RAG enhanced retrieval and the search engine Elasticsearch:

[0112] Specifically, the first formula is: Score M =αX M +βY M +γZ M ;

[0113] where M is the information in the ChromaDB vector database, Score MThe comprehensive score for the information of the ChromaDB vector database, α is the weight value of the relevance of the information, β is the weight value of the authority of the information, γ is the weight value of the timeliness of the information, and X M is the value of the relevance of the ChromaDB vector database information, and Y M is the value of the authority of the ChromaDB vector database information, and Z M is the value of the timeliness of the ChromaDB vector database information;

[0114] The second formula is:

[0115] Score N = αX N + βY N + γZ N ;

[0116] where N is the information of the external knowledge base, and Score N is the comprehensive score of the information of the external knowledge base, X N is the value of the relevance of the external knowledge base information, Y N is the value of the authority of the external knowledge base information, and Z N is the value of the timeliness of the external knowledge base information.

[0117] Step 212: According to the calculated comprehensive scores of the information of the ChromaDB vector database and the external knowledge base, determine the information with the higher comprehensive score as the main content of the candidate answer.

[0118] As Figure 3 shown, the embodiment of the present invention provides a document processing system based on a multimodal large model, and the system includes:

[0119] Processing module 301: configured to receive a target type document from the outside, and convert the target type document into a target picture based on a multimodal large model and a preset document segmentation rule, where the target type document is a document with graphic and text content;

[0120] Extraction module 302: configured to extract the structured text information of each of the target pictures of the processing module 301 by using the optical character recognition technology PaddleOCR based on a preset prompt word;

[0121] Analysis module 303: configured to analyze the target pictures of the processing module 301 and the structured text information of the extraction module 302 based on the multimodal large model, and generate a Markdown document including the picture link corresponding to the target picture and the text summary information, where the text summary information is a summary description of the structured text information;

[0122] Retrieval module 304: It is used to perform retrieval operations on the external knowledge base based on RAG enhanced retrieval and the search engine Elasticsearch, and generate candidate answers that match the user's query content. Among them, the candidate answers are generated by performing a weighted algorithm operation on the Markdown document of the parsing module 303 and the external knowledge base.

[0123] As Figure 3 shown, in the embodiment of the present invention, the processing module 301 is further configured to execute:

[0124] Receive a target type document from the outside, and determine the document format type of the target type document;

[0125] When it is determined that the document format type is PPTX, parse the target type document into single-page graphics and text based on the multimodal large model;

[0126] When it is determined that the document format type is PDF, parse the target type document according to the configuration parameters of the preset document splitting rules to generate target pictures, where the configuration parameters include: chunk_size, chunk_overlap parameter, and special custom parameters.

[0127] Based on Figure 3 shown, a document processing method based on a multimodal large model, as Figure 4 shown, after the parsing module 303 and before the retrieval module 304, it further includes: a vector conversion module 305;

[0128] The vector conversion module 305 is used to execute:

[0129] Use regular expressions to perform data cleaning operations on the Markdown document to generate a target Markdown document;

[0130] Based on the bge-m3 vector model of the RAG enhanced retrieval, convert the content of the target Markdown document into high-dimensional vectors;

[0131] Store the high-dimensional vectors in the ChromaDB vector database.

[0132] As Figure 4 shown, in the embodiment of the present invention, the retrieval module 304 is further configured to execute:

[0133] Determine at least one weighting factor, where the weighting factors include: relevance, authority, and timeliness of information;

[0134] Set corresponding weight values for each of the weighting factors;

[0135] Based on the set weight values, the first formula, and the second formula, and based on the RAG enhanced retrieval and the search engine Elasticsearch, determine the comprehensive score of the information in the ChromaDB vector database and the external knowledge base. The first formula is:

[0136] Score M = αX M + βY M + γZ M ;

[0137] where M is the information in the ChromaDB vector database, Score M is the comprehensive score of the information in the ChromaDB vector database, α is the weight value of the relevance of the information, β is the weight value of the authority of the information, γ is the weight value of the timeliness of the information, X M is the value of the relevance of the information in the ChromaDB vector database, Y M is the value of the authority of the information in the ChromaDB vector database, Z M is the value of the timeliness of the information in the ChromaDB vector database;

[0138] The second formula is:

[0139] Score N = αX N + βY N + γZ N ;

[0140] where N is the information in the external knowledge base, Score N is the comprehensive score of the information in the external knowledge base, X N is the value of the relevance of the information in the external knowledge base, Y N is the value of the authority of the information in the external knowledge base, Z N is the value of the timeliness of the information in the external knowledge base.

[0141] Based on the calculated comprehensive scores of the information in the ChromaDB vector database and the external knowledge base, determine the information with the higher comprehensive score as the main content of the candidate answer.

[0142] It can be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on an academic resource recommendation system based on a knowledge graph and a large model. In some other embodiments of the present invention, an academic resource recommendation system based on a knowledge graph and a large model may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0143] Regarding the information interaction, execution process, etc. among the various units within the above-mentioned device, since they are based on the same concept as the method embodiments of the present invention, the specific content can be referred to the description in the method embodiments of the present invention, and will not be elaborated here.

[0144] The embodiments of the present invention further provide a document processing system based on a multimodal large model, including: at least one memory and at least one processor;

[0145] At least one memory for storing machine-readable programs;

[0146] At least one processor for calling the machine-readable program to execute a document processing method based on a multimodal large model in any one of the embodiments of the present invention.

[0147] The embodiments of the present invention further provide a computer-readable medium, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor executes a document processing method based on a multimodal large model in any one of the embodiments of the present invention.

[0148] Specifically, a system or device equipped with a storage medium can be provided. On this storage medium, software program code for implementing the functions in any one of the above-mentioned embodiments is stored, and the computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.

[0149] In this case, the program code read from the storage medium itself can implement the functions of any one of the above-mentioned embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.

[0150] Embodiments of the storage medium for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer via a communication network.

[0151] In addition, it should be clear that not only can the above-described functions of any one of the above embodiments be implemented by executing the program code read by a computer, but also by the operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.

[0152] Furthermore, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion unit is made to execute part and all of the actual operations, so as to implement the functions of any one of the above embodiments.

[0153] Each embodiment of the present invention has at least the following beneficial effects:

[0154] 1. In the embodiment of the present invention, a document processing method based on a multi-modal large model is provided. The method first converts a target type document into a target picture based on a multi-modal large model according to a preset document segmentation rule to achieve in-depth parsing of a text and picture mixed document, and then based on the guiding effect of a preset prompt word and using the advanced optical character recognition technology PaddleOCR technology, structured text information is extracted from the target picture, and the multi-modal large model is guided to generate a text summary for each target picture. These text summaries will be effectively integrated with the structured text information extracted from the target picture to provide a more refined content overview for the user. The integrated text and picture information and the condensed picture content will be converted into a Markdown document including the picture link of the target picture and the text summary information. During the parsing process, the picture and text information will be sorted out and associated to clarify the corresponding relationship between the picture and the text. The existence of the picture link enables the Markdown document to directly reference the corresponding picture during display, facilitating the user to intuitively view the document content. At the same time, the structured text information is clearly and orderly presented in the form of a text summary, improving the readability and manageability of the document. Through the above steps of format recognition, text and picture separation, and content condensation, the target type document is converted into a format that is easy to manage and retrieve, and with the help of RAG retrieval enhancement, information can be quickly retrieved, and the document and information related to the user's query can be quickly found as candidate answers. When the user submits a query content, the system will generate candidate answers matching the user's query content according to the Markdown document and the information retrieved from the external knowledge base by using a weighted algorithm. Compared with the traditional OCR technology, the omission rate of key information is significantly reduced, and text and picture separation can be achieved and the structured text information in the document can be completely extracted, thus improving the accuracy of document processing;

[0155] 2. In the embodiments of the present invention, the multimodal large model can automatically determine the document format type uploaded by the user. For PPTX files, it will directly parse them into single-page pictures and texts. For PDF files, a parameter setting interface is provided, allowing users to set parameters such as chunk_size, chunk_overlap, and special custom parameters. According to these parameters, the user can choose single-page splitting or the default splitting method to meet the processing requirements of different documents. This flexible and intelligent document splitting rule makes the parsing process more in line with user needs, ensures that the parsing result is closely aligned with the original document format, completely solves the problem of format disconnection, maximally restores the original appearance of the document, improves the user's recognition and usability of the parsing result, and at the same time enables users with different professional levels to customize the parsing process according to their own needs;

[0156] 3. In the embodiments of the present invention, in order to ensure the accuracy and consistency of information, the system will use regular expressions to remove redundant information and unify the data format. At the same time, to improve the retrieval effect, the powerful bge-m3 vector model of RAG enhanced retrieval can be used to convert the document covering picture links and text summaries into high-dimensional vectors, and then the ChromaDB vector database is used to store the parsed high-dimensional vectors. The ChromaDB vector database can achieve distributed storage. Using the sharding technology, the system can efficiently store large-scale data, ensure the stability and scalability of the data, and at the same time support the dynamic addition and modification of document content, providing great convenience and flexibility for users. In addition to the traditional keyword retrieval, a composite retrieval method based on semantics, tags, and time can also be provided, so as to meet the diverse query needs of users.

[0157] It should be noted that not all steps and modules in the above-mentioned processes and system structure diagrams are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted according to needs. The system structure described in the above embodiments can be a physical structure or a logical structure. That is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities respectively, or some components in multiple independent devices may be jointly implemented.

[0158] In the above embodiments, the hardware unit may be implemented mechanically or electrically. For example, a hardware unit may include permanently dedicated circuits or logic (such as dedicated processors, FPGAs or ASICs) to perform corresponding operations. The hardware unit may also include programmable logic or circuits (such as general-purpose processors or other programmable processors), which can be temporarily set by software to perform corresponding operations. The specific implementation method (mechanical method, or dedicated permanent circuit, or temporarily set circuit) can be determined based on cost and time considerations.

[0159] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A document processing method based on a multimodal large model, characterized in that, The method includes: A1: Receive a target type document from the outside, and based on a multimodal large model and a preset document segmentation rule, convert the target type document into a target picture, where the target type document is a document with graphic and text content; A2: Based on a preset prompt, use the optical character recognition technology PaddleOCR to extract the structured text information of each target picture; A3: Based on the multimodal large model, parse the target picture and the structured text information to generate a Markdown document including the picture link corresponding to the target picture and text summary information, where the text summary information is a general description of the structured text information; A4: Based on RAG enhanced retrieval and the search engine Elasticsearch, perform a retrieval operation on an external knowledge base to generate candidate answers that match the user's query content, where the candidate answers are generated by performing a weighted algorithm operation on the Markdown document and the external knowledge base.

2. The method according to claim 1, wherein A1 includes: Receive a target type document from the outside and determine the document format type of the target type document; When it is determined that the document format type is PPTX, parse the target type document into single-page graphic and text based on the multimodal large model; When it is determined that the document format type is PDF, parse the target type document according to the configuration parameters of the preset document segmentation rule to generate a target picture, where the configuration parameters include: chunk_size, chunk_overlap parameter, and special custom parameter.

3. The method according to claim 1, wherein After A3 and before A4, it further includes: Use regular expressions to perform data cleaning operations on the Markdown document to generate a target Markdown document; Based on the bge-m3 vector model of the RAG enhanced retrieval, convert the content of the target Markdown document into a high-dimensional vector; Store the high-dimensional vector in the ChromaDB vector database.

4. The method according to claim 3, wherein A4 includes: Determine at least one weighting factor, where the weighting factors include: relevance, authority, and timeliness of information; Set corresponding weight values for each weighting factor; According to the set weight values, the first formula and the second formula, based on the RAG enhanced retrieval and the search engine Elasticsearch, determine the comprehensive scores of the information in the ChromaDB vector database and the external knowledge base. The first formula is: Score M = αX M + βY M + γZ M ; Among them, M is the information of the ChromaDB vector database, and Score M is the comprehensive score of the information of the ChromaDB vector database, α is the weight value of the relevance of the information, β is the weight value of the authority of the information, γ is the weight value of the timeliness of the information, and X M is the value of the relevance of the ChromaDB vector database information, and Y M is the value of the authority of the ChromaDB vector database information, and Z M is the value of the timeliness of the ChromaDB vector database information; The second formula is: Score N = αX N + βY N + γZ N ; Among them, N is the information of the external knowledge base, Score N is the comprehensive score of the information of the external knowledge base, X N is the numerical value of the relevance of the information of the external knowledge base, Y N is the numerical value of the authority of the information of the external knowledge base, Z N is the numerical value of the timeliness of the information of the external knowledge base. According to the calculated comprehensive scores of the information in the ChromaDB vector database and the external knowledge base, determine the information with a high comprehensive score as the main content of the candidate answer.

5. A document processing system based on a multimodal large model, characterized in that, The system includes: Processing module: used to receive a target type document from the outside, and based on a multimodal large model and a preset document segmentation rule, convert the target type document into a target picture, where the target type document is a document with graphic and text content; Extraction module: used to extract the structured text information of each target picture of the processing module by using the optical character recognition technology PaddleOCR; Parsing module: used to parse the target picture of the processing module and the structured text information of the extraction module based on the multimodal large model, and generate a Markdown document including the picture link corresponding to the target picture and the text summary information, where the text summary information is a general description of the structured text information; Retrieval module: used to perform a retrieval operation on an external knowledge base based on RAG enhanced retrieval and the search engine Elasticsearch, and generate candidate answers that match the user's query content, where the candidate answers are generated by performing a weighted algorithm operation on the Markdown document of the parsing module and the external knowledge base.

6. The system according to claim 5, wherein The processing module is further configured to execute: Receive a target type document from the outside and determine the document format type of the target type document; When it is determined that the document format type is PPTX, parse the target type document into single-page graphics and text based on the multimodal large model; When it is determined that the document format type is PDF, parse the target type document according to the configuration parameters of the preset document segmentation rule to generate a target picture, where the configuration parameters include: chunk_size, chunk_overlap parameter, and special custom parameter.

7. The method according to claim 1, wherein After the parsing module and before the retrieval module, a vector conversion module is further included; The vector conversion module is configured to execute: Use a regular expression to perform data cleaning on the Markdown document to generate a target Markdown document; Based on the bge-m3 vector model of the RAG enhanced retrieval, convert the content of the target Markdown document into a high-dimensional vector; Store the high-dimensional vector in the ChromaDB vector database.

8. The system according to claim 7, wherein The retrieval module is further configured to execute: Determine at least one weighting factor, where the weighting factors include: relevance, authority, and timeliness of information; Set corresponding weight values for each of the weighting factors; According to the set weight values, the first formula and the second formula, based on the RAG enhanced retrieval and the search engine Elasticsearch, determine the comprehensive score of the information in the ChromaDB vector database and the external knowledge base, and the first formula is: Score M = αX M + βY M + γZ M ; Among them, M is the information of the ChromaDB vector database, and Score M is the comprehensive score of the information of the ChromaDB vector database, α is the weight value of the relevance of the information, β is the weight value of the authority of the information, γ is the weight value of the timeliness of the information, and X M is the value of the relevance of the ChromaDB vector database information, and Y M is the value of the authority of the ChromaDB vector database information, and Z M is the value of the timeliness of the ChromaDB vector database information; The second formula is: Score N = αX N + βY N + γZ N ; Among them, N is the information of the external knowledge base, Score N is the comprehensive score of the information of the external knowledge base, X N is the numerical value of the relevance of the information of the external knowledge base, Y N is the numerical value of the authority of the information of the external knowledge base, Z N is the numerical value of the timeliness of the information of the external knowledge base. Based on the comprehensive scores of the information in the calculated ChromaDB vector database and the external knowledge base, determine the information with a high comprehensive score as the main content of the candidate answer.

9. A document processing system based on a multimodal large model, characterized in that, Including: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program and execute any one of the methods described in claims 1 to 4.

10. A computer-readable medium, characterized in that, Computer instructions are stored on the computer-readable medium, and when the computer instructions are executed by the processor, the processor executes any one of the methods described in claims 1 to 4.

Citation Information

Cited By

  • Document directory error repairing method and system based on multi-model collaboration

    CN120893402A

  • Interactive big and small model collaborative public service complex document analysis method and device

    CN121415423A

  • Concurrent conversion method for converting multi-type documents into target documents

    CN121636446A