Intelligent document analysis and question-answering method and system based on multi-modal large model
Through intelligent document analysis and question-and-answer methods based on multimodal large models, the problems of low document analysis efficiency and difficulty in information extraction in the existing technology are solved, and efficient analysis and question-and-answer interaction of documents are realized, improving user experience.
Patent Information
- Application Number
- CN202510083338.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
AI Technical Summary
Existing document management systems are difficult to meet complex query requirements and efficient information extraction. Traditional question-and-answer systems have limitations in document analysis, which consumes a lot of time and may not be able to parse.
Intelligent document analysis and question-answer methods based on multimodal large models are adopted to comprehensively identify and analyze documents through multimodal large models, structured data is generated, and a question-answer database is built for easy retrieval.
It realizes structured data acquisition for any document, improves the efficiency of document retrieval, and enhances the efficiency and user experience of Q&A interaction.
Smart Images

Figure CN119990096A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent document parsing and question-answering method and system based on a multimodal large model. Background Art
[0002] With the ever-increasing volume of internal enterprise data, documents in various formats, such as PDF and Word, have become fundamental to business operations. Employees frequently search and use these documents in their daily work, and their need for question-and-answer sessions regarding document content is increasing. Traditional document management systems primarily provide document storage and simple retrieval capabilities. However, faced with complex query requirements and constantly updated documents, existing technologies struggle to meet users' demands for efficient information extraction.
[0003] With the rapid development of natural language processing (NLP) technology, question-answering systems based on machine learning and deep learning are gradually being applied to various scenarios. However, existing question-answering systems still have certain limitations in document parsing. Not only do they consume a lot of time, but they may also fail to parse the document, thus affecting document-based question-answering interactions. Therefore, the present invention proposes an intelligent document parsing and question-answering method and system based on a multimodal large model. The multimodal large model is used to eliminate the limitations of document parsing, achieve comprehensive identification and parsing of documents, and enable structured data acquisition for any document, providing convenience for document retrieval. Therefore, when interacting with questions, document response information can be efficiently obtained in the document database, improving the efficiency of question-answering interactions and further enhancing the user's question-answering experience. Summary of the Invention
[0004] The purpose of the present invention is to provide an intelligent document parsing and question-answering method and system based on a multimodal large model to solve the problems raised in the above background technology.
[0005] To achieve the above objectives, the present invention provides the following technical solution: an intelligent document parsing and question-answering method based on a multimodal large model, comprising:
[0006] Acquire the document, and preprocess the document to obtain a preprocessed document;
[0007] Use the multimodal large model to parse the pre-processed documents and obtain document parsing information;
[0008] The document parsing information is stored to obtain a document database;
[0009] Use the question-answer generation model to perform semantic understanding on document parsing information and generate basic question-answer pairs;
[0010] The basic question and answer pairs are stored to obtain a question and answer database;
[0011] Obtain user questions and search the question-answering database or document database based on the user questions to obtain question answer information.
[0012] Furthermore, the multimodal large model is used to parse the preprocessed documents, including:
[0013] Detect and identify pre-processed documents to determine document layout and content;
[0014] Parsing is performed based on the document content. When the document content is text, a multimodal first parsing model is used to extract key information from the preprocessed document, and the document layout is combined to perform segmentation to obtain the first parsed information of the document. When the document content is a chart, a multimodal second parsing model is used to convert the chart information, and the converted information is formatted and segmented according to logical relationships to generate hierarchical structured data to obtain the second parsed information of the document.
[0015] Furthermore, a question-answering generation model is used to perform semantic understanding of document parsing information, including:
[0016] Obtain answers to frequently asked questions based on document parsing information, obtain answers to frequently asked questions, and match frequently asked questions with answers to obtain frequently asked questions and answers pairs;
[0017] Using the first pre-trained model to perform similarity analysis on the frequently asked question and answer pairs, identify similar questions, and obtain answers to similar questions in the document database based on the similar questions to obtain a first expanded question and answer pair;
[0018] Performing a correlation analysis on the frequently asked question and answer pairs using the second pre-trained model to determine the related questions, and obtaining answers to the related questions in the document database based on the related questions to obtain a second expanded question and answer pair;
[0019] The Embedding vector model is used to vectorize the common question and answer pairs, the first expanded question and answer pairs, and the second expanded question and answer pairs to obtain the basic question and answer pairs.
[0020] Furthermore, searching the question-answer database or document database based on the user's question includes:
[0021] Analyze the complexity of user problems to determine whether they are simple or complex, and obtain the analysis results.
[0022] Answer retrieval is performed based on the results of user question analysis. When the user question analysis result shows that the user question is a simple question, vectorization processing is performed on the user question to obtain the user question vector. Then, real-time retrieval is performed in the question and answer database based on the user question vector, and the similarity matching degree between the user question vector and the basic question and answer pair is analyzed. The best answer to the question is determined according to the similarity matching degree to obtain the question answer information. When the user question analysis result shows that the user question is a complex question, the RAG model is used to search in the document database based on the user question, filter the target data fragment, and then perform semantic analysis based on the target data fragment to determine the answer to the user question and obtain the question answer information.
[0023] Furthermore, after obtaining the question answer information, the user will provide feedback and evaluation on the user question and question answer information to obtain user question evaluation information, and then record the interaction based on the user question, question answer information and user question evaluation information. After the number of interactions meets the data flywheel mechanism, training data is obtained according to the training data screening rules for the interaction records, and model fusion training is performed based on the training data.
[0024] An intelligent document parsing and question-answering system based on a multimodal large model, comprising: a preprocessing module, a first building module, a second building module, and a question-answering interaction module;
[0025] The preprocessing module is used to obtain a document and perform preprocessing on the document to obtain a preprocessed document;
[0026] The first construction module is used to parse the pre-processed document using the multimodal large model to obtain document parsing information, and store the document parsing information to obtain a document database.
[0027] The second construction module is configured to use a question-answer generation model to perform semantic understanding on document parsing information, generate basic question-answer pairs, and store the basic question-answer pairs to obtain a question-answer database;
[0028] The question-answer interaction module is used to obtain user questions and search the question-answer database or document database based on the user questions to obtain question answer information.
[0029] Furthermore, the first building block includes: a detection and identification unit, a first processing unit, and a second processing unit;
[0030] The detection and recognition unit is used to detect and recognize the pre-processed document and determine the document layout and document content;
[0031] The first processing unit is configured to extract key information from the preprocessed document using a multimodal first parsing model when the document content is text content, and segment the document based on its layout to obtain first parsed information of the document;
[0032] The second processing unit is used to use the multimodal second parsing model to convert information on the chart when the document content is a chart, and to format and segment the converted information according to logical relationships to generate hierarchical structured data to obtain second parsed information of the document.
[0033] Furthermore, the second building module includes: a first generating unit, a second generating unit, a third generating unit and a vector processing unit;
[0034] The first generating unit is configured to obtain answers to frequently asked questions based on the document parsing information, obtain answers to frequently asked questions, and match the frequently asked questions with the answers to frequently asked questions to obtain frequently asked question-answer pairs;
[0035] The second generating unit is configured to perform similarity analysis on the frequently asked question and answer pairs using the first pre-trained model, determine similar questions, and obtain answers to similar questions in the document database based on the similar questions to obtain a first expanded question and answer pair;
[0036] The third generating unit is configured to perform a related question-answer analysis on the common question-answer pairs using the second pre-trained model, determine related questions, and obtain answers to the related questions in the document database based on the related questions to obtain second expanded question-answer pairs;
[0037] The vector processing unit is used to use the Embedding vector model to perform vectorization processing on the common question and answer pairs, the first expanded question and answer pairs, and the second expanded question and answer pairs to obtain basic question and answer pairs.
[0038] Furthermore, the question-answer interaction module includes: a question analysis unit, a first interaction unit and a second interaction unit;
[0039] The problem analysis unit is used to analyze the complexity of the user's problem, determine whether the user's problem is simple or complex, and obtain the user's problem analysis result;
[0040] The first interaction unit is configured to, when the user question analysis result indicates that the user question is a simple question, perform vectorization processing on the user question to obtain a user question vector, then perform real-time retrieval in the question-answer database based on the user question vector, analyze the similarity matching degree between the user question vector and the basic question-answer pair, and determine the best answer to the question based on the similarity matching degree to obtain question answer information;
[0041] The second interaction unit is used to use the RAG model to search in the document database according to the user question when the user question analysis result is that the user question is a complex question, filter the target data fragments, and then perform semantic analysis based on the target data fragments to determine the answer to the user question and obtain question response information.
[0042] Furthermore, the system also includes: an interaction recording module and a model training module;
[0043] The interaction recording module is used to provide feedback and evaluation on user questions and question answer information, obtain user question evaluation information, and then record the interaction based on the user questions, question answer information and user question evaluation information;
[0044] The model training module is used to obtain training data according to the training data screening rules for the interaction records after the number of interactions meets the data flywheel mechanism, and perform model fusion training based on the training data.
[0045] The present invention adopts a multimodal large model to implement document parsing, eliminating the limitations of document parsing, making it possible to obtain structured data for any document. At the same time, the document is formed into a document database with structured data, which provides convenience for document retrieval and enables document response information to be efficiently obtained in the document database when interacting with questions, thereby improving the efficiency of question-answering interaction and enhancing the user question-answering experience.
[0046] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.
[0047] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0049] Figure 1 A schematic diagram of the steps of the intelligent document parsing and question-answering method according to the present invention;
[0050] Figure 2 This is a schematic diagram of step 2 in the intelligent document parsing and question-answering method of the present invention;
[0051] Figure 3 This is a module diagram of the intelligent document parsing and question-answering system according to the present invention;
[0052] Figure 4 Schematic diagram of the first building block of the intelligent document parsing and question-answering system of the present invention;
[0053] Figure 5 Schematic diagram of the question answering module in the intelligent document parsing and question answering system according to the present invention. DETAILED DESCRIPTION
[0054] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0055] like Figure 1 As shown, an embodiment of the present invention provides an intelligent document parsing and question-answering method based on a multimodal large model, including:
[0056] Step 1: Obtain a document and preprocess it to obtain a preprocessed document;
[0057] Step 2: Use the multimodal large model to parse the pre-processed document to obtain document parsing information;
[0058] Step 3: Store the document parsing information to obtain a document database;
[0059] Step 4: Use the question-answer generation model to perform semantic understanding on the document parsing information and generate basic question-answer pairs;
[0060] Step 5: Store the basic question and answer pairs to obtain a question and answer database;
[0061] Step 6: Obtain the user's question and search the question-answering database or document database based on the user's question to obtain the answer information.
[0062] In the above technical solution, preprocessing is aimed at standardizing the document format, unifying documents of different formats into MarkDown format.
[0063] In the above technical solution, when obtaining a document, the document is uploaded to the target location.
[0064] In the above technical solution, the basic question-answer pair is a question-answer pair consisting of a common question and an answer corresponding to the common question based on the document parsing information.
[0065] The above technical solution adopts a multimodal large model to realize document parsing, eliminating the limitations of document parsing, so that structured data can be obtained for any document. At the same time, the document is formed into a document database with structured data, which provides convenience for document retrieval, so that when interacting with questions, document response information can be efficiently obtained in the document database, thereby improving the efficiency of question-answering interaction and enhancing the user's question-answering experience. By preprocessing the documents, the format of the documents is standardized, so that the preprocessed documents can be parsed by the multimodal large model, ensuring that the documents can be parsed by the multimodal large model, providing guarantee for the parsing of the multimodal large model, and using the multimodal large model to parse the preprocessed documents makes it possible to parse any document content, and structure the documents so that the document parsing information is structured and hierarchical, thereby enabling better document retrieval through document parsing information in the document database, providing convenience for document retrieval, and effectively improving the efficiency of document retrieval, thereby making it possible to consume less time when searching based on the document database according to user questions, improving the efficiency of determining question answer information, and by constructing a question and answer database, it is possible to quickly determine the question answer information for user questions, reducing the time for determining question answer information, and improving the efficiency of question and answer interaction.
[0066] like Figure 2 As shown, in one embodiment provided by the present invention, a multimodal large model is used to parse a pre-processed document, including:
[0067] S1. Detect and identify the pre-processed document to determine the document layout and document content;
[0068] S2. Parse the document according to its content. When the document content is text, use the multimodal first parsing model to extract key information from the preprocessed document, and segment it in combination with the document layout to obtain the first parsed information of the document. When the document content is a chart, use the multimodal second parsing model to convert the chart information, and format and segment the converted information according to the logical relationship to generate hierarchical structured data to obtain the second parsed information of the document.
[0069] In the above technical solution, the multimodal first parsing model uses technical solutions such as word embedding and semantic understanding to extract key information from preprocessed documents.
[0070] In the above technical solution, when segmenting in combination with the document layout, segmentation is performed according to structured information such as paragraphs, chapters, and titles.
[0071] In the above technical solution, when the multimodal second parsing model is used to convert information for a chart, the multimodal second parsing model is used to convert the chart into OCR text information and picture text description Description information.
[0072] In the above technical solution, the charts include: images, tables, charts, etc.
[0073] In the above technical solution, when the conversion information is formatted and segmented according to the logical relationship, it is segmented into text blocks Chunks of fixed size.
[0074] In the above technical solution, the multimodal large model includes: a multimodal first analytical model and a multimodal second analytical model.
[0075] The above technical solution realizes the structured processing of documents by parsing the pre-processed documents using a multimodal large model, so that the documents can exist in the form of structured data, which provides convenience for document retrieval, so that when determining the answer information based on the user's question, the answer to the question can be quickly obtained from the document database according to the structured information of the document. It can also ensure the consistency of the document parsing information in the document database, avoid the inability to retrieve due to different document contents, and thus the inability to obtain the answer information, ensuring the progress of question-answering interaction, and the multimodal parsing model adopts different multimodal parsing models for processing according to different document contents, avoiding the phenomenon that different document contents cannot be parsed, and ensuring that different document contents can be effectively parsed and stored.
[0076] In one embodiment provided by the present invention, a question-answer generation model is used to perform semantic understanding on document parsing information, including:
[0077] Obtain answers to frequently asked questions based on document parsing information, obtain answers to frequently asked questions, and match frequently asked questions with answers to obtain frequently asked questions and answers pairs;
[0078] Using the first pre-trained model to perform similarity analysis on the frequently asked question and answer pairs, identify similar questions, and obtain answers to similar questions in the document database based on the similar questions to obtain a first expanded question and answer pair;
[0079] Performing a correlation analysis on the frequently asked question and answer pairs using the second pre-trained model to determine the related questions, and obtaining answers to the related questions in the document database based on the related questions to obtain a second expanded question and answer pair;
[0080] The Embedding vector model is used to vectorize the common question and answer pairs, the first expanded question and answer pairs, and the second expanded question and answer pairs to obtain the basic question and answer pairs.
[0081] In the above technical solution, the common question and answer pairs, the first expanded question and answer pairs and the second expanded question and answer pairs after vectorization are stored to form a question and answer database.
[0082] In the above technical solution, the frequently asked question and answer pair, the first expanded question and answer pair, and the second expanded question and answer pair include questions and answers on important information in the document, frequently asked questions, and high-frequency terms.
[0083] In the above technical solution, the first pre-training model is an analysis model for similar problems, and the second pre-training model is an expansion model for related problems.
[0084] In the above technical solution, the Embedding model can be a Bert-like architecture model (such as BGE bge-large-zh-v1.5) or an improved Large Language Model (LLM) architecture dual-tower model (Qwen1.5 model).
[0085] The above technical solution obtains frequently asked question and answer pairs and expands them based on frequently asked question and answer pairs, so that the corresponding documents are summarized in advance, so that the question and answer database can contain answers to basic questions. Then, when determining question answer information based on user questions, it can directly search the question and answer database in real time, obtain question answer information in a relatively short time, and improve the interaction efficiency of user questions. The frequently asked question and answer pairs are expanded through the first pre-trained model and the second pre-trained model, so that the basic question and answer pairs contain more simple questions. In addition, through vectorization processing, calculation and analysis are performed according to vectors when similarity matching is performed, which facilitates the calculation of similarity and enables question answer information to be obtained in a relatively short time for user questions, thereby improving the efficiency of question interaction.
[0086] Furthermore, the basic question-answer pairs are stored, including: merging the basic question-answer pairs, integrating the common question-answer pairs, the first expanded question-answer pairs and the second expanded question-answer pairs to determine the information set to be stored; performing mapping relationship analysis on the information set to be stored according to the question-answer pairs to obtain a one-to-one question-and-answer mapping relationship; further analyzing the one-to-one question-and-answer mapping relationship to determine whether the same question or the same answer exists, and merging the one-to-one question-and-answer mapping relationships of the same question or the same answer to obtain a one-to-many or many-to-one question-and-answer mapping relationship; and performing structured storage in the storage unit according to the one-to-many or many-to-one question-and-answer mapping relationship to obtain a question-and-answer database.
[0087] The above technical solution uses mapping relationship analysis to enable basic question and answer pairs to be stored according to the mapping relationship, making the information in the question and answer database clearer and more convenient for subsequent retrieval in the question and answer database, thereby obtaining retrieval results in a shorter time. Moreover, through further analysis of the one-to-one question and answer mapping relationship, the mapping relationship is linked together, so that the influence relationship of the same answer or the same question is merged together. While ensuring the integrity of the stored information, it can not only reduce the storage space occupied, but also make the stored information clear and organized, reduce the degree of confusion in the question and answer database, and thus reduce the error probability of real-time retrieval in the question and answer database, and ensure the accuracy of the question response information.
[0088] In one embodiment of the present invention, searching a question-and-answer database or a document database based on a user question includes:
[0089] Analyze the complexity of user problems to determine whether they are simple or complex, and obtain the analysis results.
[0090] Answer retrieval is performed based on the results of user question analysis. When the user question analysis result shows that the user question is a simple question, vectorization processing is performed on the user question to obtain the user question vector. Then, real-time retrieval is performed in the question and answer database based on the user question vector, and the similarity matching degree between the user question vector and the basic question and answer pair is analyzed. The best answer to the question is determined according to the similarity matching degree to obtain the question answer information. When the user question analysis result shows that the user question is a complex question, the RAG model is used to search in the document database based on the user question, filter the target data fragment, and then perform semantic analysis based on the target data fragment to determine the answer to the user question and obtain the question answer information.
[0091] In the above technical solution, complexity analysis of user questions involves: optimizing the wording of the user question to remove interfering wording information in the question, thereby obtaining optimized user question information; performing semantic analysis on the optimized user question information and breaking it down based on the semantic information to determine the sentence structure of the user question; and performing sentence complexity analysis based on the sentence structure of the user question to determine whether the question is simple or complex, thereby obtaining the user question analysis result. Interfering wording information refers to expressions that have no practical meaning in the user question, such as "um," "ah," and "actually."
[0092] In the above technical solution, the best answer to the question is determined according to the similarity matching degree to obtain the question answer information, including: performing size analysis on the similarity matching degree, and arranging the sequence according to the size to determine the similarity matching degree analysis sequence; extracting the first two in the similarity matching degree analysis sequence in descending order to obtain the target analysis data; using the answer verification standard to verify the target analysis data to determine whether the target analysis data meets the answer verification standard to obtain the verification judgment result; when the verification judgment result is that the target analysis data all meet the answer verification standard, the corresponding basic question and answer pair is retrieved for the target analysis data to obtain the target basic question and answer pair, and then the answers to the target basic question and answer pair are compared and analyzed to see if they are the same. If they are the same, the question answer information is obtained based on the answer to the target basic question and answer pair. If they are different, the difference between the target analysis data is analyzed. When the difference is large, the larger difference is obtained for the target analysis data. The target analysis data is retrieved from the corresponding basic question and answer pair, and then the answer information is read from the basic question and answer pair to obtain the question answer information. When the difference is small, the association relationship analysis is performed on the answer of the target basic question and answer pair, and the information of the answer of the target basic question and answer pair is reorganized based on the association relationship to obtain the question answer information; when the inspection and judgment result is that there is a target analysis data that meets the answer inspection standard, the target analysis data that meets the answer inspection standard is screened out, and the corresponding basic question and answer pair is retrieved for the screened target analysis data, and then the answer information is read from the basic question and answer pair to obtain the question answer information; when the inspection and judgment result is that none of the target analysis data meets the answer inspection standard, the RAG model is used to search in the document database according to the user question, screen the target data fragment, and then determine the answer to the user question based on the target data fragment to obtain the question answer information.
[0093] In the above technical solution, when the RAG model is used to search in the document database according to user questions, the relevant document chunks Chunk are retrieved through multiple channels (BM25 retrieval and vector retrieval), and reranking Rerank is combined with the generation of the large model LLM.
[0094] The above technical solution realizes real-time responses to user questions and answers, so that question interaction can be carried out under any circumstances. In addition, question answer information is determined in different ways according to whether the user question is a simple question or a complex question, so that question answer information can be obtained quickly for simple user questions, and more complex and accurate answers can be generated in real time through the RAG model for more complex questions. It can significantly reduce the hallucination phenomenon of large models, so that question answer information can be obtained for complex user questions or complex documents, thereby improving the quality of question answer information.
[0095] In one embodiment provided by the present invention, after obtaining the question answer information, the user will provide feedback and evaluation on the user question and question answer information to obtain user question evaluation information, and then record the interaction based on the user question, question answer information and user question evaluation information. After the number of interactions meets the data flywheel mechanism, training data is obtained according to the training data screening rules for the interaction records, and model fusion training is performed based on the training data.
[0096] In the above technical solution, in the interaction record, classification records are made according to the user question evaluation information, and the types include: correct questions and answers and incorrect questions and answers.
[0097] In the above technical solution, a threshold of the number of interactions is pre-set in the data flywheel mechanism.
[0098] In the above technical solution, the training data screening rule refers to the data retrieval ratio of the correct question and answer set and the incorrect question and answer set when obtaining training data for interaction records.
[0099] In the above technical solution, when performing model fusion training based on training data, an efficient fine-tuning algorithm is used to fine-tune the answers to correct questions and answers, and DPO or ORPO is used to optimize based on incorrect questions and answers, thereby realizing model training.
[0100] The above technical solution obtains user question evaluation information to make it clear whether the user question answer information obtained based on the user question meets the user's needs, realizes the evaluation of question interaction, and enables the interaction record to be divided according to the user question evaluation information, providing convenience for the subsequent acquisition of training data, so that when conducting model training, data information of different evaluation types can be integrated for training, optimize the models involved in the question and answer process, improve the comprehensiveness of model training, and thereby improve the overall performance of the question and answer process and user satisfaction.
[0101] like Figure 3 As shown, an embodiment of the present invention provides an intelligent document parsing and question-answering system based on a multimodal large model, comprising: a preprocessing module, a first building module, a second building module, and a question-answering interaction module;
[0102] The preprocessing module is used to obtain a document and perform preprocessing on the document to obtain a preprocessed document;
[0103] The first construction module is used to parse the pre-processed document using the multimodal large model to obtain document parsing information, and store the document parsing information to obtain a document database.
[0104] The second construction module is configured to use a question-answer generation model to perform semantic understanding on document parsing information, generate basic question-answer pairs, and store the basic question-answer pairs to obtain a question-answer database;
[0105] The question-answer interaction module is used to obtain user questions and search the question-answer database or document database based on the user questions to obtain question answer information.
[0106] like Figure 4 As shown, in one embodiment provided by the present invention, the first building module includes: a detection and identification unit, a first processing unit and a second processing unit;
[0107] The detection and recognition unit is used to detect and recognize the pre-processed document and determine the document layout and document content;
[0108] The first processing unit is configured to extract key information from the preprocessed document using a multimodal first parsing model when the document content is text content, and segment the document based on its layout to obtain first parsed information of the document;
[0109] The second processing unit is used to use the multimodal second parsing model to convert information on the chart when the document content is a chart, and to format and segment the converted information according to logical relationships to generate hierarchical structured data to obtain second parsed information of the document.
[0110] In one embodiment provided by the present invention, the second building module includes: a first generating unit, a second generating unit, a third generating unit and a vector processing unit;
[0111] The first generating unit is configured to obtain answers to frequently asked questions based on the document parsing information, obtain answers to frequently asked questions, and match the frequently asked questions with the answers to frequently asked questions to obtain frequently asked question-answer pairs;
[0112] The second generating unit is configured to perform similarity analysis on the frequently asked question and answer pairs using the first pre-trained model, determine similar questions, and obtain answers to similar questions in the document database based on the similar questions to obtain a first expanded question and answer pair;
[0113] The third generating unit is configured to perform a related question-answer analysis on the common question-answer pairs using the second pre-trained model, determine related questions, and obtain answers to the related questions in the document database based on the related questions to obtain second expanded question-answer pairs;
[0114] The vector processing unit is used to use the Embedding vector model to perform vectorization processing on the common question and answer pairs, the first expanded question and answer pairs, and the second expanded question and answer pairs to obtain basic question and answer pairs.
[0115] like Figure 5 As shown, in one embodiment provided by the present invention, the question-answer interaction module includes: a question analysis unit, a first interaction unit and a second interaction unit;
[0116] The problem analysis unit is used to analyze the complexity of the user's problem, determine whether the user's problem is simple or complex, and obtain the user's problem analysis result;
[0117] The first interaction unit is configured to, when the user question analysis result indicates that the user question is a simple question, perform vectorization processing on the user question to obtain a user question vector, then perform real-time retrieval in the question-answer database based on the user question vector, analyze the similarity matching degree between the user question vector and the basic question-answer pair, and determine the best answer to the question based on the similarity matching degree to obtain question answer information;
[0118] The second interaction unit is used to use the RAG model to search in the document database according to the user question when the user question analysis result is that the user question is a complex question, filter the target data fragments, and then perform semantic analysis based on the target data fragments to determine the answer to the user question and obtain question response information.
[0119] In one embodiment provided by the present invention, the system further comprises: an interaction recording module and a model training module;
[0120] The interaction recording module is used to provide feedback and evaluation on user questions and question answer information, obtain user question evaluation information, and then record the interaction based on the user questions, question answer information and user question evaluation information;
[0121] The model training module is used to obtain training data according to the training data screening rules for the interaction records after the number of interactions meets the data flywheel mechanism, and perform model fusion training based on the training data.
[0122] The intelligent document parsing and question-answering system based on a multimodal large model corresponds one-to-one to the intelligent document parsing and question-answering method based on a multimodal large model. The technical effects of the intelligent document parsing and question-answering system based on a multimodal large model have been explained in the corresponding method embodiments and will not be repeated here.
[0123] Those skilled in the art should understand that the first, second and third in the present invention merely refer to different application stages.
[0124] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the disclosure herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow from the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0125] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An intelligent document parsing and question-answering method based on a multimodal large model, characterized in that: include: Acquire the document, and preprocess the document to obtain a preprocessed document; Use the multimodal large model to parse the preprocessed documents and obtain document parsing information; The document parsing information is stored to obtain a document database; The question-answer generation model is used to perform semantic understanding on document parsing information and generate basic question-answer pairs. The basic question and answer pairs are stored to obtain a question and answer database; Obtain user questions, and search in the question-answering database or document database based on the user questions to obtain question answer information.
2. The intelligent document parsing and question-answering method according to claim 1, characterized in that: Use a multimodal large model to parse preprocessed documents, including: Detect and identify pre-processed documents to determine document layout and document content; Parsing is performed according to the document content. When the document content is text content, a multimodal first parsing model is used to extract key information from the preprocessed document, and the document is segmented in combination with the document layout to obtain the first parsed information of the document. When the document content is a chart, a multimodal second parsing model is used to convert the chart information, and the converted information is formatted and segmented according to logical relationships to generate hierarchical structured data to obtain the second parsed information of the document.
3. The intelligent document parsing and question-answering method according to claim 1, characterized in that: The question-answer generation model is used to perform semantic understanding of document parsing information, including: Obtain answers to frequently asked questions based on document parsing information, obtain answers to frequently asked questions, and match frequently asked questions with answers to frequently asked questions to obtain frequently asked question-answer pairs; Perform similarity analysis on the common question-answer pairs using the first pre-trained model, determine similar questions, and obtain answers to similar questions in a document database based on the similar questions to obtain a first expanded question-answer pair; Performing a related question-answer analysis on the common question-answer pairs using the second pre-trained model to determine related questions, and obtaining answers to the related questions in the document database based on the related questions to obtain a second expanded question-answer pair; The Embedding vector model is used to vectorize the common question and answer pairs, the first expanded question and answer pairs, and the second expanded question and answer pairs to obtain basic question and answer pairs.
4. The intelligent document parsing and question-answering method according to claim 1, characterized in that: Search in the question-answering database or document database based on the user's question, including: Perform complexity analysis on user problems to determine whether the user problems are simple or complex, and obtain the user problem analysis results; Answer retrieval is performed based on the results of user question analysis. When the result of user question analysis is that the user question is a simple question, vectorization processing is performed on the user question to obtain the user question vector, and then real-time retrieval is performed in the question and answer database based on the user question vector, and the similarity matching degree between the user question vector and the basic question and answer pair is analyzed, and the best answer to the question is determined according to the similarity matching degree to obtain question answer information; when the result of user question analysis is that the user question is a complex question, the RAG model is used to search in the document database based on the user question, filter the target data fragment, and then perform semantic analysis based on the target data fragment to determine the answer to the user question and obtain question answer information.
5. The intelligent document parsing and question-answering method according to claim 4, characterized in that: After obtaining the question answer information, the user provides feedback and evaluation on the user question and question answer information to obtain user question evaluation information, and then records the interaction based on the user question, question answer information and user question evaluation information. After the number of interactions meets the data flywheel mechanism, training data is obtained according to the training data screening rules for the interaction records, and model fusion training is performed based on the training data.
6. An intelligent document parsing and question-answering system based on a multimodal large model, characterized in that: The system comprises: a preprocessing module, a first building module, a second building module and a question-answering interaction module; The preprocessing module is used to obtain a document and preprocess the document to obtain a preprocessed document; The first construction module is used to parse the pre-processed document using the multimodal large model to obtain document parsing information, and store the document parsing information to obtain a document database. The second construction module is used to use the question-answer generation model to perform semantic understanding on the document parsing information, generate basic question-answer pairs, and store the basic question-answer pairs to obtain a question-answer database; The question-answer interaction module is used to obtain user questions and search in a question-answer database or a document database according to the user questions to obtain question answer information.
7. The intelligent document parsing and question-answering system according to claim 6, characterized in that: The first building module includes: a detection and identification unit, a first processing unit and a second processing unit; The detection and identification unit is used to detect and identify the pre-processed document to determine the document layout and document content; The first processing unit is used to extract key information from the preprocessed document using the multimodal first parsing model when the document content is text content, and segment the document in combination with the document layout to obtain first parsed information of the document; The second processing unit is used to use the multimodal second parsing model to convert information on the chart when the document content is a chart, and to format and segment the converted information according to logical relationships to generate hierarchical structured data to obtain second parsed information of the document.
8. The intelligent document parsing and question-answering system according to claim 6, characterized in that: The second building module includes: a first generating unit, a second generating unit, a third generating unit and a vector processing unit; The first generating unit is used to obtain answers to frequently asked questions based on the document parsing information, obtain answers to frequently asked questions, and match the frequently asked questions with the answers to frequently asked questions to obtain frequently asked question-answer pairs; The second generating unit is used to perform similarity analysis on the common question-answer pairs using the first pre-trained model, determine similar questions, and obtain answers to similar questions in the document database based on the similar questions to obtain a first expanded question-answer pair; The third generating unit is used to perform related question and answer analysis on the common question and answer pairs using the second pre-trained model, determine related questions, and obtain answers to the related questions in the document database based on the related questions to obtain second expanded question and answer pairs; The vector processing unit is used to use the Embedding vector model to perform vectorization processing on the common question and answer pairs, the first expanded question and answer pairs, and the second expanded question and answer pairs to obtain basic question and answer pairs.
9. The intelligent document parsing and question-answering system according to claim 6, characterized in that: The question-answer interaction module includes: a question analysis unit, a first interaction unit and a second interaction unit; The problem analysis unit is used to analyze the complexity of the user's problem, determine whether the user's problem is a simple problem or a complex problem, and obtain the user's problem analysis result; The first interaction unit is used to, when the user question analysis result is that the user question is a simple question, perform vectorization processing on the user question to obtain the user question vector, and then perform real-time retrieval in the question-answer database according to the user question vector, analyze the similarity matching degree between the user question vector and the basic question-answer pair, and determine the best answer to the question according to the similarity matching degree to obtain the question answer information; The second interaction unit is used to use the RAG model to search in the document database according to the user question, filter the target data fragments, and then perform semantic analysis based on the target data fragments to determine the answer to the user question and obtain the question answer information when the user question analysis result is that the user question is a complex question.
10. The intelligent document parsing and question-answering system according to claim 9, characterized in that: The system also includes: an interaction recording module and a model training module; The interaction recording module is used to provide feedback and evaluation on user questions and question answering information, obtain user question evaluation information, and then perform interaction recording based on user questions, question answering information and user question evaluation information; The model training module is used to obtain training data according to the training data screening rules for the interaction records after the number of interactions meets the data flywheel mechanism, and perform model fusion training based on the training data.
Citation Information
Cited By
Multi-modal document information extraction method, system and equipment and storage device
CN121766326A