Document question and answer method and system based on multi-modal primitives, terminal and medium

By performing layout analysis on the document, multimodal primitives are extracted and correlation mapping tables are constructed, the problem of large language models dealing with multimodal data in document question and answer system is solved, and the accuracy and interpretability of the answers are improved.

CN119917686APending Publication Date: 2025-05-02SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411674419.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-21
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

Large language models can create hallucinations in document Q&A systems, lack interpretability, and difficulty in effectively handling multimodal data, especially in understanding and integrating chart and image information in documents.

Method used

By performing layout analysis on the documents uploaded by users, multimodal primitives such as titles, paragraphs, images and tables are extracted, and a correlation mapping table is constructed. Store text vectors into vector databases, obtain user questions and convert them into query vectors, and answer questions based on multimodal information to improve the accuracy and interpretability of answers.

Benefits of technology

The accuracy and interpretability of model output are improved. Through the processing of multimodal information, key information in the document can be extracted more accurately and rich context information can be provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917686A_ABST
    Figure CN119917686A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of large models, and particularly relates to a document question and answer method and system based on multi-modal primitives, a terminal and a medium, and the method comprises the steps: extracting title primitives, paragraph primitives, image primitives and table primitives of a plurality of documents uploaded by a user; constructing a correlation mapping table of the paragraph primitives, the image primitives and the table primitives; converting the primitives of the text modality into text vectors for storage; converting the user question into a query vector; screening out a target title primitive and a target paragraph primitive according to the query vector, and obtaining a related target image primitive and a target table primitive; and constructing the target title primitive, the target paragraph primitive, the target image primitive, the target table primitive and the user question into a cue word, inputting the cue word into a multi-modal large language model for processing, and outputting a question result. Question answering is carried out based on the multi-modal information, and the accuracy and interpretability of the question result output by the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of large models, and specifically relates to a document question-and-answer method, system, terminal and medium based on multimodal primitives. Background Art

[0002] The large language model has strong semantic understanding capabilities through training on large data sets, enabling the system to understand and answer business questions in the form of natural language. The application of large language models, especially in the retrieval-enhanced generation framework, greatly improves the processing capabilities of knowledge-intensive tasks such as question answering, text summarization, and content generation by combining information retrieval technology and language generation models. By retrieving relevant information from external knowledge bases and feeding it as context to the large language model, the model's ability to generate accurate and rich text content is enhanced.

[0003] Although large language models have performed well in document question answering systems, they also have some challenges and limitations. On the one hand, generative models may produce hallucinations, that is, generate inaccurate or fictitious information, especially perform poorly on arithmetic tasks, and lack interpretability. In addition, large models have difficulties in processing multimodal data, especially in understanding and integrating charts and images in documents. Summary of the invention

[0004] To solve the above problems, the present invention provides a document question and answer method, system, terminal and medium based on multimodal primitives, which conducts in-depth analysis on documents uploaded by users, extracts multimodal key information including text and images, retrieves multimodal information according to user questions, and then answers questions based on the multimodal information, thereby improving the accuracy and interpretability of the question results output by the model.

[0005] In a first aspect, the technical solution of the present invention provides a document question-answering method based on multimodal primitives, comprising the following steps: Perform layout analysis on several documents uploaded by users to extract target primitives of each document, wherein the target primitives include title primitives, paragraph primitives, image primitives and table primitives; Obtain the correlation between the paragraph primitive and the image primitive and the table primitive, and construct a correlation mapping table, and add the mapping relationship between the paragraph primitive and the related image primitive and table primitive in the correlation mapping table; Convert the title primitive and paragraph primitive of the text mode into text vectors through a vectorization model, and store the text vectors in a vector database; Obtain user questions and convert them into query vectors through a vectorization model; Retrieve relevant text vectors in the vector database according to the query vector, filter out the target title primitive and the target paragraph primitive, and obtain the target image primitive and the target table primitive related to the target paragraph primitive according to the correlation mapping table; The target title primitive, target paragraph primitive, target image primitive, target table primitive and user question are constructed into prompt words, the prompt words are input into the multimodal large language model for processing, and the question result is output.

[0006] In an optional implementation, the layout analysis of several documents uploaded by the user is performed to extract the target primitives of each document, specifically including: Render the document as an image; Inputting the document image into the layout analysis model to divide the document into headings, paragraphs, tables, and image areas; For the title and paragraph regions, the OCR model is used to extract text information and classify the text information into corresponding title primitives and paragraph primitives respectively; Screenshot table and image areas, as table primitive and image primitive respectively.

[0007] In an optional implementation, obtaining the correlation between the paragraph primitive and the image primitive and the table primitive specifically includes: Extract characters from paragraph primitives, image primitives, and table primitives; Compare the characters in the paragraph primitive with the title characters in the image primitive and the table primitive respectively; If the title of a paragraph primitive and an image primitive contain the same character, then the image primitive is related to the paragraph primitive; If a paragraph primitive and a table primitive's title contain the same characters, then the table primitive is related to the paragraph primitive.

[0008] In an optional implementation, storing the text vector in a vector database specifically includes: Dividing the target primitive into a plurality of primitive blocks based on the paragraph block measurement; Several text vectors corresponding to each primitive block are stored in the vector database as a storage unit.

[0009] In an optional implementation, searching for relevant text vectors in a vector database according to the query vector to screen out target title primitives and target paragraph primitives specifically includes: Through parallel processing, the query vector and the text vectors in each current storage unit are searched for similarity, and the cosine similarity is used as the measurement standard for similarity search; Obtain the N text vectors with the highest cosine similarity as the final target text vector; where N ≥ 1; The target title primitive and the target paragraph primitive are obtained according to the target text vector.

[0010] In an optional implementation, after performing layout analysis on several documents uploaded by the user to extract target primitives from each document, the following steps are also included: A position mapping table is constructed, and a mapping relationship between each target primitive and its position in the document is added to the position mapping table; the position of the target primitive in the document includes a page number and a region coordinate.

[0011] In an optional implementation, after outputting the question result, the following steps are further included: Obtaining the positions of the title primitive, the target paragraph primitive, the target image primitive, and the target table primitive in the document according to the position mapping table; Jumps in the user interface based on the acquired location and highlights relevant areas in the document.

[0012] In a second aspect, the technical solution of the present invention provides a document question-answering system based on multimodal primitives, comprising: A primitive extraction module, used to analyze the layout of several documents uploaded by users and extract target primitives of each document, wherein the target primitives include title primitives, paragraph primitives, image primitives and table primitives; A correlation building module is used to obtain the correlation between the paragraph primitive and the image primitive and the table primitive, and to build a correlation mapping table, in which a mapping relationship between the paragraph primitive and the related image primitive and table primitive is added; A vector storage module, used for converting the title primitive and paragraph primitive of the text mode into text vectors through a vectorization model, and storing the text vectors in a vector database; The query vector acquisition module is used to obtain user questions and convert them into query vectors through a vectorization model; A similar primitive query module is used to retrieve relevant text vectors in the vector database according to the query vector, filter out the target title primitive and the target paragraph primitive, and obtain the target image primitive and the target table primitive related to the target paragraph primitive according to the correlation mapping table; The question result output module is used to construct the target title primitive, the target paragraph primitive, the target image primitive, the target table primitive and the user question into prompt words, input the prompt words into the multimodal large language model for processing, and output the question result.

[0013] In a third aspect, the technical solution of the present invention provides a terminal, including: A memory for storing a document question-answering program based on multimodal primitives; A processor is used to implement the steps of the document question and answer method based on multimodal primitives as described in any of the above items when executing the document question and answer program based on multimodal primitives.

[0014] In a fourth aspect, the technical solution of the present invention provides a computer-readable storage medium, on which a document question and answer program based on multimodal primitives is stored. When the document question and answer program based on multimodal primitives is executed by a processor, the steps of the document question and answer method based on multimodal primitives as described in any of the above items are implemented.

[0015] The document question-and-answer method, system, terminal and medium based on multimodal primitives provided by the present invention have the following beneficial effects compared with the prior art: first, through layout analysis, key information in the document, including titles, paragraphs, images and tables, is accurately extracted to provide rich contextual information for subsequent questions and answers, thereby improving the accuracy of the answers; secondly, by constructing the correlation between paragraph primitives and image primitives and table primitives, image information is retrieved, and then the information of text modality and image modality is processed through a multimodal large language model, and question results are output according to the multimodal information, thereby improving the accuracy and interpretability of the question results output by the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions of the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0017] Figure 1 It is a flowchart of a document question-and-answer method based on multimodal primitives provided in an embodiment of the present invention.

[0018] Figure 2 It is a schematic block diagram of the structure of a document question-answering system based on multimodal primitives provided by an embodiment of the present invention.

[0019] Figure 3 It is a schematic diagram of the structure of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0020] In order to enable those skilled in the art to better understand the scheme of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0022] Figure 1 : is a flowchart of a document question-answering method based on multimodal primitives provided by an embodiment of the present invention. Figure 1 The execution subject may be a document question-and-answer system based on multimodal primitives. The document question-and-answer method based on multimodal primitives provided in the embodiment of the present invention is executed by a computer device, and accordingly, the document question-and-answer system based on multimodal primitives runs in the computer device. According to different requirements, the order of the steps in the flowchart can be changed, and some can be omitted.

[0023] like Figure 1 As shown, the method includes the following steps.

[0024] S1, performing layout analysis on several documents uploaded by users to extract target primitives of each document, wherein the target primitives include title primitives, paragraph primitives, image primitives and table primitives.

[0025] In this embodiment, the document uploaded by the user is first analyzed for layout, so as to automatically identify and extract key primitives such as titles, paragraphs, images and tables in the document. This process is implemented through a deep learning model, so that the system can efficiently process various complex document types, including those with complex layouts and table structures. Specifically, the following steps are included.

[0026] S1.1, render the document into an image form.

[0027] For example, in the initial stage of layout analysis, the document uploaded by the user is converted into a cv::Mat object of OpenCV. This step involves rendering the document page into an image, providing basic data for the subsequent deep learning model to further analyze and process the document.

[0028] S1.2, input the document image into the layout analysis model to divide the document into title, paragraph, table and image areas.

[0029] S1.3, for the title and paragraph areas, use the OCR model to extract text information, and classify the text information into corresponding title primitives and paragraph primitives respectively.

[0030] S1.4, screenshot table and image regions, as table primitive and image primitive, respectively.

[0031] After the document image is input into the layout analysis model, the system divides the document into different areas such as titles, paragraphs, tables and images. This process is achieved using deep learning technology. Specifically, a convolutional neural network (CNN) can be used to pre-train the layout analysis model to identify and distinguish different elements in the document. For titles and paragraphs, the OCR model is used to extract text information and classify this information into title primitives and paragraph primitives. For images and tables, the system captures the corresponding image areas as image primitives and table primitives, respectively.

[0032] When using large models for intelligent question answering, the generated question answering results often lack source information, resulting in incomplete results and uncertainty of the source. After extracting the target primitives of each document, this embodiment also establishes a location mapping table of primitives and document locations. This location mapping table adds a mapping relationship between each target primitive and its location in the document, which is used to record the specific location of each primitive in the original document, including page numbers and area coordinates. Based on this location mapping table, document traceability is achieved, allowing the relevant area in the document to be accurately located when answering user questions. S2, obtaining the correlation between the paragraph primitive and the image primitive and the table primitive, and constructing a correlation mapping table, and adding the mapping relationship between the paragraph primitive and the related image primitive and table primitive in the correlation mapping table.

[0033] In this embodiment, after layout analysis and OCR recognition, the title primitive, paragraph primitive, image primitive and table primitive in the document are successfully extracted. This embodiment further analyzes these primitives, especially the paragraph primitive and the image primitive and table primitive, to determine the correlation between them.

[0034] For image and table primitives, they may not be directly related to the previous and next paragraphs due to document layout. Therefore, this embodiment performs in-depth character analysis on the paragraph primitive and compares the extracted characters with the characters in the chart. For example, if both contain " Figure 1 character", it can be determined that the paragraph primitive is correlated with the image primitive. Specifically, the following steps are included.

[0035] S2.1, extract characters from paragraph primitives, image primitives, and table primitives.

[0036] S2.2, compare the characters in the paragraph primitive with the title characters in the image primitive and the table primitive respectively.

[0037] S2.3, if a paragraph primitive and a title of an image primitive contain the same character, then the image primitive is related to the paragraph primitive.

[0038] S2.4, if a paragraph primitive and a table primitive contain the same character in their title, then the table primitive is related to the paragraph primitive.

[0039] It should be noted that, in this embodiment, primitives of image and table image modes are stored as image files, and a mapping relationship with the corresponding paragraph primitives is established. This mapping relationship allows the system to retrieve the corresponding paragraph primitives according to the user's questions and extract the chart image information contained therein at the same time. This multimodal data processing method enables the system to make full use of image and text information and use a multimodal large model to improve the accuracy of the question-answering system.

[0040] In an optional implementation, after the correlation between primitives is determined, the primitives are sorted according to primitive analysis results and location information of the primitives in the document to ensure that the logical order of the primitives matches the physical layout of the document.

[0041] S3, converts the title primitive and paragraph primitive of the text modality into text vectors through a vectorization model, and stores the text vectors in the vector database.

[0042] For example, for primitives of text modality, such as titles and paragraphs, the bge-m3 vectorization model is used for processing. This vectorization model can convert text into vectors in a high-dimensional space, which can capture the semantic information of the text. In this way, text primitives are converted into vector form and stored in the Milvus vector database. Milvus is a high-performance, highly scalable vector database suitable for processing large-scale vector data, making the storage and retrieval of text primitives more efficient.

[0043] Considering that large language models usually have context length limitations, directly processing the entire document may result in exceeding the processing capacity of the model, thereby affecting the performance of the answer. Therefore, this embodiment processes the primitives in blocks. This process not only helps to comply with the input limitations of the model, but also improves retrieval accuracy. This embodiment adopts a paragraph-based blocking strategy because the layout analysis has identified different paragraphs, each of which contains semantically different information. This blocking strategy allows each primitive block to maintain semantic integrity. By processing the primitives in blocks, the system can use GPU parallel retrieval when retrieving document content, which can increase the retrieval speed.

[0044] Specifically, the target primitive is divided into a plurality of primitive blocks based on paragraph block measurement; and a plurality of text vectors corresponding to each primitive block are stored as a storage unit in a vector database.

[0045] S4, obtains user questions and converts them into query vectors through a vectorization model.

[0046] The user inputs a relevant question, retrieves relevant information from the vectorized database, and processes it through the multimodal large model to generate an accurate answer. This embodiment first uses the bge-m3 vector model to convert the user's input question into a query vector. By converting the text into a vector form, a fast similarity search can be performed using the vector database.

[0047] S5, retrieve relevant text vectors in the vector database according to the query vector, filter out the target title primitive and the target paragraph primitive, and obtain the target image primitive and the target table primitive related to the target paragraph primitive according to the correlation mapping table.

[0048] In the vector retrieval stage, the Faiss library is used to retrieve similar vectors in the Milvus vector database. The Faiss library is designed for efficient similarity search and dense vector clustering and can handle large-scale data sets. Cosine similarity is used as the similarity measurement standard, which is a similarity metric commonly used to measure the angle between two non-zero vectors. By calculating the cosine similarity between the query vector and the vector in the database, the Top-K similar vectors most relevant to the user's question can be quickly found. The specific steps include the following steps.

[0049] S5.1, through parallel processing, the query vector and the text vectors in each current storage unit are subjected to similarity retrieval, and the similarity retrieval uses cosine similarity as the measurement standard.

[0050] S5.2, obtain the N text vectors with the highest cosine similarity as the final target text vector; where N≥1.

[0051] S5.3, obtain the target title primitive and the target paragraph primitive according to the target text vector.

[0052] S6, constructs the target title primitive, target paragraph primitive, target image primitive, target table primitive and user question into prompt words, inputs the prompt words into the multimodal large language model for processing, and outputs the question result.

[0053] In this embodiment, once the relevant multimodal primitives are retrieved, these primitives are constructed into prompt words together with the user's question. These prompt words are then input into the multimodal large language model to output the answer to the user's question.

[0054] As an example, the open source Qwen2-VL-7B-Instruct model is used. The Qwen2-VL-7B-Instruct model can process text and image data. It converts text data into embedded representations through a standard embedding layer, and image data into a series of small block embeddings through PatchEmbed. VisionRotaryEmbedding is then used to encode the rotational position of the visual data, enhancing the model's understanding of the image space structure. The embeddings of the two different modalities are then input into the Transformer structure, and the relationship between the modalities is learned through the attention mechanism. Finally, the hidden state is converted into a predicted token ID through lm_head. These token IDs correspond to the output of the model, that is, the generated answer.

[0055] S7, obtaining the positions of the title primitive, the target paragraph primitive, the target image primitive, and the target table primitive in the document according to the position mapping table; jumping to and highlighting the relevant area in the document in the user interface according to the obtained positions.

[0056] In this embodiment, once the multimodal primitive retrieval based on the user's question is completed, the primitive and document location mapping table established previously will be used to accurately locate the specific location area of ​​the retrieved primitive in the original document. This mapping table is a key achievement in the layout analysis stage, which records the physical location information of each primitive in detail, including page numbers and coordinate ranges. Using this information, it is possible to directly jump to and highlight the relevant area in the document in the user interface, allowing the user to intuitively see the source of the answer. This ability to trace directly to the original document greatly enhances the credibility of the answer because it allows the user to personally verify the accuracy of the information. The user can trace back to a specific paragraph, chart or data in the document to verify the answer provided by the system, thereby improving the overall practicality and reliability of the system.

[0057] An embodiment of a document question and answer method based on multimodal elements is described in detail above. Based on the document question and answer method based on multimodal elements described in the above embodiment, an embodiment of the present invention also provides a document question and answer system based on multimodal elements corresponding to the method.

[0058] Figure 2 It is a schematic block diagram of the structure of a document question-answering system based on multimodal primitives provided by an embodiment of the present invention. In this embodiment, the document question-answering system 200 based on multimodal primitives can be divided into multiple functional modules according to the functions performed by them. The module referred to in the present invention refers to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, which are stored in a memory.

[0059] The primitive extraction module 210 is used to analyze the layout of several documents uploaded by the user and extract the target primitives of each document, wherein the target primitives include title primitives, paragraph primitives, image primitives and table primitives; A correlation building module 220 is used to obtain the correlation between the paragraph primitive and the image primitive and the table primitive, and to build a correlation mapping table, in which a mapping relationship between the paragraph primitive and the related image primitive and table primitive is added; A vector storage module 230, used to convert the title primitive and paragraph primitive of the text mode into text vectors through a vectorization model, and store the text vectors in a vector database; A query vector acquisition module 240 is used to acquire user questions and convert the user questions into query vectors through a vectorization model; A similar primitive query module 250 is used to retrieve relevant text vectors in the vector database according to the query vector, filter out target title primitives and target paragraph primitives, and obtain target image primitives and target table primitives related to the target paragraph primitive according to the correlation mapping table; The question result output module 260 is used to construct the target title primitive, the target paragraph primitive, the target image primitive, the target table primitive and the user question into prompt words, input the prompt words into the multimodal large language model for processing, and output the question result.

[0060] The document question and answer system based on multimodal elements of this embodiment is used to implement the aforementioned document question and answer method based on multimodal elements. Therefore, the specific implementation method of this system can be seen in the embodiment part of the document question and answer method based on multimodal elements in the previous text. Therefore, its specific implementation method can refer to the description of the corresponding embodiments of each part, which will not be introduced in detail here.

[0061] In addition, since the document question and answer system based on multimodal elements of this embodiment is used to implement the aforementioned document question and answer method based on multimodal elements, its function corresponds to that of the aforementioned method and will not be repeated here.

[0062] Figure 3 A schematic diagram of the structure of a terminal 300 provided in an embodiment of the present invention includes: a processor 310, a memory 320 and a communication unit 330. The processor 310 is used to implement the following steps when implementing a document question-and-answer program based on multimodal primitives stored in the memory 320: Perform layout analysis on several documents uploaded by users to extract target primitives of each document, wherein the target primitives include title primitives, paragraph primitives, image primitives and table primitives; Obtain the correlation between the paragraph primitive and the image primitive and the table primitive, and construct a correlation mapping table, and add the mapping relationship between the paragraph primitive and the related image primitive and table primitive in the correlation mapping table; Convert the title primitive and paragraph primitive of the text mode into text vectors through a vectorization model, and store the text vectors in a vector database; Obtain user questions and convert them into query vectors through a vectorization model; Retrieve relevant text vectors in the vector database according to the query vector, filter out the target title primitive and the target paragraph primitive, and obtain the target image primitive and the target table primitive related to the target paragraph primitive according to the correlation mapping table; The target title primitive, target paragraph primitive, target image primitive, target table primitive and user question are constructed into prompt words, the prompt words are input into the multimodal large language model for processing, and the question result is output.

[0063] The terminal 300 includes a processor 310, a memory 320 and a communication unit 330. These components communicate via one or more buses. It can be understood by those skilled in the art that the structure of the server shown in the figure does not constitute a limitation of the present invention, and it can be a bus structure or a star structure, and can also include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0064] The memory 320 can be used to store the execution instructions of the processor 310, and the memory 320 can be implemented by any type of volatile or non-volatile storage terminal or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. When the execution instructions in the memory 320 are executed by the processor 310, the terminal 300 can perform some or all of the steps in the following method embodiments.

[0065] The processor 310 is the control center of the storage terminal, and uses various interfaces and lines to connect various parts of the entire electronic terminal. It runs or executes software programs and / or modules stored in the memory 320, and calls data stored in the memory to perform various functions of the electronic terminal and / or process data. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions. For example, the processor 310 can only include a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single computing core or multiple computing cores.

[0066] The communication unit 330 is used to establish a communication channel so that the storage terminal can communicate with other terminals, receive user data sent by other terminals or send user data to other terminals.

[0067] The present invention also provides a computer storage medium, wherein the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM).

[0068] The computer storage medium stores a document question-answering program based on multimodal primitives, and the document question-answering program based on multimodal primitives implements the following steps when executed by a processor: Perform layout analysis on several documents uploaded by users to extract target primitives of each document, wherein the target primitives include title primitives, paragraph primitives, image primitives and table primitives; Obtain the correlation between the paragraph primitive and the image primitive and the table primitive, and construct a correlation mapping table, and add the mapping relationship between the paragraph primitive and the related image primitive and table primitive in the correlation mapping table; Convert the title primitive and paragraph primitive of the text mode into text vectors through a vectorization model, and store the text vectors in a vector database; Obtain user questions and convert them into query vectors through a vectorization model; Retrieve relevant text vectors in the vector database according to the query vector, filter out the target title primitive and the target paragraph primitive, and obtain the target image primitive and the target table primitive related to the target paragraph primitive according to the correlation mapping table; The target title primitive, target paragraph primitive, target image primitive, target table primitive and user question are constructed into prompt words, the prompt words are input into the multimodal large language model for processing, and the question result is output. Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on this understanding, the technical solution in the embodiments of the present invention, in essence or in other words, the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, and other media that can store program codes, including several instructions for enabling a computer terminal (which can be a personal computer, a server, or a second terminal, a network terminal, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention.

[0069] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0070] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0071] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0072] The above disclosure is only a preferred embodiment of the present invention, but the present invention is not limited thereto. Any non-creative changes that can be thought of by a person skilled in the art, as well as several improvements and modifications made without departing from the principle of the present invention, should fall within the protection scope of the present invention.

Claims

1. A document question answering method based on multimodal primitives, characterized in that: The following steps are involved: Perform layout analysis on several documents uploaded by users to extract target primitives of each document, wherein the target primitives include title primitives, paragraph primitives, image primitives and table primitives; Obtain the correlation between the paragraph primitive and the image primitive and the table primitive, and construct a correlation mapping table, and add the mapping relationship between the paragraph primitive and the related image primitive and table primitive in the correlation mapping table; Convert the title primitive and paragraph primitive of the text mode into text vectors through a vectorization model, and store the text vectors in a vector database; Obtain user questions and convert them into query vectors through a vectorization model; Retrieve relevant text vectors in the vector database according to the query vector, filter out the target title primitive and the target paragraph primitive, and obtain the target image primitive and the target table primitive related to the target paragraph primitive according to the correlation mapping table; The target title primitive, target paragraph primitive, target image primitive, target table primitive and user question are constructed into prompt words, the prompt words are input into the multimodal large language model for processing, and the question result is output.

2. The document question-answering method based on multimodal primitives according to claim 1, characterized in that: Analyze the layout of several documents uploaded by users and extract the target primitives of each document, including: Render the document as an image; Inputting the document image into the layout analysis model to divide the document into headings, paragraphs, tables, and image areas; For the title and paragraph regions, the OCR model is used to extract text information and classify the text information into corresponding title primitives and paragraph primitives respectively; Screenshot table and image areas, as table primitive and image primitive respectively.

3. The document question-answering method based on multimodal primitives according to claim 2, characterized in that: Get the correlation between paragraph primitives, image primitives, and table primitives, including: Extract characters from paragraph primitives, image primitives, and table primitives; Compare the characters in the paragraph primitive with the title characters in the image primitive and the table primitive respectively; If the title of a paragraph primitive and an image primitive contain the same character, then the image primitive is related to the paragraph primitive; If a paragraph primitive and a table primitive's title contain the same characters, then the table primitive is related to the paragraph primitive.

4. The document question-answering method based on multimodal primitives according to claim 3, characterized in that: Storing text vectors in a vector database includes: Dividing the target primitive into a plurality of primitive blocks based on the paragraph block measurement; Several text vectors corresponding to each primitive block are stored in the vector database as a storage unit.

5. The document question-answering method based on multimodal primitives according to claim 4, characterized in that: Retrieve relevant text vectors in the vector database based on the query vector and filter out the target title primitive and target paragraph primitive, including: Through parallel processing, the query vector and the text vectors in each current storage unit are searched for similarity, and the cosine similarity is used as the measurement standard for similarity search; Obtain the N text vectors with the highest cosine similarity as the final target text vector; where N ≥ 1; The target title primitive and the target paragraph primitive are obtained according to the target text vector.

6. The document question-answering method based on multimodal primitives according to any one of claims 1 to 5, characterized in that: After analyzing the layout of several documents uploaded by the user and extracting the target primitives of each document, the following steps are also included: A position mapping table is constructed, and a mapping relationship between each target primitive and its position in the document is added to the position mapping table; the position of the target primitive in the document includes a page number and a region coordinate.

7. The document question-answering method based on multimodal primitives according to claim 6, characterized in that: After outputting the problem results, the following steps are also included: Obtaining the positions of the title primitive, the target paragraph primitive, the target image primitive, and the target table primitive in the document according to the position mapping table; Jumps in the user interface based on the acquired location and highlights relevant areas in the document.

8. A document question answering system based on multimodal primitives, characterized in that: include: A primitive extraction module, used to analyze the layout of several documents uploaded by users and extract target primitives of each document, wherein the target primitives include title primitives, paragraph primitives, image primitives and table primitives; A correlation building module is used to obtain the correlation between the paragraph primitive and the image primitive and the table primitive, and to build a correlation mapping table, in which a mapping relationship between the paragraph primitive and the related image primitive and table primitive is added; A vector storage module, used for converting the title primitive and paragraph primitive of the text mode into text vectors through a vectorization model, and storing the text vectors in a vector database; The query vector acquisition module is used to obtain user questions and convert them into query vectors through a vectorization model; A similar primitive query module is used to retrieve relevant text vectors in the vector database according to the query vector, filter out the target title primitive and the target paragraph primitive, and obtain the target image primitive and the target table primitive related to the target paragraph primitive according to the correlation mapping table; The question result output module is used to construct the target title primitive, the target paragraph primitive, the target image primitive, the target table primitive and the user question into prompt words, input the prompt words into the multimodal large language model for processing, and output the question result.

9. A terminal, characterized in that: include: A memory for storing a document question-answering program based on multimodal primitives; A processor, configured to implement the steps of the document question and answer method based on multimodal primitives as described in any one of claims 1 to 7 when executing the document question and answer program based on multimodal primitives.

10. A computer-readable storage medium, characterized in that: The readable storage medium stores a document question and answer program based on multimodal primitives, and when the document question and answer program based on multimodal primitives is executed by a processor, the steps of the document question and answer method based on multimodal primitives as described in any one of claims 1-7 are implemented.

Citation Information

Cited By

  • Method for automatically extracting ecological parameters and driving factors of large language model

    CN120910562A

  • Ecological parameter of large language model and automatic extraction method of driving factor thereof

    CN120910562B