Question and answer method and device and electronic equipment

By obtaining the embedding vector and intention information of the input information, and using multiple search models and multimodal large language models to search and question-and-answer PDF documents, the problem of insufficient accuracy in image data recognition is solved, and the accuracy and matching of answers are improved.

CN120407718APending Publication Date: 2025-08-01BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510268065.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing OCR technology can only recognize the text in the image when processing image data, but cannot accurately identify the image, resulting in poor question-and-answer accuracy in PDF documents.

Method used

By obtaining the embed vector and intention information of the input information, the PDF document is searched using multiple search models, and a question-and-answer generation is combined with a multimodal large language model to obtain answer information.

Benefits of technology

It improves the matching and accuracy of answer information, reduces the situation of vague answers and inconsistent context, and reduces the probability of errors determined by answer information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407718A_ABST
    Figure CN120407718A_ABST
Patent Text Reader

Abstract

The invention provides a question and answer method and device and electronic equipment, and the method comprises the steps: obtaining an embedded vector corresponding to input information; obtaining intention information corresponding to the input information, and obtaining at least one retrieval model corresponding to the intention information; according to the at least one retrieval model and the embedded vector, retrieval is carried out in a database, a retrieval result set corresponding to the input information is obtained, the retrieval result set comprises at least one retrieval result corresponding to each retrieval model in the at least one retrieval model, and the database is obtained by processing the at least one PDF long document; the method comprises the following steps of: acquiring language category information corresponding to input information, inputting a cue word prompt corresponding to the language category information into a multi-modal large language model to generate questions and answers, and acquiring answer information corresponding to the input information. And the question and answer accuracy of the PDF document is poor due to the fact that the image cannot be accurately recognized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a question-answering method, apparatus, and electronic device. Background Art

[0002] With the development of science and technology, long-document multi-modal question-answering technology combines technologies such as natural language processing, computer vision, and information retrieval, and is a type of intelligent question-answering system that can extract, understand, and answer user questions from long documents. Among them, for example, optical character recognition (OCR) can be used for text recognition. However, for image data, the OCR technology has high requirements for images and can only recognize the text in the images, but cannot accurately recognize the images, resulting in poor question-answering accuracy for PDF documents. Summary of the Invention

[0003] This application aims to at least solve one of the technical problems in the related technologies to some extent.

[0004] To this end, the first object of this application is to propose a question-answering method to reduce the situation where only text results can be obtained, reduce the retrieval limitations of the OCR technology, improve the matching between the answer information and the input information, reduce the fuzziness of the answers and the inconsistency of the context, reduce the error probability of determining the answer information, and improve the accuracy of determining the answer information.

[0005] The second object of this application is to propose a question-answering apparatus.

[0006] The third object of this application is to propose an electronic device.

[0007] The fourth object of this application is to propose a computer-readable storage medium.

[0008] The fifth object of this application is to propose a computer program product.

[0009] To achieve the above object, the first aspect embodiment of this application proposes a question-answering method, including the following steps:

[0010] Obtain an embedding vector corresponding to the input information;

[0011] Obtain the intent information corresponding to the input information, and obtain at least one retrieval model corresponding to the intent information;

[0012] Retrieve in the database according to the at least one retrieval model and the embedding vector to obtain a set of retrieval results corresponding to the input information, where the set of retrieval results includes at least one retrieval result corresponding to each retrieval model in the at least one retrieval model, and the database is obtained by processing at least one long Portable Document Format (PDF) document;

[0013] Obtain the language category information corresponding to the input information, input the prompt corresponding to the language category information into a multi-modal large language model for question and answer generation, and obtain the answer information corresponding to the input information.

[0014] To achieve the above object, an embodiment of the second aspect of the present application proposes a question and answer device, including:

[0015] An information acquisition unit, configured to acquire an embedding vector corresponding to the input information;

[0016] A model acquisition unit, configured to acquire the intention information corresponding to the input information, and acquire at least one retrieval model corresponding to the intention information;

[0017] A result acquisition unit, configured to retrieve in the database according to the at least one retrieval model and the embedding vector to obtain a set of retrieval results corresponding to the input information, where the set of retrieval results includes at least one retrieval result corresponding to each retrieval model in the at least one retrieval model, and the database is obtained by processing at least one long Portable Document Format (PDF) document;

[0018] The information acquisition unit is further configured to acquire the language category information corresponding to the input information, input the prompt corresponding to the language category information into a multi-modal large language model for question and answer generation, and obtain the answer information corresponding to the input information.

[0019] To achieve the above object, an embodiment of the third aspect of the present application proposes an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0020] The memory stores computer execution instructions;

[0021] The processor executes the computer execution instructions stored in the memory to implement the method according to any one of the above first aspects.

[0022] To achieve the above object, an embodiment of the fourth aspect of the present application proposes a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer execution instructions, and the computer execution instructions are used to implement the method according to any one of the above first aspects when executed by a processor.

[0023] To achieve the above object, an embodiment of the fifth aspect of the present application proposes a computer program product, including a computer program, which when executed by a processor implements the method according to any one of the above first aspect.

[0024] The question-answering method, device and electronic device provided by the present application obtain an embedding vector corresponding to the input information; obtain intent information corresponding to the input information, and obtain at least one retrieval model corresponding to the intent information; retrieve in a database according to the at least one retrieval model and the embedding vector to obtain a retrieval result set corresponding to the input information, where the retrieval result set includes at least one retrieval result corresponding to each retrieval model in the at least one retrieval model, and the database is obtained by processing at least one Portable Document Format (PDF) long document; obtain language category information corresponding to the input information, input a prompt word corresponding to the language category information into a multimodal large language model for question-answering generation, and obtain answer information corresponding to the input information, which solves the problem that for image data, using OCR technology can only recognize the text in the image and cannot accurately recognize the image, resulting in poor question-answering accuracy of PDF documents. Since intent information can be obtained through semantic understanding of the input information, multiple retrieval models corresponding to the intent information are used for retrieval, which can retrieve PDF documents, reduce the situation where only text retrieval results can be obtained, reduce the retrieval limitations of OCR technology, and can obtain answer information according to language category information, which can improve the matching degree between the answer information and the input information, reduce the situation of fuzzy answers and inconsistent contexts, reduce the error probability of determining the answer information, and improve the accuracy of determining the answer information.

[0025] Additional aspects and advantages of the present application will be given in part in the following description, will become apparent in part from the following description, or will be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0027] Figure 1 is a schematic flowchart of a question-answering method provided by an embodiment of the present application;

[0028] Figure 2 is a schematic flowchart of a question-answering method provided by an embodiment of the present application;

[0029] Figure 3 is an example schematic diagram of a question-answering method provided by an embodiment of the present application;

[0030] Figure 4 Schematic diagram for an example of a database establishment method provided by an embodiment of the present application;

[0031] And Figure 5 Schematic structural diagram of a question - answering device provided by an embodiment of the present application. Detailed implementation manners

[0032] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described by referring to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.

[0033] The question - answering method and device of the embodiments of the present application will be described below with reference to the accompanying drawings.

[0034] Figure 1 Schematic flow chart of a question - answering method provided by an embodiment of the present application.

[0035] To address this issue, the embodiments of the present application provide a question - answering method to reduce the situation where only text retrieval results can be obtained, reduce the retrieval limitations of the OCR technology, and can improve the matching degree between the answer information and the input information, reduce the fuzziness of the answer and the inconsistency of the context, reduce the error probability of determining the answer information, and improve the accuracy of determining the answer information. As Figure 1 shown, the question - answering method includes the following steps:

[0036] Step 101, obtain the embedding vector corresponding to the input information;

[0037] According to some embodiments, the execution subject of the embodiments of the present application can be, for example, an electronic device. This electronic device does not specifically refer to a certain fixed device, and the name of this electronic device is not limited. This electronic device can also be called a terminal, a mobile device, etc. For example, when the device identifier of this electronic device changes, this electronic device can also change accordingly.

[0038] In some embodiments, the input information can be, for example, the information received during the question - answering process. This input information can be, for example, the question received during the question - answering operation. This input information does not specifically refer to a certain fixed information. For example, when the acquisition method of the input information changes, this input information can also change accordingly. For example, when the acquisition time point corresponding to the input information changes, this input information can also change accordingly. Among them, the acquisition method of the input information is not limited. For example, this input information can be input by clicking through the input method on the electronic device, or can be voice information input through the voice input control.

[0039] In some embodiments, an embedding vector can be, for example, a vector obtained by vectorizing input information. This embedding vector does not specifically refer to a certain fixed vector. For example, when the way of obtaining the embedding vector changes, the embedding vector can also change accordingly. For example, when the model for obtaining the embedding vector changes, the embedding vector can also change accordingly. Among them, Embedding (embedding) can be, for example, a representation method that maps high-dimensional, sparse discrete symbols (such as words in text) to a low-dimensional, dense vector space. This vector representation preserves the semantic relationships between the symbols, enabling it to be used in machine learning and deep learning models for various text analysis and processing.

[0040] In some embodiments, for example, when executing the question-answering method of the present application, an embedding vector corresponding to the input information can be obtained.

[0041] Step 102, obtain the intent information corresponding to the input information, and obtain at least one retrieval model corresponding to the intent information;

[0042] In some embodiments, the intent information can be, for example, used to indicate the intent corresponding to the input information and can be used to indicate the type of answer that the user wants to obtain in the question-answering method. This intent information does not specifically refer to a certain fixed information. For example, when the way of obtaining the intent information changes, the intent information can also change accordingly. For example, when the input information changes, the intent information can also change accordingly.

[0043] In some embodiments, the retrieval model can be, for example, a model that has been trained and can be used for retrieval. Among them, different retrieval models can be used for different types of retrieval, for example. This retrieval model does not specifically refer to a certain fixed model. For example, when the model type of the retrieval model changes, the retrieval model can also change accordingly.

[0044] According to some embodiments, the at least one retrieval model can also be referred to as a retrieval model set, for example. The at least one retrieval model does not specifically refer to a certain fixed model. For example, when the number of models corresponding to the at least one retrieval model changes, the at least one retrieval model can also change accordingly. For example, when the intent information changes, the at least one retrieval model can also change accordingly. For example, when any one of the at least one retrieval models changes, the at least one retrieval model can also change accordingly.

[0045] In some embodiments, the intent information corresponding to the input information may be obtained, and at least one retrieval model corresponding to the intent information may be obtained. The order of obtaining for step S101 and step S102 is not limited. That is, step S101 may be executed first, followed by step S102, or step S102 may be executed first, followed by step S101, or steps S101 and S102 may be executed simultaneously. The embodiments of the present disclosure do not limit this.

[0046] Step 103, retrieve in the database according to the at least one retrieval model and the embedding vector to obtain a retrieval result set corresponding to the input information, where the retrieval result set includes at least one retrieval result corresponding to each retrieval model in the at least one retrieval model, and the database is obtained by processing at least one Portable Document Format (PDF) long document;

[0047] In some embodiments, the database may be, for example, a database used to indicate question-and-answer retrieval. This database may be pre-established, and it does not specifically refer to a certain fixed database. For example, when the number of documents corresponding to the database changes, the database may also change accordingly. For example, when an update instruction for the database is received, the database may also change accordingly. Among them, the database may be obtained by processing at least one Portable Document Format (PDF) long document.

[0048] According to some embodiments, the retrieval result may be used to indicate the result obtained by retrieving in the database. Among them, the retrieval result may be the result corresponding to the retrieval model. Different retrieval models may correspond to different retrieval results. One retrieval model may correspond to at least one retrieval result or may also correspond to zero retrieval results. The embodiments of the present disclosure do not limit this.

[0049] In some embodiments, the retrieval result set may be, for example, a collective formed by converging at least one retrieval. The retrieval result set may include, for example, at least one retrieval result corresponding to each retrieval model in the at least one retrieval model. That is to say, the retrieval result set includes the results retrieved by all retrieval models. The retrieval result set does not specifically refer to a certain fixed set. For example, when any one of the at least one retrieval models changes or the database changes, the retrieval result set may also change accordingly. For example, when the number of retrieval results included in the retrieval result set changes, the retrieval result set may also change accordingly. For example, when a certain retrieval result in the retrieval result set changes, the retrieval result set may also change accordingly.

[0050] In some embodiments, the Portable Document Format (PDF) can be used, for example, to indicate the format of the document. Among them, the number of pages corresponding to each PDF long document in at least one PDF long document can be greater than a preset page threshold, that is, the technical solution of the embodiments of the present application can be applied to the Q&A scenario of long documents.

[0051] In some embodiments, according to the at least one retrieval model and the embedding vectors, a retrieval is performed in the database to obtain a set of retrieval results corresponding to the input information, where the set of retrieval results includes at least one retrieval result corresponding to each retrieval model in the at least one retrieval model, and the database is obtained by processing at least one PDF long document.

[0052] Step 104: Obtain the language category information corresponding to the input information, input the prompt corresponding to the language category information into the multi-modal large language model for Q&A generation, and obtain the answer information corresponding to the input information.

[0053] According to some embodiments, the language category information can be used, for example, to indicate the language category of the answer information to be retrieved. The language category information can be obtained, for example, by identifying the language category of the input information, or by identifying the intent information corresponding to the input information. For example, it can also be... The embodiments of the present disclosure do not limit this.

[0054] In some embodiments, the prompt can be used, for example, to prompt the context of the input information and the parameter information of the input model for the multi-modal large language model. The method for obtaining the prompt is not limited.

[0055] According to some embodiments, the multi-modal large language model can be, for example, a model that has been trained and can be used for Q&A generation. The multi-modal large language model does not specifically refer to a certain fixed model. For example, when the model parameters corresponding to the multi-modal large language model change, the multi-modal large language model can also change accordingly. For example, when the model type corresponding to the multi-modal large language model changes, the multi-modal large language model can also change accordingly.

[0056] In some embodiments, the long document multi-modal large language model can not only process pure text information, but also process various modal information such as images, charts, and tables, so as to provide more comprehensive and accurate answers.

[0057] In some embodiments, the answer information can be, for example, the answer corresponding to the input information.

[0058] In some embodiments, language category information corresponding to the input information is obtained, and a prompt corresponding to the language category information is input into a multi-modal large language model for question and answer generation, and answer information corresponding to the input information is obtained.

[0059] The question and answer method, device and electronic device provided by the present application. The method includes obtaining an embedding vector corresponding to the input information; obtaining intent information corresponding to the input information, and obtaining at least one retrieval model corresponding to the intent information; retrieving in a database according to the at least one retrieval model and the embedding vector to obtain a retrieval result set corresponding to the input information, where the retrieval result set includes at least one retrieval result corresponding to each retrieval model in the at least one retrieval model, and the database is obtained by processing at least one Portable Document Format (PDF) long document; obtaining language category information corresponding to the input information, inputting a prompt corresponding to the language category information into a multi-modal large language model for question and answer generation, and obtaining answer information corresponding to the input information. This solves the problem that for image data, using OCR technology can only recognize the text in the image and cannot accurately recognize the image, resulting in poor accuracy of question and answer for PDF documents. Since intent information can be obtained through semantic understanding of the input information, multiple retrieval models corresponding to the intent information are used for retrieval, which can retrieve PDF documents, reduce the situation of only obtaining text retrieval results, reduce the retrieval limitations of OCR technology, and can obtain answer information according to the language category information, which can improve the matching degree between the answer information and the input information, reduce the situation of fuzzy answers and inconsistent contexts, reduce the error probability of determining the answer information, and improve the accuracy of determining the answer information.

[0060] This embodiment provides another question and answer method. Figure 2 It is a schematic flowchart of a question and answer method provided by an embodiment of the present application.

[0061] As Figure 2 shown, the question and answer method may include the following steps:

[0062] Step 201, obtain an embedding vector corresponding to the input information;

[0063] The specific process is as described above and will not be elaborated here.

[0064] According to some embodiments, the execution subject of the embodiment of the present application may be an electronic device, for example. The electronic device does not specifically refer to a certain fixed device, and the name of the electronic device is not limited. The electronic device may also be referred to as a terminal, a mobile device, etc., for example. For example, when the device identifier of the electronic device changes, the electronic device may also change accordingly.

[0065] In some embodiments, the technical solutions of the embodiments of the present application can be used in fields such as enterprise document management, academic research, legal compliance, and finance, for example.

[0066] In some embodiments, when the input information is obtained, a semantic vector model can be used to generate an embedding vector of the input information. Figure 3 The following is an example schematic diagram showing a question-and-answer method according to an embodiment of the present application. As Figure 3 shown, a user query can be obtained, and an embedding vector corresponding to the user query can be obtained.

[0067] Step 202, obtain the intent information corresponding to the input information, and obtain at least one retrieval model corresponding to the intent information;

[0068] The specific process is as described above and will not be elaborated here.

[0069] According to some embodiments, when the input information is obtained, the intent information of the input information can be identified. For example, when a question is obtained, it can be identified whether the question only asks about pictures, tables, or a mixed query of pictures and texts, such as " Figure 1 How many people are there?" Among them, the intent information, that is, the recognition result, can be classified into, for example: pictures, tables, and mixed pictures and texts.

[0070] Step 203, perform a search in the database according to the at least one retrieval model and the embedding vector to obtain a first search result set;

[0071] The specific process is as described above and will not be elaborated here.

[0072] In some embodiments, the first search result set can be, for example, a collective formed by converging at least one search result directly obtained in the database using the search model and the embedding vector. The "first" in the first search result set is used to distinguish it from other search result sets and does not specifically refer to a certain fixed search result set. For example, when the number of search results included in the first search result set changes, the first search result set can also change accordingly. For example, when a certain search result in the first search result set changes, the first search result set can also change accordingly.

[0073] According to some embodiments, the database of the embodiments of the present application may be, for example, a vector database. Among them, for example, the vector representation of the user question can be used for efficient similarity retrieval in the Milvus database to find the document content most relevant to the question, that is, to obtain the answer information corresponding to the input information. Specifically, for example, it may include: retrieving in the Milvus database according to at least one determined retrieval model. Milvus is an open-source vector database management system dedicated to processing, storing, and retrieving high-dimensional vector data. It can perform similarity searches on large-scale, high-dimensional data quickly and efficiently, and is particularly suitable for application scenarios that require efficient operations on vectorized data in fields such as image retrieval, recommendation systems, and natural language processing. The design of Milvus aims to simplify and accelerate the process of processing vector data, enabling developers to conveniently build and manage complex vector retrieval systems.

[0074] Among them, the at least one retrieval model may include, for example, a hybrid type retrieval model, a graph type retrieval model, and a table type retrieval model, where:

[0075] The retrieval process of the hybrid type retrieval model is as follows:

[0076] Multi-modal information retrieval: Using the efficient retrieval mechanism of Milvus, relevant text, pictures, and table data are retrieved from the database according to the embedding vector of the question. The retrieval results will select the most relevant graphic information for organization according to the context and type of the question.

[0077] The retrieval process of the graph type retrieval model is as follows:

[0078] Multi-modal information retrieval: Using the efficient retrieval mechanism of Milvus, relevant text, pictures, and table data are retrieved from the database according to the embedding vector of the question, where the data of the filtered type is pictures.

[0079] The table type retrieval process is as follows:

[0080] Multi-modal information retrieval: Using the efficient retrieval mechanism of Milvus, relevant text, pictures, and table data are retrieved from the database according to the embedding vector of the question, where the data of the filtered type is tables.

[0081] According to some embodiments, the method further includes:

[0082] Adopting a layout extraction model based on the convolutional neural network CNN to identify the layout structure of each Portable Document Format (PDF) long document in the at least one PDF long document, and obtaining the layout information corresponding to each PDF long document;

[0083] Use a generative model to perform bounding box recognition on each of the Portable Document Format (PDF) long documents, and sort each bounding box in the bounding box set to obtain the order information of each bounding box, where the bounding box set includes at least one recognized bounding box;

[0084] According to the order information of each bounding box, use at least one recognition model corresponding to the layout information to recognize each PDF long document, and obtain the text information, title information, picture information, and table information corresponding to each PDF;

[0085] Perform word segmentation on the text information, the title information, and the page numbers corresponding to each PDF version document to obtain the word-segmented text information, word-segmented title information, and word-segmented page numbers, and use a multi-modal embedding model to vectorize the word-segmented text information, word-segmented title information, and word-segmented page numbers, and add the obtained first vector to the database;

[0086] Obtain the description information corresponding to the picture information, use the VLLM (Visual Large Language Model) to vectorize the title information and the description information corresponding to the picture information, and add the obtained second vector and the metadata corresponding to the picture information to the database, where the second vector corresponds to the title information and the description information corresponding to the picture information;

[0087] Vectorize the picture information and add the obtained third vector to the database;

[0088] Obtain the HyperText Markup Language (HTML) structure information corresponding to the table information, vectorize the title information in the table information and the HTML structure information, and add the obtained fourth vector and the metadata corresponding to the table information to the database, where the fourth vector corresponds to the title information in the table information and the HTML structure information;

[0089] Vectorize the table information and add the obtained fifth vector to the database. Thus, multimodal parsing and information structured storage of document content can be performed, improving the convenience of retrieval and the retrieval efficiency. Since VLLM can process information in text, charts, and tables and output it in a structured manner. When parsing charts, VLLM can not only extract text information but also identify and describe the numerical values, trends, titles, and annotations of the charts, generating standardized descriptions. This enables graphic and text information to be stored and subsequently processed in a structured format, reducing the situation where OCR technology can only recognize text and cannot directly extract detailed data or picture information from charts, reducing the number of cases of image data loss, reducing the situation where OCR technology cannot generate structured data, and this application can identify metadata, which can improve the accuracy of obtaining answer information and can be achieved through the VLLM model. It can not only understand the semantics of the text but also accurately extract data from the charts and provide detailed explanations, greatly reducing information misleading or loss caused by OCR recognition errors.

[0090] According to some embodiments, the method further includes:

[0091] When the layout information corresponding to each Portable Document Format (PDF) long document includes formulas, use a formula recognition model to recognize each PDF long document to obtain picture information corresponding to the formulas;

[0092] Use a Transformer conversion model to perform format conversion processing on the picture information corresponding to the formulas to obtain formulas in the target format, and add the formulas in the target format to the text information.

[0093] According to some embodiments, Figure 4 FIG. shows an example schematic diagram of a database establishment method according to an embodiment of the present application. Among them, for example, text, pictures, and table information can be extracted from a PDF document, and it is ensured that the charts and their related information (such as titles, annotations, etc.) can be accurately corresponding and associated.

[0094] According to some embodiments, parsing a PDF document can, for example, use the open-source layout analysis tool MinerU to parse the full-text layout and finally return it in the format of a dictionary Dict in a list List. Among them, MinerU: MinerU is a tool that converts PDF into a machine-readable format (such as markdown, JSON), focusing on solving the symbol conversion problem in scientific and technological literature. Among them, different types of data include different fields, which can include, for example:

[0095] Text: type(text, representing text), text (text string), text_level (whether it is a title), page_idx (page number);

[0096] Picture: type (image, indicating a picture), img_path (picture path), img_caption (picture title), img_footnote (picture footnote), page_idx (page number);

[0097] Table: type (table, indicating a table), img_path (picture path), table_caption (table title), table_footnote (table footnote), table_body (table body, recognized as HTML format through the built-in rapid_table table content extraction model of MinerU), page_idx (page number).

[0098] In addition, for formulas, the embodiments of the present application can convert formulas into LaTeX format based on the built-in formula extraction model of MinerU and embed them into strings of text type (type = text).

[0099] According to some embodiments, the database can be, for example, a vector database. The establishment of the vector database can include, for example:

[0100] a. Identify the layout structure. Use a fine-tuned convolutional neural network (CNN)-based layout extraction model for layout extraction, where the categories are set as pictures, tables, text, titles, and formulas.

[0101] b. Sort the bounding boxes. Use a sequence-to-sequence generative model to sort the recognized bounding boxes. The input is the bounding boxes (x1, y1, x2, y2) recognized on each page, where (x1, y1) and (x2, y2) represent the left values of the upper left and lower right of the bounding box respectively. Input these bounding boxes into a Transformer-based generative model to achieve bounding box sorting and output the order of each bounding box.

[0102] c. For formulas, use a formula recognition model, a YOLO series-based formula localization model, and a Transformer-based formula conversion model to convert formula pictures into LaTeX strings and embed them into text information for storage as text type.

[0103] d. Vectorize and store text types (multi-modal Embedding models such as bge-m3). Segment the text, titles, and page numbers and store them in the vector database.

[0104] e. Vectorized storage of picture types (multi-modal Embedding models such as bge-m3). Vectorize the VLLM visual model description + title, and store the title, page number, type (picture, table, text), description text of the VLLM, original picture path, and footnotes as metadata in the database; pictures can be vectorized separately and stored in the database using the same metadata.

[0105] f. Vectorized storage of table types (multi-modal Embedding models such as bge-m3). The generative model converts the table picture into an HTML structure. Vectorize the table picture separately, and store the title, page number, type (picture, table, text), HTML string, original picture path, and footnotes as metadata in the database. Vectorize the table title + HTML and store it in the database using the same metadata.

[0106] According to some embodiments, the extracted graphic and text information can be further organized and structured. Use a visual multi-modal large language model (VLLM) to describe the detailed text and numerical values in the chart, and output the processing results in a standardized data format.

[0107] Specifically, for example: for chart descriptions, for the detailed text and numerical values in the chart, use the VLLM model for in-depth analysis to generate a detailed description of the chart, including the type of the chart, numerical trends, graphic meanings, etc. At the same time, for image content, a summary description (such as objects, scenes, colors, etc. in the image) will also be generated to ensure the complete semantic transmission of chart and image information.

[0108] For example, the prompt is designed as: "Please explain in detail the content in the picture, including specific numerical values, text, symbols, and other character records that appear." After obtaining the VLLM output, it is organized as: "The following is the picture on page {page} of the retrieved PDF content, related to "{title}". This picture corresponds to the title "{caption}", and the footnote is: "{footnotes}". The image description content is: {VLLM Description}"

[0109] In some embodiments, the format conversion can be, for example, JSON formatted output: Organize the graphic and text information, chart descriptions, and table information into a JSON-formatted structured data for subsequent processing and retrieval. Each part of the information has a clear field identifier in the JSON file, including document title, chart information, table data, text paragraphs, etc., providing a standardized structure for subsequent multi-modal information processing.

[0110] According to some embodiments, by performing word segmentation on the content of the PDF document and using a semantic vector model (such as BGE) to generate corresponding embedding vectors for each document, chart, text, and table, these vectors are finally stored in the Milvus database for subsequent retrieval.

[0111] Among them, the word segmentation and semantic vector generation process may include, for example: segmenting the extracted text by sentence and using a pre-trained semantic vector model to convert the content of each document into a vector representation. For pictures and tables, corresponding embedding vectors are also generated.

[0112] Among them, the fragment metadata storage process may include, for example: each segmented clause, the description of the picture, and the table will be accompanied by metadata information such as the title and page number.

[0113] Among them, the Milvus database storage process may include, for example: storing the generated embedding vectors in the Milvus database. Each PDF document will have a unique ID, and information such as charts, tables, and text in the document will be stored separately for subsequent retrieval.

[0114] Step 204, use a relevance ranking algorithm to rank at least one retrieval result in the first retrieval result set to obtain a second retrieval result set;

[0115] The specific process is as described above and will not be elaborated here.

[0116] Among them, for example, it may also include a retrieval result screening and ranking process. Specifically, for example, through a set relevance ranking algorithm, the content that best matches the user's question is screened out, and the retrieval results are ranked according to the degree of relevance.

[0117] According to some embodiments, the second retrieval result set may be, for example, a set obtained by performing relevance ranking on the first retrieval result set. Among them, the "second" in the second retrieval result set is used to distinguish it from the rest of the retrieval result sets.

[0118] Step 205, obtain a third retrieval result set in the second retrieval result set, where the third retrieval result set includes at least one retrieval result, and the retrieval type corresponding to the at least one retrieval result is the target retrieval type;

[0119] The specific process is as described above and will not be elaborated here.

[0120] In some embodiments, the target retrieval type can be, for example, the type required for secondary retrieval. The target retrieval type does not specifically refer to a certain fixed type. The target retrieval type can include, for example, multiple retrieval types. When a modification instruction for the target retrieval type is received, the target retrieval type can be modified and the target retrieval type can change accordingly.

[0121] In some embodiments, the third retrieval result set can be, for example, a set including at least one retrieval result whose retrieval type is the target retrieval type. The third retrieval result set does not specifically refer to a certain fixed set. For example, when the number of retrieval results included in the third retrieval result set changes, the third retrieval result set can also change accordingly. For example, when any one of the retrieval results in the third retrieval result set changes, the third retrieval result set can also change accordingly. For example, when the target retrieval type changes, the third retrieval result set can also change accordingly.

[0122] According to some embodiments, obtain the third retrieval result set in the second retrieval result set, where the third retrieval result set includes at least one retrieval result, and the retrieval type corresponding to the at least one retrieval result is the target retrieval type.

[0123] Step 206, perform secondary retrieval in the database by using the retrieval models corresponding to the respective retrieval results in the third retrieval result set and the metadata corresponding to the respective retrieval results to obtain a fourth retrieval result set;

[0124] The specific process is as described above and will not be elaborated here.

[0125] In some embodiments, the fourth retrieval result set can be, for example, a set obtained by performing secondary retrieval. The fourth retrieval result set does not specifically refer to a certain fixed set. For example, when the number of retrieval results included in the fourth retrieval result set changes, the fourth retrieval result set can also change accordingly.

[0126] In some embodiments, perform secondary retrieval in the database by using the retrieval models corresponding to the respective retrieval results in the third retrieval result set and the metadata corresponding to the respective retrieval results to obtain a fourth retrieval result set.

[0127] According to some embodiments, where the target type includes a graph type and a table type; the performing secondary retrieval in the database by using the retrieval models corresponding to the respective retrieval results in the third retrieval result set and the metadata corresponding to the respective retrieval results to obtain a fourth retrieval result set includes:

[0128] Obtain a first subset of retrieval results of the target type being a graph type and a second subset of retrieval results of the target type being a table type from the third retrieval result set;

[0129] According to the metadata corresponding to each first retrieval result in the first subset of retrieval results, obtain the original image path;

[0130] Call the Visual Large Language Model (VLLM) according to the original image path, obtain the image corresponding to the original image path, and input the image corresponding to the original image path and the input information into a multi-modal large language model to obtain a third subset of retrieval results;

[0131] According to the metadata corresponding to each second retrieval result in the second subset of retrieval results, obtain the original Hypertext Markup Language (HTML) table path;

[0132] Call the VLLM according to the original HTML table path, obtain the HTML table corresponding to the original HTML table path, and input the HTML table corresponding to the original HTML table path and the input information into a multi-modal large language model to obtain a fourth subset of retrieval results;

[0133] Add the third subset of retrieval results and the fourth subset of retrieval results to the fourth retrieval result set.

[0134] Step 207, add the first retrieval result set and the fourth retrieval result set to the retrieval result set corresponding to the input information;

[0135] The specific process is as described above and will not be elaborated here.

[0136] Step 208, obtain the language category information corresponding to the input information, input the prompt corresponding to the language category information into a multi-modal large language model for question and answer generation, and obtain the answer information corresponding to the input information.

[0137] The specific process is as described above and will not be elaborated here.

[0138] According to some embodiments, identify the user's Query language: Chinese / English. The language of the Prompt component metadata (such as "the page number is" or "The page is") is mainly based on the user's Query language category.

[0139] For the mixed type, input the retrieved relevant text, image, and table information together with the user's question into a multi-modal large language model for question and answer generation. The specific steps are as follows:

[0140] Organizational Prompt: Combine the retrieved relevant content with the user's question to construct an effective Prompt, ensuring that all relevant information can be effectively conveyed to the multimodal large language model. The Prompt includes not only the question itself, but also graphic and text information, tabular data, and their related contexts.

[0141] Question and answer generation: Input the organized Prompt into the multimodal large language model for processing to generate the final answer. The multimodal large language model will generate accurate and well-founded answers based on the nature of the question, the graphic and text information in the document, and the chart data.

[0142] For the figure type:

[0143] Secondary retrieval: Obtain the original image path of the retrieved and sorted images from the metadata, and separately call the VLLM model again. The input is the image and the question, and the output is the answer to the relevant question.

[0144] Question and answer: Use the VLLM return result and the image-related metadata as input to obtain the return result of the unimodal LLM and generate accurate and well-founded answers.

[0145] For the table type:

[0146] Question and answer: Use the HTML table and the question as input to obtain the return result of the unimodal LLM and generate accurate and well-founded answers.

[0147] According to some embodiments, retrieve in the database according to the at least one retrieval model and the embedding vector to obtain a first retrieval result set; use a relevance ranking algorithm to rank at least one retrieval result in the first retrieval result set to obtain a second retrieval result set; obtain a third retrieval result set in the second retrieval result set, where the third retrieval result set includes at least one retrieval result, and the retrieval type corresponding to the at least one retrieval result is the target retrieval type; use the retrieval model corresponding to each retrieval result in the third retrieval result set and the metadata corresponding to each retrieval result to perform secondary retrieval in the database to obtain a fourth retrieval result set; add the first retrieval result set and the fourth retrieval result set to the retrieval result set corresponding to the input information. Therefore, secondary retrieval can be performed based on the initial retrieval result and metadata. The RAG (Retrieval-Augmented Generation) technology can be used to solve the problem of context length by enhancing the combination of retrieval and generation, dynamically supplement the context, and fully utilize the optimal information related to the question each time an answer is generated, reducing the situation of incorrect answers caused by information extraction errors or context breaks, improving the interpretability and credibility of the answers, and enhancing the practicality of the application of the embodiments of the present application.

[0148] To implement the above embodiments, the present application also proposes a question-and-answer device.

[0149] Figure 5 FIG. is a schematic structural diagram of a question-and-answer device provided by an embodiment of the present application.

[0150] As Figure 5 shown, the question-and-answer device includes:

[0151] An information acquisition unit 501, configured to acquire an embedding vector corresponding to the input information;

[0152] A model acquisition unit 502, configured to acquire intent information corresponding to the input information, and acquire at least one retrieval model corresponding to the intent information;

[0153] A result acquisition unit 503, configured to perform a search in a database according to the at least one retrieval model and the embedding vector, and acquire a retrieval result set corresponding to the input information, where the retrieval result set includes at least one retrieval result corresponding to each retrieval model in the at least one retrieval model, and the database is obtained by processing at least one Portable Document Format (PDF) long document;

[0154] The information acquisition unit 501 is further configured to acquire language category information corresponding to the input information, input a prompt corresponding to the language category information into a multi-modal large language model for question-and-answer generation, and acquire answer information corresponding to the input information.

[0155] Further, in a possible implementation manner of the embodiment of the present application, the information acquisition unit 501 is further configured to:

[0156] Adopt a layout extraction model based on a Convolutional Neural Network (CNN) to perform layout structure recognition on each Portable Document Format (PDF) long document in the at least one Portable Document Format (PDF) long document, and acquire layout information corresponding to each Portable Document Format (PDF) long document;

[0157] Adopt a generative model to perform bounding box recognition on each Portable Document Format (PDF) long document, and sort each bounding box in the bounding box set to acquire order information of each bounding box, where the bounding box set includes at least one recognized bounding box;

[0158] According to the order information of each bounding box, adopt at least one recognition model corresponding to the layout information to perform recognition on each Portable Document Format (PDF) long document, and acquire text information, title information, picture information, and table information corresponding to each PDF;

[0159] Tokenize the text information, the title information, and the page numbers corresponding to the PDF versions of the documents, obtain the tokenized text information, the tokenized title information, and the tokenized page numbers, and use a multi-modal Embedding model to vectorize the tokenized text information, the tokenized title information, and the tokenized page numbers, and add the obtained first vector to the database;

[0160] Obtain the description information corresponding to the picture information, use the VLLM vision model to vectorize the title information corresponding to the picture information and the description information, and add the obtained second vector and the metadata corresponding to the picture information to the database, where the second vector corresponds to the title information and the description information corresponding to the picture information;

[0161] Vectorize the picture information and add the obtained third vector to the database;

[0162] Obtain the HyperText Markup Language (HTML) structure information corresponding to the table information, vectorize the title information in the table information and the HTML structure information, and add the obtained fourth vector and the metadata corresponding to the table information to the database, where the fourth vector corresponds to the title information in the table information and the HTML structure information;

[0163] Vectorize the table information and add the obtained fifth vector to the database.

[0164] Further, in a possible implementation manner of the embodiment of the present application, the information acquisition unit 501 is further configured to:

[0165] When the layout information corresponding to each Portable Document Format (PDF) long document includes formulas, use a formula recognition model to recognize each PDF long document to obtain picture information corresponding to the formulas;

[0166] Use a Transformer conversion model to perform format conversion processing on the picture information corresponding to the formulas to obtain formulas in the target format, and add the formulas in the target format to the text information.

[0167] Further, in a possible implementation manner of the embodiment of the present application, when the result acquisition unit 503 is used to retrieve in the database according to the at least one retrieval model and the embedding vector to obtain a retrieval result set corresponding to the input information, it is specifically configured to:

[0168] Retrieve in the database according to the at least one retrieval model and the embedding vector to obtain a first retrieval result set;

[0169] Sort at least one search result in the first search result set using a relevance ranking algorithm to obtain a second search result set;

[0170] Obtain a third search result set from the second search result set, where the third search result set includes at least one search result, and the search type corresponding to the at least one search result is the target search type;

[0171] Perform a secondary search in the database using the search models corresponding to the respective search results in the third search result set and the metadata corresponding to the respective search results to obtain a fourth search result set;

[0172] Add the first search result set and the fourth search result set to the search result set corresponding to the input information.

[0173] Further, in a possible implementation manner of the embodiments of the present application, where the target type includes a graph type and a table type; when the result acquisition unit 503 is used to perform a secondary search in the database using the search models corresponding to the respective search results in the third search result set and the metadata corresponding to the respective search results to obtain a fourth search result set, it is specifically used for:

[0174] Obtain a first search result subset with the target type of graph type and a second search result subset with the target type of table type from the third search result set;

[0175] Obtain the original graph path according to the metadata corresponding to each first search result in the first search result subset;

[0176] Call the Visual Large Language Model (VLLM) according to the original graph path to obtain the graph corresponding to the original graph path, and input the graph corresponding to the original graph path and the input information into a multi-modal large language model to obtain a third search result subset;

[0177] Obtain the original HyperText Markup Language (HTML) table path according to the metadata corresponding to each second search result in the second search result subset;

[0178] Call the VLLM according to the original HTML table path to obtain the HTML table corresponding to the original HTML table path, and input the HTML table corresponding to the original HTML table path and the input information into a multi-modal large language model to obtain a fourth search result subset;

[0179] Add the third search result subset and the fourth search result subset to the fourth search result set.

[0180] Further, in a possible implementation manner of the embodiment of the present application, the information acquisition unit 501 is further specifically used for:

[0181] Format the text information, title information, picture information, and table information corresponding to each PDF to obtain text information in a preset format, title information in a preset format, picture information in a preset format, and table information in a preset format.

[0182] It should be noted that the foregoing explanation of the embodiment of the question-and-answer method also applies to the question-and-answer device of this embodiment, and will not be elaborated here.

[0183] The question-and-answer device provided in the present application includes an information acquisition unit 501 for acquiring an embedding vector corresponding to the input information; a model acquisition unit 502 for acquiring intent information corresponding to the input information and acquiring at least one retrieval model corresponding to the intent information; a result acquisition unit 503 for retrieving in a database according to the at least one retrieval model and the embedding vector to obtain a retrieval result set corresponding to the input information, where the retrieval result set includes at least one retrieval result corresponding to each retrieval model in the at least one retrieval model, and the database is obtained by processing at least one portable document format (PDF) long document; the information acquisition unit 501 is further used for acquiring language category information corresponding to the input information, inputting a prompt corresponding to the language category information into a multimodal large language model for question-and-answer generation, and acquiring answer information corresponding to the input information, which solves the problem that for image data, using OCR technology can only recognize the text in the image and cannot accurately recognize the image, resulting in poor question-and-answer accuracy of PDF documents. Since the intent information can be obtained through semantic understanding of the input information, multiple retrieval models corresponding to the intent information are used for retrieval, which can retrieve PDF documents, reduce the situation where only text retrieval results can be obtained, reduce the retrieval limitation of OCR technology, and can obtain answer information according to the language category information, which can improve the matching degree between the answer information and the input information, reduce the situation of fuzzy answers and inconsistent contexts, reduce the error probability of determining the answer information, and improve the accuracy of determining the answer information.

[0184] To implement the above embodiment, the present application also proposes an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiment.

[0185] To implement the above embodiments, the present application also provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the method provided by the foregoing embodiments when executed by a processor.

[0186] To implement the above embodiments, the present application also provides a computer program product including a computer program, which implements the method provided by the foregoing embodiments when executed by a processor.

[0187] The collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in the present application all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0188] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of such legal uses. In addition, such collection / sharing should be carried out after obtaining the informed consent of the user, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization including authorizing the relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others with access to the personal information data comply with their privacy policies and procedures.

[0189] The present application is expected to provide an implementation for users to selectively block the use or access of personal information data. That is, the present application is expected to provide hardware and / or software to prevent or block access to such personal information data. Once the personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of the user.

[0190] In the descriptions of the foregoing embodiments, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0191] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In the description of this application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined.

[0192] Any process or method description represented in a flowchart or described otherwise herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of this application includes additional implementations where functions may be executed not in the order shown or discussed, including in a substantially simultaneous manner according to the functions involved or in a reverse order, which should be understood by those skilled in the technical field to which the embodiments of this application pertain.

[0193] The logic and / or steps represented in a flowchart or described otherwise herein, for example, can be considered as an ordered listing of executable instructions for implementing a logical function and can be specifically implemented in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0194] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits having logic gate circuits for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0195] Those of ordinary skill in the art can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0196] In addition, in each embodiment of the present application, the functional units can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0197] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limitations on the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. A question-and-answer method, characterized in that, Including: Obtain the embedding vector corresponding to the input information; Obtain the intent information corresponding to the input information, and obtain at least one retrieval model corresponding to the intent information; Retrieve in the database according to the at least one retrieval model and the embedding vector to obtain a retrieval result set corresponding to the input information, wherein the retrieval result set includes at least one retrieval result corresponding to each retrieval model in the at least one retrieval model, and the database is obtained by processing at least one Portable Document Format (PDF) long document; Obtain the language category information corresponding to the input information, input the prompt corresponding to the language category information into a multi-modal large language model for question and answer generation, and obtain the answer information corresponding to the input information.

2. The method according to claim 1, wherein The method further includes: Use a layout extraction model based on a Convolutional Neural Network (CNN) to identify the layout structure of each PDF long document in the at least one PDF long document, and obtain the layout information corresponding to each PDF long document; Use a generative model to identify the bounding boxes of each PDF long document, and sort each bounding box in the bounding box set to obtain the order information of each bounding box, wherein the bounding box set includes at least one identified bounding box; According to the order information of each bounding box, use at least one recognition model corresponding to the layout information to recognize each PDF long document, and obtain the text information, title information, picture information and table information corresponding to each PDF; Segment the text information, the title information, and the page numbers corresponding to each PDF version document to obtain the segmented text information, segmented title information, and segmented page numbers, and use a multi-modal Embedding model to vectorize the segmented text information, segmented title information, and segmented page numbers, and add the obtained first vector to the database; Obtain the description information corresponding to the picture information, use a VLLM vision model to vectorize the title information and the description information corresponding to the picture information, and add the obtained second vector and the metadata corresponding to the picture information to the database, wherein the second vector corresponds to the title information and the description information corresponding to the picture information; Vectorize the picture information, and add the obtained third vector to the database; Obtain the HyperText Markup Language (HTML) structure information corresponding to the table information, vectorize the title information in the table information and the HTML structure information, and add the obtained fourth vector and the metadata corresponding to the table information to the database, wherein the fourth vector corresponds to the title information in the table information and the HTML structure information; Vectorize the table information, and add the obtained fifth vector to the database.

3. The method according to claim 2, wherein The method further includes: When the layout information corresponding to each Portable Document Format (PDF) long document includes formulas, use a formula recognition model to recognize each PDF long document to obtain the picture information corresponding to the formulas; Use a Transformer conversion model to perform format conversion processing on the picture information corresponding to the formulas to obtain formulas in the target format, and add the formulas in the target format to the text information.

4. The method according to claim 2, wherein The retrieving in the database according to the at least one retrieval model and the embedding vector to obtain a retrieval result set corresponding to the input information includes: Retrieve in the database according to the at least one retrieval model and the embedding vector to obtain a first retrieval result set; Use a relevance ranking algorithm to rank at least one retrieval result in the first retrieval result set to obtain a second retrieval result set; Obtain a third retrieval result set in the second retrieval result set, where the third retrieval result set includes at least one retrieval result, and the retrieval type corresponding to the at least one retrieval result is the target retrieval type; Use the retrieval models corresponding to the respective retrieval results in the third retrieval result set and the metadata corresponding to the respective retrieval results to perform a secondary retrieval in the database to obtain a fourth retrieval result set; Add the first retrieval result set and the fourth retrieval result set to the retrieval result set corresponding to the input information.

5. The method according to claim 4, characterized in that, Wherein, The target types include figure type and table type; the using the retrieval models corresponding to the respective retrieval results in the third retrieval result set and the metadata corresponding to the respective retrieval results to perform a secondary retrieval in the database to obtain a fourth retrieval result set includes: Obtain a first retrieval result subset in the third retrieval result set where the target type is the figure type and a second retrieval result subset where the target type is the table type; Obtain the original picture path according to the metadata corresponding to each first retrieval result in the first retrieval result subset; Call a Vision-Language Large Language Model (VLLM) according to the original picture path to obtain the picture corresponding to the original picture path, and input the picture corresponding to the original picture path and the input information into a multi-modal large language model to obtain a third retrieval result subset; Obtain the original Hypertext Markup Language (HTML) table path according to the metadata corresponding to each second retrieval result in the second retrieval result subset; Call a VLLM according to the original HTML table path to obtain the HTML table corresponding to the original HTML table path, and input the HTML table corresponding to the original HTML table path and the input information into a multi-modal large language model to obtain a fourth retrieval result subset; Add the third retrieval result subset and the fourth retrieval result subset to the fourth retrieval result set.

6. The method according to claim 2, characterized in that, The method further includes: Format the text information, title information, picture information, and table information corresponding to each of the PDFs to obtain text information in a preset format, title information in a preset format, picture information in a preset format, and table information in a preset format.

7. A question-and-answer device, characterized in that, Including: An information acquisition unit for acquiring an embedding vector corresponding to the input information; A model acquisition unit for acquiring intention information corresponding to the input information and acquiring at least one retrieval model corresponding to the intention information; A result acquisition unit for retrieving in a database according to the at least one retrieval model and the embedding vector to obtain a retrieval result set corresponding to the input information, wherein the retrieval result set includes at least one retrieval result corresponding to each retrieval model in the at least one retrieval model, and the database is obtained by processing at least one long document in a portable document format (PDF); The information acquisition unit is further configured to acquire language category information corresponding to the input information, input a prompt corresponding to the language category information into a multi-modal large language model for question and answer generation, and acquire answer information corresponding to the input information.

8. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-6.

9. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.

10. A computer program product, including a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1-6.

Citation Information

Cited By

  • Knowledge base document intention recognition method based on layout recognition

    CN121072695A

  • Technical drawing information extraction device, technical drawing information extraction method, and technical drawing information extraction program

    JP7888845B1