Question and answer method and device, equipment, storage medium and program product
By constructing a multimodal question-answering library and performing semantic segmentation, structured question data is generated, solving the illusion problem in multimodal question-answering systems and improving the accuracy and efficiency of question answering.
Patent Information
- Application Number
- CN202511699483.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-02-17
AI Technical Summary
Existing multimodal question-answering systems are prone to the "illusion" problem, where the generated question-answer content appears reasonable but is actually inconsistent with the facts or the user's requirements.
By pre-building a question-and-answer database based on multimodal documents, receiving user question data, retrieving target document fragments from the question-and-answer database, performing semantic segmentation, generating structured question data, and finally outputting multimodal question-and-answer data, the system avoids relying on question-and-answer models to directly process multimodal documents.
It improves the efficiency and accuracy of image retrieval, reduces the possibility of the model generating illusions, and enhances the accuracy of multimodal question answering.
Smart Images

Figure CN121542382A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing technology, and in particular to a question-answering method, apparatus, device, storage medium, and program product. Background Technology
[0002] Large-scale document question answering systems are an advanced technology that analyzes user-submitted questions and precisely retrieves the most relevant document fragments from a vast document library. The system then leverages powerful modeling capabilities to generate detailed answers to user questions based on these retrieved fragments.
[0003] In related large-scale document question-answering systems, a multimodal approach is used, combining images and text to provide a richer question-answering experience. However, the generated questions and answers are prone to the "illusion" problem, where the model generates content that seems reasonable but is actually inconsistent with the facts or the user's requirements. Summary of the Invention
[0004] Therefore, it is necessary to provide a question-answering method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the accuracy of multimodal question answering in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a question-and-answer method, including:
[0006] Acquire user question data, and search for the corresponding target document fragment from a preset question and answer database based on the user question data. The preset question and answer database is built based on multimodal documents.
[0007] The target document fragment is segmented to obtain multiple answer text fragments;
[0008] Based on preset question-and-answer templates, question data, and multiple answer text fragments, structured question data is generated;
[0009] Based on the structured question data, generate and output the corresponding multimodal question-and-answer data.
[0010] Secondly, this application also provides a question-and-answer device, comprising:
[0011] The data acquisition module is used to acquire user-generated question data.
[0012] The document retrieval module is used to find the corresponding target document fragments from a preset question and answer database based on user question data. The preset question and answer database is built based on multimodal documents.
[0013] The document segmentation module is used to segment the target document fragment into multiple answer text fragments;
[0014] The structured question data generation module is used to generate structured question data based on preset question and answer templates, question data, and multiple answer text fragments;
[0015] The question-and-answer module is used to generate and output corresponding multimodal question-and-answer data based on structured question data.
[0016] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in any of the above-described question-and-answer method embodiments.
[0017] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in any of the above-described question-and-answer method embodiments.
[0018] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in any of the above-described question-and-answer method embodiments.
[0019] Compared to related technologies that rely entirely on large models to generate multimodal question-answering results, the aforementioned question-answering methods, apparatus, computer devices, computer-readable storage media, and computer program products, by pre-constructing a question-answering library based on multimodal documents, receive user-sent question data during the actual multimodal question-answering process and retrieve target document fragments matching the question data from the question-answering library. This allows for simultaneous retrieval of text and images, improving image retrieval efficiency and accuracy while reducing the possibility of model "illusions." Subsequently, the target document is semantically segmented to obtain multiple answer text fragments. Then, based on a pre-set question-answering template, question data, and answer text fragments, structured question data is generated. Finally, multimodal question-answering data is output based on the structured question data. This approach directly processes multimodal documents without relying on a question-answering model. After processing the multimodal documents to construct a question-answering library, it incorporates the data into the retrieval process, generates structured question data from the retrieved answer text fragments, and outputs multimodal question-answering data, thus improving the accuracy of multimodal question-answering. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1This is a diagram illustrating the application environment of the question-and-answer method in one embodiment;
[0022] Figure 2 This is a flowchart illustrating a question-and-answer method in one embodiment;
[0023] Figure 3 This is a flowchart illustrating the question-and-answer method in another embodiment;
[0024] Figure 4 This is a flowchart illustrating the question-and-answer method in yet another embodiment;
[0025] Figure 5 This is a flowchart illustrating the question-and-answer method in a detailed embodiment;
[0026] Figure 6 This is a structural block diagram of a question-and-answer device in one embodiment;
[0027] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0029] The question-and-answer method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another network server.
[0030] Specifically, terminal 102 may initiate a question request to server 104, carrying user question data. Server 104 obtains the user question data, and then, based on the user question data, retrieves the corresponding target document fragment from a preset question-and-answer database. The preset question-and-answer database is constructed based on multimodal documents. Then, the target document fragment is semantically segmented to obtain multiple answer text fragments. After that, based on the preset question-and-answer template, question data, and answer text fragments, structured question data is generated. Finally, based on the structured question data, multimodal question-and-answer data is output.
[0031] The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0032] In one exemplary embodiment, such as Figure 2 As shown, a question-answering method is provided, which can be applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps (hereinafter referred to as S): S100 to S400. Wherein:
[0033] S100: Obtain user question data and search for the corresponding target document fragment from the preset question and answer database based on the user question data. The preset question and answer database is constructed based on multimodal documents.
[0034] The user-generated question data includes text entered by the user through the terminal.
[0035] Multimodal documents can be documents containing text paragraphs and embedded images. Sources of multimodal documents include, but are not limited to, technical documents, corporate knowledge bases, contracts and regulations, and academic materials. Technical documents may include product manuals, API documentation, and user manuals (such as "XX System Operation Guide.pdf"); multimodal documents in corporate knowledge bases may include internal FAQs, process specifications, and operation and maintenance manuals; contracts and regulations may include legal provisions, service agreements, and compliance policies; academic materials may include papers, research reports, and textbook chapters. The structure of multimodal documents may include, but is not limited to, PDF, Word (.docx), PPT, HTML web pages, Markdown, and TXT.
[0036] The preset question-and-answer database can contain multiple document fragments, which are obtained by dividing multimodal documents.
[0037] In practical applications, images can be extracted from multimodal documents in advance, then semantic analysis can be performed on the images, the semantically analyzed text can be merged into the multimodal document, and then the merged multimodal document can be segmented to obtain multiple document fragments. The document fragments can be converted into embedding vectors and stored in the library along with the document fragments to obtain a question-answering library.
[0038] In practice, a user can input question data and submit a question request through the interactive interface of a question-and-answer system on a terminal. The server then responds to the user's request by receiving the question data. Subsequently, the question data can be converted into vectors, and based on the similarity between vectors, target document fragments with high similarity to the question data can be found in the question-and-answer database. Understandably, there can be multiple target document fragments.
[0039] S200: Segment the target document fragment to obtain multiple answer text fragments.
[0040] In practice, the retrieved target document fragments are semantically segmented into multiple answer text fragments, ensuring that each fragment accurately reflects its core meaning. Specifically, this can be achieved by using a pre-trained model to semantically segment the target document fragments into multiple text fragments.
[0041]
[0042] in, It is a vector of the target document fragment retrieved based on the query data. It is a function based on semantic segmentation. It is a segmented answer text.
[0043] S300: Generate structured question data based on the preset question-and-answer template, the question data, and the multiple answer text fragments.
[0044] The preset question-and-answer templates are constructed through prompt engineering, also known as instruction engineering, prompt engineering, or prompt template engineering. This technique improves the output quality of generative artificial intelligence models by optimizing input text. Its core lies in designing precise prompts to guide the model to accurately capture user needs. Based on the working principle of pre-trained speech models, this technology constructs structured prompts through three key elements: task, instruction, and role, guiding the model to generate expected content in conjunction with the context.
[0045] The task is to clearly and concisely describe what the model is expected to generate. Specifically, the task defines the specific activities (actions, objects of operation, and target output) that the model needs to perform.
[0046] Instructions: Specific guidelines that the model must follow when generating text. Specifically, instructions refine and constrain the task, specifying the model's methods, processes, standards, and prohibitions for performing the task to ensure the quality and standardization of the output. For example, explicitly requiring the model to think or output step-by-step, specifying the form of the output, defining the writing style of the output, and imposing content constraints (such as including at least two examples in the output).
[0047] Role: Define the role the model plays in the process of generating text, thereby changing the priority of the model's "knowledge base" calls and the tone of its output, making its answers more professional and targeted.
[0048] In practice, roles, tasks, and instructions can be set according to the question-and-answer requirements to generate preset question-and-answer templates. Then, the answer text fragments and question data can be integrated into structured prompts to generate structured question data.
[0049]
[0050] Among them, prompt is the prompt project. It is structured question data.
[0051] S400 generates and outputs corresponding multimodal question-and-answer data based on structured question data.
[0052] Multimodal question-answering data can include both images and text-based responses. Images include, but are not limited to, pictures and charts.
[0053] In practice, structured question data can be input into a trained question-answering model, which outputs multimodal question-answering data that matches the question data and includes textual answers and images. . The trained question-answering model can be a deep learning model trained using a large amount of text data through unsupervised learning.
[0054] In the aforementioned question-answering method, compared to related technologies that rely entirely on large models to generate multimodal question-answering results, this method pre-builds a question-answering library based on multimodal documents. During the actual multimodal question-answering process, it receives user-sent question data and retrieves target document fragments matching the question data from the question-answering library. This allows for simultaneous retrieval of text and images, improving image retrieval efficiency and accuracy while reducing the possibility of model "illusions." Subsequently, the target document is semantically segmented to obtain multiple answer text fragments. Then, based on the pre-set question-answering template, question data, and answer text fragments, structured question data is generated. Finally, multimodal question-answering data is output based on the structured question data. This approach directly processes multimodal documents without relying on a question-answering model. After processing the multimodal documents to build a question-answering library, it incorporates the data into the retrieval process, generates structured question data from the retrieved answer text fragments, and outputs multimodal question-answering data, thus improving the accuracy of multimodal question answering.
[0055] In one exemplary embodiment, such as Figure 3 As shown, multimodal documents include document text and document images. The construction process of the preset question-answering database includes the following steps S110 to S130, wherein:
[0056] S110, perform image content semantic recognition on the document image to obtain the first semantic text, and extract the text content of the document image to obtain the second semantic text.
[0057] In this context, a multimodal document can be a document containing text paragraphs and embedded images. Images can include pictures and charts, and the format of a multimodal document can include, but is not limited to, PDF, Word, TXT, and HTML formats. Document text represents the text content within the multimodal document, while document images represent the images within the multimodal document. The first semantic text may contain the semantic content of the document images, such as their meaning and purpose. Identifying the semantic content of image content can be achieved by predicting the semantic content of the image content in the document images using a pre-trained model. The second semantic text contains the text content within the document images.
[0058] In practice, the process begins with an operator uploading a multimodal document to a server via a terminal. The server then uses a document parser that matches the document's format to extract text from formats such as PDF, Word, TXT, and HTML. This extraction process reveals the text content and location (e.g., page numbers, paragraph numbers) within the multimodal document. Next, image objects within the multimodal document are identified, and the document images, their locations within the document, and links to these images are generated.
[0059] Secondly, the document image can be input into a pre-trained image semantic recognition model. The image language recognition model understands the context of the image and generates first semantic text containing a description of the meaning of the document image (such as objects, scenes, attributes, actions, relationships, and emotions) and a description of its purpose (such as application scenarios and objectives). .
[0060] Next, the text in the document image can be extracted using a trained text recognition model to obtain the second semantic text. Alternatively, an OCR (Optical Character Recognition) engine can be used to recognize the text in the document image to obtain the second semantic text. :
[0061]
[0062] Among them, OCR is an image-to-text recognition method, Image is the image in the multimodal document, and the output is the text recognized and extracted from the image.
[0063] S120, merge the links of document text, first semantic text, second semantic text, and document image into a multimodal document.
[0064] In practice, for each document image, based on its identified position within the multimodal document, the document image can be replaced with its first semantic text, second semantic text, and a link to the document image. Subsequently, based on the identified text document's position within the multimodal document, the document text is merged into the multimodal document. In this embodiment, the extracted text content is integrated back into the original multimodal document and closely linked with the original image's link, forming a comprehensive and coherent document to ensure information integrity and traceability, as well as seamless information fusion and efficient presentation.
[0065] S130: Encode the merged multimodal document segments to obtain multiple document segment encoding features, and construct a preset question-and-answer database based on the multiple document segment encoding features.
[0066] The document fragment encoding features may include, but are not limited to, vectors or matrices of the divided document fragments.
[0067] In practice, the merged multimodal document is fragmented into multiple document fragments. This can be achieved through methods including, but not limited to, rule-based segmentation, semantic clustering-based segmentation, and machine learning model-based segmentation. Rule-based segmentation can include segmentation based on a fixed number of characters or specific characters. Semantic clustering-based segmentation involves inputting the merged multimodal document into an embedding model. The embedding model converts each sentence or paragraph in the merged multimodal document into an embedding vector, evaluates the semantic relationships between sentences or paragraphs by calculating the similarity between embedding vectors (e.g., cosine similarity), and groups text paragraphs with similarity higher than a preset similarity threshold into blocks, resulting in multiple document fragments. Machine learning model-based methods involve inputting the merged multimodal document into a trained natural language model (e.g., BERT and other Transformer models), analyzing the text through the model's attention mechanism, and dividing the merged multimodal document into multiple document fragments.
[0068] After dividing the document into multiple document segments, the document segments are encoded to obtain multiple document segment encoding features. This can be done by encoding the document segments into a word frequency matrix using a word segmentation algorithm, or by converting the document segments into vectors using word embedding, or by converting the document segments into vectors using a word frequency-inverse document frequency model, a word-vector model, or a document-vector model.
[0069] In this embodiment, integrating the extracted document text, document image first semantic text, second semantic text, and image links into a multimodal document improves data relevance and retrieval convenience. Furthermore, segmenting and encoding the merged multimodal document to construct a pre-defined question-and-answer database enhances the efficiency and accuracy of question-and-answer result retrieval.
[0070] In one exemplary embodiment, such as Figure 4 As shown, image content semantic recognition is performed on the document image to obtain the first semantic text, including S112 to S114:
[0071] S112, extract the context information of the document image to obtain the context text.
[0072] S114. Based on the context text and the preset image prompt word template, identify the semantics of the document image to obtain the first semantic text.
[0073] Contextual text includes the preceding and following text of the document image. Cue word templates refer to the input provided by the user to the large language model, used to guide the model in generating text output of a specific type, topic, or format. This input can be a question, a description, a set of keywords, or contextual information.
[0074] In specific implementation, for S122, it can be based on the image forward traversal technique to locate and extract the coherent text content of the first n characters in the document image, thus obtaining the preceding text of the image, i.e., T1:
[0075]
[0076] ExtractText is a function that takes an image and a numeric parameter, representing the length of text to be extracted from the image. For example, It is understandable that the quantity n here is just an example and is not a unique limitation.
[0077] Then, starting from the end of the document image, traverse backwards to locate and extract the continuous text content of the last n characters of the document image, obtaining the text following the image, denoted as T2:
[0078]
[0079] in, This is a function that accepts an image and a numeric parameter, representing the text to be extracted from the image for a specified length. For example, It is understandable that the quantity n here is just an example and is not a unique limitation.
[0080] For S124, after obtaining the context text T1 (preceding text) and T2 (following text) containing the document image, the context text is analyzed using a preset image prompt template and a large language model to predict the potential use and deeper meaning of the image. The preset image prompt template includes the analysis role, analysis task, and output format. The image prompt template is used to optimize the input of the large model, enabling the large model to output accurate meanings of the images. For example, the preset image prompt template prompt is as follows:
[0081] You are a professional tool for predicting the meaning of images based on context. Please predict the purpose and meaning of an image based on its context.
[0082] enter:
[0083] Image preceding text: T1.
[0084] Image text following it: T2.
[0085] Output: "...".
[0086] The contextual text is integrated into a pre-defined image cue template to generate structured image semantic recognition cue words. These structured cue words are then input into a large language model to guide it in predicting the image's purpose and meaning based on the cue words, thus obtaining the first semantic text. :
[0087]
[0088] in, It is a large language model that takes two text parameters, T1 and T2, and outputs predictions of the purpose and meaning of an image based on structured cue words.
[0089] In one exemplary embodiment, the merged multimodal document fragments are encoded to obtain multiple image-text encoding features, including S132 to S136, wherein:
[0090] S132, the merged multimodal document is segmented to obtain multiple document fragments.
[0091] S134, for each document image, divide the first semantic text, the second semantic text, and the link of the document image into the same document segment.
[0092] S136, Encode multiple document fragments to obtain multiple document fragment encoding features.
[0093] In practice, the merged multimodal document can be traversed and segmented based on sentences or paragraphs. During the segmentation process, based on the identified document image positions, the first semantic text `predict_context`, the second semantic text `predict_context`, and the links to the document images are detected. In this case, it can be a link between the first semantic text predict_context of the document image, the second semantic text predict_OCR, and the document image. To ensure the integrity of the semantic content of document images, they are divided into the same document fragment chunk (such as adjacent document fragments).
[0094] After segmenting and obtaining multiple document fragments, it can be done through... The model converts each document fragment into an embedding vector to obtain the document fragment encoding features corresponding to each document fragment. Then, the document fragment encoding features of each document fragment are stored in a pre-set question-answering database. :
[0095]
[0096] In this embodiment, by dividing the relevant text of the document image into the same document segment, not only is the information organization structure optimized, but the efficiency and accuracy of document management are also improved. Furthermore, it is beneficial to improve the relevance of data and the convenience of retrieval.
[0097] In an exemplary embodiment, searching for the corresponding target document fragment from a preset question-and-answer database includes:
[0098] The question data is encoded to obtain question encoding features.
[0099] The question encoding features can include vectors or matrices of question data.
[0100] In specific implementation, this can be achieved by encoding the question data into a word frequency matrix using a word segmentation algorithm to obtain the question encoding features corresponding to the question data; alternatively, it can be achieved by converting the question data into a vector using word embedding to obtain the question encoding features corresponding to the question data; or it can be achieved by converting the question data into a vector using a word frequency-inverse document frequency model, a word-vector model, or a document-vector model to obtain the question encoding features corresponding to the question data. For example, it can be achieved through... The model converts the question data into embedding vectors, obtaining the corresponding question encoding features, vectors. .
[0101] A similarity analysis was performed on the coding features of the questions and the coding features of each document fragment in the preset question-and-answer database to obtain the similarity analysis results.
[0102] The similarity analysis results include the similarity between the question coding features and the coding features of each document fragment in the preset question-and-answer database.
[0103] In practice, similarity algorithms can be used to encode features for the query. and preset question and answer database Similarity calculations are performed on the encoding features of each document fragment to determine the query encoding features. Question and Answer Database The similarity between the encoded features of each document fragment is used to obtain the similarity analysis results. The similarity algorithms include, but are not limited to, Euclidean distance, Manhattan distance, and cosine similarity.
[0104] Based on the similarity analysis results, the target document fragment corresponding to the query data is determined.
[0105] In practice, the similarity can be compared with a preset similarity threshold, and document fragment encoding features with similarity higher than the preset similarity threshold can be selected. The document fragments corresponding to the selected document fragment encoding features are then identified as the target document fragments corresponding to the query data. :
[0106]
[0107] in, This refers to the retrieval method, where `query` represents the user's question data, and `embed` is a pre-defined question-and-answer database. It is the retrieved target document fragment.
[0108] In other implementations, the document fragment encoding features can be sorted in descending order according to similarity, a preset number of document fragment encoding features can be filtered, and the document fragments corresponding to the filtered document fragment encoding features can be determined as the target document fragments corresponding to the question data.
[0109] In this embodiment, by using a preset question-and-answer database and similarity retrieval between encoded features, text and image retrieval can be performed simultaneously on the question data, thereby improving retrieval efficiency and accuracy.
[0110] In one exemplary embodiment, the answer text fragment is configured with an index identifier; based on the structured question data, corresponding multimodal question-and-answer data is generated and output, including:
[0111] Select the target answer text fragment from the answer text fragments of structured question data.
[0112] Based on the target answer text fragment, generate question and answer text, and obtain question and answer images based on the index identifier of the target answer text fragment. Output question and answer text, index identifier, and question and answer images. Multimodal question and answer data includes question and answer text, index identifier, and question and answer images.
[0113] In this embodiment, after segmenting the target text fragment into multiple answer text fragments, the method further includes assigning a unique index identifier (ID) to each answer text fragment:
[0114]
[0115] in, This function adds an index ID to each segmented answer text fragment. It is a text snippet of the answer with an index ID added.
[0116] When assigning an index identifier to each answer text fragment, if the answer text fragment contains at least one of the first semantic text, the second semantic text, and the image link of the document image, then the index identifier of the answer text fragment is mapped to the document image to facilitate the output of the document image.
[0117] In this embodiment, the structured question data is generated based on a preset question-and-answer template containing tasks, instructions, and roles, question data, and target text fragments, and also includes the structure of tasks, instructions, and roles. The tasks in the structured question data are generated based on the question data and target text fragments, while the instructions constrain the scope of data analysis, the readability of the question-and-answer results, and the output format of the results. For example, in this embodiment, the preset question-and-answer template prompt is as follows:
[0118] You are a professional customer service assistant. Please carefully review the following document and, based on its content, provide a clearly structured and logically sound answer, along with the reference index ID. Service requirements:
[0119] - Please answer based on the content of the document; do not use knowledge beyond the content of the document to answer.
[0120] - The sentences should be fluent, well-organized, and easy to read;
[0121] - If no article answers this question, please answer "I don't know" and output "none" for the reference ID;
[0122] Remember, you must return the answer first, then the citation ID of the article. For each citation in all articles that supports the answer, return the citation ID; the answer may contain multiple citations. If there are no citations, output no citation ID. Please be sure to output in the following format:
[0123] Answer: Output the following: [id:x]、[id:x]
[0124] In practice, structured question data can be input into a trained question-answering model to guide the model in selecting target answer text fragments for use in multimodal question-answering data, based on the question data and answer text fragments in the structured question data. This forms the data foundation for generating multimodal question-answering data. Subsequently, the question-answering model generates and outputs question-answer text based on the target answer text fragments, resulting in text-based question-answering results. The output of the question-answer text also includes the index identifier of the referenced target answer text fragment. Afterward, a mapping relationship can be established between the index identifier of the target answer text fragment and a document image. If a mapping relationship exists, the mapped document image is identified as the question-answer image and output. Thus, the output contains multimodal question-answering data including question-answer text, the index identifier of the referenced target answer text fragment, and the question-answer image.
[0125] For example, the generated structured question data and the output multimodal question-answer data are as follows:
[0126] Filename: Origin of the Festival
[0127] Title: Spring Festival and Mid-Autumn Festival
[0128] Document Content: [id:1] The Mid-Autumn Festival originated from ancient moon worship and sacrificial activities, first mentioned in the *Zhou Li* (Rites of Zhou), and gradually spread from the court to the common people. [id:2] It is also related to celebrating the autumn harvest; the fifteenth day of the eighth lunar month falls in the middle of autumn, hence the name "Mid-Autumn." [id:3] The Spring Festival, also known as the Lunar New Year, is one of China's most important traditional festivals. [id:4] Its origins can be traced back to ancient year-end sacrificial activities, first mentioned in ancient books such as the *Shang Shu* (Book of Documents). The Spring Festival marks the beginning of the Lunar New Year and is closely related to the calendar of agricultural society. On this day, people celebrate the harvest and pray for good fortune and auspiciousness in the coming year.
[0129] The question is: What is the origin of the Mid-Autumn Festival?
[0130] Answer: The Mid-Autumn Festival originated from the ancient worship and sacrificial activities for the moon, first mentioned in the "Rites of Zhou," and gradually spread from the court to the common people. It is also related to celebrating the autumn harvest, as the fifteenth day of the eighth lunar month falls in the middle of autumn, hence the name "Mid-Autumn Festival." [id:1], [id:2].
[0131] In this embodiment, the retrieved target text fragments are pre-segmented with finer granularity and labeled with index identifiers. Structured question data is then generated by combining the question data and a preset question-and-answer template. This results in the output of the answer along with a reference index ID. This method not only adds image retrieval functionality to the document question-and-answer system but also effectively avoids the problem of image illusions generated by the model, and fully utilizes the textual information in the images to provide users with a completely new interactive experience.
[0132] To provide a clearer explanation of the question-and-answer method provided in this application, a specific embodiment and appendix are described below. Figure 5 The specific embodiment includes the following steps:
[0133] Before question answering, the following steps S101 to S105 are performed to build the question-and-answer library for indexing:
[0134] S101, Text Extraction and Image Parsing: Extract the document text and document images from the multimodal document respectively.
[0135] S102, extract the context information of the document image to obtain the context text, and based on the context text and the preset image prompt word template, identify the semantics of the document image to obtain the first semantic text.
[0136] S103, extract the text content of the document image to obtain the second semantic text.
[0137] S104, Preset Question-Answer Base Construction: Merge the document text, first semantic text, second semantic text, and document image links into a multimodal document. Segment the merged multimodal document to obtain multiple document fragments. For each document image, divide the first semantic text, second semantic text, and document image links into the same document fragment. Encode the multiple document fragments to obtain multiple document fragment encoding features. Construct a preset question-answer base based on the multiple document fragment encoding features.
[0138] During the specific question and answer process, execute steps S1 to S5 to obtain the question and answer results:
[0139] S1, Obtain user question data.
[0140] S2, Query and Retrieval Fragments: The query data is encoded to obtain the query encoding feature query. The query encoding feature is compared with the encoding features of each document fragment in the preset question and answer database to obtain the similarity analysis results. Based on the similarity analysis results, the target document fragment corresponding to the query data is determined.
[0141] S3, the target document fragment is segmented to obtain multiple answer text fragments, each of which carries an index identifier.
[0142] S4, prompt project: Generates structured question data based on preset question and answer templates, question data, and multiple answer text fragments with index identifiers.
[0143] S5, Answer: Text + Image: Filter the target answer text fragment from the answer text fragments of the structured question data, generate question and answer text based on the target answer text fragment, and obtain the question and answer image based on the index identifier of the target answer text fragment. Output the question and answer text, index identifier, and question and answer image. Multimodal question and answer data includes question and answer text, index identifier, and question and answer image.
[0144] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0145] In one exemplary embodiment, such as Figure 6 As shown, a question-answering device 600 is provided, including: a data receiving module 610, a document retrieval module 620, a semantic segmentation module 630, a structured question data generation module 640, and a question-answering module 650, wherein:
[0146] Data receiving module 610 is used to acquire user question data;
[0147] The document retrieval module 620 is used to search for the corresponding target document fragment from a preset question and answer database based on user question data. The preset question and answer database is constructed based on multimodal documents.
[0148] The semantic segmentation module 630 is used to segment the target document fragment to obtain multiple answer text fragments;
[0149] The structured question data generation module 640 is used to generate structured question data based on a preset question and answer template, question data, and multiple answer text fragments;
[0150] The question-and-answer module 650 is used to generate and output corresponding multimodal question-and-answer data based on structured question data.
[0151] In an exemplary embodiment, the question-answering device 600 further includes a question-answering library construction module 660, which is used to perform image content semantic recognition on the document image to obtain first semantic text, and extract the text content of the document image to obtain second semantic text, merge the document text, the first semantic text, the second semantic text, and the link of the document image into a multimodal document, encode the merged multimodal document into segments to obtain multiple document segment encoding features, and construct a preset question-answering library based on the multiple document segment encoding features.
[0152] In an exemplary embodiment, the question-answering library construction module 660 is further configured to extract contextual information of the document image to obtain contextual text, and to identify the semantics of the document image based on the contextual text and a preset image prompt word template to obtain first semantic text.
[0153] In an exemplary embodiment, the question-answering library construction module 660 is further configured to segment the merged multimodal documents to obtain multiple document fragments, and for each document image, to divide the first semantic text, the second semantic text, and the link of the document image into the same document fragment, and to encode the multiple document fragments to obtain multiple document fragment encoding features.
[0154] In an exemplary embodiment, the document retrieval module 620 is further configured to encode the question data to obtain question encoding features, perform similarity analysis on the question encoding features and the encoding features of each document fragment in the preset question-and-answer database to obtain similarity analysis results, and determine the target document fragment corresponding to the question data based on the similarity analysis results.
[0155] In an exemplary embodiment, the question-answering module 650 is further configured to filter out a target answer text fragment from the answer text fragments of the structured question data, generate question-answering text based on the target answer text fragment, obtain a question-answering image based on the index identifier of the target answer text fragment, and output the question-answering text, the index identifier, and the question-answering image. The multimodal question-answering data includes the question-answering text, the index identifier, and the question-answering image.
[0156] Each module in the aforementioned question-and-answer device 600 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0157] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a question-and-answer method.
[0158] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0159] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in any of the above-described question-and-answer method embodiments.
[0160] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in any of the above-described question-and-answer method embodiments.
[0161] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in any of the above-described question-and-answer method embodiments.
[0162] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0163] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0164] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0165] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A question and answer method, characterized by, The method comprises: acquiring user question data, and searching for a corresponding target document segment from a preset question and answer library according to the user question data, wherein the preset question and answer library is constructed based on a multi-modal document; segmenting the target document segment to obtain a plurality of answer text segments; generating structured question data according to a preset question and answer template, the question data and the plurality of answer text segments; generating and outputting corresponding multi-modal question and answer data according to the structured question data.
2. The method of claim 1, wherein, The multi-modal document comprises document text and document images; the construction process of the preset question and answer library comprises: performing image content semantic recognition on the document images to obtain first semantic text, and extracting text content of the document images to obtain second semantic text; merging the document text, the first semantic text, the second semantic text and links of the document images into the multi-modal document; performing encoding processing on the merged multi-modal document segment by segment to obtain a plurality of document segment encoding features, and constructing the preset question and answer library according to the plurality of document segment encoding features.
3. The method of claim 2, wherein, The image content semantic recognition on the document images to obtain first semantic text comprises: extracting context information of the document images to obtain context text; recognizing semantics of the document images according to the context text and a preset image prompt word template to obtain first semantic text.
4. The method of claim 2, wherein, The encoding processing on the merged multi-modal document segment by segment to obtain a plurality of document segment encoding features comprises: segmenting the merged multi-modal document respectively to obtain a plurality of document segments; for each of the document images, dividing the first semantic text of the document image, the second semantic text and the links of the document image into the same document segment; performing encoding processing on a plurality of the document segments to obtain a plurality of document segment encoding features.
5. The method of claim 2, wherein, The searching for a corresponding target document segment from a preset question and answer library comprises: performing encoding processing on the question data to obtain question encoding features; performing similarity analysis on the question encoding features and document segment encoding features in the preset question and answer library to obtain a similarity analysis result; determining a target document segment corresponding to the question data according to the similarity analysis result.
6. The method according to any one of claims 1 to 5, characterized in that, The answer text segment is configured with an index identifier; the generating and outputting of corresponding multi-modal question and answer data according to the structured question data comprises: filtering a target answer text segment from the answer text segment of the structured question data; generating question and answer text according to the target answer text segment, obtaining question and answer images according to the index identifier of the target answer text segment, and outputting the question and answer text, the index identifier and the question and answer images; The multi-modal question and answer data comprises the question and answer text, the index identifier and the question and answer images.
7. A question and answer apparatus characterized by comprising: The device comprises: a data acquisition module configured to acquire user question data; a document retrieval module configured to search for a corresponding target document segment from a preset question and answer library according to the user question data, wherein the preset question and answer library is constructed based on a multi-modal document; a document segmentation module, configured to segment the target document segment to obtain a plurality of answer text segments; a structured question data generation module, configured to generate structured question data according to a preset question and answer template, the question data, and the plurality of answer text segments; a question and answer module, configured to generate and output corresponding multi-modal question and answer data according to the structured question data. 8.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-7. The processor executes the computer program to implement the steps of the method in any one of claims 1 to 6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.
Citation Information
Patent Citations
Question and answer method and device, equipment and medium
CN118445395A
RAG-based construction technology multi-mode intelligent knowledge base question and answer processing method, medium and equipment
CN119938817A
Hierarchical retrieval method and device based on multi-modal questions and answers and computer equipment
CN120045683A
Process for delivering responses to queries expressed in natural language based on a dynamic document corpus
WO2024228712A1