Multi-mode question answering system, reply information generation method and electronic equipment
By constructing a multimodal knowledge base and employing a question-answering model for dual-path retrieval of text and images, the problem of insufficient image information recognition in multimodal question-answering systems is solved, enabling the generation of more accurate and comprehensive response information and enhancing the system's adaptability in complex scenarios.
Patent Information
- Application Number
- CN202511654531.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-02-17
AI Technical Summary
Existing multimodal question-answering systems cannot effectively recognize image information when processing non-text data, have limited room for improvement in the correlation between search results and question information, and have low accuracy in answer information, thus failing to meet users' multimodal search and question-answering needs.
A multimodal knowledge base integrating knowledge images and their context is constructed. A question-answering model is used for dual-path retrieval of text and images to generate response information. Multimodal and text resources are integrated, and the knowledge information stored in the knowledge base and the feature extraction capabilities of the question-answering model are optimized.
It improves the accuracy and comprehensiveness of responses during multimodal question answering, enhances the adaptability of intelligent question answering systems to complex scenarios, and ensures the integrity of text-image association and the retention of information.
Smart Images

Figure CN121543716A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of information technology, and more specifically, to a multimodal question-and-answer system, a method for generating response information, and an electronic device. Background Technology
[0002] With the advent of large-scale models in the multimodal era, retrieval augmentation generation (RAG) technology has emerged. However, while RAG can improve question-answering quality by leveraging external knowledge bases, existing models still rely heavily on text processing. When processing non-text data such as question images or retrieval images, they are often converted into text descriptions and processed based on text processing logic. This results in poor recognition of image information, room for improvement in the relevance between search results and question information, and unsatisfactory accuracy of the responses generated based on search results, failing to meet users' multimodal retrieval and question-answering needs.
[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0004] The purpose of this disclosure is to provide a multimodal question-answering system, a method for generating response information, and an electronic device to improve the accuracy of response information generated during the multimodal question-answering process.
[0005] According to a first aspect of the present disclosure, a multimodal question-answering system is provided, comprising: a multimodal knowledge base including a first image, the first image including a knowledge image and a context corresponding to the knowledge image; a text knowledge base including multiple knowledge texts; and a question-answering model for receiving question information and generating response information based on the question information and a target first image in the multimodal knowledge base that matches the question information and / or a target knowledge text in the text knowledge base that matches the question information, wherein the response information includes the target first image and / or the target knowledge text, and the question information includes text question information.
[0006] According to a second aspect of the present disclosure, a method for generating response information for a multimodal knowledge base is provided, applied to a multimodal question-answering system as described in any of the preceding claims, comprising: responding to a question information input message; determining a prompt text feature matrix corresponding to the question information, wherein the question information includes text question information; acquiring an image feature matrix of a first image in the multimodal knowledge base and a knowledge text feature matrix of a knowledge text corresponding to the text knowledge base, wherein the first image includes a knowledge image and a context corresponding to the knowledge image; determining the image feature matrix with the highest first similarity to the prompt text feature matrix as a target image feature matrix; determining the knowledge text feature matrix with the highest second similarity to the prompt text feature matrix as a target knowledge text feature matrix; and merging and deduplicating the target image feature matrix and the target knowledge text feature matrix to form response information corresponding to the question information.
[0007] According to a third aspect of this disclosure, an electronic device is provided, comprising: a memory storing either a multimodal knowledge base or a question-answering model in a multimodal question-answering system as described in any of the preceding claims, wherein the question-answering model is configured to perform the response information generation method as described above based on instructions stored in the memory.
[0008] According to a fourth aspect of this disclosure, a computer-readable storage medium is provided having a program stored thereon that, when executed by a processor, implements the system or method as described in any of the preceding claims.
[0009] According to a fifth aspect of this disclosure, a computer program product is provided, comprising a computer program, characterized in that, when executed by a processor, the computer program implements the steps of the method as described in any of the preceding claims, or implements the system described above.
[0010] This disclosure, through the construction of a multimodal knowledge base including a first image (integrating the knowledge image and its context), avoids the fragmentation of image-text association caused by the traditional separation of text and image processing, fully preserves modal information, and solves the problem of information loss in non-text information processing. In addition, the question-answering model performs dual-path retrieval of text and image based on the question information, recalling the target first image of the multimodal knowledge base and / or the target knowledge text of the text knowledge base to generate response information. By integrating multimodal and text resources, it can effectively make up for the deficiency of insufficient multimodal matching ability, improve the comprehensiveness and accuracy of retrieval, and thus make the response information more accurate, better meet the user's question-answering needs, and enhance the adaptability of the intelligent question-answering system to complex scenarios.
[0011] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.
[0013] Figure 1 This is a flowchart of a multimodal question-answering system in an exemplary embodiment of this disclosure.
[0014] Figure 2 This is a flowchart of the formation process of a multimodal knowledge base in an exemplary embodiment of this disclosure.
[0015] Figure 3 This is a schematic diagram of the first image in an exemplary embodiment of this disclosure.
[0016] Figure 4 This is a flowchart of the training process of the question-answering model in an exemplary embodiment of this disclosure.
[0017] Figure 5 This is a flowchart illustrating the generation of response information by question-answering model 3 in an exemplary embodiment of this disclosure.
[0018] Figure 6 This is a sub-flowchart of step S54 in an embodiment of this disclosure.
[0019] Figure 7 This is a schematic diagram of the operation process of system 100 in this embodiment of the present disclosure.
[0020] Figure 8 This is a flowchart of a response information generation method based on a multimodal knowledge base in an exemplary embodiment of this disclosure.
[0021] Figure 9 This is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0022] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.
[0023] Furthermore, the accompanying drawings are merely illustrative of this disclosure, and the same reference numerals in the drawings denote the same or similar parts, thus repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.
[0024] The exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0025] Figure 1 This is a flowchart of a multimodal question-answering system in an exemplary embodiment of this disclosure.
[0026] refer to Figure 1 The multimodal question-answering system 100 may include: Multimodal knowledge base 1 includes multiple multimodal documents, each multimodal document including a first image, the first image including a knowledge image and the context corresponding to the knowledge image; Text knowledge base 2 includes multiple knowledge texts; Question-answering model 3 is used to receive question information and generate response information based on the question information and a target first image matching the question information in the multimodal knowledge base and / or a target knowledge text matching the question information in the text knowledge base. The response information includes the target first image and / or the target knowledge text, and the question information includes text question information.
[0027] The multimodal question-answering system 100 provided in this embodiment includes multimodal knowledge storage, text knowledge storage, and intelligent question-answering functions. It can receive user question information, generate retrieval results including images and text by searching the multimodal knowledge base 1 and the text knowledge base 2 based on the question information, and generate response information to the question information based on the retrieval results.
[0028] The multimodal knowledge base 1 is used to store multimodal knowledge, which includes knowledge types other than pure text knowledge, such as images and videos. This multimodal knowledge is stored in the form of multimodal documents. In this embodiment, the multimodal knowledge base 1 includes multiple multimodal documents, each including a first image. The first image is not an independent knowledge image, but a composite information carrier that integrates a knowledge image (such as a portrait of a historical figure or a data chart) with its contextual text (such as image descriptions or related interpretive text). For example, a first image may include a photograph of a Great Wall beacon tower (a knowledge image) and corresponding text explaining the defensive functions of the Ming Dynasty beacon tower (the context corresponding to the knowledge image). This integration method can fully preserve the semantic relationship between the image and text, avoid the information fragmentation caused by traditional separate storage, and provide a comprehensive knowledge source that integrates visual and textual information for retrieval.
[0029] Text knowledge base 2 is a knowledge base that stores plain text knowledge. The stored knowledge text can include various structured and unstructured text data, such as academic paper abstracts, encyclopedia entries, policy document fragments, etc. These texts can independently carry knowledge without relying on images. They can quickly match when users ask questions that only require text information, and can also complement the information in multimodal knowledge base 1.
[0030] In some embodiments, multimodal knowledge base 1 and text knowledge base 2 can also be implemented through a single knowledge base. That is, both multimodal knowledge base 1 and text knowledge base 2 are logical concepts, not bound to independent physical storage units. They can be implemented through one or more logical modules according to actual needs. For example, in a single database, multimodal knowledge and knowledge text, including the first image (including knowledge image and context), can be classified and stored by field tags (such as "modal type: image / text"). Subsequent retrieval only requires filtering by field to distinguish and call the relevant information, without the need to deploy two separate physical databases.
[0031] Furthermore, regardless of whether the multimodal knowledge base 1 and the text knowledge base 2 are implemented through a single knowledge base or multiple knowledge bases, from a software implementation perspective, the logical division can be accomplished through a unified database management system (such as MySQL or MongoDB) or a distributed file system (such as HDFS). If it is a single software module, the two types of data can be integrated and stored through data structure design, and the target data can be quickly located through conditional query statements (such as the WHERE clause in SQL) during retrieval. If it is multiple software modules, data interaction between modules can be established through interface calls to ensure that when the question-answering model 3 initiates a retrieval request, the matching results of the two types of knowledge bases can be obtained synchronously.
[0032] From a hardware implementation perspective, a centralized or distributed deployment architecture can be flexibly chosen. In a centralized deployment, the data of both types of knowledge bases can be stored on the same physical server's storage medium (such as hard drives or SSDs), relying on the server's CPU and memory resources to complete data reading, writing, and retrieval calculations. This is suitable for scenarios with small data volumes and low access frequencies. In a distributed deployment, multimodal data (the first image) and text data can be stored on the physical hardware of different nodes (e.g., storing the first image, which occupies more storage space, in a distributed storage cluster, and storing plain text on high-performance computing nodes). Load balancing technologies (such as Nginx) can be used to distribute retrieval tasks, improving response speed and system stability in large-scale data scenarios.
[0033] Question answering model 3 belongs to the category of retrieval-enhanced generative models, and can be implemented, for example, using a specially trained large language model. Question answering model 3 is associated with multimodal knowledge base 1 and text knowledge base 2. It can receive question information, retrieve knowledge from the contents of multimodal knowledge base 1 and text knowledge base 2 based on the question information, and then generate and output answer information based on the retrieval results. Question information includes, but is not limited to, text question information and text + multimodal question information (e.g., text + image).
[0034] This disclosure, through the construction of a multimodal knowledge base including a first image (integrating the knowledge image and its context), avoids the fragmentation of image-text association caused by the traditional separation of text and image processing, fully preserves modal information, and solves the problem of information loss in non-text information processing. In addition, the question-answering model performs dual-path retrieval of text and image based on the question information, recalling the target first image of the multimodal knowledge base and / or the target knowledge text of the text knowledge base to generate response information. By integrating multimodal and text resources, it can effectively make up for the deficiency of insufficient multimodal matching ability, improve the comprehensiveness and accuracy of retrieval, and thus make the response information more accurate, better meet the user's question-answering needs, and enhance the adaptability of the intelligent question-answering system to complex scenarios.
[0035] The following is a detailed description of each part of the multimodal question-answering system 100.
[0036] To address the issue of inaccurate response information generated by existing retrieval enhancement technologies, this disclosure mainly focuses on three aspects: optimizing the knowledge information stored in the knowledge base to provide more comprehensive multimodal reference information; improving retrieval accuracy by enhancing the question-answering model's understanding of question and knowledge information; and improving the accuracy of response information generation by adjusting the question-answering model's response information generation method.
[0037] First, the embodiments of this disclosure optimize the knowledge information stored in the knowledge base.
[0038] In the initial stage of knowledge base construction, traditional methods typically input images and text separately into their respective representation models to generate their own embedding vectors. However, this separate processing approach may cause the inherent connection between images and text to disappear, affecting the data's expressive power; furthermore, since RAG multimodal retrieval requires understanding the features of three modalities, this further increases the difficulty of model understanding.
[0039] To effectively overcome this problem, this disclosure proposes a method of deep stitching and integration of images and their corresponding text content. Specifically, in each multimodal knowledge base entry, the image and text content are integrated to generate a single image (the first image). In this way, the previously scattered image information and text descriptions are unified into a new multimodal data unit, thus laying a solid data foundation for subsequent model training and knowledge base construction. This integration method preserves the integrity and relevance of the image and text semantics to the greatest extent possible, and also simplifies the features of multimodal data, helping the question-answering model 3 to better understand and utilize the rich information in the knowledge base.
[0040] Figure 2 This is a flowchart of the formation process of a multimodal knowledge base in an exemplary embodiment of this disclosure.
[0041] refer to Figure 2 In this embodiment of the disclosure, the process of forming a multimodal knowledge base includes: Step S201: Obtain multimodal documents and identify knowledge images in the multimodal documents; Step S202: Identify the context corresponding to the knowledge image in the multimodal document; Step S203: Based on the context corresponding to the knowledge image and the position of the knowledge image in the multimodal document, take a screenshot of the multimodal document to form the first image corresponding to the knowledge image; Step S204: Form the multimodal document based on the first image.
[0042] Multimodal documents refer to composite data carriers that simultaneously contain multiple modalities of information, such as text, images, and tables, rather than a single text format. Examples include document pages containing portraits of historical figures and corresponding textual descriptions, and research report excerpts with accompanying data charts and interpretive text. Their core characteristic lies in the interconnectedness of different modalities of information, which together carry complete knowledge, making them the basic data source for constructing multimodal knowledge bases.
[0043] In step S201, raw data containing multiple modal information can be collected to obtain multimodal documents. The sources of the raw data can include electronic documents, scanned documents, web page content, etc. Next, through image detection and content analysis technology, images with knowledge-carrying value (such as historical images, technical diagrams, data charts, etc.) are selected from the multimodal documents, i.e., knowledge images, while irrelevant and redundant images (such as decorative icons) are excluded, thus locking in the core visual information for subsequent knowledge integration. The method for selecting knowledge images can be carried out using general methods, and this disclosure does not impose any special restrictions on it.
[0044] In step S202, image analysis technology is first used to determine the information corresponding to the selected knowledge images in the multimodal document, and to determine the description of the knowledge image. Next, text recognition, semantic parsing, and other methods are used to identify the correspondence between text information and the description of the knowledge image within a preset range (e.g., 5 lines around the knowledge image). That is, text content directly related to the information corresponding to the knowledge image is identified in the multimodal document, such as explanatory text below the knowledge image, interpretive statements about the knowledge image in the paragraph containing the knowledge image, etc. Its purpose is to supplement the background information and semantic logic that the knowledge image cannot fully convey, and to ensure the integrity of the knowledge expression.
[0045] In step S203, after determining the location of the knowledge image (e.g., coordinates or paragraph) and the location of its context (e.g., coordinates or paragraph), the knowledge image and its context text are cropped as a whole according to their spatial positional relationship in the original multimodal document (e.g., the context text is located above / below / around the image), and finally the first image is formed. The first image is not a single visual element, but an integrated multimodal information unit that combines the knowledge image (visual modality) and the corresponding context (text modality).
[0046] In step S204, the process of forming a multimodal document based on the first image can either involve storing the first image in association with its corresponding multimodal document, or it can involve saving the first image as a separate multimodal document.
[0047] Figure 3 This is a schematic diagram of the first image in an exemplary embodiment of this disclosure.
[0048] refer to Figure 3In the multimodal document 30, the knowledge image 31 is located in region A, and the location of the context 32 corresponding to the knowledge image 31 is identified in region B. Then, screenshots are taken of regions A and B to generate the first image 33.
[0049] The first image 33 not only contains the knowledge image 31, but also the context 32 that provides supplementary explanations for the knowledge image 31. Thus, it can integrate both textual and image information to provide more complete reference knowledge.
[0050] Therefore, by filtering knowledge images, matching related text, and integrating screenshots, the scattered visual and textual knowledge in the original multimodal documents is transformed into a unified and closely related first image. This not only avoids the information fragmentation problem caused by the separate storage of traditional multimodal data, but also lays a data foundation for the subsequent question-answering model to accurately understand knowledge images and match multimodal knowledge. This ensures that the multimodal knowledge base can provide high-quality multimodal knowledge information, optimize the retrieval results of the question-answering model, and thus improve the accuracy of the response information generated by the question-answering model.
[0051] Even after integrating a multimodal knowledge base, existing retrieval models or multimodal representation models still have shortcomings in Chinese document retrieval, especially when handling scenarios containing both text and images. Against this backdrop, this disclosure proposes an efficient method for fine-tuning models, capable of rapidly adjusting large models in other languages using limited domain data, thereby enabling them to possess superior Chinese multimodal representation capabilities.
[0052] Figure 4 This is a flowchart of the training process of the question-answering model in an exemplary embodiment of this disclosure.
[0053] refer to Figure 4 In an exemplary embodiment, the training process of question-answering model 3 includes: Step S401: Input N questions and M training documents into the question answering model to obtain the query feature matrix of the question information and the document feature matrix of the training documents. The i-th question information corresponds to the i-th query feature matrix, and the j-th training document corresponds to the j-th document feature matrix. 0≤i<N, 0≤j<M. The training documents include the first image. Step S402: Calculate the similarity between the i-th query feature matrix and the j-th document feature matrix as the similarity of the i-th sample; Step S403: Calculate the M-1 similarities between the i-th query feature matrix and the M-1 document feature matrices other than the j-th document feature matrix, and determine the maximum similarity among the M-1 similarities as the i-th hard negative sample similarity. Step S404: Calculate the sum of the difference between the similarity of the i-th hard negative sample and the similarity of the i-th sample and the interval parameter as the i-th value, and set the training loss corresponding to the i-th question information to the maximum value between 0 and the i-th value; Step S405: Determine the current loss function value of the question answering model based on the average of the N training loss values corresponding to the N question information; Step S406: Adjust the parameters and interval parameters of the question-answering model according to the loss function value until the preset training stopping condition is reached.
[0054] Before training, a multimodal document retrieval dataset can be constructed as a training dataset for training the question-answering model 3. This dataset can contain a large number of high-quality Chinese question-answer pairs, representing samples from a Chinese multimodal document knowledge base. Each multimodal knowledge base entry includes not only concatenated and integrated image and text data (i.e., the first image) but also integrated plain text data. This diverse data design significantly enhances the ability of the trained question-answering model 3 to process both image and text data simultaneously.
[0055] In step S401, the question information can be based on the question portion of question-answer pairs in the training database. The number of question information N and the number of documents M used for training are both positive integers and can be dynamically adjusted according to the model training objective.
[0056] In some embodiments, the question information may be plain text information, while in other embodiments, the question information may be a combination of text information and multimodal information, such as including text information and images.
[0057] In the process of feature extraction to form the query feature matrix and the text feature matrix, the query input (question information) and training documents can be processed as tokens, thereby forming a token embedding matrix of the question information as the question feature matrix and a token embedding matrix of the training documents as the document feature matrix. This part can demonstrate the model's feature extraction capability for multimodal information.
[0058] In an exemplary embodiment, the training data may use the target language, thereby enhancing the question-answering model 3's understanding of the target language and its feature extraction capabilities. In this embodiment, the N textual question messages are all textual question messages corresponding to the target language, and the contextual information contained in the N first images is all textual information in the target language, including Chinese.
[0059] In step S402, the similarity between the i-th text feature matrix and the i-th image feature matrix can be calculated using common similarity calculation methods such as cosine similarity.
[0060] In step S403, the difficult negative sample refers to the negative sample that has little semantic difference from the positive sample and is easily misjudged as a match by the model. By focusing on such samples, the ability of question answering model 3 to distinguish similar information can be improved, and the problem of low discrimination of negative samples in traditional training can be solved.
[0061] This embodiment of the disclosure trains the question-answering model 3 by using hard negative samples, thereby training the question-answering model 3 to produce as large a difference as possible in the feature extraction of two similar samples (e.g., two question messages). In this way, by optimizing the feature extraction capability of the question-answering model 3, the probability of feedback error information in subsequent retrieval is reduced, and the retrieval accuracy of the question-answering model 3 is improved.
[0062] Specifically, this embodiment of the disclosure uses hard-to-bear sample similarity to improve the feature extraction capability of question-answering model 3. The method for screening hard-to-bear samples in this embodiment of the disclosure is to calculate the M-1 similarities between the i-th question feature matrix and the M-1 document feature matrices other than the j-th document feature matrix, and determine the maximum similarity among the M-1 similarities as the i-th hard-to-bear sample similarity.
[0063] In step S404, the interval parameter is a learnable model parameter (initial value can be set to 0.5) used to dynamically adjust the decision boundary of positive and negative sample similarity. The i-th value is obtained by calculating the difference between the similarity of the difficult negative sample and the sample similarity, and then superimposing the interval parameter. This i-th value can amplify the difference between positive and negative samples. The training loss is set to the maximum of 0 and the i-th value, meaning that loss is only generated and model optimization is driven when the similarity of the difficult negative sample is close to or exceeds the sample similarity, avoiding meaningless loss calculations and improving training efficiency.
[0064] In step S405, the loss function value is a global loss index obtained by averaging the training loss of N text question information. This index comprehensively reflects the overall matching performance of question answering model 3 on the current batch of training samples, provides a unified basis for adjusting model parameters, and ensures that the model optimization direction conforms to the global training objective.
[0065] In step S406, the model parameters and interval parameters are adjusted based on the loss function value. The parameters of the feature extraction and similarity calculation modules of the question-answering model are updated through the backpropagation algorithm, while the interval parameter is optimized to adapt to the needs of different training stages. The preset training stopping conditions can be set as the loss function value converging to a preset threshold, the number of training iterations reaching a set upper limit, or the model's performance on the validation set no longer improving, ensuring that the model is trained sufficiently and avoiding overfitting.
[0066] Figure 4 The steps shown can be implemented using the following forms and formulas.
[0067] After the knowledge base is built, an open-source pre-trained retrieval model can be introduced. Then, fine-tuning can be performed using Chinese training data to form the feature extraction module (e.g., a multimodal feature extraction module) in question-answering model 3, thereby endowing it with powerful Chinese embedding capabilities. During fine-tuning, this embodiment applies an adaptive interval loss function to guide the model in learning the discriminative relationship between positive and negative sample pairs. Compared to the traditional contrastive learning loss, the specific definition of this loss function is as follows: (1) The symbols are defined as follows: Batch size; The token embedding matrix of the i-th query input (i.e., the question information) is the i-th question feature matrix, containing... One query token; The token embedding matrix of the j-th training document is also the feature matrix of the j-th document, containing: A document token; : Query input and The similarity score between documents is calculated as follows: (2) Here, cos represents cosine similarity, and a and b respectively iterate through all tokens of the query input and document.
[0068] m: The learnable interval parameter, initially set to 0.5 and automatically optimized through backpropagation; : Select the highest score among all negative samples (excluding positive samples) in the current batch to form a hardnegative sample.
[0069] This loss function improves the model's discriminative ability through the following mechanism: (1) Dynamically select the most difficult negative sample within the batch. (3) Introduce learnable interval parameters to adaptively adjust the decision boundary between positive and negative samples, thereby further enhancing the generalization and discrimination capabilities of the model.
[0070] Understandably, the above training process requires traversing all N question messages and all M training documents to enable the model to distinguish features between similar samples (similar question messages or similar training documents).
[0071] The training process of the aforementioned question-answering model 3 involves feature quantization, positive and negative polarity similarity calculation, dynamic loss construction, and parameter iterative optimization. On the one hand, it enhances the ability of question-answering model 3 to distinguish similar information through hard negative sample mining; on the other hand, it uses learnable interval parameters and dynamic loss functions to continuously optimize the matching strategy during training, ultimately improving the accuracy of question-answering model 3 in understanding and retrieving multimodal knowledge, laying the model foundation for subsequent accurate response generation based on multimodal knowledge bases and text knowledge bases.
[0072] After the aforementioned fine-tuning and training, the information preservation modality integration question-answering model 3 has achieved excellent multimodal representation capabilities for Chinese documents. By processing the multimodal knowledge base using the trained multimodal model, more accurate multimodal features can be generated.
[0073] In some embodiments, to fully leverage the advantages of text, a text representation module can be included in the question-answering model 3 to represent the text knowledge base using a text representation model (e.g., BGE). This allows multimodal data and text data to be transformed into their respective knowledge base representation vectors, facilitating efficient subsequent retrieval and utilization.
[0074] For user query input, i.e., question information, it is also input into these two representation models in the same way (text question information input text representation function module, multimodal question information input). Figure 4 The trained multimodal feature extraction module obtains its vector representation in the representation space. This method ensures that the multimodal knowledge base and the text knowledge base achieve dual-path recall within the same representation space, providing a reliable foundation for subsequent retrieval.
[0075] After optimizing the feature extraction capability of question-answering model 3, in this embodiment of the disclosure, question-answering model 3 is also set to realize multimodal retrieval through dual-path recall of text retrieval results and multimodal retrieval results, and then form response information based on the multimodal retrieval results.
[0076] Figure 5 This is a flowchart illustrating the generation of response information by question-answering model 3 in an exemplary embodiment of this disclosure.
[0077] refer to Figure 5 After receiving a question, the question-answering model 3 can generate a response by including: Step S51: Generate the question information feature matrix corresponding to the question information, and obtain the multimodal document feature matrix of multimodal documents in the multimodal knowledge base and the knowledge text feature matrix of knowledge text in the text knowledge base; Step S52: Calculate the first similarity between the multimodal document features and the question information feature matrix, and determine the image feature matrix with the largest first similarity as the target document feature matrix; Step S53: Calculate the second similarity between the knowledge text feature matrix and the question information feature matrix, and determine the knowledge text feature matrix with the largest second similarity as the target knowledge text feature matrix; Step S54: Merge and deduplicate the target multimodal document feature matrix and the target knowledge text feature matrix to generate response information.
[0078] In an exemplary embodiment, the question information includes text question information and image question information. Step S51 includes: generating a question text feature matrix corresponding to the text question information and a question image feature matrix corresponding to the image question information.
[0079] At this point, calculating the first similarity between the multimodal document feature matrix and the question information feature matrix in step S52 includes: calculating the first sub-similarity between the multimodal document feature matrix and the question text feature matrix, and the second sub-similarity between the multimodal document feature matrix and the question image feature matrix, and forming the first similarity based on the first sub-similarity and the second sub-similarity.
[0080] The first similarity is formed based on the first and second sub-similarity, and a weighted sum can be calculated by setting fixed or dynamic weights for the first and second sub-similarity, respectively. In some embodiments, for text and image question information, the importance (e.g., information entropy) of the text and image question information can be calculated in real time, thereby dynamically setting the weights of the corresponding first and second sub-similarity.
[0081] Correspondingly, after generating the question image feature matrix corresponding to the question information, when calculating the second similarity between the knowledge text feature matrix and the question information feature matrix in step S53, the third sub-similarity between the knowledge text feature matrix and the question text feature matrix, as well as the fourth sub-similarity between the knowledge text feature matrix and the question image feature matrix, can be calculated first, and the second similarity is formed based on the third sub-similarity and the fourth sub-similarity.
[0082] The methods for generating the first and second similarities based on self-similarity described above are the same and will not be repeated here.
[0083] Figure 6 This is a sub-flowchart of step S54 in an embodiment of this disclosure.
[0084] refer to Figure 6 In an exemplary embodiment, step S54 may include: Step S541: Based on the target multimodal document feature matrix and the target knowledge text feature matrix, form a temporary answer set corresponding to the question information; Step S542: Calculate the similarity between the target multimodal document feature matrix and the target knowledge text matrix in the temporary answer set as the recall similarity. Step S543: When the recall similarity is greater than a preset threshold, delete the target knowledge text matrix in the temporary answer set, or delete the part of the target knowledge text matrix that overlaps with the target multimodal document feature matrix, and then merge the remaining information with the target multimodal document feature matrix to store it in the temporary answer set. Step S544: Generate the response information corresponding to the question information based on the feature matrix in the temporary answer set corresponding to the question information.
[0085] exist Figure 5 In the illustrated embodiment, a temporary answer set corresponding to the question information is first formed based on the target multimodal document feature matrix and the target knowledge text feature matrix of the dual-path recall. Since each document corresponds to a feature matrix, the multimodal document corresponding to the target multimodal document feature matrix is called the target multimodal document, and the knowledge text document corresponding to the target knowledge text feature matrix is called the target knowledge text document.
[0086] Then, the similarity between the target multimodal document feature matrix and the target knowledge text matrix is compared, for example, by cosine similarity.
[0087] If the similarity between the two is greater than the preset threshold, it indicates that the multimodal retrieval is effective. Only the target multimodal document feature matrix with richer information is retained. Alternatively, the overlapping parts of the target knowledge text matrix and the target multimodal document feature matrix are deleted, and the remaining information is merged with the target multimodal document feature matrix and stored in the temporary answer set.
[0088] In some embodiments, the target multimodal document and the target knowledge text document may also be stored in the temporary answer set. Thus, in step S543, after deleting the overlapping parts in the target knowledge text document and the target multimodal document, the remaining information in the target knowledge text document is merged with the target multimodal document and stored in the temporary answer set.
[0089] In this embodiment of the disclosure, the temporary answer set can store both the target multimodal document and the target multimodal document feature matrix, or it can store only one of them. Similarly, it can store both the target knowledge text and the target knowledge text feature matrix, or it can store only one of them.
[0090] If the similarity between the target multimodal document feature matrix and the target knowledge text feature matrix is not greater than the preset threshold, it indicates an error in the multimodal retrieval, and a more reliable plain text content will be used to generate the response information. That is, if the similarity between the two is not greater than the preset threshold, the target multimodal document feature matrix will be deleted from the temporary answer set, and / or the target multimodal document will be deleted.
[0091] During the answer generation phase, the fragments obtained from multimodal retrieval are merged and deduplicated to ensure the integrity and uniqueness of the information. Subsequently, the processed content (i.e., the remaining content in the temporary answer set) is concatenated with the user's question (question information) and input into the generation module of Question Answering Model 3 (e.g., implemented through a large language model) for answer generation. Thus, Question Answering Model 3 can generate rich, semantically accurate, and highly relevant answers to users' multimodal queries, fully meeting users' needs in multimodal knowledge retrieval and question answering scenarios.
[0092] Figure 7 This is a schematic diagram of the operation process of system 100 in this embodiment of the present disclosure.
[0093] In this example, question-answering model 3 may include text representation model 31, multimodal representation model 32, and large language model 33.
[0094] refer to Figure 7 From left to right, firstly, a text knowledge base 2 is constructed using text data; then, a first image is formed by splicing text and images; and finally, a multimodal knowledge base 1 is formed using the first image.
[0095] Question answering model 3 receives user questions and uses the trained text representation model 31 and multimodal representation model 32 (through...) Figure 4 (The method shown is used for training) to extract features from the text and image information in the question information to form the question text feature matrix and the question image feature matrix, respectively. In addition, features are extracted from the multimodal documents in the multimodal knowledge base 1 and the knowledge text in the text knowledge base 2 to form the multimodal document feature matrix and the knowledge text feature matrix.
[0096] Next, according to Figure 5 The method of the illustrated embodiment performs similarity comparison and filtering on each of the above feature matrices, and finally forms a target multimodal document feature matrix and a target knowledge text feature matrix, and / or, the target multimodal document and the target knowledge text document, and stores them in the temporary answer set corresponding to the question information.
[0097] In the temporary answer set, as follows Figure 6The embodiment shown uses merging and deduplication to update the temporary answer set. Then, the information in the temporary answer set is concatenated with the question information and input into the large language model, which outputs the response information.
[0098] In summary, the embodiments of this disclosure explore a deep retrieval method for training multimodal large models for non-textual information such as images, charts, and screenshots, which can significantly enhance the RAG system's ability to understand and match complex Chinese documents.
[0099] First, by addressing the challenges of complex matching of query information (query) and text / image (text blocks containing images, i.e., the first image), such as high training difficulty and corpus collection difficulties, the features of the matching model are simplified. The query information (query) and text / image matching are simplified into a text / image matching paradigm, which improves matching efficiency and ensures that modal information is not lost. It also solves the problems of high training difficulty and difficulty in constructing multimodal samples in one fell swoop.
[0100] Secondly, by achieving deep concatenation and integration of original text and image data, and combining it with multimodal representation training methods, the knowledge base can efficiently carry and express multi-source heterogeneous information, providing a solid data foundation for subsequent intelligent question answering and content generation. Specifically, through adaptive contrastive learning training, and on this basis, a multimodal dual-path recall retrieval scheme is constructed, enabling Question Answering Model 3 to fully integrate various information carriers such as text, images, and tables. When facing complex Chinese document retrieval tasks, the recall results are richer and more relevant, significantly outperforming the traditional RAG system that relies solely on text, and greatly improving the retrieval performance of Question Answering Model 3 in multimodal scenarios.
[0101] Finally, the system 100 of this embodiment supports users to search and ask questions in multiple ways, such as text and images, which can better meet the industry's needs for understanding and integrating multimodal information, and improve the practicality and user experience of the Chinese intelligent question answering system in real business.
[0102] To systematically evaluate the effectiveness of the proposed Chinese multimodal knowledge base construction and retrieval method, comparative experiments were conducted on the constructed dataset to evaluate different model schemes. The evaluation metrics included Top-K accuracy and Normalized Discounted Cumulative Gain (NDCG), with K values of 3, 5, and 10, comprehensively measuring the model's performance in multimodal retrieval tasks. Experimental results are shown in Table 1.
[0103] Table 1:
[0104] Table 1. Performance of different models on multimodal retrieval tasks
[0105] Experimental results show that the multimodal knowledge base construction and adaptive fine-tuning method proposed in this disclosure significantly improves retrieval performance. Compared with the baseline model that relies solely on text retrieval, the scheme using ColQwen2.5 and introducing contrastive learning loss improves performance by 8.75% and 11.03% in Top-3 Accuracy and Top-3 NDCG, respectively, and by 2.91% and 8.90% in Top-10 Accuracy and NDCG, respectively. This indicates that the introduction of multimodal information and the representation optimization through contrastive learning effectively enhance the model's ability to capture complex semantics and cross-modal associations.
[0106] Furthermore, after training the model with Adaptive Margin Loss, the model achieved further improvements across all evaluation metrics. For example, Top-3 Accuracy reached 81.43%, and Top-3 NDCG reached 67.21%, representing improvements of 1.06% and 1.22% respectively compared to the contrastive learning loss scheme. This improvement is attributed to the Adaptive Margin Loss dynamically adjusting the decision boundary between positive and negative samples during training, combined with hardnegative sampling and token-level max pooling, effectively enhancing the model's ability to distinguish difficult samples and its generalization capabilities.
[0107] Furthermore, the improvement in the NDCG metric is particularly significant, indicating that the improved model not only recalls more relevant results but also performs better in relevance ranking, thus better meeting users' actual retrieval needs.
[0108] In summary, the experimental results fully verify the effectiveness of the multimodal knowledge base construction, question-answering model training method, and dual-path recall retrieval method proposed in this disclosure. Through multimodal data integration, adaptive interval loss optimization training, and dual-path recall mechanism, System 100 achieves leading performance in Chinese multimodal retrieval tasks, laying a solid foundation for the subsequent research and application of multimodal knowledge base systems.
[0109] Figure 8 This is a flowchart of a response information generation method based on a multimodal knowledge base in an exemplary embodiment of this disclosure.
[0110] refer to Figure 8 In an exemplary embodiment, the response information generation method 800 based on a multimodal knowledge base may include: Step S81: Respond to the question information input message and determine the question information feature matrix corresponding to the question information; Step S82: Obtain the multimodal document feature matrix of the multimodal document in the multimodal knowledge base and the knowledge text feature matrix of the knowledge text in the text knowledge base. The multimodal document includes a first image, and the first image includes a knowledge image and the context corresponding to the knowledge image. Step S83: Determine the multimodal document feature matrix with the highest first similarity to the query information feature matrix as the target multimodal document feature matrix; Step S84: The knowledge text feature matrix with the highest second similarity to the question information feature matrix is determined as the target knowledge text feature matrix; Step S85: Merge and deduplicate the target multimodal document feature matrix and the target knowledge text feature matrix to form the response information corresponding to the question information.
[0111] Figure 8 The method shown can be used to implement the response information generation process of question-and-answer system 3 in system 100. Please refer to the above description for related details. Figure 5 and Figure 6 The description of the embodiments shown will not be repeated here.
[0112] In an exemplary embodiment of this disclosure, an electronic device capable of implementing the above-described method is also provided. The electronic device includes: a memory storing either a multimodal knowledge base or a question-answering model in the multimodal question-answering system of any of the above embodiments, wherein the question-answering model is configured to execute, based on instructions stored in the memory, instructions such as those described above. Figure 8 The embodiment shown illustrates a method for generating response information.
[0113] The following reference Figure 9 To describe an electronic device 900 according to this embodiment of the present invention. Figure 9 The electronic device 900 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0114] like Figure 9 As shown, the electronic device 900 is presented in the form of a general-purpose computing device. The components of the electronic device 900 may include, but are not limited to: at least one processor 910, at least one memory 920, and a bus 930 connecting different system components (including memory 920 and processor 910).
[0115] The memory stores program code that can be executed by the processor 910, causing the processor 910 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of the present invention. For example, the processor 910 can perform methods as shown in embodiments of this disclosure.
[0116] The memory 920 may include a readable medium in the form of volatile memory, such as random access memory (RAM) 9201 and / or cache 9202, and may further include read-only memory (ROM) 9203.
[0117] The memory 920 may also include a program / utility 9204 having a set (at least one) of program modules 9205, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.
[0118] Bus 930 can represent one or more of several types of bus structures, including a memory bus or memory controller, peripheral bus, graphics acceleration port, processor, or a local bus using any of the various bus structures.
[0119] Electronic device 900 can also communicate with one or more external devices 700 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 900, and / or with any device that enables electronic device 900 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 950. Furthermore, electronic device 900 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 960. As shown, network adapter 960 communicates with other modules of electronic device 900 via bus 930. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 900, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0120] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0121] In exemplary embodiments of this disclosure, a computer-readable storage medium is also provided, having stored thereon a program product capable of implementing the systems or methods described above. In some possible implementations, various aspects of the invention may also be implemented as a program product comprising program code that, when run on a terminal device, causes the terminal device to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of the invention.
[0122] The program product for implementing the above-described method according to embodiments of the present invention may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited thereto. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0123] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0124] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.
[0125] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0126] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0127] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0128] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0129] Those skilled in the art will understand that various aspects of the present invention can be implemented as systems, methods, or program products. Therefore, various aspects of the present invention can be specifically implemented in the following forms: entirely in hardware, entirely in software (including firmware, microcode, etc.), or in a combination of hardware and software, collectively referred to herein as “circuit,” “module,” or “system.”
[0130] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and concept of this disclosure are indicated by the claims.
Claims
1. A multi-modal question answering system, characterized in that, The method comprises the steps of: a multi-modal knowledge base comprising a plurality of multi-modal documents, the multi-modal documents comprising a first image, the first image comprising a knowledge image and a context corresponding to the knowledge image; a text knowledge base comprising a plurality of knowledge texts; a question and answer model configured to receive query information, generate reply information according to the query information and a target first image in the multi-modal knowledge base matched with the query information and / or a target knowledge text in the text knowledge base matched with the query information, the reply information comprising the target first image and / or the target knowledge text, the query information comprising text query information.
2. The multi-modal question answering system of claim 1, wherein, The training process of the question and answer model comprises the steps of: inputting N pieces of query information and M training documents into the question and answer model to obtain a query feature matrix of the query information and a document feature matrix of the training documents, the i-th piece of query information corresponding to the i-th query feature matrix, the j-th training document corresponding to the j-th document feature matrix, 0≤i<N, 0≤j<M, the training documents comprising the first image; calculating the similarity between the i-th query feature matrix and the j-th document feature matrix as the i-th sample similarity; calculating the similarity between the i-th query feature matrix and M-1 document feature matrices other than the j-th document feature matrix, determining the maximum similarity among the M-1 similarities as the i-th hard negative sample similarity; calculating the sum of the difference between the i-th hard negative sample similarity and the i-th sample similarity and a margin parameter as the i-th value, and setting the training loss degree corresponding to the i-th piece of query information as the maximum value between 0 and the i-th value; determining the current loss function value of the question and answer model according to the average value of N training loss degrees corresponding to the N pieces of query information; adjusting the parameters of the question and answer model and the margin parameter according to the loss function value until a preset training stop condition is reached.
3. The multi-modal question answering system of claim 2, wherein, The N pieces of text query information are text query information in a target language, the context information contained in the N first images is text information in the target language, and the target language comprises Chinese.
4. The multi-modal question answering system of claim 1, wherein, The method of generating reply information according to query information and a target first image in the multi-modal knowledge base matched with the query information and / or a target knowledge text in the text knowledge base matched with the query information comprises the steps of: generating a query information feature matrix corresponding to the query information, obtaining a multi-modal document feature matrix of the multi-modal documents in the multi-modal knowledge base and a knowledge text feature matrix of the knowledge texts in the text knowledge base; calculating the first similarity between the multi-modal document feature matrix and the query information feature matrix, and determining the multi-modal document feature matrix corresponding to the maximum first similarity as a target multi-modal document feature matrix; calculating the second similarity between the knowledge text feature matrix and the query information feature matrix, and determining the knowledge text feature matrix corresponding to the maximum second similarity as a target knowledge text feature matrix; Merging and deduplication processing is performed on the target multi-modal document feature matrix and the target knowledge text feature matrix to generate the reply information.
5. The multi-modal question answering system of claim 4, wherein, Merging and deduplication processing is performed on the target multi-modal document feature matrix and the target knowledge text feature matrix to generate the reply information includes: forming a temporary answer set corresponding to the query information according to the target multi-modal document feature matrix and the target knowledge text feature matrix; calculating the similarity of the target multi-modal document feature matrix and the target knowledge text matrix in the temporary answer set as a recall similarity; when the recall similarity is greater than a preset threshold, deleting the target knowledge text matrix in the temporary answer set, or deleting the part of the target knowledge text matrix that overlaps with the target multi-modal document feature matrix, and then merging the remaining information with the target multi-modal document feature matrix to store in the temporary answer set; generating the reply information corresponding to the query information according to the feature matrix in the temporary answer set corresponding to the query information.
6. The multi-modal question answering system of claim 4, wherein, The query information includes text query information and image query information, and the generation of the query information feature matrix corresponding to the query information includes generating a query text feature matrix corresponding to the text query information and a query image feature matrix corresponding to the image query information.
7. The multi-modal question answering system of claim 6, wherein, The calculation of the first similarity between the image feature matrix and the query information feature matrix includes the calculation of a first sub-similarity between the image feature matrix and the query text feature matrix, and a second sub-similarity between the image feature matrix and the query image feature matrix, and the formation of the first similarity according to the first sub-similarity and the second sub-similarity.
8. The multi-modal question answering system of claim 1, wherein, The formation process of the multi-modal knowledge base includes: acquiring a multi-modal document, and identifying a knowledge image in the multi-modal document; identifying a context corresponding to the knowledge image in the multi-modal document; According to the context corresponding to the knowledge image and the position of the knowledge image in the multi-modal document, the multi-modal document is screenshot to form a first image corresponding to the knowledge image; forming the multi-modal document according to the first image. 9.A method for generating reply information based on a multi-modal knowledge base, characterized by, Applied to the multi-modal question answering system of any one of claims 1-8, comprising: in response to the query information input message, determining the query information feature matrix corresponding to the query information; acquiring the multi-modal document feature matrix of the multi-modal document in the multi-modal knowledge base and the knowledge text feature matrix of the knowledge text in the text knowledge base, the multi-modal document including a first image, the first image including a knowledge image and a context corresponding to the knowledge image; determining the multi-modal document feature matrix with the largest first similarity with the query information feature matrix as the target multi-modal document feature matrix; determining the knowledge text feature matrix with the largest second similarity with the query information feature matrix as the target knowledge text feature matrix; Merging and deduplication processing is performed on the target multi-modal document feature matrix and the target knowledge text feature matrix to form the reply information corresponding to the query information.
10. An electronic device, comprising: including: A memory storing any one of the multi-modal knowledge base or the question answering model in the multi-modal question answering system according to any one of claims 1-8, the question answering model being configured to perform the method of claim 9 for generating an answer based on instructions stored in the memory.
Citation Information
Cited By
Multi-modal data-based reply method, electronic equipment and readable storage medium
CN121811423A