Multimodal fusion question and answer method and device, electronic equipment and storage medium
By receiving multimodal input problems, image annotation, table linearization and intermediate inference are performed, and unified text input is generated, which solves the problem of insufficient information interaction during modal fusion of existing multimodal question-and-answer systems, and improves the accuracy and generalization of the system.
Patent Information
- Application Number
- CN202510093798.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing multimodal question and answer system is difficult to achieve deep-level information interaction during modal fusion, resulting in information loss and semantic relationship loss, affecting the accuracy and generalization of the system.
By receiving multimodal input questions, extracting questions, auxiliary text, tables and images, and performing image annotation, table linearization, and intermediate reasoning. After generating inference text, it is spliced with the original question and table text, and inputting a pre-trained large language model to generate answers.
Through diversified image annotation and position-enhanced table linearization, multimodal data is converted into a unified text format, promoting interactive reasoning between modal information, and improving the accuracy and generalization of the question-and-answer system.
Smart Images

Figure CN119938859A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of question answering technology, and in particular to a multimodal fusion question answering method, device, electronic device and storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, question-answering system, as one of the important applications of natural language processing, has become one of the core technologies in the field of human-computer interaction. The goal of the question-answering system is to obtain answers from knowledge bases or other information sources by semantically understanding and reasoning about natural language questions. This technology has wide application value in the fields of intelligent customer service, intelligent search, educational question-answering, e-commerce, etc. However, traditional question-answering systems usually rely only on single-modal data sources (such as text, images, or tables), which severely limits their accuracy and flexibility in complex scenarios.
[0003] In recent years, with the widespread application of multimodal technology and deep learning, multimodal question answering systems have gradually become a research hotspot. Multimodal question answering systems aim to integrate information from multiple modalities (such as text, tables, images, etc.) to provide accurate answers to complex questions. Compared with unimodal question answering systems, multimodal systems can handle more complex scenarios, such as the need to extract information from images, combine data in tables with text descriptions for reasoning. However, existing multimodal question answering technologies still face the following major challenges:
[0004] 1) Current multimodal question-answering technologies usually process different modalities separately. For example, the text modality is modeled using natural language processing technology, the image modality is processed using computer vision technology, and the table modality is linearized using table parsing technology. Although this separate processing method can extract information from each modality separately, it is often difficult to achieve deep information interaction when the modalities are fused. This method is not only inefficient, but may also lead to information loss or lack of semantic relationships between modalities, thus affecting the overall performance of the question-answering system.
[0005] 2) In multimodal question answering, converting non-textual modalities (such as tables and images) into textual representations has become a mainstream trend. However, information loss is very common during the modality conversion process. For example, for tabular data, traditional linearization methods may lose the structural relationship between rows and columns in the table; for image data, simple image descriptions or single image caption generation may ignore important visual details in the image. This information loss directly affects the question answering system's ability to understand and reason about complex questions.
[0006] 3) Multimodal question-answering systems need to extract information from data sources of different modalities and obtain answers through cross-modal reasoning. However, most existing technologies only focus on simple interactions between a single modality or two modalities, and lack the ability to model deep associations in multimodal data. For example, some current methods simply concatenate information from different modalities into a unified input text sequence so that it can be input into a pre-trained language model for processing, but this method ignores the complex semantic relationships and interactive reasoning processes between modalities. Summary of the invention
[0007] The present invention provides a multimodal fusion question-answering method, device, electronic device and storage medium, which are used to solve the technical problems of low accuracy and poor generalization of existing multimodal question-answering systems.
[0008] The present invention provides a multi-modal fusion question-answering method, comprising:
[0009] Accepting multimodal input problems;
[0010] extracting questions, supporting text, tables, and images from the multimodal input question;
[0011] Annotating the image to obtain annotated text;
[0012] Performing text mode conversion on the table by table linearization to obtain table text;
[0013] Performing intermediate reasoning using the auxiliary text, the image and the table text to obtain a reasoning text;
[0014] Concatenate the question, the reasoning text, the annotation text and the table text to obtain an input text;
[0015] The input text is input into a pre-trained large language model, and an answer to the multimodal input question is output.
[0016] Optionally, the annotated text includes a first annotated text and a second annotated text; and the step of annotating the image to obtain the annotated text includes:
[0017] Recognize explicit text from the image as first annotated text by optical character recognition;
[0018] A preset joint probability sampling strategy is adopted to filter the second annotation text of the image from a preset image annotation set.
[0019] Optionally, the step of converting the table into text mode by table linearization to obtain table text includes:
[0020] The elements in the same row of the table are spliced by using a preset first separator to obtain a row paragraph;
[0021] Each row paragraph is spliced by the preset second separator to obtain the table text.
[0022] Optionally, the step of performing intermediate reasoning using the auxiliary text, the image and the table text to obtain the reasoning text includes:
[0023] Input the image into a CLIP visual encoder and output visual features;
[0024] Input the auxiliary text and the table text into a language encoder and output language features;
[0025] The visual features and the language features are integrated to obtain an inference text.
[0026] The present invention also provides a multi-modal fusion question-answering device, comprising:
[0027] A multimodal input question receiving module, used for receiving multimodal input questions;
[0028] An extraction module, for extracting questions, auxiliary texts, tables and images from the multimodal input questions;
[0029] An image annotation module, used to annotate the image to obtain annotated text;
[0030] A table linearization module, used for performing text mode conversion on the table by table linearization to obtain table text;
[0031] An intermediate reasoning module, used to perform intermediate reasoning using the auxiliary text, the image and the table text to obtain a reasoning text;
[0032] A splicing module, used for splicing the question, the reasoning text, the annotation text and the table text to obtain an input text;
[0033] The question-answering module is used to input the input text into a pre-trained large language model and output an answer to the multimodal input question.
[0034] Optionally, the annotated text includes a first annotated text and a second annotated text; and the image annotating module includes:
[0035] A first annotation text generating submodule, used for recognizing explicit text from the image as the first annotation text through optical character recognition;
[0036] The second annotation text generation submodule is used to select the second annotation text of the image from a preset image annotation set by adopting a preset joint probability sampling strategy.
[0037] Optionally, the table linearization module includes:
[0038] A row paragraph generation submodule, used for splicing elements in the same row of the table through a preset first separator to obtain a row paragraph;
[0039] The table text generation submodule is used to splice each line paragraph through a preset second separator to obtain a table text.
[0040] Optionally, the intermediate reasoning module includes:
[0041] A visual feature output submodule, used for inputting the image into a CLIP visual encoder and outputting visual features;
[0042] A language feature output submodule, used for inputting the auxiliary text and the table text into a language encoder and outputting language features;
[0043] The inference text generation submodule is used to fuse the visual features and the language features to obtain the inference text.
[0044] The present invention also provides an electronic device, the device comprising a processor and a memory:
[0045] The memory is used to store program code and transmit the program code to the processor;
[0046] The processor is used to execute the multimodal fusion question-answering method as described in any one of the above items according to the instructions in the program code.
[0047] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to execute the multimodal fusion question-answering method as described in any one of the above items.
[0048] It can be seen from the above technical scheme that the present invention has the following advantages: the present invention adopts a multimodal fusion question-answering method, and specifically discloses: receiving a multimodal input question; extracting a question, auxiliary text, table and image from the multimodal input question; annotating the image to obtain annotated text; using table linearization to convert the table into text mode to obtain table text; using auxiliary text, image and table text to perform intermediate reasoning to obtain reasoning text; splicing the question, reasoning text, annotated text and table text to obtain input text; inputting the input text into a pre-trained large language model to output the answer to the multimodal input question. The present invention uses a diversified image annotation method and position-enhanced table linearization to convert multimodal input data containing text, table and image into a unified text format; then uses a multimodal intermediate reasoning generator to create text describing cross-modal relationships to promote interactive reasoning between different modal information; finally, uses a pre-trained language model to perform text-to-text generation tasks, and outputs the answer that best fits the question based on the information processed in the previous steps, thereby providing accurate and efficient answers to multimodal source questions. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0050] Figure 1 A flowchart of the steps of a multi-modal fusion question-answering method provided by an embodiment of the present invention;
[0051] Figure 2 A flowchart of the steps of a multi-modal fusion question-answering method provided by another embodiment of the present invention;
[0052] Figure 3 Schematic diagram of input for auxiliary text, table and image in multimodal input problem;
[0053] Figure 4 A schematic diagram of a question-answering system based on multimodal unified representation and interactive reasoning provided by an embodiment of the present invention;
[0054] Figure 5 A structural block diagram of a multi-modal fusion question-answering device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0055] The embodiments of the present invention provide a multimodal fusion question-answering method, device, electronic device and storage medium, which are used to solve the technical problems of low accuracy and poor generalization of existing multimodal question-answering systems.
[0056] In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0057] See also Figure 1 , Figure 1 A flowchart of the steps of a multimodal fusion question-answering method provided in an embodiment of the present invention.
[0058] The present invention provides a multi-modal fusion question-answering method, which may specifically include the following steps:
[0059] Step 101, receiving a multimodal input question;
[0060] Step 102, extracting questions, auxiliary texts, tables and images from the multimodal input questions;
[0061] In an embodiment of the present invention, a multimodal input question may include a question, auxiliary text, a table, and an image associated with the question.
[0062] Among them, questions can be asked in response to auxiliary questions, tables, and images.
[0063] Step 103, annotating the image to obtain annotated text;
[0064] Image annotation is an important task in the field of computer vision, which involves adding labels to objects, scenes, or features in images so that machine learning algorithms can understand and process this information. The quality of image annotation directly affects the performance of the trained model, so it is crucial when building and optimizing computer vision applications.
[0065] The present invention can obtain the annotation text representing the meaning of the image by annotating the image.
[0066] Step 104, using table linearization to perform text mode conversion on the table to obtain table text;
[0067] Table linearization is the process of converting a two-dimensional table structure into a one-dimensional text sequence, which makes the table easier to process and understand, especially when the content needs to be accessed through assistive technologies such as screen readers. This process is very important for ensuring information accessibility, and is also applicable to data transmission, storage optimization, and certain types of analytical tasks.
[0068] In the specific implementation, for the table in the multimodal input problem, the position-enhanced coding linearization method can be used to convert the structured table into graphics and text.
[0069] Step 105, using the auxiliary text, the image and the table text to perform intermediate reasoning to obtain the reasoning text;
[0070] Intermediate Inference refers to the process of gradually deriving partial conclusions based on known information and rules when solving complex problems or making decisions. Instead of jumping directly from the premise to the final answer, it builds understanding through a series of logical steps and draws small verifiable conclusions at each step. This approach helps improve the transparency, explainability and accuracy of the reasoning process.
[0071] In the embodiment of the present invention, the connection and relationship between different modes can be sought by generating intermediate reasoning descriptions through auxiliary text, images and table text as a principle.
[0072] In one example, a generative language model can be fine-tuned as an intermediate inference generator to generate inference results based on visual features and text input. The specific model training process can refer to the generation method in the prior art and is not specifically limited here.
[0073] Step 106, concatenating the question, the reasoning text, the annotation text and the table text to obtain the input text;
[0074] After completing the text conversion of each modality, the question, reasoning text, annotation text, table text and other related contexts can be spliced into a unified input sequence to obtain the input text.
[0075] Step 107, input the input text into the pre-trained large language model, and output the answer to the multimodal input question.
[0076] Large Language Models (LLMs) are deep learning models that have been trained on a large scale and have a large number of parameters, which can understand and generate human language. Such models are usually based on neural network architectures, such as Transformers, and are pre-trained on massive text data to capture complex patterns and structures in language.
[0077] In an embodiment of the present invention, a large language model may be pre-trained as a question-answering model to analyze the input text and output answers to multimodal input questions.
[0078] The present invention utilizes diversified image annotation methods and position-enhanced table linearization to convert multimodal input data containing text, tables and images into a unified text format; then a multimodal intermediate reasoning generator is used to create text describing cross-modal relationships to promote interactive reasoning between different modal information; finally, a pre-trained language model is used to perform text-to-text generation tasks, and the answer that best suits the question is output based on the information processed in the previous steps, thereby providing accurate and efficient answers to questions of multimodal origin.
[0079] See also Figure 2 , Figure 2 A flowchart of a multi-modal fusion question-answering method provided by another embodiment of the present invention. Specifically, the following steps may be included:
[0080] Step 201, receiving a multimodal input question;
[0081] Step 202, extracting questions, auxiliary texts, tables and images from the multimodal input questions;
[0082] Steps 201-202 are the same as steps 101-102. For details, please refer to the description of steps 101-102, which will not be repeated here.
[0083] In one example, auxiliary text, tables, and images in a multimodal input question can be as follows: Figure 3 shown.
[0084] Step 203, identifying explicit text from the image as first annotated text through optical character recognition;
[0085] Step 204, using a preset joint probability sampling strategy to select a second annotation text of the image from a preset image annotation set;
[0086] In actual scenarios, images may contain both text and patterns. Therefore, text extraction from images can be done through two image-to-text conversion strategies: optical character recognition (OCR) and image annotation. The annotation text may include a first annotation text and a second annotation text.
[0087] Firstly, all explicit text in the image is directly extracted by OCR method as the first annotation text.
[0088] Secondly, the image annotation model is used to annotate images into text from noisy network data, which is used as the image annotation set for subsequent text sampling. It should be noted that multiple image annotation sets can be generated in different ways, and then multiple image annotation sets can be sampled to make the annotation generation more diverse to enhance the generalization of the system.
[0089] Among them, for the sampling of the image annotation set, a joint probability sampling strategy combining Top-K sampling and Top-p sampling can be adopted.
[0090] Top-K sampling refers to selecting the top k highest scoring or most relevant items from a sorted list.
[0091] Top-p sampling, also known as Nucleus Sampling, is a decoding strategy used in natural language generation tasks, especially in text generation models, such as the GPT series, BERT and other pre-trained language models. It aims to solve the balance problem between traditional greedy search (such as maximum likelihood) and pure random sampling (such as top-K sampling).
[0092] The sampling strategy of the present invention selects words according to the conditional probability distribution:
[0093]
[0094] in, Represents the word segmentation in the sampling pool, and t represents the next sampling process. Specifically, by combining Top-K and Top-p sampling, K word segmentations with the highest probability are selected in the sampling pool, and the probability mass is redistributed according to their cumulative probability as the evaluation indicator for the next word segmentation screening:
[0095]
[0096] Where V represents the sampling pool, It represents the conditional probability of the tth word after filtering the first t-1 words. In Top-p sampling, the minimum word set with cumulative probability exceeding probability p is selected to select the next word, as shown in the following formula:
[0097]
[0098] Top-p sampling will only filter a few words in the sampling pool when the next word segmentation prediction is more certain, which means that it is more sensitive to word segmentations with similar probabilities than Top-K sampling. To solve this problem, Top-p is combined with Top-K as a decoding strategy. When the sum of the probabilities of the current K candidate word segmentations is lower than p, the number of word segmentations is increased to ensure that the sum of probabilities is greater than or equal to p; when the sum of the candidate word segmentation probabilities is greater than or equal to p but the number of candidate word segmentations is too small, the number of candidate word segmentations is increased to K. This strategy can avoid the impact of too few candidate word segmentations on the diversity and robustness of the model output. The strategy is as follows:
[0099]
[0100] Here, x represents the number of words with the highest probability, and x ≥ K is guaranteed to prevent too few words from being filtered in the sampling pool.
[0101] Step 205, using table linearization to perform text mode conversion on the table to obtain table text;
[0102] In a specific implementation, for tables in multimodal input problems, a position-enhanced encoding linearization method can be used to convert structured tables into text.
[0103] In one example, step 205 may include the following sub-steps:
[0104] S51, concatenating elements in the same row of the table using a preset first separator to obtain a row paragraph;
[0105] S52, splicing each line paragraph by using a preset second separator to obtain a table text.
[0106] In a specific implementation, all elements in the same row of the table can be concatenated, and different elements can be separated by a first separator, such as "|", and all rows including the header can be concatenated into a long paragraph, separated by a predefined second separator, such as "header:" or "row:x", where x represents the row number. This linearization of the table can be specifically expressed as:
[0107]
[0108] Among them, M and N represent the number of rows and columns respectively.
[0109] Step 206, using the auxiliary text, the image and the table text to perform intermediate reasoning to obtain a reasoning text;
[0110] In the embodiment of the present invention, the connection and relationship between different modes can be sought by generating intermediate reasoning descriptions through auxiliary text, images and table text as a principle.
[0111] In one example, step 206 may include the following sub-steps:
[0112] S61, input the image into the CLIP visual encoder and output the visual features;
[0113] S62, inputting the auxiliary text and the table text into a language encoder and outputting language features;
[0114] S63, integrates visual features and language features to obtain inference text.
[0115] In the specific implementation, the image can be input into the CLIP visual encoder to extract visual features, and the text and linearization table can be input into the language encoder to extract language features. The corresponding visual feature and language feature extraction formulas are as follows:
[0116]
[0117] in, represents the visual encoder, Denotes a language encoder based on Transformer, I denotes image input, and P denotes text input. Then, the visual features and language representations are fused together to encode their joint features. The joint features are then input into the decoder to generate an intermediate inference describing the relationship between different modalities. The intermediate inference can be obtained by the following formula:
[0118]
[0119] Among them, R represents intermediate reasoning, Represents an intermediate inference generator.
[0120] Step 207, concatenating the question, the reasoning text, the annotation text and the table text to obtain the input text;
[0121] After incorporating intermediate reasoning of natural language into text input to enhance cross-modal interaction, it is then concatenated with the question and context into a unified input sequence.
[0122] Step 208, input the input text into the pre-trained large language model, and output the answer to the multimodal input question.
[0123] In an embodiment of the present invention, a large language model may be pre-trained as a question-answering model to analyze the input text and output answers to multimodal input questions.
[0124] In the embodiment of the present invention, the large language model is defined as follows:
[0125]
[0126] in, represents the question-answering model, Indicates the predicted answer. During the fine-tuning phase, the average negative log-likelihood loss of each batch of word segmentation is minimized As training goals:
[0127]
[0128] in, is the maximum length of the output sequence, and Represent the first In addition, the embodiment of the present invention also adopts prefix adjustment as a fine-tuning strategy, that is, adding a prefix of a specific task as a prompt word in the input sequence.
[0129] Figure 4 A schematic diagram of a question-answering system based on multimodal unified representation and interactive reasoning provided in an embodiment of the present invention.
[0130] The present invention utilizes diversified image annotation methods and position-enhanced table linearization to convert multimodal input data containing text, tables and images into a unified text format; then a multimodal intermediate reasoning generator is used to create text describing cross-modal relationships to promote interactive reasoning between different modal information; finally, a pre-trained language model is used to perform text-to-text generation tasks, and the answer that best suits the question is output based on the information processed in the previous steps, thereby providing accurate and efficient answers to questions of multimodal origin.
[0131] See also Figure 5 , Figure 5 A structural block diagram of a multi-modal fusion question-answering device provided in an embodiment of the present invention.
[0132] The embodiment of the present invention provides a multi-modal fusion question-answering device, including:
[0133] A multi-modal input question receiving module 501 is used to receive a multi-modal input question;
[0134] An extraction module 502, for extracting questions, auxiliary text, tables and images from the multimodal input questions;
[0135] The image annotation module 503 is used to annotate the image to obtain annotated text;
[0136] A table linearization module 504 is used to perform text mode conversion on the table by table linearization to obtain table text;
[0137] An intermediate reasoning module 505 is used to perform intermediate reasoning using auxiliary text, image and table text to obtain a reasoning text;
[0138] A splicing module 506 is used to splice the question, the reasoning text, the annotation text and the table text to obtain the input text;
[0139] The question-answering module 507 is used to input the input text into a pre-trained large language model and output the answer to the multi-modal input question.
[0140] In the embodiment of the present invention, the annotated text includes a first annotated text and a second annotated text; the image annotating module 503 includes:
[0141] A first annotation text generation submodule is used to recognize explicit text from the image as the first annotation text through optical character recognition;
[0142] The second annotation text generation submodule is used to select the second annotation text of the image from the preset image annotation set by adopting a preset joint probability sampling strategy.
[0143] In the embodiment of the present invention, the table linearization module 504 includes:
[0144] A row paragraph generation submodule is used to splice the elements in the same row of the table through a preset first separator to obtain a row paragraph;
[0145] The table text generation submodule is used to splice each line paragraph through a preset second separator to obtain a table text.
[0146] In this embodiment of the present invention, the intermediate reasoning module 505 includes:
[0147] The visual feature output submodule is used to input the image into the CLIP visual encoder and output the visual features;
[0148] A language feature output submodule, used for inputting auxiliary text and table text into a language encoder and outputting language features;
[0149] The inference text generation submodule is used to fuse visual features and language features to obtain inference text.
[0150] An embodiment of the present invention further provides an electronic device, the device comprising a processor and a memory:
[0151] The memory is used to store the program code and transmit the program code to the processor;
[0152] The processor is used to execute the multimodal fusion question-answering method of the embodiment of the present invention according to the instructions in the program code.
[0153] An embodiment of the present invention further provides a computer-readable storage medium, which is used to store program code, and the program code is used to execute the multi-modal fusion question-answering method of the embodiment of the present invention.
[0154] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0155] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0156] It will be appreciated by those skilled in the art that the embodiments of the present invention may be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0157] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0158] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0159] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0160] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0161] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.
[0162] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal fusion question-answering method, characterized in that: include: Accepting multimodal input problems; extracting questions, supporting text, tables, and images from the multimodal input question; Annotating the image to obtain annotated text; Performing text mode conversion on the table by table linearization to obtain table text; Performing intermediate reasoning using the auxiliary text, the image and the table text to obtain a reasoning text; Concatenate the question, the reasoning text, the annotation text and the table text to obtain an input text; The input text is input into a pre-trained large language model, and an answer to the multimodal input question is output.
2. The method according to claim 1, characterized in that The annotated text includes a first annotated text and a second annotated text; the step of annotating the image to obtain the annotated text includes: Recognize explicit text from the image as first annotated text by optical character recognition; A preset joint probability sampling strategy is adopted to filter the second annotation text of the image from a preset image annotation set.
3. The method according to claim 1, characterized in that The step of converting the table into text mode by table linearization to obtain table text includes: The elements in the same row of the table are spliced by using a preset first separator to obtain a row paragraph; Each row paragraph is spliced by the preset second separator to obtain the table text.
4. The method according to claim 1, characterized in that: The step of performing intermediate reasoning using the auxiliary text, the image and the table text to obtain the reasoning text comprises: Input the image into a CLIP visual encoder and output visual features; Input the auxiliary text and the table text into a language encoder and output language features; The visual features and the language features are integrated to obtain an inference text.
5. A multi-modal fusion question-answering device, characterized in that: include: A multimodal input question receiving module, used for receiving multimodal input questions; An extraction module, for extracting questions, auxiliary texts, tables and images from the multimodal input questions; An image annotation module, used to annotate the image to obtain annotated text; A table linearization module, used for performing text mode conversion on the table by table linearization to obtain table text; An intermediate reasoning module, used to perform intermediate reasoning using the auxiliary text, the image and the table text to obtain a reasoning text; A splicing module, used for splicing the question, the reasoning text, the annotation text and the table text to obtain an input text; The question-answering module is used to input the input text into a pre-trained large language model and output an answer to the multimodal input question.
6. The device according to claim 5, characterized in that The annotated text includes a first annotated text and a second annotated text; The image annotation module comprises: A first annotation text generating submodule, used for recognizing explicit text from the image as the first annotation text through optical character recognition; The second annotation text generation submodule is used to select the second annotation text of the image from a preset image annotation set by adopting a preset joint probability sampling strategy.
7. The device according to claim 5, characterized in that The table linearization module comprises: A row paragraph generation submodule, used for splicing elements in the same row of the table through a preset first separator to obtain a row paragraph; The table text generation submodule is used to splice each line paragraph through a preset second separator to obtain a table text.
8. The device according to claim 5, characterized in that The intermediate reasoning module comprises: A visual feature output submodule, used for inputting the image into a CLIP visual encoder and outputting visual features; A language feature output submodule, used for inputting the auxiliary text and the table text into a language encoder and outputting language features; The inference text generation submodule is used to fuse the visual features and the language features to obtain the inference text.
9. An electronic device, characterized in that: The device comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the multimodal fusion question-answering method according to any one of claims 1 to 4 according to the instructions in the program code.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program code, and the program code is used to execute the multimodal fusion question-answering method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Question and answer method and device based on multi-modal information and application of question and answer method and device
CN117828142A
Multi-modal question and answer method and device, electronic equipment and computer readable storage medium
CN119047583A
Cited By
Data processing method and electronic equipment
CN121010837A