Conversational intelligent question and answer method and system for multi-modal data of power grid
By encoding multimodal data using BERT and ResNet101 models, and combining efficient retrieval and adaptive answer acquisition strategies, the shortcomings of existing models in multiple rounds of dialogue and multimodal information processing are solved, and higher accuracy and comprehensiveness of multimodal dialogue-based Q&A is achieved.
Patent Information
- Application Number
- CN202510216151.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
Smart Images

Figure CN120146062A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence and language processing technology, and specifically relates to a conversational intelligent question-answering method and system for multimodal data of power grids. Background Art
[0002] Multimodal conversational question answering (MMCoQA) is an emerging task that aims to answer user questions through multi-round dialogues using multiple modal knowledge sources such as text, tables, and images. This task brings challenges such as multimodal knowledge priority, consistency, and complementarity. For specific questions, how to correctly determine the appropriate knowledge modality, how to use the consistency between different modalities to verify the answer, and how to reason across modalities to get the final answer are all urgent issues to be addressed. Unlike previous conversational question answering tasks, questions in MMCoQA do not give the most appropriate answer modality, requiring the model to have a deep multimodal understanding and reasoning ability. As the conversation progresses, these challenges are intertwined, placing higher demands on the models; published datasets have promoted the development of text- and knowledge-based conversational QA, but existing methods still only involve a single source of knowledge and ignore visual clues. The ManymodalQA challenge introduced by Hannan et al. and the complex problem scenarios introduced by Talmor further emphasize the importance of multimodal QA; MMConvQA proposes to decompose complex problems into conversational questions to gradually meet users' complex information needs; the existing ORConvQA and ManyModalQA baseline models have obvious shortcomings in handling MMCoQA tasks; as a multimodal conversational question-answering task, MMCoQA requires the model to be able to handle multiple rounds of conversations and integrate information from multiple modalities such as text, tables, and images; however, the baseline models of ORConvQA and ManyModalQA either only support single-modal input and cannot fully utilize the advantages of multimodal knowledge sources; or are only applicable to single-round conversations and cannot capture key information in the conversation context. Summary of the invention
[0003] The purpose of the present invention is to provide a conversational intelligent question-answering method and system for multimodal data of power grids to solve the problems raised in the above-mentioned background technology.
[0004] In order to achieve the above object, the present invention provides the following technical solution: a conversational intelligent question-answering method for multimodal data of power grid, the specific use steps of the method are as follows:
[0005] S1: Data preparation and preprocessing: Collect multimodal data including questions, texts, tables, and images. Preprocess texts and tables to fit the BERT model, and keep images in their original format for processing by ResNet101.
[0006] S2: Model Selection and Initialization: Select BERT series and ResNet101 models, load pre-trained weights, and initialize the projection matrix and encoder;
[0007] S3: Multimodal Data Representation: Use the BERT series model to encode the question to obtain embeddings, and different encoders to process the knowledge embeddings of text, tables, and images;
[0008] S4: Multimodal Evidence Extraction: Construct and load an efficient retriever to quickly retrieve the most relevant multimodal evidence from the knowledge base for the question;
[0009] S5: Adaptive Answer Obtaining: The modality detector predicts the answer modality, and the answer extractor extracts the answer from the evidence according to the modality;
[0010] S6: Model Training and Optimization: Define the overall loss function, fuse the losses of evidence extraction, modality detection, and answer obtaining, and train with the AdamW optimizer until convergence;
[0011] S7: Evaluation of the MMConvQA Dataset Based on ChatGPT: Call the ChatGPT API to evaluate the performance of MMConvQA and explore directions for improvement.
[0012] Preferably, the data preparation and preprocessing in S1 refer to that when constructing a multimodal conversational question-answering system, it is necessary to widely collect data, including information in multiple modalities such as questions, text paragraphs, tables, and images. After the data collection is completed, the data preprocessing work is carried out next. For text and table data, specific formatting methods are adopted to make them adapt to the input requirements of the BERT series model. When processing long texts, the sliding window advanced technology is used to ensure the integrity and accuracy of the text information. For image data, its original format is kept unchanged so that it can be directly processed by the ResNet101 model later.
[0013] Preferably, the model selection and initialization in S2 refer to that for the natural language encoding part, bert-base-uncased and albert-base-v2 in the BERT series models are selected. These models have excellent performance in the field of natural language processing. For the image encoding part, the ResNet101 model is adopted, which is famous for its strong image feature extraction ability. When initializing the model, the weights of these pre-trained models are first loaded to ensure that the model can learn useful knowledge from a large amount of data. At the same time, the necessary projection matrix and encoder are also initialized to map data in different modalities to the same semantic space.
[0014] Preferably, the specific usage steps of the multimodal data representation in S3 are as follows:
[0015] Step 1: Problem Encoding
[0016] Select a model: Select the bert-base-uncased model from the BERT series of models for problem encoding;
[0017] Input processing: Preprocess the problem text, including tokenization, adding special tokens, and adjusting the length to meet the input requirements of the BERT model;
[0018] Generate embeddings: Input the processed problem text into the BERT model to obtain the problem embedding representation; Step 2: Text Paragraph and Table Encoding
[0019] Select a knowledge encoder: For text paragraphs and tables, select the same BERT series of models as for problem encoding, or select other suitable text encoders according to specific circumstances;
[0020] Input formatting: Convert the text paragraph and table data into a format that the encoder can understand, and convert the table data into a series of text fragments or structured representations;
[0021] Generate embeddings: Input the formatted text paragraph and table data into the knowledge encoder to generate the corresponding knowledge embedding representations;
[0022] Step 3: Image Encoding
[0023] Select an image encoder: Select ResNet101 as the image encoder, which can effectively extract image features;
[0024] Preprocess the image: Perform necessary preprocessing on the image, such as resizing and normalizing, to meet the input requirements of ResNet101;
[0025] Generate embeddings: Input the preprocessed image into the ResNet101 model to obtain the image embedding representation; This is achieved by extracting the feature vectors in the output layer of the model.
[0026] Preferably, the specific steps of multi-modal evidence extraction in S4 are as follows:
[0027] Step 1: Train an embedding model:
[0028] Use a large amount of text data to train a machine learning model to map queries and document content to a common vector embedding space; Ensure that the embedding model can capture the semantic similarity between texts, so that the distance between similar texts in the embedding space is closer;
[0029] Step 2: Construct database embeddings:
[0030] Preprocess all documents to be retrieved, and use the trained embedding model to convert them into vector representations;
[0031] Step 3: Query Vector Generation and Retrieval:
[0032] When the user enters a query, use the embedding model to convert the query text into a vector representation;
[0033] Perform a maximum inner product search in the database embedding set to find the database embedding with the largest inner product with the query vector, that is, the closest embedding;
[0034] The formula for the maximum inner product search is:
[0035] p = argmax x∈X x T q#(6)
[0036] where X is the set of embedding vectors, q is the vector to be queried, and the vector p with the largest inner product in the set is to be found;
[0037] Extraction and Evaluation of Results:
[0038] Based on the retrieved closest embedding, accurately locate the original document or paragraph. Subsequently, carefully extract the information related to the query from these documents as answers or evidence. To ensure accuracy, strictly evaluate the extraction results, and integrate or sort multiple results to provide the user with a comprehensive and accurate answer.
[0039] Preferably, the specific usage steps of the adaptive answer acquisition in S5 are as follows:
[0040] Step 1: Modal Detection: First, input the user's question into the modal classifier; the classifier encodes the question and predicts the probabilities of three modalities, and selects the modality with the highest probability as the most appropriate answer modality;
[0041] The modal classifier is expressed as:
[0042] s b = f(W c F c (q k ′))#(8)
[0043] where f() represents the softmax function, F c is the question encoder, s b ∈R 3
[0044] Step 2: Answer Acquirer Selection and Application:
[0045] If the modality is text, use the text answer acquirer, input the question and the retrieved text paragraph into the machine reading comprehension model, and the machine reading comprehension model calculates the possibility of each token as the start and end of the answer, and outputs the best answer passage;
[0046] If the modality is a table, then select the table answer extractor, combine the question text with the linearized table data, encode it through BERT, and use a linear classifier to determine the starting position of the answer in the table, thereby extracting the answer;
[0047] If the modality is an image, then use the image answer extractor, combine the question text with the image descriptions in the preset answer set, process it through BERT and combine visual features, and use a linear classifier to find the best answer start token and end token in the description as the answer;
[0048] Answer integration and output: Based on the results output by the selected answer extractor, input the answer text sequence into the BERT model to obtain the deep representation of each token. Subsequently, these text representations are cleverly combined with the visual feature v i Combined. Through the integration process, the final answer not only maintains a close correlation with the user's query question but also can provide accurate and useful information;
[0049] The answer extraction score s of the candidate answers predicted by the above three extractors c Is defined as the average of the probabilities of the start and end tokens. For each candidate answer, its final score is calculated as the evidence extraction score s a , the modality detection score s b And the answer extraction score s c The sum of. The answer with the highest score is selected as the final answer output, where the modality detection loss L md Is the cross-entropy loss commonly used for training multi-class classification tasks. The answer extraction loss for training the three answer extractors is defined as
[0050]
[0051]
[0052] Where y 1 And y 2 Represent whether the token is the answer start token and the answer end token respectively. The final loss is defined as the sum of the above three losses of evidence extraction, modality detection, and answer extraction.
[0053] Preferably, the specific usage steps of the model training and optimization in S6 are as follows:
[0054] Step 1: Model initialization and retriever training
[0055] Use the albert-base-v2 model in the BERT series to encode natural language knowledge, and at the same time use the Resnet101 model pre-trained on the ImageNet dataset to encode image data; set the index vector mapping dimension to 128, and determine the number of retrieved evidences to be 2000. Then, use bert-base-uncased as the basis of the machine reading comprehension model and initialize it; based on the checkpoint model pre-trained on CoQA; run a dedicated training script to train the retriever model and save its checkpoint after training is completed;
[0056] Step 2: Evidence Embedding Generation and Vector Retrieval Library Construction
[0057] Load the previously saved retriever model checkpoint and use this model to generate the embedding representations of all evidences; these embedding representations will be used as feature vectors for subsequent similarity calculation and evidence retrieval; pass the generated embedding representations into the faiss vector retrieval library to quickly and efficiently retrieve evidences related to the query;
[0058] Step 3: Pipeline Model Training and Overall Model Saving:
[0059] Based on the faiss vector retrieval library, further train the pipeline model. The pipeline model includes a question-answering module and an answer generation module, and uses the retrieved evidences to reason and generate answers; during the training process, use the AdamW optimizer to optimize the model parameters; after training is completed, save the overall model, including the retriever and the pipeline model, for subsequent result evaluation and practical application.
[0060] Preferably, the specific steps for evaluating based on the MMConvQA dataset of ChatGPT in S7 are as follows:
[0061] Step 1: Call the ChatGPT API to obtain answers:
[0062] Use the gpt-3.5-turbo model API provided by OpenAI to submit requests in the form of a conversation according to the given prompts and original questions and obtain the answers generated by ChatGPT;
[0063] Step 2: Evaluate the performance of ChatGPT on the MMConvQA dataset:
[0064] Compare the answers generated by ChatGPT with the standard answers in the MMConvQA dataset, calculate the accuracy metric to evaluate its performance; pay special attention to questions with the answer type of yes / no and make manual corrections to ensure the accuracy of the evaluation results; at the same time, record and analyze the performance differences of ChatGPT on different types of questions;
[0065] Step 3: Analyze and discuss the evaluation results:
[0066] Conduct in-depth analysis and discussion on the evaluation results, exploring the advantages and disadvantages of ChatGPT in multi-modal conversational question-answering tasks; analyze ChatGPT's performance in understanding questions, generating answers, and processing multi-modal information, and propose potential improvement directions to optimize the model structure and enhance the training data to further improve its performance in multi-modal conversational question-answering tasks.
[0067] Preferably, the system includes a retriever and a reader; the retriever includes a question encoder and a multi-modal knowledge encoder; the reader includes a table answer acquirer, a text answer acquirer, and an image answer acquirer.
[0068] The beneficial effects of the present invention are as follows:
[0069] 1. By deeply integrating multi-modal data representation theory, evidence extraction technology, and adaptive answer acquisition strategy, the present invention aims to accurately retrieve and integrate data from multiple modalities such as text, tables, and pictures to comprehensively answer questions in the conversation. By introducing an advanced adaptive extractor, the model can intelligently identify and extract multi-modal information closely related to the question, thereby improving the accuracy and comprehensiveness of the answer. To verify the performance of the model, this application chooses to evaluate on the ChatGPT platform to test the model's comprehensive capabilities in multi-modal data processing, conversation context understanding, and answer generation through actual conversation scenarios; this innovative attempt not only provides a new solution for the processing of the MMConvQA dataset but also opens up a new path for the research in the field of multi-modal conversational question-answering.
[0070] 2. By using a pre-trained model to deeply encode questions and multi-modal knowledge and adopting the maximum inner product search strategy, the present invention efficiently extracts evidence closely related to the question from three modalities of text, tables, and pictures. Subsequently, an advanced machine reading comprehension model is built to accurately analyze these evidences to obtain the answer to the question. The entire architecture is designed based on an adaptive extractor, which can flexibly handle changing questions and multi-modal knowledge environments to achieve efficient and accurate multi-modal conversational question-answering.
[0071] 3. The present invention designs a targeted prompt and an experimental scheme for calling the API based on ChatGPT to evaluate the performance of the constructed multi-modal conversational question-answering model on the MMConvQA dataset. Through a clearly defined question-answering accuracy metric, a detailed evaluation was carried out, and an in-depth discussion was conducted on the experimental results. The reasons for the model's performance were analyzed and revealed, and the existing error types were pointed out, such as inappropriate modality selection and insufficient information integration. Potential research directions were also explored, such as optimizing the multi-modal fusion strategy and enhancing the model's ability to understand complex conversation contexts, etc., in order to further improve the performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 It is a schematic diagram of the overall usage process of the present invention;
[0073] Figure 2 It is a schematic diagram of the model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0075] As Figures 1 to 2 shown, the embodiments of the present invention provide a conversational intelligent question-answering method for power grid multi-modal data. The specific usage steps of this method are as follows:
[0076] S1: Data preparation and preprocessing: Collect multi-modal data of questions, texts, tables, and images. Preprocess the texts and tables to adapt to the BERT model, and keep the images in the original format for ResNet101 processing;
[0077] S2: Model selection and initialization: Select the BERT series and ResNet101 models, load the pre-trained weights, and initialize the projection matrix and the encoder;
[0078] S3: Multi-modal data representation: Use the BERT series model to encode the questions to obtain embeddings, and use different encoders to process the texts, tables, and images to obtain knowledge embeddings;
[0079] S4: Multi-modal evidence extraction: Construct and load an efficient retriever, and quickly retrieve the most relevant multi-modal evidence from the knowledge base for the questions;
[0080] S5: Adaptive answer acquisition: The modality detector predicts the answer modality, and the answer extractor extracts the answer from the evidence according to the modality;
[0081] S6: Model Training and Optimization: Define the overall loss function, integrate the losses of evidence extraction, modality detection, and answer acquisition, and train with the AdamW optimizer until convergence;
[0082] S7: Evaluation of the MMConvQA Dataset Based on ChatGPT: Call the ChatGPT API to evaluate the performance of MMConvQA and explore directions for improvement.
[0083] Among them, the data preparation and preprocessing in S1 refer to that when constructing a multi-modal conversational question-answering system, it is necessary to widely collect data, including information in multiple modalities such as questions, text paragraphs, tables, and images. After the data collection is completed, the next step is to perform data preprocessing. For text and table data, specific formatting methods are adopted to make them adapt to the input requirements of the BERT series models. When processing long texts, the sliding window advanced technology is used to ensure the integrity and accuracy of the text information. For image data, its original format is kept unchanged so that it can be directly processed by the ResNet101 model later. This series of preprocessing steps lay a solid foundation for the subsequent model training and the construction of the question-answering system.
[0084] The "specific formatting method" refers to converting text and table data into a format that the BERT series models can understand and process. This usually involves the following steps:
[0085] Text Tokenization and Encoding: For text data, first, tokenization needs to be performed to split the text into word or sub-word units. Then, using the vocabulary provided by the BERT model, these units are mapped to the corresponding IDs to form a string of digital encodings. This is the basis for the BERT model to recognize and process text.
[0086] Input Sequence Construction: The BERT model requires the input data to have a specific format, usually including the input text, special tokens (such as [CLS] and [SEP]), and possibly segment identifiers. Therefore, it is necessary to construct the text and table data according to this format.
[0087] Length Adjustment and Padding: Since the BERT model has a limit on the length of the input sequence, it is necessary to truncate overly long texts and pad overly short texts to ensure that all input sequences have the same length.
[0088] Table Data Conversion: For table data, it may be necessary to convert it into text form, or treat each cell in the table as a separate text segment for processing. Then, encoding and formatting are performed in the same way as text data.
[0089] Among them, the model selection and initialization in S2 refer to, for the natural language encoding part, selecting bert-base-uncased and albert-base-v2 from the BERT series of models, which have excellent performance in the field of natural language processing. For the image encoding part, the ResNet101 model is adopted, which is famous for its strong image feature extraction ability. When initializing the model, the weights of these pre-trained models are first loaded to ensure that the model can learn useful knowledge from a large amount of data. At the same time, the necessary projection matrices and encoders are also initialized to map data of different modalities into the same semantic space, so as to achieve the effective fusion and utilization of multimodal information.
[0090] Among them, the specific usage steps of the multimodal data representation in S3 are as follows:
[0091] Step 1: Question encoding
[0092] Select model: Select the bert-base-uncased model from the BERT series of models for question encoding;
[0093] Input processing: Preprocess the question text, including word segmentation, adding special tokens, and adjusting the length to meet the input requirements of the BERT model;
[0094] Generate embedding: Input the processed question text into the BERT model to obtain the question embedding representation; this is usually achieved by extracting the vector corresponding to the [CLS] token in the output layer of the model.
[0095] Step 2: Text paragraph and table encoding
[0096] Select knowledge encoder: For text paragraphs and tables, the same BERT series of models as used for question encoding can be selected, or other suitable text encoders can be selected according to specific circumstances.
[0097] Input formatting: Convert the text paragraph and table data into a format that the encoder can understand, such as converting the table data into a series of text fragments or a structured representation;
[0098] Generate embedding: Input the formatted text paragraph and table data into the knowledge encoder to generate the corresponding knowledge embedding representation;
[0099] Step 3: Image encoding
[0100] Select image encoder: Select ResNet101 as the image encoder, which can effectively extract image features;
[0101] Preprocess the image: Perform necessary preprocessing on the image, such as resizing, normalizing, etc., to meet the input requirements of ResNet101;
[0102] Generate embeddings: Input the preprocessed image into the ResNet101 model to obtain the image embedding representation; This is achieved by extracting the feature vectors in the output layer of the model.
[0103] The BERT series of models will be described below
[0104] According to the processing of the question by ORConvQA, the reformulated question is,
[0105] q′ k =[CLS]q 1 [SEP]q k-w [SEP]···[SEP]q k-1 [SEP]q k [SEP]#(1)
[0106] where [CLS] and [SEP] are special tokens introduced by BERT, and w is the size of the sliding window.
[0107] Then the question representation v q is
[0108] v q =W q F q (q′ k )#(2)
[0109] where F q is the question encoder based on ALBERT, and W q is the question projection matrix.
[0110] For different modality terms in C, they are passed to different knowledge encoders. For each paragraph p in C p , the paragraph representation j is obtained as follows as follows
[0111]
[0112] where F p is the paragraph encoder based on ALBERT, and W p is the paragraph projection matrix.
[0113] According to the work of TaPaS, the table is linearlyized row by row into t′ j to obtain their representations. The representation of the table is calculated by the following formula
[0114]
[0115] Among them, F t is a table encoder based on ALBERT, and W t is a table projection matrix.
[0116] For each image i in C i , its representation is obtained as follows j . As follows
[0117]
[0118] Among them, F i is a Resnet network pre-trained on ImageNet. Among them, d q , d p , d t , d i have the same dimension.
[0119] Among them, the specific steps of multi-modal evidence extraction in S4 are as follows:
[0120] Step 1: Train the embedding model:
[0121] Use a large amount of text data to train a machine learning model to map queries and document content to a common vector embedding space; ensure that the embedding model can capture the semantic similarity between texts, so that the distance between similar texts in the embedding space is closer;
[0122] Step 2: Construct the database embedding:
[0123] Preprocess all documents to be retrieved, and use the trained embedding model to convert them into vector representations; store these vectors to construct a database embedding set for subsequent nearest neighbor search;
[0124] Step 3: Query vector generation and retrieval:
[0125] When the user inputs a query, use the embedding model to convert the query text into a vector representation;
[0126] Perform a maximum inner product search in the database embedding set to find the database embedding with the maximum inner product with the query vector, that is, the closest embedding.
[0127] Maximum Inner Product Search (MIPS); expressed by the formula as:
[0128] p = argmax x∈X x T q#(6)
[0129] Where X is a set of embedding vectors, q is the vector to be queried, and p, the vector with the maximum inner product, is to be found in the set;
[0130] Extraction and evaluation results:
[0131] Based on the retrieved closest embedding, accurately locate the original document or paragraph. Subsequently, carefully extract the information relevant to the query from these documents as answers or evidence. To ensure accuracy, strictly evaluate the extraction results and integrate or sort multiple results to provide users with comprehensive and accurate answers.
[0132] There are three knowledge encoders F p 、F t 、F i independent of the question above to achieve pre-computed multimodal encoding and perform effective maximum inner product search. After pre-training the question encoder and the knowledge encoder, all items in C are input into the knowledge encoder to obtain their representations, and the parameters of the knowledge encoder are frozen in the following training stage. Thanks to this, the similarity s q between the given question embedding v a and all knowledge item embeddings is effectively calculated by inner product, and the top N r items I r are selected as evidence, where N r is the number of retrieved items.
[0133] The loss of evidence retrieval is defined as
[0134]
[0135] where is the evidence retrieval score for an item in I r , and y indicates whether it is a positive example.
[0136] Internally, by training the embedding model and constructing a database embedding set, efficient query vector generation and retrieval are achieved, which can capture the semantic similarity between texts and improve retrieval accuracy. At the same time, this method introduces independent knowledge encoders for pre-computation, effectively performs maximum inner product search, quickly locates relevant evidence, and strictly evaluates and integrates the extraction results to provide users with comprehensive and accurate answers, improving retrieval efficiency and user experience.
[0137] Among them, the specific usage steps of adaptive answer acquisition in S5 are as follows:
[0138] Step 1: Modal Detection: First, input the user's question into the modal classifier; this classifier encodes the question and predicts the probabilities of three modalities (text, table, image), and selects the modality with the highest probability as the most suitable answer modality;
[0139] The modal classifier is expressed as:
[0140] s b = f(W c F c (q′ k ))#(8)
[0141] where f() represents the softmax function, and F c is the question encoder,
[0142] Step 2: Answer Extractor Selection and Application:
[0143] If the modality is text, use the text answer extractor, input the question and the retrieved text paragraphs into the machine reading comprehension model, the model calculates the likelihood of each token as the start and end of the answer, and outputs the best answer paragraph;
[0144] If the modality is table, select the table answer extractor, combine the question text with the linearized table data, encode through BERT, and use a linear classifier to determine the start position of the answer in the table, so as to extract the answer;
[0145] If the modality is image, use the image answer extractor, combine the question text with the image descriptions in the preset answer set, process through BERT and combine with visual features, use a linear classifier to find the best start token and end token in the description as the answer;
[0146] Answer Integration and Output: Based on the results output by the selected answer extractor, input the answer text sequence into the BERT model to obtain the deep representation of each token. Subsequently, these text representations are cleverly combined with the visual feature v i Combined, through integration processing, the final answer not only maintains a close correlation with the user's query question, but also can provide accurate and useful information.
[0147] The answer extraction score s c of the candidate answers predicted by the above three extractors is defined as the average of the probabilities of the start and end tokens. For each candidate answer, its final score is calculated as the evidence extraction score s a , the modal detection score s b and the answer extraction score s cFor the sum, select the answer with the highest score as the final answer output, where the modality detection loss L md is the cross-entropy loss commonly used for training multi-class classification tasks. The answer acquisition loss for training the three answer acquirers is defined as
[0148]
[0149] where y 1 and y 2 represent whether the token is the start token and end token of the answer respectively. The final loss is defined as the sum of the above three losses of evidence extraction, modality detection, and answer acquisition.
[0150] Among them, the specific usage steps of model training and optimization in S6 are as follows:
[0151] Step 1: Model initialization and retriever training
[0152] Use the albert-base-v2 model in the BERT series to encode natural language knowledge, and at the same time use the Resnet101 model pre-trained on the ImageNet dataset to encode image data; set the index vector mapping dimension to 128, and determine the number of retrieved evidences to be 2000. Then, use bert-base-uncased as the basis of the machine reading comprehension model and initialize it; based on the checkpoint model pre-trained on CoQA; run a dedicated training script to train the retriever model and save its checkpoint after training is completed;
[0153] Step 2: Evidence embedding generation and vector retrieval library construction
[0154] Load the retriever model checkpoint saved previously and use this model to generate the embedding representations of all evidences; these embedding representations will be used as feature vectors for subsequent similarity calculation and evidence retrieval; pass the generated embedding representations into the faiss vector retrieval library to quickly and efficiently retrieve evidences related to the query;
[0155] Step 3: Pipeline model training and overall model saving:
[0156] Based on the faiss vector retrieval library, further train the pipeline model, which includes a question-answering module and an answer generation module, and use the retrieved evidences (the number is set to 10) for inference and answer generation; during the training process, use the AdamW optimizer to optimize the model parameters; after training is completed, save the overall model, including the retriever and the pipeline model, for subsequent result evaluation and practical application.
[0157] Among them, the specific steps for evaluating the MMConvQA dataset based on ChatGPT in S7 are as follows:
[0158] Step 1: Call the ChatGPT API to obtain answers:
[0159] Use the gpt-3.5-turbo model API provided by OpenAI to submit a request in the form of a conversation according to the given prompt and the original question, and obtain the answer generated by ChatGPT. Ensure that the input and output formats of the API are correctly processed during the call, as well as handle possible exceptions and errors.
[0160] The message format for calling the API is as follows;
[0161]
[0162] Among them, the "content" of the system role in the first line is the prompt content, then enter the first question of the multi-round conversation, and "completion.choices.message.content" is the response of ChatGPT. When entering subsequent questions, expand this answer into the message to provide the conversation context for subsequent Q&A. Do not use the correct answer and the gold question to correct the context in consecutive conversations.
[0163] Step 2: Evaluate the performance of ChatGPT on the MMConvQA dataset:
[0164] Compare the answers generated by ChatGPT with the standard answers in the MMConvQA dataset, and calculate the accuracy metric to evaluate its performance; pay special attention to questions with answer types of yes / no, and perform manual correction to ensure the accuracy of the evaluation results; at the same time, record and analyze the performance differences of ChatGPT on different types of questions;
[0165] Step 3: Analyze and discuss the evaluation results:
[0166] Conduct in-depth analysis and discussion on the evaluation results, and explore the advantages and disadvantages of ChatGPT in multi-modal conversational Q&A tasks; analyze ChatGPT's performance in understanding questions, generating answers, and processing multi-modal information, and propose potential improvement directions, such as optimizing the model structure and enhancing the training data, to further improve its performance in multi-modal conversational Q&A tasks.
[0167] Among them, the system includes a retriever and a reader; the retriever includes a question encoder and a multi-modal knowledge encoder; the reader includes a table answer acquirer, a text answer acquirer, and an image answer acquirer.
[0168] The retriever consists of a question encoder and a multi-modal knowledge encoder, and has powerful information processing capabilities. Among them, questions, text knowledge, and table knowledge are efficiently encoded by two independent Albert-base-v2 pre-trained models respectively, while image knowledge is accurately encoded by the Resnet101 model pre-trained on the ImageNet dataset. This design ensures that various types of knowledge can be accurately and quickly represented. In the retrieval stage, the question of the current conversation turn is reformulated and processed through the encoder together with the evidence items. Subsequently, using the Faiss vector retrieval library, the cosine similarity between the given question embedding and all knowledge item embeddings is calculated, and the top 10 items with the highest similarity are accurately selected as key evidence and sent to the reader for further analysis.
[0169] The reader part uses a machine reading comprehension model with the Bert-base-uncased architecture, which has excellent text understanding and analysis capabilities. The model first receives the question as input and predicts the probabilities of three modalities (the modality is a table, the modality is text, and the modality is an image) through an internal mechanism, providing important clues for subsequent processing. Then, the model carefully calculates two scores for each token in the paragraph as the start token and the end token, so as to accurately predict the paragraph where the answer is located. For each candidate answer, we comprehensively consider its evidence extraction score, modality detection score, and answer acquisition score to obtain a final comprehensive score. Based on this score, we select the answer with the highest score as the final answer for output, ensuring that users can obtain the most accurate and valuable information.
[0170] Positive effects
[0171] 1. Comparative experiments
[0172] The present invention uses Recall and Normalized Discounted Cumulative Gain (NDCG) to evaluate the inclusion and ranking position of retrieval evidence, and uses macro-average F1 and Exact Match (EM) at the word level to evaluate the performance of answer extraction. The specific description of the evaluation metrics for this task is as follows.
[0173] Recall: It refers to the proportion of relevant evidence included in the retrieved evidence items. The higher the Recall, the better the performance of retrieving relevant items.
[0174] NDCG: A metric used to measure the quality of retrieval ranking. It combines evidence relevance and ranking position and is calculated by normalizing the discounted cumulative gain of relevance. The higher the NDCG, the higher the quality of the retrieval ranking.
[0175] F1: An indicator for evaluating the performance of answering questions. First, decompose the gold answer and the answer given by the model for each question into individual words; then calculate the precision and recall rate for each answer, where precision represents the proportion of words in the answer given by the model that exactly match the words in the gold answer, and recall rate represents the proportion of words in the reference answer that are included in the answer given by the model, thereby calculating the F1 for each question; finally, calculate the weighted average of the F1 values for each question. The higher the F1, the better the performance of the model in answering questions.
[0176] EM: Refers to the proportion of the answer given by the model that exactly matches the gold answer. The higher the EM, the higher the performance of the model in accurately answering questions.
[0177] Table 1 Experimental Results
[0178]
[0179] The experimental results are shown in Table 1, and the baseline models used in the experiment are selected as follows.
[0180] ORConvQA: The baseline model used in the ORConvQA dataset, which is a text modality, multi-turn conversation dataset;
[0181] ManyModelQA: The baseline model used in the ManyModelQA dataset, which is a multi-modal, single-turn conversation dataset;
[0182] MAE(EG): EG (Evidence Given) means that if the relevant evidence for the question is not retrieved, it is supplemented to the retrieved item list;
[0183] MAE: The model used in the present invention, which is the abbreviation of Multimodal Conversational QA-system with Adaptive Extractor.
[0184] As can be seen from Table 1, the existing baseline models of ORConvQA and ManyModalQA cannot handle MMCoQA well because they are either unimodal or single-turn conversation. The performance of the MAE model is significantly better than these two models; when supplementing the relevant evidence of the question to the retrieval list, it is much better than the standard MAE model, indicating that evidence extraction is one of the difficulties in this task.
[0185] 2. Evaluation of the MMConvQA Dataset Based on ChatGPT
[0186] Table 2 Evaluation Results of the MMConvQA Dataset Based on ChatGPT
[0187]
[0188] Table 3 Examples of ChatGPT Error Types
[0189]
[0190]
[0191] Table 2 shows the evaluation results of the MMConvQA dataset based on ChatGPT, and Table 3 lists examples of three common types of errors. The summary and discussion of the experimental results are as follows:
[0192] (1) Compared with the EM results of the multi-modal conversational question-answering model based on the adaptive extractor, the accuracy of ChatGPT is much higher, and it can answer nearly one-third of the questions on both the test set and the development set. This is related to the types of questions in the dataset, indicating that ChatGPT can answer some world questions and knowledge questions in professional fields, which is related to its large pre-trained corpus.
[0193] (2) For the question in the first row of Table 3, the knowledge modality required is an image. Due to the limitation of single-text modality input and the lack of the ability to retrieve knowledge bases, ChatGPT lacks the rich information brought by the image modality, resulting in incorrect answers; the second row is an example that requires specific domain knowledge, and the knowledge modality required for the question is a table. The lack of such semi-structured knowledge leads to the need for more supplementary information in its response; the third row is the second question in a multi-turn conversation. Due to the incorrect answer to the first question, the subsequent semantic reference is unclear, which belongs to a context understanding error.
[0194] (3) The experimental results show that in tasks such as MMConvQA that require retrieving multi-modal knowledge, there is still room for improvement in the performance of ChatGPT, such as bridging models of other modalities, using the Chain of Thought (CoT) or Self-Consistency (SC) strategies, and leveraging external retrieval models to supplement the model's knowledge.
[0195] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0196] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A conversational intelligent question-answering method for multimodal data of power grid, characterized by: The specific steps of using this method are as follows: S1: Data preparation and preprocessing: Collect multimodal data including questions, texts, tables, and images. Preprocess texts and tables to fit the BERT model, and keep images in their original format for processing by ResNet101. S2: Model selection and initialization: Select the BERT series and ResNet101 models, load pre-trained weights, initialize the projection matrix and encoder; S3: Multimodal data representation: Using the BERT series model to encode the problem embedding, different encoders to process text, tables, and images to obtain knowledge embedding; S4: Multimodal evidence extraction: Build and load an efficient retriever to quickly retrieve the most relevant multimodal evidence from the knowledge base for the question; S5: Adaptive answer acquisition: modality detection predicts the answer modality, and the answer extractor extracts the answer from the evidence based on the modality; S6: Model training and optimization: define the overall loss function, integrate the evidence extraction, modality detection and answer acquisition losses, and train with AdamW optimizer until convergence; S7: Evaluation of MMConvQA dataset based on ChatGPT: Call ChatGPT API to evaluate MMConvQA performance and explore improvement directions.
2. The method of conversational intelligent question-answering for multimodal data of power grid according to claim 1, characterized in that: The data preparation and preprocessing in S1 refers to the need to collect data extensively when building a multimodal conversational question-answering system, including information in multiple modes such as questions, text paragraphs, tables, and images. After the data collection is completed, data preprocessing is performed next. For text and table data, a specific formatting method is used to adapt it to the input requirements of the BERT series models. When processing long texts, the advanced sliding window technology is used to ensure the integrity and accuracy of the text information. For image data, its original format is kept unchanged so that it can be directly processed using the ResNet101 model later.
3. The method of conversational intelligent question-answering for multimodal data of power grid according to claim 1, characterized in that: The model selection and initialization in S2 refers to the selection of bert-base-uncased and albert-base-v2 in the BERT series of models for the natural language encoding part. These models have excellent performance in the field of natural language processing, and for the image encoding part, the ResNet101 model is used, which is well-known for its powerful image feature extraction capabilities. When initializing the model, the weights of these pre-trained models are first loaded to ensure that the model can learn useful knowledge from a large amount of data. At the same time, the necessary projection matrix and encoder are initialized to map data of different modalities into the same semantic space.
4. The method of conversational intelligent question-answering for multimodal data of power grid according to claim 1, characterized in that: The specific steps for using the multimodal data representation in S3 are as follows: Step 1: Problem coding Select model: Select the bert-base-uncased model from the BERT series of models for question encoding; Input processing: preprocess the question text, including word segmentation, adding special tags, and adjusting the length to meet the input requirements of the BERT model; Generate embedding: Input the processed question text into the BERT model to obtain the question embedding representation; Step 2: Encoding text paragraphs and tables Select knowledge encoder: For text paragraphs and tables, choose the same BERT series model as question encoding, or choose other appropriate text encoders according to the specific situation; Input formatting: converting text paragraphs and tabular data into a format that the encoder can understand, and converting tabular data into a series of text fragments or structured representations; Generate embedding: Input the formatted text paragraphs and table data into the knowledge encoder to generate the corresponding knowledge embedding representation; Step 3: Image encoding Select image encoder: Select ResNet101 as the image encoder, which can effectively extract image features; Preprocess the image: perform necessary preprocessing on the image, resize and normalize it to meet the input requirements of ResNet101; Generate embedding: Input the preprocessed image into the ResNet101 model to obtain the image embedding representation; This is achieved by extracting the feature vector in the output layer of the model.
5. The method of conversational intelligent question and answer for multimodal data of power grid according to claim 1, characterized in that: The specific steps of multimodal evidence extraction in S4 are as follows: Step 1: Train the embedding model: Use a large amount of text data to train machine learning models to map queries and document content into a common vector embedding space; ensure that the embedding model can capture the semantic similarity between texts so that similar texts are closer in the embedding space; Step 2: Build database embedding: Preprocess all documents to be retrieved and convert them into vector representations using the trained embedding model; Step 3: Query vector generation and retrieval: When a user enters a query, the query text is converted into a vector representation using an embedding model; Perform a maximum inner product search on the database embedding set to find the database embedding with the largest inner product with the query vector, i.e. the closest embedding; The maximum inner product search formula is: p=argmax x∈X x T q#(6) Where X is the set of embedded vectors, q is the vector to be queried, and the vector p with the largest inner product is to be found in the set; Extraction and evaluation results: Based on the closest embedding retrieved, the original document or paragraph is accurately located. Subsequently, information related to the query is carefully extracted from these documents as answers or evidence. To ensure accuracy, the extracted results are strictly evaluated, and multiple results are integrated or ranked to provide users with comprehensive and accurate answers.
6. The method of conversational intelligent question and answer for multimodal data of power grid according to claim 1, characterized in that: The specific steps for obtaining the adaptive answer in S5 are as follows: Step 1: Modality detection: First, the user’s question is input into the modality classifier; the classifier encodes the question and predicts the probabilities of the three modalities, and selects the modality with the highest probability as the most appropriate answer modality; The modality classifier is expressed as: Where f() represents the softmax function, F c is the question encoder, s b ∈R 3 Step 2: Answer acquisition device selection and application: If the modality is text, use the text answer acquirer to input the question and the retrieved text paragraph into the machine reading comprehension model. The machine reading comprehension model calculates the probability of each token being the beginning and end of the answer and outputs the best answer paragraph. If the modality is a table, select the table answer obtainer, combine the question text with the linearized table data, encode it through BERT, and use a linear classifier to determine the starting position of the answer in the table to extract the answer; If the modality is an image, the image answer getter is used to combine the question text with the image description in the preset answer set, and then processed by BERT and combined with visual features, and a linear classifier is used to find the best answer start and end tokens in the description as the answer; Answer integration and output: Based on the output of the selected answer acquirer, the answer text sequence is input into the BERT model to obtain the deep representation of each token. Subsequently, these text representations are cleverly combined with the visual features v i Combined with the integrated processing, the final answer not only maintains close relevance to the user's query, but also provides accurate and useful information; The answer extraction scores s of the candidate answers predicted by the above three getters c Defined as the average of the probabilities of the start and end tokens, for each candidate answer, its final score is calculated as the evidence extraction score s a , modal detection score s b and answer to get score c The answer with the highest score is selected as the final answer output, where the modality detection loss L md is the cross entropy loss commonly used to train multi-class classification tasks. The answer acquisition loss used to train the three answer acquirers is defined as, Among them, y1 and y2 represent whether the word is the answer start word and the answer end word respectively. The final loss is defined as the sum of the three losses mentioned above: evidence extraction, modality detection and answer acquisition.
7. The method of conversational intelligent question and answer for multimodal data of power grid according to claim 1, characterized in that: The specific steps for model training and optimization in S6 are as follows: Step 1: Model initialization and retriever training Use the albert-base-v2 model in the BERT series to encode natural language knowledge, and use the Resnet101 model pre-trained on the ImageNet dataset to encode image data; Set the index vector mapping dimension to 128 and determine the number of retrieval evidences to 2000. Then, use bert-base-uncased as the basis of the machine reading comprehension model and initialize it; Based on the CoQA pre-trained checkpoint model; run a dedicated training script to train the retriever model and save its checkpoint after training is complete; Step 2: Evidence embedding generation and vector retrieval library construction Load the previously saved retriever model checkpoint and use the model to generate embedding representations of all evidence; these embedding representations will be used as feature vectors for subsequent similarity calculations and evidence retrieval; pass the generated embedding representations into the Faiss vector retrieval library to quickly and efficiently retrieve evidence related to the query; Step 3: Pipeline model training and overall model saving: Based on the faiss vector retrieval library, the pipeline model is further trained. The pipeline model includes a question-answering module and an answer generation module, which use the retrieved evidence to reason and generate answers. During the training process, the AdamW optimizer is used to optimize the model parameters. After the training is completed, the overall model, including the retriever and the pipeline model, is saved for subsequent result evaluation and practical application.
8. The method of conversational intelligent question and answer for multimodal data of power grid according to claim 1, characterized in that: The specific steps of the ChatGPT-based MMConvQA dataset evaluation in S7 are as follows: Step 1: Call ChatGPT API to get the answer: Use the gpt-3.5-turbo model API provided by OpenAI to submit requests in the form of dialogues and obtain answers generated by ChatGPT based on the given prompts and original questions; Step 2: Evaluate the performance of ChatGPT on the MMConvQA dataset: Compare the answers generated by ChatGPT with the standard answers in the MMConvQA dataset and calculate the accuracy metric to evaluate its performance; Pay special attention to questions with yes / no answers and make manual corrections to ensure the accuracy of the evaluation results. At the same time, record and analyze the performance differences of ChatGPT on different types of questions. Step 3: Analyze and discuss the evaluation results: Conduct an in-depth analysis and discussion on the evaluation results to explore the strengths and weaknesses of ChatGPT in multimodal conversational question-answering tasks; analyze ChatGPT's performance in understanding questions, generating answers, and processing multimodal information, propose potential improvement directions, optimize the model structure, and enhance training data to further improve its performance in multimodal conversational question-answering tasks.
9. A conversational intelligent question-answering system for multimodal data of power grid, characterized by: The system includes a retriever and a reader; the retriever includes a question encoder and a multimodal knowledge encoder; the reader includes a table answer acquirer, a text answer acquirer, and an image answer acquirer.