Implementation method of visual question-answering system
Through the methods of retrieval enhancement and multimodal large model data enhancement, the problem of the inability to generate better models in the prior art is solved, a more accurate and comprehensive question-and-answer system model is realized, and the processing capability of visual question-and-answer tasks is improved.
Patent Information
- Application Number
- CN202510049986.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art cannot generate better models, resulting in the inability to effectively implement the Q&A system, especially when dealing with visual Q&A tasks, where accuracy and factuality are limited.
By obtaining the knowledge base, including knowledge points related to picture data, search enhancement and multimodal big model data enhancement, combining the retrieval enhancement text description data and the enhanced data output from the multimodal big model, modal fusion is carried out and the question-and-answer system model is trained.
It improves the comprehensiveness and accuracy of the learnable content of the model, effectively improves the performance of the model, and enables the Q&A system to handle visual Q&A tasks more accurately.
Smart Images

Figure CN119988659A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for implementing a visual question answering system. Background Art
[0002] In recent years, with the development of large-scale language model technology, for example, the language model GPT (Generative Pre-training Model) series based on artificial intelligence technology, the language representation model BERT (Bidirectional Encoder Representation), etc., it can significantly improve the ability of natural language processing.
[0003] However, although the corresponding large-scale language model technology can generate fluent natural language and answer simple questions, it still has certain limitations in terms of accuracy and factuality.
[0004] At present, visual question answering is a task that tests the capabilities of various large models. The ability of related models to answer questions has been unprecedentedly improved and has been applied to various fields of life, such as education question answering, medical question answering, video question answering, etc. With the growing demand for corresponding natural language processing applications, especially the increasing demand for question answering systems that can provide accurate and comprehensive answers, existing language models are unable to adapt to the growth of this demand.
[0005] Moreover, while the rapid development of question-answering systems has resulted in language models being unable to adapt to the corresponding growth, information retrieval technology has also developed rapidly, laying the foundation for the emergence of RAG (retrieval-augmented generation). The corresponding RAG technology combines the functions of information retrieval and generation models, so that relevant external information can be introduced through RAG technology during the generation of the model, so as to answer relevant questions more accurately and comprehensively. However, the existing technology accurately retrieves relevant content but cannot use this information in an effective way, resulting in the inability to generate a better model, which further leads to the inability to implement the corresponding question-answering system based on the corresponding model.
[0006] In view of this, the present invention is proposed. Summary of the invention
[0007] The purpose of the present invention is to provide a method for implementing a visual question answering system to solve the problem in the prior art that a better model cannot be generated, thereby well implementing the corresponding question answering system.
[0008] The objective of the present invention is achieved through the following technical solutions:
[0009] A method for implementing a visual question answering system, comprising:
[0010] Acquire a knowledge base, wherein the knowledge base includes knowledge points related to the image data;
[0011] Searching the knowledge base based on the question text related to each image contained in the original training data set to obtain question enhanced text description data after retrieval enhancement;
[0012] The image data set in the original training data set is combined with the corresponding text prompts through the multimodal large model to perform data enhancement, and the enhanced data output by the multimodal large model is obtained as the image text description data;
[0013] Modal fusion is performed based on the question-enhanced text description data after retrieval enhancement, the enhanced data output by the multimodal large model and the original training data set, and the constructed question-answering system model is trained based on the modal fusion data, and a visual question-answering system for processing question-answering tasks is established based on the trained question-answering system model.
[0014] The knowledge base includes: text data obtained by manually constructing a knowledge outline and expanding a generative large model, and before performing the retrieval enhancement process, it also includes: removing redundant data in the knowledge base.
[0015] The process of obtaining enhanced data further includes:
[0016] The output length and format of the enhanced data are restricted, and the content and format requirements of the output picture text description are specified.
[0017] The retrieval enhancement adopts a serial retrieval enhancement processing method, and the corresponding processing process includes:
[0018] In the knowledge base, questions related to the image data are input for retrieval to obtain retrieval result information, the retrieval result information includes: knowledge points related to the image data;
[0019] The available search result information is determined in the knowledge base as the question enhanced text description data after the search enhancement.
[0020] The process of determining the available search result information includes:
[0021] The retrieval results are processed using a model that can capture contextual information and semantic relationships in the language to obtain usable retrieval result information.
[0022] The model capable of capturing contextual information and semantic relations in a language includes an all-MiniLM-L6-v2 model, and the step of obtaining available search result information includes:
[0023] The text encoder of the all-MiniLM-L6-v2 model is used to encode the question text and the knowledge base, and the cosine similarity between each question and all knowledge points in the knowledge base is calculated, that is:
[0024]
[0025] Among them, Q E and K E denote the embedding of all questions and the embedding of all answers respectively; n and m denote the number of questions and the number of answers respectively; q i Indicates Q E The i-th question in k j K E The jth piece of knowledge in
[0026] An optimal predetermined number of retrieval result information is selected as the available retrieval result information according to the calculated cosine similarity.
[0027] The constructed question-answering system model includes an input layer, a modality fusion layer and an output layer, wherein:
[0028] The input layer is used to extract features and includes:
[0029] Use the pre-trained bidirectional encoder representation BERT model to encode the question text and image text description in the input data and the retrieved knowledge points to obtain the corresponding hidden information representation H T and the hidden information representation of image text description and retrieval knowledge H D ; The visual features of the associated picture data are extracted through an image recognition network, and the linear transformation function is used to project the visual features into the same feature space as the picture text description to obtain the corresponding visual information representation H V ;
[0030] The modality fusion layer is used to perform multimodal fusion processing on the extracted features, and includes:
[0031] The above H D , H V Input the cross-attention module to obtain a visual representation H that enhances data awareness D→V , and then H D , H T Input the cross attention module to obtain enhanced data-aware text representation H D→T , then H D→V and H D→T After connecting together, we get the multimodal information representation H, namely:
[0032]
[0033] Among them, Cross-ATT is the processing performed by the cross-attention module;
[0034] The output layer is used to perform self-attention calculation and includes:
[0035] The obtained multimodal information representation H is input into the multimodal encoder for self-attention calculation. The multimodal data representation H is input into the self-attention layer ATT and the residual layer output to obtain Afterwards Input a standard feedforward network with GELU as activation function and residual layer to obtain H', and finally get the output calculated by the multimodal encoder: the multimodal information represents the first tag H of H' 0 , input it into a normalized exponential function softmax layer for output:
[0036]
[0037] Among them, W M ∈R d×2 represents the weight matrix, and y represents the true answer label of the question.
[0038] The visual representation H V Obtained by the following formula:
[0039] H v =W v ResNet(I),W v ∈R d×2048 ;
[0040] Among them, W v ∈R d×2048 is a learnable parameter, and the output of the convolutional layer is calculated as:
[0041] ResNet(I)={r j |r j ∈R 2048 ,j=1,2,...,49};
[0042] In the formula, when the original image is divided into 7×7 regions, each region is represented by a 2048-dimensional vector r j Indicates that the pixels of the image are 224×224.
[0043] The implementation of the cross attention module includes:
[0044] use and Represents the data of two modes α and β respectively, where T represents the sequence length and d represents the feature dimension; and the query Querys is defined using the desired mode of interest. Another mode defines keys and values, specifically in, There are three weight matrices, corresponding to the query matrix, key matrix and value matrix respectively;
[0045] Then, the potential adjustment from modality β to modality α is expressed as cross-modal attention, namely:
[0046]
[0047] Among them, Y α The length and Q α equal;
[0048] Based on the above formula, a residual link is added, and another position feed-forward sublayer is added to form a cross-modal attention block Cross-ATT.
[0049] The process of training the constructed question answering system model includes a process of updating model parameters, and the process includes:
[0050] Use the cross-loss entropy function as the training objective, minimize the loss function, and update the model parameters according to the gradient of the loss function to the model parameters through the back-propagation algorithm:
[0051]
[0052] Among them, H 0 represents the answer predicted by the model, y i Indicates the true answer label of the question.
[0053] Compared with the prior art, the implementation method of the visual question answering system provided by the present invention uses a retrieval enhancement method to enhance the image text data related to the image data, thereby obtaining additional knowledge. Compared with the traditional question answering model, the learnable content of the model is more comprehensive and accurate, and the performance of the model can be effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0055] Figure 1 A schematic diagram of the implementation process of the method provided in the embodiment of the present invention;
[0056] Figure 2 A schematic diagram of the model framework structure provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0057] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention; it is obvious that the described embodiments are only part of the embodiments of the present invention, not all of the embodiments, which does not constitute a limitation of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of the present invention.
[0058] First, the terms that may be used in this article are explained as follows:
[0059] The term “and / or” means that either or both of them can be realized at the same time. For example, X and / or Y means both “X” or “Y” and “X and Y”.
[0060] The terms "include", "comprises", "contains", "has" or other descriptions with similar semantics should be interpreted as non-exclusive inclusion. For example, including certain technical feature elements (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or products, etc.) should be interpreted as including not only certain technical feature elements explicitly listed, but also other technical feature elements known in the art that are not explicitly listed.
[0061] The term "consisting of..." means excluding any technical feature elements not explicitly listed. If this term is used in a claim, it will make the claim closed, so that it does not contain technical feature elements other than the technical feature elements explicitly listed, except for the conventional impurities related to them. If this term only appears in a clause of a claim, it only limits the elements explicitly listed in the clause, and the elements recorded in other clauses are not excluded from the overall claim.
[0062] The term "parts by mass" refers to the mass ratio relationship between multiple components. For example, if it is described that component X is x parts by mass and component Y is y parts by mass, then the mass ratio of component X to component Y is x:y. 1 part by mass can represent any mass, for example, 1 part by mass can be represented as 1 kg or 3.1415926 kg. The sum of the parts by mass of all components is not necessarily 100 parts, but can be greater than 100 parts, less than 100 parts or equal to 100 parts. Unless otherwise specified, the parts, proportions and percentages described herein are all measured by mass.
[0063] Unless otherwise specified or limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example: it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in this article can be understood according to specific circumstances.
[0064] When concentration, temperature, pressure, size or other parameters are expressed in the form of a numerical range, the numerical range should be understood to specifically disclose all ranges formed by the pairing of any upper limit, lower limit, and preferred value in the numerical range, regardless of whether the range is explicitly stated; for example, if a numerical range of "2 to 8" is stated, the numerical range should be interpreted as including ranges such as "2 to 7", "2 to 6", "5 to 7", "3 to 4 and 6 to 7", "3 to 5 and 7", "2 and 5 to 7", etc. Unless otherwise specified, the numerical ranges stated herein include both their end values and all integers and fractions within the numerical range.
[0065] The orientation or position relationship indicated by terms such as "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", etc. are based on the orientation or position relationship shown in the drawings and are only for the convenience and simplification of description, and do not explicitly or implicitly indicate that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation of this document.
[0066] The following is a detailed description of a method for implementing a visual question-answering system provided by the present invention. The contents not described in detail in the embodiments of the present invention belong to the prior art known to professional and technical personnel in the field. If no specific conditions are specified in the embodiments of the present invention, the conditions are carried out according to the conventional conditions in the field or the conditions recommended by the manufacturer. The reagents or instruments used in the embodiments of the present invention, if the manufacturer is not specified, are all conventional products that can be purchased commercially.
[0067] In the implementation scheme of the present invention, it is specifically a visual question answering system based on retrieval enhancement technology, which mainly combines information in an external knowledge base to improve the performance of the question answering system, thereby solving some limitations of the existing technology.
[0068] The purpose of the implementation scheme of the present invention is to provide a technical solution based on retrieval enhancement technology that can be used to answer relevant questions in the field of computer science, thereby making full use of image datasets and related external knowledge, combined with retrieval enhancement technology to answer simple questions related to computer science knowledge.
[0069] Specifically, a method for implementing a visual question answering system provided by an embodiment of the present invention can be implemented by the following technical solution, that is, the corresponding processing process may include:
[0070] (1) A specific knowledge base is constructed by combining manual and generative large models, and the knowledge base mainly includes subject knowledge related to the image data (i.e., knowledge points related to the image data) and is saved in the form of a knowledge graph;
[0071] For example, according to the image dataset used, the corresponding knowledge base may include the following contents: concepts, structures, characteristics, advantages and disadvantages, and applications of lists, linked lists, trees, binary trees, directed graphs, undirected graphs, stacks, queues, flow charts, network topologies, logic circuit diagrams, and deadlocks;
[0072] (2) using the constructed knowledge base to enhance the description of problems related to the image data;
[0073] Specifically, a retrieval enhancement method can be used to implement the corresponding enhancement processing, that is, the question text related to each image (i.e., picture) in the picture data set contained in the training data set is searched in the knowledge base to obtain the most relevant knowledge point, that is, to obtain the corresponding question enhanced text description data (i.e., knowledge point text) after retrieval enhancement;
[0074] (3) Generate a picture text description (or picture text description data) related to the picture data using the parameter-frozen GPT-4. Specifically, using the prompt engineering, the picture data and the text prompt are used as the input of GPT-4, and the GPT-4 model generates the picture text description related to the picture data according to the output constraints given by the text prompt.
[0075] The corresponding generation process based on the output constraints given by the prompt project can specifically include: encoding the image data into Base64 format and inputting it into the generation model, providing the output content requirements to the large model in the form of text, and obtaining the image text description that meets the constraints. The output result can then be used as the enhanced data of the model for subsequent use;
[0076] (4) The question-enhanced text description data after the retrieval enhancement (i.e., the knowledge point text obtained in the above process (2)), the image text description data (i.e., the image text description obtained in the above process (3)) and the original image data and its corresponding question text (i.e., the original data set used for training) are combined as the input data of the model to train the model, that is, to train and construct the corresponding question-answering system model, and to establish a question-answering system for processing question-answering tasks based on the question-answering system model.
[0077] After completing the above model training, you can also use the accuracy of the test data to evaluate the performance of the model to determine the processing effect of the corresponding question-answering system model on the visual question-answering task.
[0078] It can be seen from the technical solution provided by the present invention that the retrieval enhancement method is used to enhance the text data related to the picture. Compared with the traditional question-answering model, the learnable content of the model is more comprehensive and accurate, which can effectively improve the performance of the model.
[0079] The specific implementation methods of the present invention will be described in detail below in conjunction with the accompanying drawings. That is, the technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of the present invention.
[0080] An implementation method of a visual question answering system provided by an embodiment of the present invention is specifically implemented as follows: Figure 1 As shown, it can mainly include the following steps:
[0081] Step 11, constructing a knowledge base, wherein the knowledge base includes subject knowledge related to the image data, mainly text descriptions of knowledge points;
[0082] This step can be implemented by combining manual and large language models, which can ensure the integrity of the data to a certain extent, while also preventing the existence of excessive redundant information, thus ensuring the quality of the data as a whole;
[0083] Specifically, in order to solve the problem of scarce computer science training data sets, this step uses a generative large language model to assist in building a computer science knowledge base, thereby obtaining a knowledge base containing richer data. Different from the previous manual construction of data sets, using a large model to build a knowledge base will make the processing more efficient.
[0084] Furthermore, in order to avoid a gap between the quality of the data in the corresponding knowledge base and the quality of the manually constructed data set, it is necessary to process the data in the knowledge base to improve the quality of the data.
[0085] Step 12, using the data in the knowledge base to perform retrieval enhancement;
[0086] Using the knowledge base obtained in step 11, a retrieval method is used to retrieve the most relevant knowledge points for each question in the data set to facilitate understanding of the semantics of the question;
[0087] As an artificial intelligence technology, retrieval enhancement combines the advantages of retrieval models and generation models, aiming to improve the accuracy and context relevance of the generation model when generating text; however, there is more than one method for retrieval enhancement. According to the combination of retrieval and generation, it includes parallel retrieval enhancement, serial retrieval enhancement and hybrid retrieval enhancement; in order to ensure that knowledge highly relevant to the query is retrieved, the embodiment of the present invention can adopt a serial retrieval enhancement method. Specifically, in the retrieval enhancement process, it is necessary to focus on retrieving knowledge in the first stage, and then perform answer prediction in the second stage on the premise of ensuring the correctness of the information. The specific retrieval enhancement process may include:
[0088] First, the model searches for the input query or question and finds relevant information in the document;
[0089] Later, after the retrieval is complete and the external information to be used is determined, the answer prediction is performed based on the query and the retrieved information:
[0090] In the above-mentioned retrieval enhancement process, the corresponding retrieved information can play a great role in the performance of the model, so it is necessary to ensure that there is a high correlation between the retrieved information and the query in order to effectively improve the capabilities of the model.
[0091] Step 13, use the image data combined with the prompt project to perform data enhancement and obtain the text description of the image;
[0092] Taking the data structure in the field of education as an example, there are fewer relevant data sets, that is, the available data sets are small in size and contain a small amount of data, so data enhancement processing can be performed on the data they contain;
[0093] The biggest difference between the data that can be used in the embodiments of the present invention and the previous data is that the image content is simple but contains a lot of information; that is, the image in the data structure is highly abstract, consisting of geometric shapes such as rectangles and circles and logical symbols such as arrows and broken lines; it is very simple visually, but it contains rich semantic information; the relationship between objects is no longer limited to spatial relationships, but includes complex logical relationships; therefore, how to make the model understand the content of the image is a problem that needs to be solved;
[0094] Specifically, in order to obtain the data required in the embodiment of the present invention, on the one hand, more and more effective image features can be extracted through neural networks, and on the other hand, the corresponding text description of the image can also help the model better understand the image. Therefore, better linking the image and its text description will be able to obtain more comprehensive information than single-modal data, that is, more image data information; for this reason, the embodiment of the present invention uses a multimodal large model to describe and expand the content of the image. Since the ability of the large model to process long texts needs to be improved, when outputting, a prompt project will be used to limit its output length and format, and specify the description content that must be output.
[0095] Step 14: Use the attention mechanism to perform modal fusion on the enhanced data;
[0096] Whether it is the image features recorded in the image or the text features recorded in the text description, the information contained in the data of a single modality is always not as comprehensive as the information contained in the multimodal data. Therefore, in the embodiment of the present invention, it is necessary to obtain multimodal information to achieve multimodal information fusion; the multimodal information fusion refers to combining information from different senses or data sources to obtain a richer and more accurate representation than using any modal information alone;
[0097] Specifically, in an embodiment of the present invention, data of different modalities are first extracted through separate image feature extraction models and text feature extraction models, and then the attention mechanism is used to fuse the data of the two modalities in order to better obtain more comprehensive and detailed information.
[0098] Step 15, performing model training based on the data processed by modal fusion to obtain a corresponding question-answering system model to establish a question-answering system for processing question-answering tasks.
[0099] The above-mentioned embodiment of the present invention uses the method of retrieval enhancement to help the model acquire more knowledge, and adopts the method of independent input query and related knowledge to solve the problem of semantic interaction between the two. Compared with the traditional retrieval method, the model understands the input information more clearly. Moreover, it can be seen from the obtained test results that it has significantly improved the evaluation indicators.
[0100] To facilitate understanding of the present invention, the processing process provided by the above-mentioned embodiment of the present invention will be further described in detail below.
[0101] Reference Figure 2 As shown, the processing process provided by the embodiment of the present invention includes the processing of acquiring a knowledge base, preprocessing data, retrieval enhancement, data enhancement, building a model, and training a model. The specific implementation methods of each processing process will be described in detail below.
[0102] (1) Acquisition of knowledge base
[0103] A knowledge outline is constructed according to the categories of images in the data set. The categories of images in the data set include lists, linked lists, trees, binary trees, directed graphs, undirected graphs, stacks, queues, logic circuit diagrams, network topologies, flow charts, and deadlocks. The corresponding knowledge outline includes the concept, structure, characteristics, advantages and disadvantages, and applications of each category of images. The content of the knowledge outline is input into the generative large model to obtain the original knowledge base, which can be used as an external knowledge base for retrieval enhancement after being sorted (i.e., data preprocessing). The knowledge base contains subject knowledge related to the image data, mainly text description data of the knowledge.
[0104] (2) Data preprocessing
[0105] For the original data in the acquired knowledge base, in order to ensure the effect of the model, it is necessary to perform some preprocessing on the original data to improve the quality of the data. Specifically, the corresponding preprocessing may include:
[0106] (21) Remove redundant knowledge
[0107] Due to the particularity of image data, many data structures have some identical descriptions, such as structural descriptions, application scenarios, etc. These duplicate data need to be removed;
[0108] (22) Remove incorrect knowledge
[0109] Although the large model has outstanding capabilities, it still has some problems, such as hallucinations. Therefore, it is necessary to eliminate errors in the knowledge it outputs. The four structures of logic circuit diagrams, network topology structures, flow charts, and deadlocks are, in a sense, application scenarios of certain data structures. Therefore, the information contained in the outline does not exist, such as the advantages and disadvantages of deadlocks. However, the large model will generate some non-existent content based on the outline, so the corresponding knowledge needs to be removed.
[0110] (23) Processing the text content length
[0111] When the model processes long text, information loss may occur. Therefore, overly long text descriptions will not only fail to improve the performance of the model, but will reduce its capabilities. In addition, when performing retrieval, overly long texts will also contain redundant information for images. Therefore, the length of the output text content needs to be constrained to meet application needs.
[0112] (24) Storage of knowledge base;
[0113] The acquired unstructured knowledge base can be saved in a csv file, and neo4j can be used to build a knowledge graph related to the data structure; the name of the data structure is used as the central node, and the related knowledge in it is used as the surrounding nodes to form a simple data structure knowledge graph.
[0114] The text in the knowledge base obtained after the above processing is specific and concise, which is more conducive to model use and understanding learning.
[0115] (3) Retrieval enhancement
[0116] After obtaining the required data through the above preprocessing, the preprocessed data (i.e. the knowledge base obtained through the above construction, which carries the text information of the corresponding image or the text data of the image) can be subjected to retrieval enhancement processing to obtain the corresponding question enhanced text description data, i.e. the corresponding knowledge point text description;
[0117] In the retrieval enhancement process, a serial retrieval enhancement can be specifically adopted. Its biggest advantage is that there is a clear division of steps, and the quality of retrieval and generation can be evaluated separately, which is helpful for debugging and optimization.
[0118] The corresponding retrieval enhancement processing may include:
[0119] In the knowledge base, a question is input to search and obtain relevant search result information, wherein the search result information includes: a text description of a knowledge point in the knowledge base;
[0120] Determining available search result information in the knowledge base to obtain search enhancement data; wherein the process of determining available search result information includes: processing the knowledge base using a model capable of capturing context information and semantic relationships in a language to obtain available search result information;
[0121] Specifically, based on the Sentence Transformers (a Python library for calculating embedding vectors of sentences, texts, and images) open source library, the questions in the image text description corresponding to each image (i.e., picture) can be extracted and put into a document, and encoded separately from the external knowledge base. Then, the semantic similarity between each of the questions and all knowledge points in the knowledge base (i.e., subject knowledge related to the picture in the knowledge base) is calculated, and finally the most similar knowledge point is output as the retrieval enhanced processed data corresponding to the question, i.e., the corresponding question enhanced text description data.
[0122] Since Sentence Transformers uses pre-trained Transformer models, and these models have been trained on large-scale corpora and can capture contextual information and semantic relationships in language, you need to load a pre-trained model when encoding text data. You can choose different models, such as all-MiniLM-L6-v2, paraphrase-Mini-LM-L3-v2, etc.
[0123] In the embodiment of the present invention, all-MiniLM-L6-v2 can be used as a text encoder to encode the questions and knowledge points in the text description of the picture. After encoding all questions and knowledge points at the sentence level, the cosine similarity between each question and all knowledge points is calculated as follows:
[0124] Q E =Encoder(Q)
[0125] K E =Encoder(K)
[0126]
[0127] Among them, Q E and K E They represent the embedding of all questions and the embedding of all knowledge points respectively; n and m represent the number of questions and the number of knowledge points respectively; q i Indicates Q E The i-th question in k j K E After calculating the relevance score of each question with all the knowledge, select the K most relevant knowledge points before output as external knowledge to help the model understand the relevant issues of the image.
[0128] Before outputting knowledge points, you can set a similarity threshold to control the specific content of the output knowledge points to prevent too many knowledge points from bringing noise data or too few knowledge points from insufficient information, so as to judge the impact of knowledge points on the model.
[0129] (4) Data enhancement
[0130] The data enhancement process of the embodiment of the present invention mainly uses a multimodal large model (such as a GPT-4 model, etc.) to generate text descriptions related to the image (i.e., generate image text description data); specifically, the image data and related text prompts (text prompts constructed by oneself according to different categories of images) can be used as inputs of the multimodal large model to enable the multimodal large model to output the content of the image text description related to the image data; when generating the corresponding content, different types of image data need to use different prompts; in the implementation process, the image needs to be first converted into a base64 encoded format, and then input into the multimodal large model in combination with the corresponding text prompt, and the obtained output is as follows:
[0131] The output template can be: picture:{name}; describe:{xxx}; that is, it includes the picture name and the corresponding picture text description;
[0132] The output content requirements include: for different data structure types of images, the output content requirements are also different; specifically, if it is a list or linked list, it is necessary to describe the number and order of elements; if it is a tree structure, it is necessary to describe the number of nodes in the tree, the degree of the tree, the number of layers and other basic information; if it is a graph structure, it is necessary to describe the number of nodes and edges in the graph, the degree of the nodes and other basic information; if it is a stack, it is necessary to describe the number of elements, the top and bottom elements of the stack; if it is a queue, it is necessary to describe the number of elements, the head and tail elements;
[0133] The output text length requirements include: the maximum output text length does not exceed a predetermined length, such as 100 tokens (i.e., words);
[0134] The output format requirements include: outputting the required content in JSON format;
[0135] After the multimodal large model generates a description related to the image, the generated content (i.e., the description related to the image, or the descriptive text or image text description data of the corresponding image) can be input into the subsequent model as enhanced data to help the corresponding model understand the content in the image.
[0136] (5) Model building
[0137] (51) Input layer
[0138] The input layer is used to extract features from the data (knowledge points) obtained after the retrieval enhancement, the image-related description obtained in the data enhancement step, and the original training data as input data, and the specific processing process of the input data may include:
[0139] (511) For the text data in the input data, a pre-trained BERT (Bidirectional Encoder Representation) model is used to extract text features. BERT is essentially a multi-layer bidirectional transformer encoder (i.e., the encoder of the Transformer model); in order to obtain global information, a multi-head self-attention layer is first used to convert each position in the input sequence into a weighted sum of the input layer; specifically, for the i-th head attention layer, the input X∈R d×N The conversion is based on the dot product attention mechanism (where d represents the embedding dimension of the feature and N represents the maximum length of the sequence), namely:
[0140]
[0141] in, They are learnable parameters of queries, keys and values.
[0142] Afterwards, the outputs of the m multi-head self-attention layers are concatenated together and linearly transformed:
[0143] MATT(X)=W m [ATT1(X),...,ATT m (X)] T ;
[0144] Among them, W m is a learnable parameter. Based on the output of the self-attention layer, BERT adds a residual layer between the input and output, namely:
[0145]
[0146] Afterwards, a standard feed-forward network with GeLU as activation function and another residual connection with layer norm is stacked on top to generate the output of the BERT layer:
[0147]
[0148] The sentence S input to the model consists of two parts: the first part: the question in the dataset; the second part: the text description of the image + the retrieved knowledge points (i.e. the retrieval results); in form, it can be expressed as X = (X1, X2, ...X N ) represents the transformed input sequence, where X i ∈R dIt is the sum of word embeddings, segment embeddings and position embeddings, d is the embedding dimension of the feature, and N is the maximum length of the sequence. The two parts are encoded using the pre-trained BERT model to obtain the hidden information representation H of the problem. T and the corresponding image text description and the hidden information representation H of the retrieved knowledge D ;
[0149] (512) For the associated image in the input data, i.e., the picture data, the features are extracted using the image recognition network ResNet-152; first, the image size is adjusted to 224×224 pixels, and then the output of the last convolutional layer of the ResNet-152 network is obtained:
[0150] ResNet(I)={r j |r j ∈R 2048 ,j=1,2,...,49};
[0151] It divides the original image into 7×7 regions, each of which is represented by a 2048-dimensional vector r j express.
[0152] Next, a linear transformation function is used to project the visual features into the same feature space as the text description of the image:
[0153] H v =W v ResNet(I),W v ∈R d×2048 ;
[0154] Where Wv∈R d×2048 is a learnable parameter.
[0155] (52) Modal Fusion Layer
[0156] The modality fusion layer is used to perform multimodal fusion processing on the features extracted and processed in the above-mentioned processing processes (511) and (512). The specific implementation process may include:
[0157] After the above processing (51) extracts the unimodal features of the image data and the text data related to the image, in order to obtain more comprehensive information, the cross attention block is used to fuse the information of the two modes. and Represents the data of two modes α and β respectively, where T represents the sequence length and d represents the feature dimension. Queries can be defined using the desired mode of interest, specifically: Use another mode to define Keys and Values, specifically in, are three weight matrices corresponding to the query matrix, key matrix, and value matrix respectively; the potential adjustment from modality β to modality α is expressed as cross-modal attention, namely:
[0158]
[0159] Among them, Y α The length and Q α equal; then add a residual link, and finally, add another position feed-forward sublayer to form a cross-modal attention block Cross-ATT.
[0160] Specifically, based on the above cross-modal attention block Cross-ATT, H D , H V Input the cross-attention module to obtain a visual representation H that enhances data awareness D→V , and then H D , H T Input the cross attention module to obtain enhanced data-aware text representation H D→T , then H D→V and H D→T After connecting together, we get the multimodal output representation H, namely:
[0161] H D→V = Cross-ATT(H D ,H V )
[0162] H D→T = Cross-ATT(H D ,H T )
[0163] H=H D→T +H D→V
[0164] (53) Output layer
[0165] The output layer is used to perform self-attention calculation, which may specifically include: for the multimodal information representation H obtained above, H can be input into the multimodal encoder for self-attention calculation, and the self-attention calculation process includes: inputting the multimodal data representation H into the self-attention layer ATT and the residual layer output to obtain Afterwards Input a standard feedforward network with GELU as activation function and residual layer to obtain H', and finally get the output calculated by the multimodal encoder: the multimodal information represents the first tag H of H' 0, input it into a softmax (normalized exponential function) layer for output, and obtain:
[0166]
[0167] H'=LN(H+MLP(H))
[0168]
[0169] Among them, W M ∈R d×2 represents the weight matrix, and y represents the true answer label of the question.
[0170] (6) Model training
[0171] This step mainly trains all the parameters in the model established in the previous step, using the cross-entropy loss function as the training target. During the model training process, the goal is to minimize the loss function. Through the back-propagation algorithm, the loss calculated by the cross-entropy loss function will be back-propagated along the model's computational graph, and the model parameters will be updated according to the gradient of the loss to the model parameters.
[0172]
[0173] In the formula, H 0 represents the answer predicted by the model, y i Indicates the true answer label of the question.
[0174] Through the above model training process, a trained visual question answering system model can be obtained. Based on the trained question answering system model, a question answering system for processing question answering tasks is established to process the visual question answering tasks.
[0175] After the above visual question answering system model is trained, the training results of the model can also be tested using a test set that the model has not learned. Specifically, the test data is input into the trained model to obtain the output result, which includes the question ID and the corresponding answer. Then the correct answer and the model predicted answer are extracted and evaluated using the accuracy, that is:
[0176]
[0177] Based on the above accuracy calculation results, we can know the performance of the trained visual question answering system model.
[0178] In summary, the embodiment of the present invention uses a parameter-frozen generative large model to perform corresponding enhancement processing on the text description of the picture, and at the same time builds a detailed and comprehensive knowledge base, and uses a search enhancement method to search the relevant external knowledge base to obtain the knowledge text description related to the picture. Then, by using data from multiple modalities to capture the deep semantic relationship between pictures, related questions and answers, the model can learn professional subject knowledge, improve the teaching efficiency of educators and learners in practical applications, and help build a better smart education project.
[0179] In summary, the embodiments of the present invention can make full use of questions, images and knowledge texts to solve the question-answering system problems of specific subjects in the field of education; at the same time, the retrieval enhancement method can be combined with a large model to perform data enhancement, and a cross-modal attention mechanism is used to solve the problem of multimodal data fusion, so that more abundant and comprehensive information can be used in the process of training the model. Therefore, compared with the traditional model training method, the embodiments of the present invention use more information, and the predicted results are significantly improved in multiple evaluation indicators.
[0180] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with the technical field within the technical scope disclosed in the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims. The information disclosed in the background technology section of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or in any form that the information constitutes prior art known to those skilled in the art.
Claims
1. A method for implementing a visual question answering system, characterized in that: include: Acquire a knowledge base, wherein the knowledge base includes knowledge points related to the image data; Searching the knowledge base based on the question text related to each image contained in the original training data set to obtain question enhanced text description data after retrieval enhancement; The image data set in the original training data set is combined with the corresponding text prompts through the multimodal large model to perform data enhancement, and the enhanced data output by the multimodal large model is obtained as the image text description data; Modal fusion is performed based on the question-enhanced text description data after retrieval enhancement, the enhanced data output by the multimodal large model and the original training data set, and the constructed question-answering system model is trained based on the modal fusion data, and a visual question-answering system for processing question-answering tasks is established based on the trained question-answering system model.
2. The method according to claim 1, characterized in that The knowledge base includes: text data obtained by manually constructing a knowledge outline and expanding a generative large model, and before performing the retrieval enhancement process, it also includes: removing redundant data in the knowledge base.
3. The method according to claim 1, characterized in that: The process of obtaining enhanced data further includes: The output length and format of the enhanced data are restricted, and the content and format requirements of the output picture text description are specified.
4. The method according to claim 1, characterized in that The retrieval enhancement adopts a serial retrieval enhancement processing method, and the corresponding processing process includes: In the knowledge base, questions related to the image data are input for retrieval to obtain retrieval result information, the retrieval result information includes: knowledge points related to the image data; The available search result information is determined in the knowledge base as the question enhanced text description data after the search enhancement.
5. The method according to claim 4, characterized in that The process of determining the available search result information includes: The retrieval results are processed using a model that can capture contextual information and semantic relationships in the language to obtain usable retrieval result information.
6. The method according to claim 5, characterized in that The model capable of capturing contextual information and semantic relations in a language includes an all-MiniLM-L6-v2 model, and the step of obtaining available search result information includes: The text encoder of the all-MiniLM-L6-v2 model is used to encode the question text and the knowledge base, and the cosine similarity between each question and all knowledge points in the knowledge base is calculated, that is: Among them, Q E and K E denote the embedding of all questions and the embedding of all answers respectively; n and m denote the number of questions and the number of answers respectively; q i Indicates Q E The i-th question in k j K E The jth piece of knowledge in An optimal predetermined number of retrieval result information is selected as the available retrieval result information according to the calculated cosine similarity.
7. The method according to any one of claims 1 to 6, characterized in that: The constructed question-answering system model includes an input layer, a modality fusion layer and an output layer, wherein: The input layer is used to extract features and includes: Use the pre-trained bidirectional encoder representation BERT model to encode the question text and image text description in the input data and the retrieved knowledge points to obtain the corresponding hidden information representation H T and the hidden information representation of image text description and retrieval knowledge H D ; The visual features of the associated picture data are extracted through an image recognition network, and the linear transformation function is used to project the visual features into the same feature space as the picture text description to obtain the corresponding visual information representation H V ; The modality fusion layer is used to perform multimodal fusion processing on the extracted features, and includes: The above H D , H V Input the cross-attention module to obtain a visual representation H that enhances data awareness D→V , and then H D , H T Input the cross attention module to obtain enhanced data-aware text representation H D→T , then H D→V and H D→T After connecting together, we get the multimodal information representation H, namely: H D→V =Cross-ATT(H D ,H V ) H D→T =Cross-ATT(H D ,H T ); H=H D→T +H D→V Among them, Cross-ATT is the processing performed by the cross-attention module; The output layer is used to perform self-attention calculation and includes: The obtained multimodal information representation H is input into the multimodal encoder for self-attention calculation. The multimodal data representation H is input into the self-attention layer ATT and the residual layer output to obtain Afterwards Input a standard feedforward network with GELU as activation function and residual layer to obtain H', and finally get the output calculated by the multimodal encoder: the multimodal information represents the first tag H of H' 0 , input it into a normalized exponential function softmax layer for output: Among them, W M ∈R d×2 represents the weight matrix, and y represents the true answer label of the question.
8. The method according to claim 7, characterized in that The visual representation H V Obtained by the following formula: H v =W v ResNet(I),W v ∈R d×2048 ; Among them, W v ∈R d×2048 is a learnable parameter, and the output of the convolutional layer is calculated as: ResNet(I)={r j |r j ∈R 2048 ,j=1,2,...,49}; In the formula, when the original image is divided into 7×7 regions, each region is represented by a 2048-dimensional vector r j Indicates that the pixels of the image are 224×224.
9. The method according to claim 7, characterized in that: The implementation of the cross attention module includes: use and Represents the data of two modes α and β respectively, where T represents the sequence length and d represents the feature dimension; and the query Querys is defined using the desired mode of interest. Another mode defines keys and values, specifically in, There are three weight matrices, corresponding to the query matrix, key matrix and value matrix respectively; Then, the potential adjustment from modality β to modality α is expressed as cross-modal attention, namely: Among them, Y α The length and Q α equal; Based on the above formula, a residual link is added, and another position feed-forward sublayer is added to form a cross-modal attention block Cross-ATT.
10. The method according to claim 7, characterized in that The process of training the constructed question answering system model includes a process of updating model parameters, and the process includes: Use the cross-loss entropy function as the training objective, minimize the loss function, and update the model parameters according to the gradient of the loss function to the model parameters through the back-propagation algorithm: Among them, H 0 represents the answer predicted by the model, y i Indicates the true answer label of the question.
Citation Information
Cited By
Large model training data generation method and device, medium and electronic equipment
CN120913010A