Multi-modal perception problem generation method and system based on large language model, and medium

By extracting image features and text interactions, a comprehensive representation of multimodal semantics is obtained and converted into input representations that can be understood by large language models, solving the problem of large language models understanding complex multimodal inputs in multimodal problem generation tasks, and achieving more efficient multimodal problem generation.

CN119942300APending Publication Date: 2025-05-06XI AN JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510009573.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize large language models to understand complex multimodal inputs, resulting in the inability to fully utilize model performance in multimodal problem generation tasks.

Method used

By extracting image features and interacting with input text, a multimodal semantic comprehensive representation is obtained and converted into input representations that are understandable by large language models to guide the model generation problem.

Benefits of technology

It improves the understanding of multimodal inputs by large language models, enables it to handle more complex multimodal inputs and generates effective problems, making full use of the performance of large language models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942300A_ABST
    Figure CN119942300A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal perception problem generation method and system based on a large language model and a medium, and belongs to the technical field of computer semantic analysis, and the method comprises the steps: extracting image features, and extracting features most related to an input text from the image features; converting a text background in the input content into word embedding, and interacting with features most relevant to the input text extracted from the image features to obtain image representation most relevant to the text content and text representation most relevant to the image content; performing semantic alignment on the image representation most relevant to the text content and the text representation most relevant to the image content to obtain a multi-modal semantic comprehensive representation of the visual and text information, and converting the multi-modal semantic comprehensive representation into an input representation which can be understood by a large language model; a big language model is guided to generate a problem based on input characterization which can be understood by the big language model. According to the method, the large language model can be fully utilized to understand more complex multi-modal input and generate an effective problem.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of computer semantic analysis, and specifically relates to a method, system and medium for generating multimodal perception questions based on a large language model. Background Art

[0002] With the development of large language models, modern intelligent agents such as chatbots and dialogue systems have achieved human-like conversation capabilities. As the study of vision and language fusion deepens, researchers have begun to turn more to visual dialogue systems. Such systems are designed to understand and interpret multimodal scenarios and be able to interact with users about these scenarios. They not only need to answer questions accurately, but also be able to recognize their own knowledge blind spots and actively seek additional information by asking questions about multimodal content. Therefore, multimodal question generation has become an emerging research direction. Although the demand for multimodal question generation continues to grow, related research is still in its early stages. Some researchers have made some attempts on simple multimodal inputs, for example, using short text information such as knowledge triples in text form or answers to guide images to generate questions, and preliminarily achieving the task of generating questions from multimodal inputs. However, the focus of these studies is still on images, which is far from the complexity of actual multimodal scenarios. The reason is that these methods can only understand short texts and images due to the limitation of the number of parameters of the model, and cannot encode more complex multimodal inputs.

[0003] With the improvement of computing power, the powerful understanding ability of large language models has become an effective means to understand complex multimodal inputs. Researching how to apply large language models to multimodal question generation tasks has become a mainstream trend. However, large language models can only understand text content, and how to make large language models understand multimodal inputs is a difficult problem. Some researchers have tried to use additional expert models to translate visual images into text subtitles, thereby using large language models to achieve multimodal question generation tasks. The definition of multimodal question generation tasks is that the model can receive complex multimodal inputs containing visual images and long background texts, and ask questions based on these multimodal information. In order to solve the cross-modal semantic gap, some researchers have introduced image description generation models as expert models to generate text descriptions of image content, that is, converting visual content into text, and then inputting the image description and background text into the large language model, so that the text-centric large language model introduces visual understanding capabilities to generate questions. Although this expert model-based strategy enables the large language model to generate questions based on multimodal input to a certain extent, when the expert model translates images, it ignores the interaction between visual semantics and textual semantics, making it difficult to fully understand the key content in the modal input and unable to fully utilize the performance of the large language model, which has become the bottleneck of such methods. Summary of the invention

[0004] The purpose of the present invention is to address the problems in the above-mentioned prior art and provide a method, system and medium for generating multimodal perception questions based on a large language model, so as to improve the large language model's ability to understand multimodal inputs and make more effective use of the large language model to enable it to understand more complex multimodal inputs and generate effective questions.

[0005] In order to achieve the above object, the present invention has the following technical solutions:

[0006] In a first aspect, a method for generating multimodal perceptual questions based on a large language model is provided, comprising:

[0007] Extract image features, and extract the features most relevant to the input text from the image features;

[0008] The text background in the input content is converted into word embedding, and interacted with the features most relevant to the input text extracted from the image features to obtain the image representation most relevant to the text content and the text representation most relevant to the image content;

[0009] Semantically align the image representation that is most relevant to the text content and the text representation that is most relevant to the image content to obtain a multimodal semantic comprehensive representation of visual and text information, and convert it into an input representation that can be understood by the large language model;

[0010] Guide the large language model to generate problems based on the input representation that the large language model can understand.

[0011] As a preferred solution, the step of extracting image features is performed using a ViT visual encoder, and the image is input into the ViT visual encoder to obtain an embedding representation matrix Where N is the length of the image feature vector sequence, which is determined by the number of image blocks, and d is the dimension of each image block vector. For a given image V, the image features are extracted by the ViT visual encoder as follows:

[0012] I v =ViT(V;θ)

[0013] Where θ is the parameter of the ViT visual encoder.

[0014] As a preferred solution, the features most relevant to the input text are extracted from the image features by a visual sensor;

[0015] The image features are extracted through the ViT visual encoder, and the following embedded representation is obtained:

[0016] I v =[a 1 ,a 2 ,...,an ]

[0017] In the formula, a i is the vector extracted by the visual encoder for each block of the image;

[0018] Embedded Representation I v Perform self-attention calculation, for each vector a i ∈I v , calculate its relationship with any other vector a as follows j The alignment score e is:

[0019] e=h T tanh(W k a j +W q a i )

[0020] Where W q , W k are all trainable parameter matrices, h T is the weight vector;

[0021] The vector a is converted into i Normalize the alignment scores with each other vector to get the probability vector p i :

[0022]

[0023] Vector a i After a layer of self-attention operation, we get The calculation expression is as follows:

[0024]

[0025] Where W v is a trainable weight matrix;

[0026] For I v Any eigenvector a in i All are obtained through the self-attention layer From this we get

[0027]

[0028] As a preferred solution, through the cross-attention layer, the visual perceptron will Each element in I is represented by text information t Attention calculation is performed on the related elements in the text to capture the semantic association between the two; text information representation I t For a vector sequence, the expression is as follows:

[0029] I t =[x 1 ,x 2 ,...,x n ]

[0030] In the formula, x i Represents the embedding of each word in the sentence, where the sentence contains n words in total;

[0031] Press the formula to I t and Any vector in Calculate the attention score matrix score:

[0032]

[0033] Where W s is a trainable weight matrix;

[0034] Normalize the attention scores through the softmax function to get the attention weight matrix:

[0035] w i =softmax(score)

[0036] Using the original image region representation vector Perform attention weighting to obtain the intermediate state vector c k , after a layer of feedforward neural network calculation, the final image region representation vector is obtained:

[0037]

[0038] for Any feature vector in All are calculated through the self-attention layer From this we get

[0039]

[0040] As a preferred solution, the step of converting the text background in the input content into word embedding includes:

[0041] The input background text T is divided into sentences, and the sentences are concatenated as follows:

[0042] Seq=[<cls>,t 1 ,<sep>,t 2 ,…, <sep>,t n ]

[0043] Where, t i represents the i-th sentence in the text, <cls> represents the start marker, <sep> represents the concatenation marker, Seq represents the sentence sequence after the sentence segmentation and concatenation markers, and Seq is mapped to a sentence-level vector sequence as follows:

[0044] E text =BERT(Seq)

[0045] In the formula, BERT is used to map each sentence in Seq into a corresponding text embedding vector sequence, and all text embedding vector sequences are concatenated to obtain E text ; After embedding the text, the position information of the sentence is encoded as follows:

[0046] P text =pos(Seq)×W p

[0047] In the formula, the position information of the sentence in the text segment is obtained through the pos function, W p is a learnable weight matrix;

[0048] P text With E text Sum and get the final text representation I t :

[0049] I t =P text +E text

[0050] The text representation obtained by embedding the input text Where M is the length of the text representation and d is the dimension of each embedding vector.

[0051] As a preferred solution, the text representation I is obtained by receiving the input text through text perception and embedding the text t , using self-attention mechanism to capture text representation I t The key semantics in the image are represented by cross-attention layer and image embedding. v Interact to extract the text representation most relevant to the image content; finally, the feedforward neural network layer outputs the text representation most relevant to the image content

[0052] As a preferred solution, the semantic alignment of the image representation most relevant to the text content and the text representation most relevant to the image content includes:

[0053] Embedding the image into a representation matrix and the text representation of the input text As input, a set of condensed fixed-length context representations is generated. In the formula, N is a customizable hyperparameter determined by the number of blocks specified in the image encoder, depending on the feature richness of the image, and M is a dynamic value determined by the length of the input text. A fixed number size is set to The learnable embedding H Text and H Img , where L satisfies the following formula:

[0054] L<M+N

[0055] For text, we can learn to embed H Text , the semantic representation after condensed alignment is obtained as follows

[0056]

[0057] For the visual learnable embedding H Img , the semantic representation after condensed alignment is obtained as follows

[0058]

[0059] Will and Connect them together to obtain a multimodal semantic comprehensive representation of visual and textual information

[0060]

[0061] As a preferred solution, the multimodal semantics of visual and textual information is comprehensively represented After a linear fully connected layer, it is converted into an input representation that can be understood by the large language model as follows:

[0062]

[0063] Where W and b are learnable matrices and vectors, respectively. It is a sequence of multimodal semantic representations that is input into a large language model;

[0064] In the step of guiding the large language model to generate questions based on the input representation that the large language model can understand, the text tag of the prompt instruction sentence is embedded in Add to multimodal semantic representation sequence At the beginning, beam search is used to select the best generated output; Vicuna-7B is used as the core of the question generation decoder; the expression of question generation decoding is as follows:

[0065]

[0066] In a second aspect, a multimodal perceptual question generation system based on a large language model is provided, comprising:

[0067] An image feature extraction module is used to extract image features and extract features most relevant to the input text from the image features;

[0068] A feature interaction module is used to convert the text background in the input content into word embedding, and interact with the features most relevant to the input text extracted from the image features to obtain the image representation most relevant to the text content and the text representation most relevant to the image content;

[0069] The semantic alignment module is used to semantically align the image representation that is most relevant to the text content and the text representation that is most relevant to the image content, obtain a multimodal semantic comprehensive representation of visual and text information, and convert it into an input representation that can be understood by the large language model;

[0070] The question generation module is used to guide the large language model to generate questions based on the input representation that the large language model can understand.

[0071] According to a third aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for generating multimodal perception questions based on a large language model is implemented.

[0072] Compared with the prior art, the present invention has at least the following beneficial effects:

[0073] By exploring the associations between different modal inputs and strengthening the interaction of semantic information between different modalities, the key semantic representations in each modality can be obtained. At the same time, in order to bridge the semantic gap between different modalities, the multimodal representation is aligned to the text representation space of the large language model to obtain a comprehensive multimodal semantic representation of visual and textual information, and converted into an input representation that can be understood by the large language model, thereby successfully utilizing the powerful language understanding ability of the large language model to generate questions. The present invention can enhance the large language model's ability to understand multimodal inputs, thereby making more full use of the performance of the large language model in question generation tasks, enabling the large language model to understand more complex multimodal inputs and generate effective questions.

[0074] Furthermore, the multimodal perception question generation method based on the large language model of the present invention encodes the semantic features of the visual image through the pre-trained visual encoder, and uses the word embedding technology to encode the semantic features of the input text, and uses the attention-based perception module for interaction to obtain accurate text embedding representation and image embedding representation. Combining the visual perception module and the text perception module, and then further optimizing through the semantic alignment module, using attention weighting and splicing operations, a multimodal semantic comprehensive representation of visual and text information is obtained, which provides a solid semantic foundation for subsequent cross-modal semantic alignment training.

[0075] Furthermore, the present invention proposes a question generation decoder based on a two-stage training strategy for a multimodal perception question generation method based on a large language model, which semantically aligns the rich embedded representation obtained by the multimodal perception encoder with the large language model, and uses a linear layer alignment dimension to enable the embedded representation extracted from the multimodal encoder to be seamlessly connected to the input part of the large language model. The first stage of the two-stage training strategy uses an image subtitle dataset to align cross-modal semantic representations, and the second stage of training embeds the instructions guiding question generation into the front end of the sequence, which together with the multimodal semantic representation serve as soft prompt inputs for the large language model, thereby guiding the large language model to generate questions. This fully utilizes the powerful understanding ability of the large language model, enabling the model to understand more complex multimodal inputs and generate effective questions. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0077] Figure 1 The architecture diagram of the multimodal perception question generation method based on a large language model in an embodiment of the present invention;

[0078] Figure 2 A schematic diagram of the structure of the attention module according to an embodiment of the present invention;

[0079] Figure 3 A schematic diagram of the structure of a semantic alignment module according to an embodiment of the present invention;

[0080] Figure 4 Schematic diagram of a two-stage training strategy according to an embodiment of the present invention. DETAILED DESCRIPTION

[0081] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, ordinary technicians in this field can also obtain other embodiments without making creative work.

[0082] The embodiment of the present invention proposes a method for generating multimodal perception questions based on a large language model, comprising the following steps:

[0083] Extract image features, and extract the features most relevant to the input text from the image features;

[0084] The text background in the input content is converted into word embedding, and interacted with the features most relevant to the input text extracted from the image features to obtain the image representation most relevant to the text content and the text representation most relevant to the image content;

[0085] Semantically align the image representation that is most relevant to the text content and the text representation that is most relevant to the image content to obtain a multimodal semantic comprehensive representation of visual and text information, and convert it into an input representation that can be understood by the large language model;

[0086] Guide the large language model to generate problems based on the input representation that the large language model can understand.

[0087] The method of the embodiment of the present invention can promote the understanding of multimodal content by the large language model, make full use of the powerful semantic understanding and generation capabilities of the large model, generate better questions, and provide auxiliary support for question demanders represented by teachers.

[0088] See also Figure 1 In one possible implementation, the multimodal perception question generation method based on a large language model in an embodiment of the present invention is based on an encoder-decoder architecture and includes two main modules: a multimodal perception encoder and a question generation decoder module. The multimodal perception question generation method based on a large language model in an embodiment of the present invention includes the following two processes:

[0089] Step 1: Design a multimodal perceptual encoder, which includes 2 modules:

[0090] 1) Visual perception module

[0091] In the unimodal question generation task, the encoder mainly focuses on understanding the input content and providing a semantic basis for subsequent question generation. For multimodal question generation tasks involving images and texts, the first thing to do is to understand the input content. Unlike the question generation task with unimodal input, this process not only requires understanding the content of each modality separately, but also involves the interaction between different modalities to deeply explore the semantic connection between them and extract their respective key contents. In order to deeply understand the semantics of the visual modality, the present invention uses a pre-trained visual encoder to extract image features, preliminarily understand the image and obtain the original representation of the image. In order to further strengthen the interaction between different modal contents, the present invention proposes a novel design: a visual perceptron. The visual perceptron receives the visual features extracted by the visual encoder and promotes the interaction between the visual features and the text context features by introducing an attention module. Doing so not only deepens the comprehensive understanding of visual and textual information, but also can more effectively extract feature embeddings that are highly relevant to the text from the visual content.

[0092] (1) Visual Encoder

[0093] In order to effectively process multimodal input visual images, the embodiment of the present invention uses the open source pre-trained model ViT as a visual encoder to extract image features. Unlike traditional visual tasks, the present invention requires not only an understanding of the features of the image itself, but also the combination of these visual information with relevant background text. The ViT visual encoder is specially trained for image and text alignment tasks. Therefore, the image features extracted by it are initially similar to text features in distribution. This feature enables the visual information encoded by ViT to interact more easily with text features and facilitates subsequent alignment with the input of large language models. The embedded representation matrix obtained after the image passes through ViT Where N is the length of the image feature vector sequence, which is determined by the number of image blocks, and d is the dimension of each specified image block vector. For a given image V, the process of extracting image features through ViT can be expressed as follows:

[0094] I v =ViT(V;θ)

[0095] Where θ is the parameter of the pre-trained ViT model.

[0096] (2) Visual sensor

[0097] After the image passes through the visual encoder, it obtains an initial representation, but this initial representation lacks relevance to the text content, so a visual perceptron module is needed. The main purpose of the visual perceptron module is to extract the features that are most relevant to the text input. The core module of the visual perceptron is an attention module based on two-stage attention. It starts with a self-attention layer to capture the intrinsic visual relationship, and then uses cross-attention to interact with the text tag embedding from the text encoder, allowing the model to enrich its understanding by integrating information from the text domain. Finally, the output vector sequence after interaction is output through the feedforward layer to obtain the final visual semantic representation. Figure 2 As shown, the attention module structure of the embodiment of the present invention has three layers of superimposed attention modules in the visual sensor, text sensor and semantic alignment module respectively. The attention module structures in the three sensors are similar, but the input and output are not the same and the parameters are not shared.

[0098] The image is embedded in the visual encoder and represented as I v :

[0099] I v =[a 1 ,a 2 ,...,a n ]

[0100] In the formula, a i is the vector feature extracted by the visual encoder for each block of the image. In order to further extract the key visual information in the image, the first stage of attention first needs to v Perform self-attention calculation. For each vector a i ∈I v , calculate its difference with any other vector a j The alignment score e is calculated as follows:

[0101] e=h T tanh(W k a j +W q a i )

[0102] Where W q , W k is the trainable parameter matrix, h T is the weight vector.

[0103] Next, vector a i Normalize the alignment scores with each other vector to get the probability vector p i :

[0104]

[0105] Finally, according to the formula, vector a i After a layer of self-attention, we finally get Compared to a i , The feature information of the represented image area is more focused.

[0106]

[0107] Where W v is the trainable weight matrix.

[0108] For I v Any eigenvector a in i , all need to go through this layer of self-attention to get Based on this, we can get

[0109]

[0110] Compared to the initial output I of the visual encoder v , The key information in the image is highlighted, making it easier to interact with the text embedding in the next step. The text features I encoded by the text perceptron t Interact based on the cross-attention layer. Through the cross-attention layer, the visual perceptron can Each element in I t The attention calculation is performed on the related elements in the image, so as to better capture the semantic relationship between them. This helps to ensure the correct information transfer in the subsequent problem tasks. t is also a vector sequence, where x i Represents the embedding of each word in the sentence, and the sentence contains n words in total.

[0111] I t =[x 1 ,x 2 ,...,x n ]

[0112] Through I t and Any vector in Calculate the attention score matrix score:

[0113]

[0114] Where W s is a trainable weight matrix.

[0115] Then, the attention scores are normalized by the softmax function to obtain the attention weight matrix:

[0116] w i =softmax(score)

[0117] Using the original image region representation vector Perform attention weighting to obtain the intermediate state vector c k , and finally a layer of feedforward neural network is used to calculate the final image region representation vector:

[0118]

[0119] for Any feature vector in All of these must be calculated through this self-attention layer Based on getting

[0120]

[0121] In the above process, I v After the self-attention layer focuses on its own important information, a cross-attention layer is used to represent the text information. t Interact and finally output the final representation through the feedforward neural network This is the complete process of the attention module in the model. The overall representation of this process is shown in the following formula, where θ represents all the parameters of the attention module.

[0122]

[0123] Visual representation of the output of the visual encoder I v The visual perception device interacts with the background text representation to obtain the visual representation most relevant to the text content. But it is essentially a visual representation encoded by ViT, so in order to align it to the text representation space, the final output of the perceptron It is further transformed as the initial input vector to the semantic alignment module.

[0124] 2) Text perception module

[0125] The text perception module is designed to enable the model to understand text information. Specifically, it first converts the text context in the input content into word embeddings so that the model can process text data. In order to understand the visual information in the input content at the same time, the module interacts with the visual representation of the image through a dedicated text perception. This means that the text perception module not only handles the understanding of text data, but also involves cross-modal interaction with visual data. Therefore, from a broader perspective, the text perception module and the visual perception module present a symmetrical relationship: the text representation processed by the text perception module is used to interact with the image representation in the visual perception module to obtain the image representation most relevant to the text content, and the image representation processed by the visual perception module is also used to interact with the text representation in the text perception module to obtain the text representation most relevant to the image content. Such a design promotes the model's ability to extract key information when processing multimodal inputs. The text perception module is divided into text embedding and text perception.

[0126] (1) Text Embedding

[0127] For the input background text T, since the text is long, it must first be divided into sentences, and the entire text is organized in sentences. Then, the marks are used to splice the sentences. The mathematical expression is as follows:

[0128] Seq=[<cls>,t 1 ,<sep>,t 2 ,…,>sep>,t n ]

[0129] In the formula, t i represents the i-th sentence in the text, >cls> represents the start marker, >sep> represents the concatenation marker, and Seq represents the sentence sequence after the sentence segmentation and concatenation markers. Then map Seq to a sentence-level vector sequence:

[0130] E text =BERT(Seq)

[0131] In the formula, BERT is used to map each sentence in the text sequence Seq into a corresponding text embedding vector sequence, and all sequences are concatenated to obtain E text . The embodiment of the present invention uses a frozen pre-trained BERT to obtain an embedded representation. One of its advantages is that it uses the bidirectional characteristics of the Transformer to simultaneously consider the left and right contextual information of each word in the sentence. This all-round contextual understanding enables BERT to more accurately capture the relationship between words and the semantic information of the sentence. This method can more comprehensively capture semantic and syntactic relationships, so that word vectors can better express the context of words in context. After the text is embedded, the position information of the sentence is encoded to better represent the semantics of the text segment.

[0132] P text =pos(Seq)×W p

[0133] In the formula, the position information of the sentence in the text segment is obtained through the pos function, W p is a learnable weight matrix. Next, P text With E text Sum and get the final text representation I t .

[0134] I t =P text +E text

[0135] Text representation of input text obtained by text embedding Where M is the length of the text representation, which is determined by the length of the original text, and d is the dimension of each specified embedding vector. The entire process of text embedding is represented as follows:

[0136] I t =E(text)

[0137] (2) Text Perceptron

[0138] The text perceptron accepts text token embeddings I t As input, these embeddings encapsulate the language content of the text. Its internal structure is similar to that of a visual perceptron. It first passes through a self-attention layer, which uses the self-attention mechanism to capture the key semantics in the text embedding. Then, it passes through a cross-attention layer and the visual embedding I v Interact to extract the text semantic representation most relevant to the image. Finally, after a layer of feedforward neural network, the final output vector is generated. This process is symmetrical with the visual perceptron, and the overall mathematical expression of this process is as follows, where θ represents all the parameters of the attention module.

[0139]

[0140] This module is in a dual relationship with the visual perceptron. Similarly, the text representation I obtained by text embedding t After the text perceptron interacts with the visual representation of the input image, the text representation most relevant to the image content is extracted. In the second stage of training, the final output of the text perceptron is The initial input representation used as the semantic alignment module is further transformed to feed the large language model with the semantic content of greatest interest in question generation.

[0141] (3) Semantic alignment module

[0142] See also Figure 3 The semantic alignment module of the embodiment of the present invention adopts a method similar to Blip-2. It customizes a set of fixed-length learnable feature embeddings, interacts with the semantic representations from the visual perception module and the text perception module through the attention module, and outputs a set of fixed-length feature representations. It reduces the dimensionality of the rich multimodal representations and condenses the key semantics, which is more convenient for the subsequent first-stage alignment training.

[0143] The semantic representations after multimodal interaction are obtained through the visual perception module and the text perception module respectively. These representations already contain all the semantic information needed for subsequent question generation. However, since the decoder is based on the large language model pre-trained with text input, the decoder cannot directly understand this representation and needs to be trained in cross-modal alignment. The purpose of the semantic alignment module is to better align the multimodal representation with the input of the downstream decoder module based on the large language model in subsequent training, so that the large language model can understand the multimodal input more easily. Due to the richness of image features and the length of text materials, the amount of information contained in the two representations is dynamic. If the traditional fusion method is used, it is difficult to uniformly map it to the text representation space of the large language model. In recent multimodal task research based on large language models, the method based on semantic alignment has become the mainstream. The Blip-2 model proposes an innovative Q-form adapter, which converts the visual representation into a form that is easier for the language model to understand through a set of fixed-length learnable vector sets, thereby facilitating the training of subsequent visual features mapped to the text semantic space of the large language model. Based on this, the semantic alignment module of the present invention adopts the same design.

[0144] Based on the steps described above, the features output by the visual perception module are obtained And the features output by the text perception module Where N is a customizable hyperparameter, which is determined by the number of blocks specified in the image encoder and depends on the feature richness of the image. M is a dynamic value determined by the length of the input text. One of the tasks of the semantic alignment module is to take these two long outputs from different modal perceptrons as input and produce a set of condensed fixed-length context representations. A fixed number size is specified as The learnable embedding H Text and H Img , where L satisfies the following formula to achieve the effect of dimensionality reduction and concentration of multimodal contextual semantics.

[0145] L<M+N

[0146] Similar to the visual perceptron, these learnable embeddings are processed through a self-attention layer, followed by a cross-attention layer interacting with the feature matrix from the visual perceptron or the language perceptron, and finally output through a feed-forward layer. Text , as shown in the following formula, after the attention module, the concentrated and aligned semantic representation is obtained Where θ represents all the parameters in the attention module.

[0147]

[0148] Similarly, for the visual learnable embedding H Img , as shown in the following formula, after the attention module, the concentrated and aligned semantic representation is obtained

[0149]

[0150] Then, as shown in the following formula, and Connect together to produce the context representation of the final output of the multimodal perceptual encoder

[0151]

[0152] The second step is to design the problem generation encoder, which is divided into two parts: model design and model training.

[0153] 1) Model design

[0154] Whether in real life or in the field of education, multimodal scenarios are often accompanied by complex and long texts, and the content is broader. Traditional end-to-end models can only understand the representation of simple text and image mixtures, and perform poorly in dealing with problems involving complex text and image mixture representations.

[0155] With the development and popularization of large-scale language models, current research on multimodal text generation increasingly uses these advanced models to perform related tasks. These large language models, due to their deep language understanding and generation capabilities, provide new possibilities for understanding complex multimodal scenarios where images and texts are mixed. The present invention also constructs a question generation decoder based on a frozen large language model. The large language model has been pre-trained on a large scale and performs well in traditional text input question generation tasks. They have learned rich language knowledge and semantic representations, and can capture the structure and context information of the language well. The frozen large language model is used as the architecture and training process of the decoding simplified model. There is no need to calculate and train a special decoder model, while reducing the complexity of the model and the difficulty of training, and can cope with complex multimodal scenarios. The embodiment of the present invention uses Vicuna-7B as the core of the question generation decoder. Vicuna is an open source lightweight large language model that uses fewer parameters to achieve more than 90% of ChatGPT's performance. The selection of this model can reduce the requirements for the performance of the experimental platform and improve the performance of question generation.

[0156] In the first step, the representation of the preliminarily aligned multimodal context has been obtained through the multimodal perceptual encoder. In order to facilitate the subsequent ablation experiment of the decoder's large language model, the present invention designs a linear layer to align the dimensionality gap between the encoder output and the input representations of different large language models. As mentioned above, the visual encoder ViT based on image-text matching pre-training has aligned the image representation and text representation to a certain extent. In addition, the semantic alignment module uses two custom learnable embeddings H Text and H Img The semantic gap between multimodal representation and textual representation is further aligned, so in the decoder part, we only need to simply After a linear fully connected layer, it can be well converted into an input representation that can be understood by a large language model. As shown in the following formula, W and b are respectively the learnable matrix and vector, It is the multimodal semantic representation sequence that will eventually be input into the large model.

[0157]

[0158] In addition to inputting the learned multimodal semantic representation into the large language model In the model reasoning stage, task instruction statements need to be input so that the large model can understand the question generation task. It should be noted that different prompt instruction statements will produce different results for question generation in a specific multimodal context. Due to the uninterpretability of the large language model, it is difficult to directly design the most appropriate instruction statement. However, when conducting experiments on the entire dataset, as long as the large language model can understand and perform the question generation task, there will not be much difference in the overall question evaluation index. Embed the text markup of the prompt instruction statement into Add multimodal semantic representation sequence to At the beginning, this integrated data is used as input to a large language model, and beam search is used to select the best generated output.

[0159] The entire process of the question generation decoder is as follows:

[0160]

[0161] 2) Model training

[0162] The embodiment of the present invention adopts a two-stage training scheme widely used in the multimodal model (VLM) based on the large language model. Figure 4 As shown in the figure, in the pre-training stage, the goal is to align the large language model with visual information using image-text pairs from the image captioning dataset, which provides captions describing images and carefully depicts the details of the image content. After the first stage of training, the large language model is already familiar with visual embedding and can understand the content of the image. In the second stage, the model is trained using a dataset for question generation tasks so that the model can successfully perform question generation tasks.

[0163] Since the semantic alignment module of the embodiment of the present invention adopts a method similar to the Q-Form module of Blip-2, Blip-2 uses a 129M image-text matching dataset. The purpose of using this large-scale dataset for pre-training is to better fine-tune various downstream tasks in the second stage. Different from this, the second stage of the embodiment of the present invention only needs to fine-tune the question generation task, and only pre-trains on the 0.5M image title pair dataset obtained after filtering the CC3M dataset. In the first training stage, the text perception module and the text semantic alignment module are hidden. The main purpose of training in this stage is to achieve cross-modal semantic alignment from image representation to text representation. Since the semantic representation of the background text is also in the text representation space and there is no cross-modal semantic gap, there is no need for separate training. It can be aligned to the semantic space of the large language model in the second training stage.

[0164] After the first stage of training, the decoder with the large language model as the core has been able to convert visual representations into a text representation space that the large language model can understand, so the task of the second stage is mainly to guide the model to generate problems. In the second stage of training, the instructions to guide the generation of problems are also embedded in the front end of the encoder output sequence, and used together with the multimodal semantic representation as the soft prompt input of the large language model to guide the large language model to generate problems. It is worth noting that the second stage of training in the embodiment of the present invention still freezes the large language model, so the model in the embodiment of the present invention is essentially a prompt tuning that depends on the input, and the weights of the large language are not updated. This avoids the catastrophic forgetting problem caused by fine-tuning the large language model, and this method also reduces the requirements for hardware resources. When trying to unfreeze the large language model for fine-tuning, the model achieved worse results, that is, when there is less data, fine-tuning is often unnecessary. During the training process, the difference between the model output problem and the true value problem is compared, and the model is trained by minimizing the following language modeling loss:

[0165]

[0166] In the formula, θ←(θ vp ,θ lp ,θ sa ,H Text ,H Img ) is the trainable parameter of the model, y i is the target truth value at the current time step, Yes i The i-1 word tokens output previously.

[0167] Another embodiment of the present invention further provides a multimodal perceptual question generation system based on a large language model, comprising:

[0168] An image feature extraction module is used to extract image features and extract features most relevant to the input text from the image features;

[0169] A feature interaction module is used to convert the text background in the input content into word embedding, and interact with the features most relevant to the input text extracted from the image features to obtain the image representation most relevant to the text content and the text representation most relevant to the image content;

[0170] The semantic alignment module is used to semantically align the image representation that is most relevant to the text content and the text representation that is most relevant to the image content, obtain a multimodal semantic comprehensive representation of visual and text information, and convert it into an input representation that can be understood by the large language model;

[0171] The question generation module is used to guide the large language model to generate questions based on the input representation that the large language model can understand.

[0172] Another embodiment of the present invention further provides an electronic device, comprising:

[0173] A memory storing at least one instruction; and a processor executing the instruction stored in the memory to implement the multimodal perception question generation method based on a large language model.

[0174] Another embodiment of the present invention further proposes a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method for generating multimodal perception questions based on a large language model is implemented.

[0175] Exemplarily, the instructions stored in the memory may be divided into one or more modules / units, which are stored in a computer-readable storage medium and executed by the processor to complete the multimodal perception question generation method based on a large language model of the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program in the server.

[0176] The electronic device may be a computing device such as a smart phone, a notebook, a PDA, and a cloud server. The electronic device may include, but is not limited to, a processor and a memory. Those skilled in the art will appreciate that the electronic device may also include more or fewer components, or a combination of certain components, or different components, for example, the electronic device may also include an input / output device, a network access device, a bus, etc.

[0177] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0178] The memory may be an internal storage unit of the server, such as a hard disk or memory of the server. The memory may also be an external storage device of the server, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the server. Furthermore, the memory may include both an internal storage unit of the server and an external storage device. The memory is used to store the computer-readable instructions and other programs and data required by the server. The memory may also be used to temporarily store data that has been output or is to be output.

[0179] It should be noted that the information interaction, execution process and other contents between the above-mentioned module units are based on the same concept as the method embodiment. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0180] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0181] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the camera device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a disk or an optical disk.

[0182] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0183] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.< / sep>

Claims

1. A method for generating multimodal perceptual questions based on a large language model, characterized in that: include: Extract image features, and extract the features most relevant to the input text from the image features; The text background in the input content is converted into word embedding, and interacted with the features most relevant to the input text extracted from the image features to obtain the image representation most relevant to the text content and the text representation most relevant to the image content; Semantically align the image representation that is most relevant to the text content and the text representation that is most relevant to the image content to obtain a multimodal semantic comprehensive representation of visual and text information, and convert it into an input representation that can be understood by the large language model; Guide the large language model to generate problems based on the input representation that the large language model can understand.

2. The method for generating multimodal perceptual questions based on a large language model according to claim 1, characterized in that: The step of extracting image features is performed using a ViT visual encoder, and the image is input into the ViT visual encoder to obtain an embedding representation matrix. Where N is the length of the image feature vector sequence, which is determined by the number of image blocks, and d is the dimension of each image block vector. For a given image V, the image features are extracted by the ViT visual encoder as follows: I v =ViT(V;θ) Where θ is the parameter of the ViT visual encoder.

3. The method for generating multimodal perceptual questions based on a large language model according to claim 2, characterized in that: Extract the features most relevant to the input text from the image features through the visual perceptron; The image features are extracted through the ViT visual encoder, and the following embedded representation is obtained: I v =[a1,a2,...,a n ] In the formula, a i is the vector extracted by the visual encoder for each block of the image; Embedded Representation I v Perform self-attention calculation, for each vector a i ∈I v , calculate its relationship with any other vector a as follows j The alignment score e is: e=h T tanh(W k a j +W q a i ) Where W q , W k are all trainable parameter matrices, h T is the weight vector; The vector a is converted into i Normalize the alignment scores with each other vector to get the probability vector p i : Vector a i After a layer of self-attention operation, we get The calculation expression is as follows: Where W v is a trainable weight matrix; For I v Any eigenvector a in i All are obtained through the self-attention layer From this we get 4. The method for generating multimodal perceptual questions based on a large language model according to claim 3, characterized in that: Through the cross-attention layer, the visual perceptron will Each element in the text information representation I t Attention calculation is performed on the related elements in the text to capture the semantic association between the two; text information representation I t For a vector sequence, the expression is as follows: I t =[x1,x2,...,x n ] In the formula, x i Represents the embedding of each word in the sentence, where the sentence contains n words in total; Press the formula to I t and Any vector in Calculate the attention score matrix score: Where W s is a trainable weight matrix; Normalize the attention scores through the softmax function to get the attention weight matrix: w i =softmax(score) Using the original image region representation vector Perform attention weighting to obtain the intermediate state vector c k , after a layer of feedforward neural network calculation, the final image region representation vector is obtained: for Any feature vector in All are calculated through the self-attention layer From this we get 5. The method for generating multimodal perceptual questions based on a large language model according to claim 1, characterized in that: The step of converting the text background in the input content into word embedding comprises: The input background text T is divided into sentences, and the sentences are concatenated as follows: Seq=[<cls>,t1,<sep>,t2,...,<sep>,t n ] Where, t i represents the i-th sentence in the text, <cls> represents the start marker, <sep> represents the concatenation marker, Seq represents the sentence sequence after the sentence segmentation and concatenation markers, and Seq is mapped to a sentence-level vector sequence as follows: E text =BERT(Seq) In the formula, BERT is used to map each sentence in Seq into a corresponding text embedding vector sequence, and all text embedding vector sequences are concatenated to obtain E text ; After embedding the text, the position information of the sentence is encoded as follows: P text =pos(Seq)×W p In the formula, the position information of the sentence in the text segment is obtained through the pos function, W p is a learnable weight matrix; P text With E text Sum and get the final text representation I t : I t =P text +E text The text representation obtained by embedding the input text Where M is the length of the text representation and d is the dimension of each embedding vector.

6. The method for generating multimodal perceptual questions based on a large language model according to claim 5, characterized in that: The text representation I obtained by receiving the input text through text perception and text embedding t , using self-attention mechanism to capture text representation I t The key semantics in the image are represented by cross-attention layer and image embedding. v Interact to extract the text representation most relevant to the image content; finally, the feedforward neural network layer outputs the text representation most relevant to the image content 7. The method for generating multimodal perceptual questions based on a large language model according to claim 6, characterized in that: The semantically aligning the image representation most relevant to the text content and the text representation most relevant to the image content comprises: Embedding the image into a representation matrix and the text representation of the input text As input, a set of condensed fixed-length context representations is generated. In the formula, N is a customizable hyperparameter determined by the number of blocks specified in the image encoder, depending on the feature richness of the image, and M is a dynamic value determined by the length of the input text. A fixed number size is set to The learnable embedding H Text and H Img , where L satisfies the following formula: L <M+N For text, we can learn to embed H Text , the semantic representation after condensed alignment is obtained as follows For the visual learnable embedding H Img , the semantic representation after condensed alignment is obtained as follows Will and Connect them together to obtain a multimodal semantic comprehensive representation of visual and textual information 8. The method for generating multimodal perceptual questions based on a large language model according to claim 7, characterized in that: Comprehensively represent the multimodal semantics of visual and textual information After a linear fully connected layer, it is converted into an input representation that can be understood by the large language model as follows: Where W and b are learnable matrices and vectors, respectively. It is a sequence of multimodal semantic representations that is input into a large language model; In the step of guiding the large language model to generate questions based on the input representation that the large language model can understand, the text tag of the prompt instruction sentence is embedded in Add to multimodal semantic representation sequence At the beginning, beam search is used to select the best generated output; Vicuna-7B is used as the core of the question generation decoder; the expression of question generation decoding is as follows:

9. A multimodal perceptual question generation system based on a large language model, characterized in that: include: An image feature extraction module is used to extract image features and extract features most relevant to the input text from the image features; A feature interaction module is used to convert the text background in the input content into word embedding, and interact with the features most relevant to the input text extracted from the image features to obtain the image representation most relevant to the text content and the text representation most relevant to the image content; The semantic alignment module is used to semantically align the image representation that is most relevant to the text content and the text representation that is most relevant to the image content, obtain a multimodal semantic comprehensive representation of visual and text information, and convert it into an input representation that can be understood by the large language model; The question generation module is used to guide the large language model to generate questions based on the input representation that the large language model can understand.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for generating multimodal perceptual questions based on a large language model as described in any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Water conservancy unmanned aerial vehicle inspection method based on visual language action multi-modal model

    CN120704350A