Method and device for processing multi-modal data by using multi-modal large model
By using multiple attention headers and different mask matrices for attention processing in the multimodal large model, the problem of insufficient acquisition of context information caused by causal masks is solved, and the accuracy of multimodal data processing is improved.
Patent Information
- Application Number
- CN202510225325.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
When the existing multimodal large model processes multimodal data, since the large language model of the decoder only uses a causal mask, the context information of the multimodal data is insufficiently obtained, which affects the model's understanding of the data.
In the multimodal large model, multiple attention heads are used to perform attention processing on the multimodal data, and mask the attention matrix through different mask matrices, allowing the information of the rear position to be considered during attention calculation, thereby improving the acquisition of context information.
Through this method, the accuracy of the processing results of multimodal data can be improved and the context information understanding of images and text can be enhanced by the large language model.
Smart Images

Figure CN120068940A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the technical field of data processing, and in particular, to a method and device for processing multimodal data using a multimodal large model. Background Art
[0002] With the rapid development of large language models, more and more multimodal large models based on large language models have been developed to better understand multimodal data.
[0003] Currently, most of the large language models used in the relatively high-performance multimodal large models in the industry are decoder-only large language models (Decoder-only LLMs). In this type of large language model, a causal mask is used as the mask for the self-attention mechanism. Here, the causal mask is a backward-invalid mask, that is, when performing self-attention calculation, only the information at the previous positions is considered, which will lead to insufficient acquisition of the context information of multimodal data, and thus affect the understanding of multimodal data by the multimodal large model. Summary of the Invention
[0004] One or more embodiments of this specification describe a method for processing multimodal data using a multimodal large model, which can improve the accuracy of the processing results of multimodal data.
[0005] In a first aspect, a method for processing multimodal data using a multimodal large model is provided. The multimodal large model includes a large language model, and the large language model includes multiple attention heads, and the multiple attention heads correspond to different mask matrices; the method includes:
[0006] Performing attention processing on multiple representation vectors using a target attention head among the multiple attention heads to obtain an initial attention matrix, where the multiple representation vectors include several image representations corresponding to the input image and several text representations corresponding to the input text;
[0007] Performing mask processing on the initial attention matrix using a target mask matrix corresponding to the target attention head to obtain an updated attention matrix, where the target mask matrix has valid values at several target positions where the row number is less than the column number.
[0008] In a second aspect, a device for processing multimodal data using a multimodal large model is provided. The multimodal large model includes a large language model, and the large language model includes multiple attention heads, and the multiple attention heads correspond to different mask matrices; the device includes:
[0009] A processing unit, configured to perform attention processing on a plurality of representation vectors by using a target attention head among the plurality of attention heads to obtain an initial attention matrix, where the plurality of representation vectors include a plurality of image representations corresponding to an input image and a plurality of text representations corresponding to an input text;
[0010] A masking unit, configured to perform masking processing on the initial attention matrix by using a target masking matrix corresponding to the target attention head to obtain an updated attention matrix, where the target masking matrix has valid values at a plurality of target positions where the row number is less than the column number.
[0011] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method of the first aspect.
[0012] In a fourth aspect, a computing device is provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method of the first aspect is implemented.
[0013] The method for processing multimodal data by using a multimodal large model provided by one or more embodiments of this specification masks the attention matrices corresponding to respective attention heads by using different masking matrices in a large language model. Among them, these different masking matrices at least include a target masking matrix, which has valid values at a plurality of target positions where the row number is less than the column number. It should be understood that when using the target masking matrix to perform masking processing on the attention matrix, it is possible to consider the information at the subsequent positions during attention calculation, that is, in this solution, it is possible to consider the context information at the current position during attention calculation, which helps to improve the accuracy of the processing result of multimodal data. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] To more clearly illustrate the technical solutions of the embodiments of this specification, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0015] Figure 1 It is a schematic diagram of an implementation scenario of an embodiment disclosed in this specification;
[0016] Figure 2 It shows a flowchart of a training method for a multimodal large model according to an embodiment of this specification;
[0017] Figure 3 It shows a schematic diagram of a method for processing a sequence of representation vectors in an example of this specification;
[0018] Figure 4 Schematic diagrams of each mask matrix in an example of this specification are shown;
[0019] Figure 5 Schematic diagrams of the setting method of the mask matrix in an example of this specification are shown;
[0020] Figure 6 Schematic diagrams of a device for processing multimodal data using a multimodal large model according to an embodiment of this specification are shown. Detailed implementation manners
[0021] The solutions provided in this specification will be described below with reference to the accompanying drawings.
[0022] In this specification, the above-mentioned multimodal large model may include a vision-language large model, an audio-language large model, etc., and the multimodal data processed by using this multimodal large model may include image-text pairs, audio-text pairs, etc. Taking the vision-language large model as an example, it may include: an image encoder, an adapter (or connector), and a large language model.
[0023] For the above-mentioned vision-language large model, the current training method is as follows: First, based on a plurality of pre-collected image-text pairs, the adapter is trained to enable it to perform alignment processing on the input image and text, that is, to perform textification processing on the input image. Then, the entire vision-language large model is fine-tuned to enable it to understand the content of the input image and text. However, since the adapter is usually implemented based on two-layer MLP and has a small number of parameters, it is difficult to make the large language model fully understand the image content by only focusing on the adapter.
[0024] For the above problems, the following several improvement solutions are proposed:
[0025] First, use a Q-Former network based on the Transformer architecture to implement the adapter. A cross-attention mechanism layer is added to this Transformer architecture, which uses a fixed number of learnable query vectors to compress the input image, and these learnable query vectors are simultaneously concatenated with the input text, and a fused representation of the input text is obtained based on the self-attention mechanism. In addition, during the training process, a text-image matching loss, a text-image contrast loss, and a text generation loss are used in combination, and only the parameters of the adapter are adjusted.
[0026] Second, implement an image encoder based on Vision Transformers (abbreviated as ViT), and implement an adapter based on a resampler network. It should be noted that since the number of image representations extracted by the ViT-based image encoder is usually in the hundreds to thousands, this will increase the inference speed and latency of the large language model. Therefore, an adapter is implemented based on the resampler network. Specifically, this resampler network is a structure based on the multi-head attention mechanism of Transformer, in which the image representations to be compressed are used as keys and values, and a fixed number of learnable query vectors are used to compress the image representations. Finally, the image representations are compressed to the same number as the learnable query vectors, which can greatly alleviate the inference speed and latency of the large language model.
[0027] In the improved version of the second solution above, the above adapter can also be implemented as a multi-layer perceptron, which uses this multi-layer perceptron to compress the image representations. For example, compress the image representations at a compression ratio of 4:1.
[0028] Third, implement an image encoder based on Vision Transformers (abbreviated as ViT), and implement an adapter based on a language middleware (QLLaMA), where QLLaMA uses the weights of the pre-trained multilingual LLaMA and adds 96 learnable query vectors and cross-attention layers. During the training process, a combination of image-text matching loss, image-text contrastive loss, and text generation loss is used to adjust the parameters of the adapter and the image encoder.
[0029] In the improved version of the third solution above, in the image encoder, a dynamic high-resolution mechanism is used, that is, the input image is first divided into blocks of different ratios, and then the image encoder is used to extract the image representations of these blocks. In addition, an adapter is implemented based on a multi-layer perceptron to minimize the loss of information in the image representations.
[0030] In summary, the above various improvement solutions all focus on how to design a good adapter to connect the image encoder and the large language model. However, as mentioned before, compared with the large language model and the image encoder, the number of parameters of the adapter is relatively small. For example, in the Scholar Vision-Language Model, the maximum number of parameters of the adapter is only 172M, while the number of parameters of the large language model reaches 103B, and the number of parameters of the image encoder also reaches 5.5B. Therefore, simply using an adapter with such a small number of parameters to textify the input image will result in poor processing results, which will affect the large language model's understanding of the image content.
[0031] Therefore, in this solution, it is proposed to use the large language model to undertake the task of fusing image representations and text representations to relieve the pressure on the adapter.
[0032] However, as mentioned above, in currently commonly used large language models, when performing self-attention calculation, only the information at the previous positions is considered, which will affect the large language model's understanding of image content. Therefore, this solution further proposes to use different mask matrices to perform mask processing on the attention matrices corresponding to each attention head in the large language model. Among them, these different mask matrices at least include a target mask matrix, which has valid values at several target positions where the row number is less than the column number. It should be understood that when using this target mask matrix to perform mask processing on the attention matrix, it is possible to consider the information at the subsequent positions when performing attention calculation, that is, this solution can consider the context information of the current position when performing attention calculation, which helps to improve the accuracy of the processing results of multimodal data.
[0033] Figure 1 Schematic diagram of the implementation scenario of an embodiment disclosed in this specification. Figure 1 In it, the multimodal large model may include an image encoder, an adapter, and a large language model.
[0034] In one embodiment, the above image encoder may be implemented as a Visual Geometry Group (VGG) model or a Residual Network (ResNet) model, etc.
[0035] In another embodiment, the above image encoder is also implemented as a neural network model based on Transformer, such as Vision Transformers (abbreviated as ViT).
[0036] In a more specific embodiment, the above ViT may be selected from CLIP (a model trained based on the ViT architecture). In CLIP, it includes an image encoder (including ViT) and a text encoder, and is trained by the method of contrastive learning to establish a corresponding relationship between images and texts in the same vector space.
[0037] In another more specific embodiment, the above ViT may be selected from SigLip (built on the basis of CLIP), and SigLip uses a loss function of the sigmoid function to alleviate the influence of the batch size on the training effect when using the contrastive loss.
[0038] The above adapter is used to connect the image encoder and the large language model, and it can be implemented as a multi-layer perceptron for aligning the image representation extracted by the image encoder to the text representation space.
[0039] The above-mentioned large language model may include multiple attention heads, and the multiple attention heads correspond to different mask matrices. Among them, each of the mask matrices corresponding to the multiple attentions may include a target mask matrix, and the target mask matrix has valid values at a number of target positions where the row number is less than the column number.
[0040] Specifically, the input image can be subjected to feature extraction through an image encoder to obtain a number of initial representations of the input image. Then, the number of initial representations are input into an adapter for text processing to obtain a number of image representations aligned to the text representation space. Finally, the number of image representations and the input text are input into the large language model. In the large language model, the input text is first subjected to embedding processing to obtain corresponding number of text representations, and the number of text representations and the number of image representations together form a representation vector sequence. Then, for the representation vector sequence, any one of the multiple attention heads can be used to perform attention processing on the multiple representation vectors in the representation vector sequence to obtain an initial attention matrix. Then, the initial attention matrix can be masked using the mask matrix corresponding to the any one attention head, and further an updated attention matrix can be obtained. Finally, a target text can be generated according to each of the updated attention matrices corresponding to the multiple attention heads.
[0041] It should be understood that in the case where the input text is a question about the input image, the target text is the answer corresponding to the input text.
[0042] It should also be understood that Figure 1 This is only an exemplary illustration. In practice, the above image encoder can also be replaced by a feature extraction algorithm, such as a feature extraction algorithm based on Scale-Invariant Feature Transform (SIFT), a feature extraction algorithm based on Histogram of Oriented Gradients (HOG), etc. This specification does not make any limitation in this regard.
[0043] As mentioned above, in this solution, the large language model undertakes the task of fusing image representations and text representations, so there is no need to pre-train the adapter, and only one overall fine-tuning of the multi-modal large model is required. The following describes this fine-tuning process.
[0044] Figure 2 The flowchart of the training method of a multi-modal large model according to an embodiment of this specification is shown. This method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. It should be noted that this method includes multiple rounds of iteration. Figure 2 The method steps included in the t-th (t is a positive integer) round of iteration are shown. It can be understood that by repeatedly executing the steps shown therein, multi-round iterative updates of the multi-modal large model can be achieved, and then the multi-modal large model updated in the last round is used as the finally used multi-modal large model. As Figure 2 shown, this method may include the following steps:
[0045] Step S202: Obtain a sample set, which includes a number of matching image-text pairs.
[0046] Taking any matching image-text pair as an example, the text therein (hereinafter referred to as the input text) can be a question asked about the image (hereinafter referred to as the input image). For Figure 1 example, the input text can be, for example: "What is the person in the picture doing?"
[0047] It should be understood that when the text in the image-text pair is a question asked about the image, this question can have a corresponding standard answer.
[0048] Step S204: Use an image encoder to extract features from the input image in any image-text pair to obtain a number of initial representations of the input image.
[0049] As mentioned above, the image encoder can be implemented as a Visual Geometry Group (VGG) model, a Residual Network (ResNet) model, etc. It can also be implemented as a neural network model based on Transformer, such as Vision Transformers (abbreviated as ViT).
[0050] When the image encoder is implemented based on a neural network model of Transformer, since the neural network model based on Transformer processes sequence data, the input image can be first divided into multiple patches. For example, Figure 1 in, the input image can be divided into 9 patches. Then, these multiple patches can be tiled to obtain the input sequence of the image encoder. Of course, in practice, the position information of each patch can also be added to the input sequence to facilitate the image encoder to distinguish patches at different positions.
[0051] Next, in the image encoder, based on the attention mechanism, features of multiple patches in the input sequence can be extracted, thereby obtaining a number of initial representations of the input image.
[0052] Step S206: Input a number of initial representations of the input image into an adapter for text processing to obtain a number of image representations aligned to the text representation space.
[0053] It should be noted that in this solution, each image representation output by the adapter has text semantics and has the same dimension as the input representation of the large language model.
[0054] Step S208, input a number of image representations of the input image and the input text matching the input image into the large language model.
[0055] In the large language model, the input text can be first embedded to obtain a number of text representations of the input text. Then, a number of image representations of the input image and a number of text representations of the input text can be combined to obtain a sequence of representation vectors. That is to say, the multiple representation vectors in the sequence of representation vectors include a number of image representations of the input image and a number of text representations of the input text.
[0056] In one example, a number of image representations among the multiple representation vectors are sorted before a number of text representations.
[0057] Of course, in practice, the above-mentioned number of text representations can also be sorted before the number of image representations.
[0058] Figure 3 The schematic diagram of the method for processing the sequence of representation vectors shown in an example of this specification is as follows Figure 3 As shown, the method may include the following steps:
[0059] Step S302, use any attention head headi among the multiple attention heads to perform attention processing on the multiple representation vectors in the sequence of representation vectors to obtain the initial attention matrix A corresponding to the attention head headi, which contains the attention coefficients between the representation vectors. Step S304 uses the mask matrix corresponding to the attention head headi to perform mask processing on the initial attention matrix A to obtain the updated attention matrix A' corresponding to the attention head headi, so that the updated attention matrices corresponding to the multiple attention heads can be obtained. Step S306, determine the target text according to the updated attention matrices corresponding to the multiple attention heads.
[0060] Specifically, the above-mentioned attention processing may include: based on the parameter matrices (including the Q parameter matrix, the K parameter matrix, and the V parameter matrix) corresponding to the attention head headi, map the multiple representation vectors into respective query vectors, respective key vectors, and respective value vectors. Then, according to the respective query vectors and respective key vectors, the initial attention matrix A corresponding to the attention head headi is obtained.
[0061] Taking the length of the sequence of representation vectors as N as an example, that is, the sequence of representation vectors includes N representation vectors, the size of the above-mentioned initial attention matrix is N×N, and it can be expressed as follows:
[0062]
[0063] In Formula 1, D is the dimension of the representation vector, Q and K are the query vector and the key vector respectively, where Q = XW Q, K = XW K , where X is the above-mentioned sequence of representation vectors, and W Q and W K are the Q parameter matrix and the K parameter matrix, respectively.
[0064] It should be understood that each row from left to right (or each column from top to bottom) in the initial attention matrix A of size N×N corresponds to each representation vector in the sequence of representation vectors from left to right.
[0065] And the corresponding masking process may include: performing element-wise multiplication on the initial attention matrix A and the masking matrix corresponding to the attention head headi, so as to obtain the updated attention matrix A′ corresponding to the attention head headi, which can be specifically expressed by the following formula:
[0066]
[0067] Among them, the definitions of D, Q, and K can be referred to in Formula 1, and M is the masking matrix (described later).
[0068] In this solution, multiple attention heads in the large language model correspond to different masking matrices. Taking any attention head headi as an example, each row and each column of the masking matrix corresponding to it correspond to each representation vector in the above-mentioned sequence of representation vectors. Among them, when several image representations are sorted before several text representations in each representation vector, the masking matrix corresponding to the attention head headi can be selected from Figure 4 each of the masking matrices shown.
[0069] Figure 4 Among them, the masking matrix M C , which has valid values only at the base positions where the row number is greater than the column number (shown by the gray box), and the values at other positions are zero, and can be specifically expressed by the following formula:
[0070]
[0071] where i is the row number and j is the column number.
[0072] The masking matrix M V2V , which has valid values at the above-mentioned base positions and the first target positions (shown by the gray dotted box), and the values at other positions are zero. Among them, any first target position satisfies that the row number is less than the column number, and the representation vectors corresponding to the row and the column are both image representations, or in other words, both the row number and the column number belong to the index range (also called the position range) of several image representations in the sequence of representation vectors, and can be specifically expressed by the following formula:
[0073]
[0074] where i is the row number, j is the column number, and vs and v e are respectively the lower and upper limits of the index ranges of a number of image representations.
[0075] Mask matrix M V2Q , which has valid values at the above-mentioned base position and the second target position (shown by the gray dashed box), and the values at other positions are zero; wherein, any second target position satisfies that the row number is less than the column number, and the representation vector corresponding to the row is an image representation, and the row number belongs to the index range of a number of image representations in the representation vector sequence; the representation vector corresponding to the column is a text representation, and the column number belongs to the index range of a number of text representations in the representation vector sequence, and can be specifically expressed as the following formula:
[0076]
[0077] wherein, i is the row number, j is the column number, v s and v e are respectively the lower and upper limits of the index ranges of a number of image representations, q s and q e are respectively the lower and upper limits of the index ranges of a number of text representations.
[0078] Mask matrix M Q2Q , which has valid values at the above-mentioned base position and the third target position (shown by the gray dashed box), and the values at other positions are zero. Among them, any third target position satisfies that the row number is less than the column number, and the representation vectors corresponding to both the row and the column are text representations, or rather, both the row number and the column number belong to the index range of a number of text representations in the representation vector sequence, and can be specifically expressed as the following formula:
[0079]
[0080] wherein, i is the row number, j is the column number, q s and q e are respectively the lower and upper limits of the index ranges of a number of text representations.
[0081] It should be noted that in practice, among the representation vectors corresponding to each row and each column of the mask matrix corresponding to any attention head headi, it is also possible that a number of text representations are sorted before a number of image representations. At this time Figure 4 the mask matrix M V2Q in needs to be replaced with M′. After replacement, M′ has valid values at the above-mentioned base position and the second target position, and the values at other positions are zero; wherein, any second target position satisfies that the row number is less than the column number, and the representation vector corresponding to the row is a text representation, and the row number belongs to the index range of a number of text representations in the representation vector sequence; the representation vector corresponding to the column is an image representation, and the column number belongs to the index range of a number of image representations in the representation vector sequence.
[0082] From Figure 4 it can be seen that the mask matrices M V2V , M V2Q , M Q2Q not only have valid values at positions where the row number is greater than or equal to the column number, but also have valid values at several target positions where the row number is less than the column number. Therefore, when using these mask matrices to perform mask processing on the initial attention matrix, it is possible to consider not only the information before the current position but also the information after the current position during attention calculation, that is, it is possible to consider the context information of the current position.
[0083] Specifically, based on the mask matrix M V2V it is possible to consider all the image representations during attention calculation for the image representation, that is, it is possible to obtain the complete context information of the image. Based on the mask matrix M V2Q it is possible to consider all the text representations during attention calculation for the image representation, thereby effectively establishing the correlation relationship from the image to the text. Based on the mask matrix M Q2Q it is possible to consider all the text representations during attention calculation for the text representation, that is, it is possible to obtain the complete context information of the text.
[0084] It should be understood that in the case where the image representation is sorted before the text representation, due to the mask matrix M C , the correlation relationship from the text to the image will be effectively established, so there is no need to separately set the mask matrix.
[0085] All in all, in this solution, based on the above four mask matrices, it is possible to obtain different context information respectively.
[0086] Returning to Figure 2 , after obtaining the updated attention matrix A′ corresponding to the attention head headi, multiple updated vectors corresponding to multiple representation vectors can be obtained according to the updated attention matrix A′ and the above value vectors. Similarly, multiple updated vectors corresponding to each attention head can be obtained.
[0087] Next, the multiple updated vectors corresponding to multiple attention heads can be fused correspondingly, so as to obtain the output corresponding to the attention layer to which these multiple attention heads belong, which can be specifically expressed by the following formula:
[0088] MultiHead(Q,K,V)=Concat(head1,…,headn)W o (Formula 7)
[0089] where n is the number of attention heads, and W o is the linear transformation matrix.
[0090] It should be noted that large language models usually include multiple attention layers. In this solution, different masking matrices can be set for multiple attention heads in the first several layers of the multiple attention layers, while the same masking matrix is set for multiple attention heads in other layers. For example, all are set to the masking matrix M. C 。
[0091] In a specific embodiment, setting different masking matrices for multiple attention heads can specifically include: multiple attention heads can be sequentially divided into several groups, and each group corresponds to a different masking matrix, and each attention head in the same group corresponds to the same masking matrix.
[0092] Figure 5 Schematic diagram showing the setting method of the masking matrix in an example of this specification. Figure 5 In, the large language model includes multiple attention layers, in which different masking matrices are set for multiple attention heads in the first two layers of the multiple attention layers, while the same masking matrix is set for multiple attention heads in other layers. For example, all are set to the masking matrix M. C 。
[0093] Specifically, assuming that each of the first two layers has 12 attention heads, where the 1st and 2nd attention heads can be divided into group 1, and each attention head in this group 1 corresponds to the masking matrix M V2Q correspondingly; the 3rd and 4th attention heads are divided into group 2, and each attention head in this group 2 corresponds to the masking matrix M V2V correspondingly; the 5th and 6th attention heads are divided into group 3, and each attention head in this group 3 corresponds to the masking matrix M Q2Q correspondingly, and the remaining 6 attention heads are divided into group 4, and each attention head in this group 4 corresponds to the masking matrix M C correspondingly.
[0094] It should be understood that Figure 5 This is only an exemplary illustration. In practice, a Scale layer, a Softmax layer, etc. can also be set between adjacent attention layers. This specification does not limit this.
[0095] In summary, after the processing of each attention layer (including attention processing and masking processing), the target text can be generated based on the output of the last attention layer.
[0096] It should be understood that in the case where the input text in the image-text pair is a question for the input image, the target text can be the answer corresponding to the input text.
[0097] Step S210: Calculate the classification loss based on the probabilities predicted by the large language model for each word in the target text, and adjust the parameters of the large language model according to this classification loss.
[0098] Specifically, the cross-entropy loss function can be used to calculate the classification loss based on the probabilities predicted by the large language model for each word in the target text and the actual values determined according to the standard answers corresponding to the input text. After calculating the classification loss, the update gradients corresponding to the large language model, the image encoder, and the adapter are calculated respectively using the backpropagation method, and the parameters of the large language model, the image encoder, and the adapter are updated based on them.
[0099] Of course, in practice, the above cross-entropy loss function can also be replaced by the mean squared error loss function, etc., and this specification does not limit this.
[0100] Thus, one round of iterative update of the multi-modal large model is completed. After multiple rounds of iterative update, the multi-modal large model updated in the last round can be used as the final multi-modal large model for use, and it is used to process multi-modal data.
[0101] It should be understood that the above is the training method proposed for the vision-language large model. Similarly, it is also possible to train multi-modal large models such as the audio-language large model, and this solution will not be elaborated here.
[0102] In summary, the training method for the multi-modal large model provided by this solution does not require pre-training for the adapter in advance, but only needs to perform a fine-tuning on the multi-modal large model as a whole. And after fine-tuning, not only can the adapter perform text-based processing on the image representation, but also the large language model itself can fully understand the content of the image, thereby relieving the pressure on the adapter.
[0103] In addition, by using different mask matrices for multiple attention heads in this solution, the deficiency in the ability to obtain context information of images or texts using only the mask matrix M C and the problem of being unable to effectively establish the association relationship from images to texts can be alleviated. Different from the mask matrix M C , the mask matrices M V2V , M V2Q , M Q2Q constructed in this solution can enable the large language model to effectively establish the association relationship from images to texts, and at the same time can fully explore the context information of images and texts, which helps to improve the accuracy of the processing results of multi-modal data.
[0104] The main innovation points of this solution are summarized as follows:
[0105] 1. A variety of masking strategies are proposed, enabling large language models to fully explore the context information of images and texts, and establish the correlation between images and texts.
[0106] 2. A decoupled masking strategy is proposed, that is, different masking matrices are set for different attention heads to mix a variety of masking strategies into large language models.
[0107] Corresponding to the above method of using a multimodal large model to process multimodal data, an embodiment of this specification also provides a device for using a multimodal large model to process multimodal data, as Figure 6 shown, the device may include:
[0108] A processing unit 602, configured to perform attention processing on a plurality of representation vectors by using a target attention head among a plurality of attention heads to obtain an initial attention matrix, where the plurality of representation vectors include several image representations corresponding to the input image and several text representations corresponding to the input text.
[0109] A masking unit 604, configured to perform masking processing on the initial attention matrix by using a target masking matrix corresponding to the target attention head to obtain an updated attention matrix, where the target masking matrix has valid values at several target positions where the row number is less than the column number.
[0110] In one embodiment, each row and each column of the above target masking matrix corresponds to each representation vector and is selected from one of the following:
[0111] A first masking matrix, which has valid values at a base position and a first target position, and the values at other positions are zero; where the row number of the base position is greater than or equal to the column number, and any first target position satisfies that the row number is less than the column number, and the representation vectors corresponding to the row and the column are both image representations;
[0112] A second masking matrix, which has valid values at a base position and a second target position, and the values at other positions are zero; where any second target position satisfies that the row number is less than the column number, and the representation vectors corresponding to the row and the column include image representations and text representations;
[0113] A third masking matrix, which has valid values at a base position and a third target position, and the values at other positions are zero; where any third target position satisfies that the row number is less than the column number, and the representation vectors corresponding to the row and the column are both text representations.
[0114] In one embodiment, several image representations among the above plurality of representation vectors are sorted before several text representations;
[0115] The second target position satisfies that the row number is less than the column number, the row number belongs to the index range of several image representations, and the column number belongs to the index range of several text representations.
[0116] In one embodiment, the mask unit 604 is specifically configured to:
[0117] Perform element-wise multiplication on the target mask matrix and the initial attention matrix to obtain an updated attention matrix.
[0118] In one embodiment, the above-mentioned multiple attention heads are sequentially divided into several groups, and each group corresponds to a different mask matrix. Each attention head in the same group corresponds to the same mask matrix.
[0119] In one embodiment, the above-mentioned multiple attention heads include a second attention head, which corresponds to a fourth mask matrix. The fourth mask matrix has valid values only at the base positions where the row number is greater than or equal to the column number, and the values at other positions are zero.
[0120] In one embodiment, the multimodal large model further includes an image encoder and an adapter; the device further includes:
[0121] An extraction unit 606, configured to perform feature extraction on the input image through the image encoder to obtain a number of initial representations of the input image;
[0122] An input unit 608, configured to input the number of initial representations into the adapter for text processing to obtain a number of image representations aligned to the text representation space.
[0123] In one embodiment, the extraction unit 606 includes:
[0124] A slicing sub-module 6062, configured to slice the input image into multiple tiles;
[0125] An input sub-module 6064, configured to input the multiple tiles and their position information into the image encoder. In the image encoder, based on the attention mechanism, perform feature extraction on the multiple tiles to obtain a number of initial representations of the input image.
[0126] In one embodiment, the input text is a question asking about the input image, and the device further includes:
[0127] An acquisition unit 610, configured to acquire the target text generated by the large language model according to the updated attention matrices corresponding to the multiple attention heads;
[0128] A determination unit 612, configured to determine the target text as the answer corresponding to the input text.
[0129] In one embodiment, the device further includes:
[0130] A calculation unit 614, configured to calculate the classification loss according to the probabilities predicted by the large language model for each word in the target text;
[0131] An adjustment unit 616 for adjusting the parameters of the large language model according to the classification loss.
[0132] The functions of the functional units of the device in the above embodiments of this specification can be implemented by the steps of the above method embodiments. Therefore, the specific working process of the device provided in an embodiment of this specification will not be repeated here.
[0133] The device for processing multimodal data using a multimodal large model provided in an embodiment of this specification can improve the accuracy of the processing results of multimodal data.
[0134] According to an embodiment of another aspect, there is also provided a computer-readable storage medium having a computer program stored thereon, which when executed on a computer, causes the computer to execute in combination with Figure 2 or Figure 3 the method described.
[0135] According to an embodiment of still another aspect, there is also provided a computing device including a memory and a processor, where an executable code is stored in the memory, and when the processor executes the executable code, it implements in combination with Figure 2 or Figure 3 the method described.
[0136] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the medium or device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0137] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0138] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of this specification. It should be understood that the above is only the specific embodiments of this specification and is not used to limit the protection scope of this specification. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of this specification shall be included within the protection scope of this specification.
Claims
1. A method for processing multimodal data using a multimodal large model, wherein the multimodal large model includes a large language model, wherein the large language model includes a plurality of attention heads, and the plurality of attention heads correspond to different mask matrices; the method comprises: Performing attention processing on multiple representation vectors using a target attention head among the multiple attention heads to obtain an initial attention matrix, wherein the multiple representation vectors include multiple image representations corresponding to the input image and multiple text representations corresponding to the input text; The initial attention matrix is masked using the target mask matrix corresponding to the target attention head to obtain an updated attention matrix, wherein the target mask matrix has valid values at several target positions whose row numbers are less than their column numbers.
2. The method according to claim 1, wherein: Each row and column of the target mask matrix corresponds to each characterization vector and is selected from one of the following: A first mask matrix having valid values at a base position and a first target position and zero values at other positions; wherein the row number of the base position is greater than or equal to the column number, any first target position satisfies that the row number is less than the column number, and the representation vectors corresponding to the rows and columns are both image representations; A second mask matrix having valid values at the base position and the second target position and zero values at other positions; wherein any second target position satisfies that the row number is less than the column number, and the representation vectors corresponding to the row and column include image representation and text representation; A third mask matrix has valid values at the base position and the third target position, and values at other positions are zero; wherein any third target position satisfies that the row number is less than the column number, and the representation vectors corresponding to the rows and columns are both text representations.
3. The method according to claim 2, wherein: The plurality of image representations in the plurality of representation vectors are sorted before the plurality of text representations; The second target position satisfies that the row number is smaller than the column number, the row number belongs to the index range of the plurality of image representations, and the column number belongs to the index range of the plurality of text representations.
4. The method according to claim 1, wherein: The mask processing includes: The target mask matrix and the initial attention matrix are bitwise multiplied to obtain the updated attention matrix.
5. The method according to claim 1, wherein: The multiple attention heads are divided into a number of groups in sequence, and each group corresponds to a different mask matrix, and each attention head in the same group corresponds to the same mask matrix.
6. The method according to claim 1, wherein: The multiple attention heads include a second attention head, which corresponds to a fourth mask matrix, and the fourth mask matrix has valid values only at basic positions where the row number is greater than or equal to the column number, and the values of other positions are zero.
7. The method according to claim 1, wherein: The multimodal large model also includes an image encoder and an adapter; the plurality of image representations corresponding to the input image are obtained by the following steps: Extracting features from the input image using the image encoder to obtain a number of initial representations of the input image; The plurality of initial representations are input into the adapter for text processing to obtain a plurality of image representations aligned to the text representation space.
8. The method according to claim 7, wherein: The step of extracting features from the input image by using the image encoder comprises: Dividing the input image into a plurality of blocks; The multiple image blocks and their position information are input into the image encoder, and in the image encoder, features of the multiple image blocks are extracted based on an attention mechanism to obtain several initial representations of the input image.
9. The method according to claim 1, wherein: The input text is a question asked about the input image; the method further includes: Obtain target text generated by the large language model according to each updated attention matrix corresponding to the multiple attention heads; The target text is determined as the answer corresponding to the input text.
10. The method according to claim 9, further comprising: Calculating classification loss according to the predicted probability of each word in the target text by the large language model; According to the classification loss, parameters of the large language model are adjusted.
11. A device for processing multimodal data using a multimodal large model, wherein the multimodal large model includes a large language model, wherein the large language model includes a plurality of attention heads, and wherein the plurality of attention heads correspond to different mask matrices; the device comprises: a processing unit, configured to perform attention processing on a plurality of representation vectors using a target attention head among the plurality of attention heads to obtain an initial attention matrix, wherein the plurality of representation vectors include a plurality of image representations corresponding to the input image and a plurality of text representations corresponding to the input text; A mask unit is used to mask the initial attention matrix using a target mask matrix corresponding to the target attention head to obtain an updated attention matrix, wherein the target mask matrix has valid values at several target positions whose row numbers are less than the column numbers.
12. The device according to claim 11, wherein Each row and column of the target mask matrix corresponds to each characterization vector and is selected from one of the following: A first mask matrix having valid values at a base position and a first target position and zero values at other positions; wherein the row number of the base position is greater than or equal to the column number, any first target position satisfies that the row number is less than the column number, and the representation vectors corresponding to the rows and columns are both image representations; A second mask matrix having valid values at the base position and the second target position and zero values at other positions; wherein any second target position satisfies that the row number is less than the column number, and the representation vectors corresponding to the row and column include image representation and text representation; A third mask matrix has valid values at the base position and the third target position, and values at other positions are zero; wherein any third target position satisfies that the row number is less than the column number, and the representation vectors corresponding to the rows and columns are both text representations.
13. The device according to claim 12, wherein: The plurality of image representations are sorted before the plurality of text representations in the plurality of representation vectors; The second target position satisfies that the row number is smaller than the column number, the row number belongs to the index range of the plurality of image representations, and the column number belongs to the index range of the plurality of text representations.
14. The device according to claim 11, wherein: The mask unit is specifically used for: The target mask matrix and the initial attention matrix are bitwise multiplied to obtain the updated attention matrix.
15. The device according to claim 11, wherein The multiple attention heads are divided into a number of groups in sequence, and each group corresponds to a different mask matrix, and each attention head in the same group corresponds to the same mask matrix.
16. The device according to claim 11, wherein The multiple attention heads include a second attention head, which corresponds to a fourth mask matrix, and the fourth mask matrix has valid values only at basic positions where the row number is greater than or equal to the column number, and the values of other positions are zero.
17. The device according to claim 11, wherein: The multimodal large model also includes an image encoder and an adapter; the device also includes: an extraction unit, configured to perform feature extraction on the input image through the image encoder to obtain a plurality of initial representations of the input image; The input unit is used to input the plurality of initial representations into the adapter for text processing to obtain a plurality of image representations aligned to the text representation space.
18. The device according to claim 17, wherein: The extraction unit comprises: A sub-slicing module, used for dividing the input image into a plurality of image blocks; The input submodule is used to input the multiple image blocks and their position information into the image encoder, and in the image encoder, feature extraction is performed on the multiple image blocks based on the attention mechanism to obtain several initial representations of the input image.
19. The device according to claim 11, wherein: The input text is a question asked about the input image; the device also includes: An acquisition unit, configured to acquire a target text generated by the large language model according to each updated attention matrix corresponding to the plurality of attention heads; A determination unit is used to determine the target text as an answer corresponding to the input text.
20. The apparatus according to claim 19, further comprising: A calculation unit, configured to calculate the classification loss according to the probability predicted by the large language model for each word in the target text; An adjustment unit is used to adjust the parameters of the large language model according to the classification loss.
21. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed in a computer, the computer is caused to execute the method according to any one of claims 1 to 10.
22. A computing device comprising a memory and a processor, wherein: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 10 is implemented.