Multi-modal large model method based on layered visual injection and mixed attention mechanism
Through the multimodal large-model method of layered visual injection and mixed attention mechanism, the problem of high computational complexity in the prior art is solved, and the calculation cost is significantly reduced and high performance is maintained.
Patent Information
- Application Number
- CN202510124361.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-27
AI Technical Summary
The existing multimodal large models have significantly increased computational complexity when processing high-resolution images or video scenes, limiting the application scenarios of the model.
Hierarchical visual injection and mixed attention mechanisms are adopted to efficiently integrate visual features with text features, and to avoid forward transmission of visual features in language decoders, thereby reducing overall computational complexity.
It significantly reduces computing costs, maintains high performance, enhances the scalability of the model, and has a wider range of applications.
Smart Images

Figure CN120047785A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology. Specifically, it relates to a multimodal large model method based on hierarchical visual injection and hybrid attention mechanism. This method can effectively integrate visual and text information, and significantly improve the computational efficiency of multimodal tasks. Background Art
[0002] In recent years, multimodal large models have attracted extensive attention from researchers. By integrating multimodal information, such models can achieve more complex and high-level semantic understanding, and are widely used in fields such as image captioning, video analysis, and visual question answering.
[0003] Currently, most methods directly connect visual and text features and input them into a large language model (LLM) for fusion. This approach has improved the processing ability of multimodal information to a certain extent, but there are still significant challenges: Visual sequences are usually longer than language sequences, especially when dealing with high-resolution images or introducing video scenarios, resulting in a sharp increase in the number of input visual tokens. The computational cost of large language models grows quadratically with the increase in the number of input tokens. Existing methods will lead to a significant increase in the computational complexity of the model, limiting the application scenarios of the model.
[0004] To address the above challenges, current research mainly adopts two strategies: One is to use lightweight large language models to reduce the computational complexity by reducing the parameter scale, but this often leads to a significant decline in performance; the other is to compress visual tokens to reduce the input length and computational amount. However, these methods can only relieve the computational pressure to a certain extent, and the compressed visual features still need to participate in the forward calculation of the entire model, and cannot fundamentally solve the problem of computational overhead. Summary of the Invention
[0005] The present invention aims to provide an innovative multimodal large model method to solve the problem of high computational complexity in existing methods. Different from existing methods, the present invention enables the efficient fusion of visual and text features through hierarchical visual injection and hybrid attention mechanism, while avoiding the forward transmission of visual features in the language decoder, thereby reducing the overall computational complexity.
[0006] The technical solution of the present invention includes the following two key modules:
[0007] 1. Hierarchical Visual Injection Module: This module provides visual features for each layer, avoiding interference between visual and text information, and enriching the visual representation in the model.
[0008] 2. Hybrid Attention Mechanism Module: This module realizes the efficient interaction between visual and text features through the attention mechanism.
[0009] The present invention includes the following steps:
[0010] Step 1: Data preparation
[0011] For the input image I, use a pre-trained visual encoder to extract the visual feature sequence X v = f VE (I), where represents the input image, H and W are the image height and width respectively, and C is the number of channels. f VE (·) represents the feature extraction function of the Visual Encoder. represents the set of real numbers, N represents the length of the visual feature sequence, and d v is the dimension of the visual feature. For the input text T, extract the feature X l = f +E (T), where f +E (·) represents the feature extraction function of the Textual Encoder. M represents the length of the text feature sequence, d is the hidden layer dimension of the large language model, a fixed hidden layer dimension determined by the design of the large language model, and is usually preset according to the scale of the model.
[0012] Step 2: Obtain hierarchical visual features
[0013] After obtaining the visual feature sequence, the present invention projects the visual features to the same hidden layer dimension d as the large language model. To achieve this mapping and avoid interference between visual features and text features, the present invention sets a visual projection matrix and in each layer of the large language model to convert the visual feature X v into key-value pairs:
[0014]
[0015] where respectively represent the key and value corresponding to the visual feature. The projection matrices and are randomly initialized during training and are iteratively updated through subsequent optimization algorithms. The keys K v and values V v obtained here will be used in the subsequent attention mechanism, which not only realizes the dimension alignment of visual and text features but also provides more diverse visual feature inputs for the model.
[0016] Step 3: Visual-text feature fusion under the hybrid attention mechanism
[0017] To achieve visual-text feature fusion in each layer of the large language model, the present invention proposes a hybrid attention mechanism. In the attention sublayer of each layer, the text feature sequence X l is mapped to query, key, and value features:
[0018] Q l = W l Q X l , K l = W l K X l , V l = W l V X l
[0019] where represent the query, key, and value representations of the text at this layer, respectively. is the text feature projection matrix. At the same time, to enhance the model's perception of the positional relationship between sequence data, position information is additionally added to Q l , K l in each layer:
[0020] Q l′ = Q l + P l , K l′ = K l + P l
[0021] Here represents the position encoding information, and Q l′ and K l′ are the text query and key features with added position information. Since the visual sequence has already had position information added during visual feature extraction, no additional addition is made here.
[0022] Subsequently, the key-value features of the visual sequence and the language sequence are concatenated into a complete KV sequence:
[0023] K vl = [K v ; K l′
[0024] V vl = [V v ; V l
[0025] where [;] represents concatenation in the sequence dimension.
[0026] Subsequently, attention scores are calculated through the attention calculation formula:
[0027]
[0028] where represents the matching score between the query and the key, d is the hidden layer dimension of the large language model, which is the same as described in step 1, is a scaling factor introduced to prevent the dot product of the query Q l′ and the key K vl from being too large, which may lead to an overly flat gradient of the subsequent Softmax output and affect the learning efficiency of the model. T represents the transpose.
[0029] During the training of each multi-modal task, to prevent future information leakage between texts, a partial causal mask is added to S, so that only past text information can be seen at the current moment, but visual features can be used globally. Finally, the attention scores are normalized by Softmax and multiplied by V vl , to obtain the output of the attention layer:
[0030] A = Softmax(S)V vl
[0031] Here, Softmax normalizes the attention scores for each row, which is the output of the current attention layer and represents the result of the fusion of visual and text features.
[0032] Step 4: Feed-forward layer calculation
[0033] After the visual and text fusion in step 3, the present invention only passes the text part to the feed-forward layer for calculation. The feed-forward layer contains 2 projection matrices and 1 activation function. The specific calculation process is as follows:
[0034] Y = W 2 (Act(W 1 A))
[0035] where and are projection matrices, which are used for dimension elevation and dimension reduction of features respectively. Act is a non-linear activation function, which is the output obtained by the current feed-forward layer.
[0036] Step 5: Obtain the output
[0037] The large language model has a total of L layers, where L is a hyperparameter determined according to the model architecture. For small-scale models, L is usually between 12 and 24 layers. The number of layers of the model directly affects the computational complexity and performance. Each layer includes two operations: the attention mechanism and the feed-forward layer. Finally, after L layers of calculation, the final output Y of the model can be obtained a .
[0038] Step 6: Model Optimization
[0039] In a multi-modal task, the input to the model includes visual feature X v and text feature X l , and its goal is to generate an answer Y that conforms to the actual semantics a . The model is optimized by maximizing the probability of generating the answer Y a , and the specific definition is as follows:
[0040]
[0041] where is the generated answer sequence, with a length of y <i represents all the words before the i-th word currently generated. θ are the parameters that need to be updated during the training of the model, specifically including: visual projection module parameters text module parameters W l Q ,W l K ,W l V , and the weights W 1 ,W 2 of the feed-forward layer. p(y i |X v ,X l ,y <i ; θ) represents the conditional probability of generating the i-th word based on the multi-modal input and the previously generated words.
[0042] To train the model, the negative log-likelihood is used as the loss function, and its definition is as follows:
[0043]
[0044] The training objective is to minimize the loss function which is equivalent to maximizing the conditional probability p(y i |X v ,X l ,y <i ; θ) of the generated text. By optimizing the parameter θ, the model can gradually learn how to better generate sequences that conform to the distribution of the training data.
[0045] Compared with existing methods, the present invention has the following obvious advantages and beneficial effects:
[0046] 1. Significantly reduce the computational cost: Through hierarchical visual injection and hybrid attention mechanism, the forward transmission of visual features in the language decoder is avoided, reducing the repeated participation of visual sequences in multi-layer calculations, and significantly reducing the overall computational complexity.
[0047] 2. Maintain high performance: While reducing the computational cost, it ensures the efficient fusion of visual and text features, and the model can still maintain a high level of performance in multi-modal tasks.
[0048] 3. Enhance scalability: Due to the lower computational complexity of the optimized model, it improves the model's ability to process larger-scale multi-modal data and has a wider range of applications.
[0049] In summary, the present invention proposes a multi-modal large model method based on hierarchical visual injection and hybrid attention mechanism, which greatly reduces the computational complexity of the model while maintaining the model performance. Brief Description of the Drawings
[0050] Figure 1 It is a flow chart of the present invention
[0051] Figure 2 It is a schematic diagram of the overall model structure of the present invention
[0052] Figure 3 It is a schematic diagram of the hybrid attention part of the present invention Detailed Embodiment
[0053] To verify the technical solution of the present invention, this embodiment is carried out in the following experimental environment: The hardware environment is an A100-80G graphics card, the software environment is CUDA 11.8, Python 3.10.14 and PyTorch 2.0.1, and the experiment is based on the TinyLLaVA framework.
[0054] The details of the specific experimental steps are as follows:
[0055] Step 1: Data Preparation
[0056] The first stage is the pre-training stage, and the training data uses the LLaVA-1.5-558k dataset. This dataset contains 558,000 multi-modal samples, each sample consists of an image and the corresponding text pair, and the text content is the description of the image. The second stage is the supervised fine-tuning stage, using the LLaVA-1.5-mix-665k dataset, which contains 665,000 multi-modal samples. The specific content is task instructions and answers based on images, which are used to guide the model to generate more target-semantic-compliant answers in multi-modal tasks. In terms of the visual encoder, the pre-trained Siglip SoViT-400m / 14 model is used, the feature sequence length of a single image is 728, the dimension is 1152, and the text decoder selects Qwen2 and Llama3.2.
[0057] Step 2: Obtain hierarchical visual features
[0058] During the hierarchical visual feature extraction process, project the visual features into the hidden layer space of the language model. A single projection layer is used to complete the key-value mapping, with its input dimension being the output dimension of the visual encoder and the output dimension being the hidden layer dimension of the language model, ensuring that the dimensions of visual and text features are consistent. In practice, the hidden layer dimension of Qwen2-0.5B is 512, and that of Llama3.2-1B is 2048.
[0059] Step 3: Visual-Text Feature Fusion under the Hybrid Attention Mechanism
[0060] In the attention calculation of each layer, introduce a hybrid attention mechanism to achieve efficient interaction between visual and text features. The specific steps are as follows:
[0061] Map the text feature X l to queries, keys, and values, and at the same time add rotary position encoding to Q l , K l :
[0062] Q l′ = Q l + P l 、K l′ = K l + P l
[0063] where Q l′ , K l′ represent the query and key features with added position information, and concatenate the visual and language key and value features into a complete K vl , V vl sequence.
[0064] According to the attention mechanism, calculate the attention scores:
[0065]
[0066] S is the matching score matrix between the query and the key. After softmax normalization of the attention scores, the output of the attention layer is obtained:
[0067] A = Softmax(S)V vl
[0068] Step 4: Feed-Forward Layer Calculation
[0069] In the design of the present invention, only the text features will be passed to the feed-forward layer for calculation, thereby reducing the repeated participation of visual features in subsequent layers. In this embodiment, the activation function used in the feed-forward layer is GeLU.
[0070] Step 5: Obtain the Output
[0071] After finally repeating the calculations of the L-layer attention and the feed-forward layer, the output Y of the model is obtained. a 。
[0072] Step 6: Model optimization
[0073] Calculate the cross-entropy loss between the generated answer and the target answer, and then perform backpropagation to optimize the network parameters.
[0074] To verify the effectiveness of the present invention, the present invention was tested on different language models, and the specific results are shown in Table 1.
[0075]
[0076]
[0077] Table 1: Comparison of model performance and model efficiency of the algorithm on different LLMs
[0078] GQA, MM-Vet, POPE, and MMMU are all evaluation metrics for multi-modal tasks. Params(M) represents the parameter scale of the model (unit: million). FLOPs(GB) represents the amount of floating-point operations required for the inference or training process (unit: GigaFLOPs). The "(9%)" in parentheses indicates that this scheme only occupies about 9% of the computational amount compared with the original model.
[0079] From the experimental results, it can be seen that the present invention can significantly reduce the computational amount of the model on different language models (Qwen2-0.5B and Llama-3.2-1B):
[0080] For Qwen2-0.5B, after adopting the present invention, the computational amount is reduced from 830 GFLOPs to 72 GFLOPs, which is only about 9% of the original model;
[0081] For Llama-3.2-1B, the computational amount is reduced from 2030 GFLOPs to 179 GFLOPs, which is also about 9% of the original model.
[0082] At the same time, the present invention still maintains the same or higher performance on multi-modal task metrics such as GQA, MM-Vet, POPE, and MMMU as the original model, fully demonstrating its superior scalability and effectiveness. Therefore, the method described in the present invention significantly reduces the computational complexity while maintaining high performance on multi-modal tasks in different test sets, and the technology is reasonable and reliable.
Claims
1. A multimodal large model method based on hierarchical visual injection and hybrid attention mechanism, characterized in that: The following steps are involved: Step 1: Data preparation; For the input image, use the pre-trained visual encoder to extract the visual feature sequence X v ,in For the input text, extract the text feature X through the text encoder in the large language model l ,in Step 2: Obtain hierarchical visual features; The visual feature X v Project to the same dimension as the large language model embedding space, setting the visual projection matrix for each layer and The visual features are converted into key-value pairs, and the calculation formula is: Said And update the projection matrix through back propagation during training and Step 3: Visual-text feature fusion under hybrid attention mechanism; In each layer of the large language model, the text feature sequence X l is mapped to query Q l , key K l Sum value V l Features, and add location information to get Q l′ and K l′ , concatenate the key-value features of vision and text into a complete key-value sequence, and calculate the attention score, which is calculated as follows: According to the attention score, the output A of the attention layer is calculated as follows: A=Softmax(S)V vl Step 4: Feedforward layer calculation; Only the text features are passed to the feed-forward layer for calculation, which contains two projection matrices W1 and W + , an activation function Act(·), whose calculation formula is: Y=W + (Act(W1A)) Step 5: Generate model output; After repeatedly executing the attention mechanism of the L layer and the feedforward layer calculation, the final output Y of the model is obtained a ; Step 6: Model optimization; The generated answer is compared with the target answer to calculate the cross entropy loss, and backpropagation is performed to optimize the network parameters.
2. According to claim 1, a multimodal large model method based on layered visual injection and hybrid attention mechanism is characterized in that: In step 1, the visual encoder is the pre-trained Siglip visual encoder SoViT-400m / 14, and the large language model is Qwen2 or Llama3.
2.
3. According to claim 1, a multimodal large model method based on layered visual injection and hybrid attention mechanism is characterized in that: In the hybrid attention mechanism described in step 3, the query is text information, and the key value is a sequence of visual and text concatenation.
4. The multimodal large model method based on hierarchical visual injection and hybrid attention mechanism according to claim 1, characterized in that: The activation function of the feedforward layer in step 4 is GeLU.
Citation Information
Patent Citations
Medical visual question and answer method, device and equipment and storage medium
CN118467707A
Multi-modal data fusion control method and device, equipment and medium
CN118734250A
Cited By
Dynamic confrontation cross-domain time sequence anomaly detection method and system
CN121302205A
A dynamic confrontation cross-domain time series anomaly detection method and system
CN121302205B