A multimodal large model system integrating vision and language

By combining DINOv2 and SigLIP vision encoder, using MLP projector and Mamba backbone model to build a multimodal large model system, the inefficiency of MLLMs in high real-time performance and edge deployment scenarios is solved, and the training speed and performance improvement is achieved.

CN119227744BActive Publication Date: 2025-07-04EAST CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411736605.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-07-04
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

Existing multimodal large language models (MLLMs) are less efficient in high real-time performance and edge deployment scenarios, and traditional methods improve efficiency by reducing model capacity or compressing visual context length, but usually at the cost of performance.

Method used

The combination of DINOv2 and SigLIP vision encoder is adopted to reduce the dimensionality of visual features through a multi-layer perceptron (MLP) projector module, and a multimodal large model system based on state space model is constructed using the Mamba backbone model.

Benefits of technology

It significantly improves training and inference speed, improves the overall performance of the model, especially in handling long sequences and fast inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119227744B_ABST
    Figure CN119227744B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal large model system that integrates vision and language, including: a vision encoder that integrates DINOv2 and SigLIP, used to collect low-level spatial attributes and semantic attributes; a multi-layer perceptron projector, used to map visual features to the language embedding space and the Mamba backbone model network based on the state space model. Compared with the multimodal large language model that relies on the Transformer network as the basic model, the large model system of the present invention has improved in indicators such as inference speed and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal language models, and in particular to a multimodal large model system that integrates vision and language. Background Art

[0002] Recently, multimodal large language models (MLLMs) have achieved remarkable results in various downstream tasks, including multimodal content generation, vision-based question answering, and embodied intelligence. These advancements not only demonstrate the versatility of MLLMs but also pave the way for further research into more nuanced and complex applications. For example, as MLLMs continue to evolve, significant improvements are possible in areas such as real-time interaction in dynamic environments, cross-modal retrieval tasks, and seamless integration of language and vision processing in everyday technologies.

[0003] MLLMs typically rely on the well-known Transformer network as the base model for many downstream tasks. However, due to its quadratic computational complexity, the Transformer network is often inefficient, making it difficult to meet the requirements in application scenarios that demand high real-time performance and are suitable for edge deployment.

[0004] Traditional methods mainly improve the efficiency of multimodal large language models (MLLMs) by reducing model capacity or compressing the length of visual context, while usually keeping the Transformer architecture unchanged in the language model. Although this method improves efficiency, it often comes at the cost of significantly reducing model performance. Summary of the Invention

[0005] The purpose of the present invention is to solve the above problems existing in the prior art and provide a multimodal system based on the Mamba language model.

[0006] The specific technical solution adopted by the present invention is as follows:

[0007] A multimodal large model system that integrates vision and language, the system includes:

[0008] (1) An encoder module, including a visual encoder and a text encoder, the visual encoder includes a visual encoder DINOv2 and a visual encoder SigLIP. Among them, DINOv2 captures low-level spatial information in the image, and SigLip captures semantic information in the image; First, image chunking: Scale and crop all images into a unified format. For a given image , where C is the number of channels of the image, H and W respectively represent the height and width of the image, and R represents the tensor space; Divide the image into the same size patches, where P is the side length of each patch; both the vision encoder DINOv2 and the vision encoder SigLIP take the image after patching as the input sequence, and extract the channel concatenation of the outputs of the two vision encoders as the compact vision tokens ; where and represent the feature dimensions of the outputs of the two vision encoders respectively is expressed as: where and represent the feature extraction operations of DINOv2 and SigLIP on the input image respectively. Concatenate the outputs of the two vision encoders to form output of dimension; The text encoder selects GPT-NeoXTokenizer to convert text into text tokens;

[0009] (2) Projector module, using a multi-layer perceptron MLP to reduce the dimensionality of the vision tokens output by the vision encoder module to the vision tokens of the dimension input to the next module , the formula is , represents a multi-layer perceptron; The purpose of this step is to convert the dimension of the vision tokens into the dimension input to the next module, and the output result of the projector module will be sent to the next module for processing;

[0010] (3) Mamba backbone model, select the version of the Mamba backbone model with 2.8B parameters pre-trained on the SlimPa-jama dataset or the 7B version pre-trained on the RefinedWeb dataset, and then concatenate the output result of the projector module with the output result of the text encoder and denote it as , divided into two parts: the representation of text tokens , the representation of vision tokens , convert the input token sequence into the output target tokens sequence in an autoregressive manner , where L is the length of the output sequence, represents the i-th element in the output sequence, and the formula is:

[0011]

[0012] where is the text token representation, is the vision token representation.

[0013] Further, the image is scaled and cropped into a unified format, cropped at a resolution of 384*384, and the pixel values are normalized.

[0014] Further, the multi-layer perceptron MLP consists of two fully connected layers, and the activation function is GELU.

[0015] The present invention has the following beneficial effects compared with the prior art:

[0016] 1) The present invention adopts the Mamba model based on the state space model, which has a linear computational complexity and performs excellently in efficiently processing long sequences, fast inference, and linear scalability of sequence length;

[0017] 2) The present invention has conducted a large number of experiments on 3 benchmark datasets, and it is proved that the method of the present invention can significantly improve the training speed, inference speed, and overall performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a schematic structural diagram of the present invention;

[0019] Figure 2 is a flowchart of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific embodiments of the present invention will be made with reference to the accompanying drawings. Many specific details are set forth in the following description in order to fully understand the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. The technical features in the various embodiments of the present invention can be combined correspondingly without conflict.

[0021] The specific structure of the present invention is as follows:

[0022] Encoder module: DINOv2 and SigLIP are combined as the visual encoder. Among them, DINOv2 is used to capture the low-level spatial attributes in the input image, and SigLIP is used to capture the semantic attributes in the input image. The feature vectors output by the two are combined and sent to the projection layer for further processing.

[0023] Projector module: A multi-layer perceptron (MLP) is used as the projection layer to reduce the dimension of the feature vectors output from the visual encoder, so that the dimension of the visual feature vectors is consistent with the dimension of the feature vectors of the text tokens.

[0024] The Mamba backbone model uses the Mamba backbone network as the model architecture, which consists of a short convolutional layer, an SSM module, residual connections, and the RM-SNorm regularization method.

[0025] Embodiment

[0026] In this embodiment, a multimodal large model system based on the extension of the Mamba language model is proposed to solve the following problems: Multimodal large models based on the Transformer network are often inefficient, making it difficult to meet the requirements in application scenarios that require high real-time performance and are suitable for edge deployment.

[0027] As Figure 1 shown, this embodiment specifically includes:

[0028] Encoder module: DINOv2 and SigLIP are combined as the visual encoder. Among them, DINOv2 is used to capture the low-level spatial attributes in the input image, and SigLIP is used to capture the semantic attributes in the input image. The feature vectors output by both are combined and sent to the projection layer for further processing.

[0029] It should be noted that in the encoder module, considering the input image , the visual encoder divides the image into equally sized blocks, being the scale of the blocks. Both of these visual encoders take the patched image as the input token sequence and extract the channel connection of the outputs of the two encoders as a compact visual representation :

[0030]

[0031] Projector module: The multi-layer perceptron MLP is used to reduce the dimensionality of the visual tokens output by the visual encoder module to visual tokens, and the formula is , representing the multi-layer perceptron; the purpose of this step is to convert the dimensionality of the visual tokens to the dimensionality required for the input of the next module, and the output result of the projector module will be sent to the next module for processing;

[0032] Mamba backbone model: The 2.8B parameter version pre-trained on the SlimPa-jama dataset or the 7B version pre-trained on the RefinedWeb dataset of the Mamba backbone model is selected, and then the output result of the projector module is concatenated with the output result of the text encoder and denoted as , which is divided into two parts: the representation of text tokens , the representation of visual tokens , converting the input token sequence into the output target token sequence in an autoregressive manner , where L is the length of the output sequence, represents the i-th element in the output sequence, and the formula is:

[0033]

[0034] where is the text token representation, is the visual token representation.

[0035] Refer to Figure 2 , the training and experimental process of the multimodal large model Cobra extended based on the Mamba language model is as follows:

[0036] 1) Dataset

[0037] The following datasets are used: (1) The mixed dataset used in LLaVA v1.5, which contains a total of 655K multi-turn visual-text mixed conversations including academic VQA. (2) The LVIS-Instruct-4V dataset, which contains 220K images generated by GPT-4V, and these images have visually aligned and context-aware internal structures. (3) The LRV-Instruct dataset, which contains 400K visual instructions covering 16 visual language tasks and aims to reduce hallucinations. Overall, the entire dataset contains approximately 1.2 million images, corresponding multi-turn conversation data, and pure text conversation data.

[0038] 2) Evaluation metrics

[0039] Experiments were conducted on nine different benchmarks, including: (1) Four open-ended visual question answering tasks, namely VQA-v2, GQA, VizWiz, and TextVQA. (2) Two closed-ended visual question answering tasks, namely VSR and POPE. (3) Three visual grounding tasks, namely RefCOCO, RefCOCO+, and RefCOCOg. VQA-v2 (Goyal et al., 2017) is used to evaluate the model's understanding and reasoning ability of images and questions. GQA (Hudson and Manning, 2019) focuses on spatial understanding and multi-dimensional reasoning and is tested on real-world images. VizWiz (Gurari et al., 2018) is similar to VQA-v2 but adds some questions without answers to test the model's coping ability. TextVQA (Singh et al., 2019) focuses on reasoning from text information in images. VSR (Liu, Emerson, and Collier, 2023) consists of a series of yes / no questions to examine the model's ability to judge spatial relationships in complex scenes, which poses a challenge to multi-modal large language models (MLLMs). POPE (Li et al., 2023b) consists of specific yes / no questions to evaluate the model's tendency to generate hallucinated content. RefCOCO mainly deals with short descriptions with spatial anchors, RefCOCO+ relies on appearance-based descriptions, and RefCOCOg emphasizes the use of long and rich descriptions (Kazemzadeh et al., 2014; Yu et al., 2016).

[0040] 3) Baseline methods

[0041] Cobra was compared with various algorithms of different scales, including: (1) Large-scale multi-modal large language models (MLLMs): OpenFlamingo (Awadalla et al., 2023), BLIP-2 (Li et al., 2023a), MiniGPT-4 (Zhu et al., 2023), InstructBLIP (Dai et al., 2023), Shikra (Chen et al., 2023a), IDEFICS (Laurençon et al., 2023), Qwen-VL (Bai et al., 2023), and LLaVA v1.5 (Liu et al., 2023b). (2) Small-scale MLLMs: MoE-LLaVA (Lin et al., 2024), LLaVA-Phi (Zhu et al., 2024), and MobileVLM v2 (Chu et al., 2024). Cobra led in all the above-mentioned metrics and inference speed.

[0042] 4) Implementation details

[0043] The training process includes multi-modal instruction tuning, during which the multi-modal projector and the Mamba large language model (LLM) are fine-tuned. The model is trained using 8 NVIDIA A100 80GB GPUs. A variety of open-source model weights are selected, including Mamba with 2.8 billion and 7 billion parameters as the LLM backbone of the model.

[0044] The above-described embodiments are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical fields can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.

Claims

1. A multimodal large model system that integrates vision and language, characterized in that, The system includes: (1) Encoder module, including a visual encoder and a text encoder. The visual encoder includes a visual encoder DINOv2 and a visual encoder SigLIP. Among them, DINOv2 captures low-level spatial information in the image, and SigLIp captures semantic information in the image. First, image chunking: Scale and crop all images into a unified format. For a given image x v ∈R C×H×W , where C is the number of channels of the image, H and W represent the height and width of the image respectively, and R represents the tensor space; Divide the image into N v =(H + W) / P 2 small patches Patch, where P is the side length of each small patch; Both the visual encoder DINOv2 and the visual encoder SigLIP take the chunked image as the input sequence and extract the channel concatenation of the outputs of the two visual encoders as the compact visual where D DINOv2 and D SigLIP represent the feature dimensions of the outputs of the two visual encoders respectively; R v is expressed as: where and represent the feature extraction operations of DINOv2 and SigLIP on the input image respectively. Concatenate the outputs of the two visual encoders to form an output of dimension N v ×(D DINOv2 +D SigLIP ); The text encoder selects GPT-NeoXTokenizer to convert the text into text tokens; (2) Projector module, which uses a multi-layer perceptron (MLP) to reduce the dimensionality of the visual tokens R output by the visual encoder module to the dimensionality of the visual tokens H input to the next module. v The formula is H q = φ(R q ), where φ() represents a multi-layer perceptron; v ​ (3) Mamba backbone model, select the version of the Mamba backbone model with 2.8B parameters pre-trained on the SlimPa-jama dataset or the 7B version pre-trained on the RefinedWeb dataset. Then, concatenate the result output by the projector module and the result output by the text encoder and denote it as H, which is divided into two parts: the representation of text tokens H v , and the representation of visual tokens H q . Convert the input token sequence into the output target tokens sequence in an autoregressive manner where L is the length of the output sequence, and y i represents the i-th element in the output sequence, and the formula is: Among them, H v is the text token representation, and H q is the visual token representation.

2. The multimodal large model system according to claim 1, wherein The image is scaled and cropped into a unified format, cropped at a resolution of 384*384, and the pixel values are normalized.

3. The multimodal large model system according to claim 1, wherein The multi-layer perceptron MLP consists of two fully connected layers, and the activation function is GELU.

Citation Information

Patent Citations

  • Image processing method

    CN116958766A

  • Context awareness medical vision language model pre-training method, system and application

    CN118039056A