Multi-modal visual understanding model based on double-path visual coding, training method, reasoning method and equipment

Through the dual-channel vision encoder structure, combined with global and local visual features, the single-channel vision encoder lacks understanding on high-resolution and dense text images, and achieves stronger image resolution and semantic understanding capabilities.

CN120339798APending Publication Date: 2025-07-18SHANG HAI JIE YUE XING CHEN ZHI NENG KE JI YOU XIAN GONG SI
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510475698.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing single-channel vision encoders have problems with insufficient understanding when processing high-resolution images, dense text images and chart images, especially in high-precision visual tasks such as image comprehension, document analysis and chart interpretation.

Method used

Using a dual-channel vision encoder structure, the first visual encoder extracts the global visual features of natural universal images and freezes the pre-training weights. The second visual encoder processes the local details of high-resolution images, splices global and local features in the channel dimension through the feature fusion layer, and converts them into a large language model input format through the linear layer, and enhances multi-scale feature expression by combining convolutional neural networks and feature pyramid networks.

Benefits of technology

The multimodal visual understanding model's ability to understand high-resolution images, dense text and graph data is improved, the alignment of visual features and large language models is optimized, and the application performance of the model in specific scenarios is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339798A_ABST
    Figure CN120339798A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal visual understanding model based on double-path visual coding, a training method, a reasoning method and equipment. The model comprises a first visual encoder used for extracting global visual features of a natural general image and outputting first image features, and weight freezing of the first visual encoder; the input of the second visual encoder is an image of which the size is adjusted to a preset high resolution, and the second visual encoder is used for extracting local detail information of the image and outputting a second image feature; the feature fusion layer is used for splicing the first image feature and the second image feature in a channel dimension to form a fused visual feature; the linear layer is used for converting the fused visual features into input dimensions required by a large language model; the large language model is used for generating natural language answers based on the fused visual features after dimension conversion and text input. According to the method, a double-path visual coding structure is adopted, the image analysis capability of the multi-modal visual understanding model is improved, and the alignment mode of the visual features and the large language model is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of visual understanding artificial intelligence technology, and in particular to a multimodal visual understanding model, training method, reasoning method and device based on dual-path visual coding. Background Art

[0002] In recent years, multimodal visual understanding models have made significant progress in the fields of computer vision and natural language processing. Existing large multimodal models widely use single-channel visual encoder Clip-ViT, such as BLIP-2, LLaVA, Mini-GPT4, mPLUG-OWL, etc. Clip-ViT pre-trains data on large-scale images and texts, and has strong image-text alignment capabilities, especially in natural general images (such as landscapes, people, objects, city street scenes, etc.). Therefore, Clip-ViT can effectively reduce the training difficulty and computational cost of large multimodal models, allowing large language models to efficiently learn image-text association information, thereby achieving tasks such as text generation, image description, and visual question answering (VQA).

[0003] However, although Clip-ViT performs well in general visual scenarios, existing multimodal large models still have significant limitations in certain types of image understanding tasks, such as high-resolution images, dense text images, and chart images. In these scenarios, the existing models' understanding ability is insufficient, which affects their application in image understanding, document parsing, chart interpretation, and high-precision visual tasks.

[0004] The main reason for the above problems is that the single-channel visual encoder Clip-ViT has insufficient upper limits, which is specifically reflected in the following two aspects:

[0005] On the one hand, the input resolution is limited: Clip-ViT uses a fixed input resolution of 224×224. When processing large-resolution images, the image needs to be scaled, which may lead to the loss of key information. If the input resolution is increased, the number of image tokens will increase significantly, causing the computational overhead to grow exponentially, thus affecting the model's reasoning efficiency.

[0006] Another aspect is the limitation of training data: Clip-ViT is mainly trained on text and images of natural scenes, but less trained on dense text images (such as document scans, invoices, bills) and chart images (such as line graphs, bar graphs, pie charts), which makes it difficult to accurately parse high-density text information and chart structures. Directly fine-tuning Clip-ViT with new training data may lead to catastrophic forgetting, that is, the loss of the original general visual knowledge. Summary of the invention

[0007] In view of the deficiencies in the prior art, the present application provides a multi-modal visual understanding model, training method, inference method and device based on dual-path visual coding, so as to at least solve the problem of insufficient ability of existing single-path visual encoders in high-resolution image parsing.

[0008] To achieve the above objectives and other advantages, the present application is implemented by the following technical solutions:

[0009] In a first aspect, the present application provides a multi-modal visual understanding model based on dual-path visual coding, including:

[0010] A first visual encoder, configured to extract global visual features of natural general images and output first image features, and the weights of the first visual encoder are frozen to keep the pre-training results unchanged;

[0011] A second visual encoder, the input of the second visual encoder is an image whose size is adjusted to a preset high resolution, and is configured to extract local detail information of the input image and output second image features;

[0012] A feature fusion layer, configured to splice the first image features and the second image features in the channel dimension to form fused visual features;

[0013] A linear layer, configured to convert the fused visual features into the input dimension required by the large language model;

[0014] A large language model, configured to generate natural language answers based on the fused visual features and text inputs after dimension conversion.

[0015] According to the multi-modal visual understanding model based on dual-path visual coding provided by the present application, the number of Tokens of the first image features and the second image features is the same to ensure matching with the input dimension of the feature fusion layer;

[0016] The first image features are composed of a first preset number of first Tokens, and each of the first Tokens corresponds to a first feature dimension;

[0017] After being flattened, the second image features are composed of second Tokens with the same number as the first preset number to ensure that the number of Tokens output by the first visual encoder and the second visual encoder is consistent, and each of the second Tokens corresponds to a second feature dimension;

[0018] The first image features and the second image features are spliced in the channel dimension to form the fused visual features obtained by multiplying the first preset number by the target feature dimension, where the target feature dimension is the sum of the first feature dimension and the second feature dimension.

[0019] A multimodal visual understanding model based on dual - path visual coding provided by the present application, wherein the second visual encoder includes a convolutional neural network and a feature pyramid network;

[0020] The convolutional neural network adopts a deep convolutional neural network structure to extract multi - scale image features and provide a basic feature map for the feature pyramid network;

[0021] The feature pyramid network adopts a bottom - up, lateral connection, and top - down feature fusion method to enhance the multi - scale feature expression ability and generate multiple feature levels including low - level, intermediate - level, and high - level;

[0022] The output of the high - level feature level is obtained by downsampling the feature map output by the previous feature level to extract global semantic information and enhance the parsing ability of high - resolution images.

[0023] A multimodal visual understanding model based on dual - path visual coding provided by the present application, the input of the second visual encoder is an image with a resolution adjusted to 1024×1024, the downsampling ratio of the high - level feature level is 64, the shape of the feature map output by the high - level feature level is 16×16×1024, and the feature map output by the high - level feature level is flattened through a flattening operation, flattening the 16×16 dimension to 256 dimensions to form a 256×1024 feature representation.

[0024] A multimodal visual understanding model based on dual - path visual coding provided by the present application, the first visual encoder adopts a visual coding network based on the Transformer structure and converts the input image into the first image feature through the following steps:

[0025] Receive an input image with a size of 224×224, and perform a fixed - size segmentation process on the image, dividing the image into 16×16 patches to form 196 patch units;

[0026] Perform a linear mapping on the patch units to transform each patch into a 768 - dimensional feature vector to generate an initial visual feature representation of 196×768 dimensions;

[0027] On the basis of the initial visual feature representation, add an additional Token as a global feature identifier to form a feature sequence of 197×768;

[0028] Through the Transformer encoding module, perform multi - layer attention calculations on the 197×768 feature sequence to extract global visual features;

[0029] Using the interpolation expansion method, the visual features of 197×768 after Transformer encoding calculation are adjusted to the Token representation of 256×768 to match the Token structure output by the second visual encoder;

[0030] The first image features of 256×768 are used to concatenate with the second image features of 256×1024 output by the second visual encoder in the channel dimension to form fused visual features.

[0031] According to a multi-modal visual understanding model based on dual-path visual encoding provided by the present application, the linear layer includes a multi-layer perceptron network for transforming the dimension of the fused visual features from 1792 to 4096 to match the input format of the large language model.

[0032] In a second aspect, the present application provides a training method for a multi-modal visual understanding model based on dual-path visual encoding, and the training method includes:

[0033] Initializing the weights of the second visual encoder pre-trained based on the ImageNet dataset, and connecting the second visual encoder to a small language model for warm-up training;

[0034] After completing the warm-up training, connecting the second visual encoder to the complete multi-modal visual understanding model, and using the image-text pair dataset to perform end-to-end pre-training on the multi-modal visual understanding model. In the end-to-end pre-training stage, only optimize the parameters of the linear layer and the second visual encoder, and freeze the weights of the first visual encoder and the large language model;

[0035] Based on the task-specific supervised fine-tuning dataset, only perform fine-tuning training on the linear layer and the large language model.

[0036] In a third aspect, the present application provides an inference method for a multi-modal visual understanding model based on dual-path visual encoding. Deploy the trained multi-modal visual understanding model to an inference device, and the inference method includes:

[0037] Receiving the multi-modal data input by the user, where the multi-modal data includes image data and text data related to the image;

[0038] Inputting the multi-modal data into the multi-modal visual understanding model for inference calculation, and obtaining the inference output of the multi-modal visual understanding model, where the inference output includes a natural language answer related to the multi-modal data;

[0039] Transmitting the inference output to the user interaction interface to support natural language interaction, and performing multi-round inference processing based on the subsequent input of the user.

[0040] Fourthly, the present application provides an electronic device, which includes:

[0041] One or more processors; and a memory storing computer program instructions, which when executed, cause the processors to execute the training method of the multi-modal visual understanding model based on dual-path visual coding as described above, or execute the inference method of the multi-modal visual understanding model based on dual-path visual coding as described above.

[0042] Fifthly, the present application provides a computer-readable storage medium, on which computer programs / instructions are stored, and when the computer programs / instructions are executed by a processor, the training method of the multi-modal visual understanding model based on dual-path visual coding as described above is implemented, or the inference method of the multi-modal visual understanding model based on dual-path visual coding as described above is implemented.

[0043] Sixthly, the present application provides a computer program product, including computer programs / instructions, characterized in that when the computer programs / instructions are executed by a processor, the training method of the multi-modal visual understanding model based on dual-path visual coding as described above is implemented, or the inference method of the multi-modal visual understanding model based on dual-path visual coding as described above is implemented.

[0044] The multi-modal visual understanding model, training method, inference method and device based on dual-path visual coding provided by the present application extract global and local visual features by combining two different visual encoders and fuse them in the channel dimension to enhance the model's understanding ability of high-resolution images, dense texts and chart data. The first visual encoder maintains pre-trained general visual features to ensure that the model has the cognitive ability of natural images, while the second visual encoder focuses on extracting details of high-resolution inputs to make up for the limitations of single-path encoders in specific scenarios. The fused visual features are linearly mapped to adapt to the input requirements of large language models, thereby optimizing the utilization efficiency of visual information in multi-modal tasks. Finally, the model can reason based on the input images and texts and generate natural language answers that conform to semantic understanding. The present application adopts a dual-path visual coding structure, improves the image parsing ability of the multi-modal visual understanding model, optimizes the alignment method of visual features with large language models, and enhances the application performance of the model in scenarios of high resolution, dense text and chart data. Description of the Drawings

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other implementation manners can be obtained based on these drawings without creative efforts.

[0046] Figure 1 It is a logical schematic diagram of a multi-modal visual understanding model based on dual-path visual coding provided by an embodiment of the present application;

[0047] Figure 2 It is a flowchart of a training method for a multi-modal visual understanding model based on dual-path visual coding provided by an embodiment of the present application;

[0048] Figure 3 It is a flowchart of an inference method for a multi-modal visual understanding model based on dual-path visual coding provided by an embodiment of the present application;

[0049] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0050] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the drawings, details are described as follows.

[0051] It should be noted that those of ordinary skill in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments without conflict. Unless otherwise defined, the technical terms or scientific terms involved in the present application should be the general meanings understood by those with ordinary skills in the technical field to which the present application belongs. The terms "a", "one", "kind", "the" and other similar words involved in the present application do not represent a limitation in quantity and can represent a single or plural. The terms "including", "comprising", "having" and any variations thereof involved in the present application are intended to cover non-exclusive inclusion; the terms "first", "second", "third", etc. involved in the present application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0052] For the convenience of understanding the embodiments of the present application, the following are the explanations of the key terms / technical abbreviations in the present application:

[0053] Clip-ViT refers to using Vision Transformer (ViT) in the CLIP (Contrastive Language-Image Pretraining) model as a visual encoder. It is a single-path visual encoder, that is, only ViT is used as an image feature extractor, without an additional high-resolution processing module or multi-channel fusion mechanism.

[0054] CNN (Convolutional Neural Network, Convolutional Neural Network), CNN is a deep learning model used for image processing and feature extraction. It can extract features such as edges, textures, and shapes in images layer by layer through operations such as convolution (Convolution) and pooling (Pooling).

[0055] FPN (Feature Pyramid Network, Feature Pyramid Network), FPN is a multi-scale feature extraction network that can combine feature maps of different resolutions to improve the model's ability to recognize objects of different scales.

[0056] The linear layer (fully connected layer) is a neural network layer used for feature transformation, usually used to adjust the dimension of the input features to meet the requirements of subsequent calculation modules.

[0057] Large language model (LLM, Large Language Model), a large language model is a deep learning model based on the Transformer structure and trained on a large-scale corpus, capable of performing tasks such as text generation, question answering, and reasoning.

[0058] Warm-up Training, Warm-up Training is a training method that gradually adjusts the model parameters, usually used to reduce the training calculation overhead and improve the training stability.

[0059] End-to-End Training means optimizing multiple modules of the entire model in the same training process, without the need to train each part separately and then combine them.

[0060] Supervised Fine-Tuning (SFT), SFT is a strategy for fine-tuning the model based on a task-specific dataset, usually used to improve the performance of the pre-trained model on specific tasks.

[0061] In a deep learning model, Token refers to the smallest processing unit of the input data, which can be used to represent text (such as words or sub-words in NLP tasks) or image features (such as ViT divides an image into patches and converts them into Tokens).

[0062] Refer to Figure 1 As shown, the embodiments of the present application provide a multi-modal visual understanding model based on dual-channel visual coding, including:

[0063] The first visual encoder is used to extract the global visual features of natural general images and output the first image features. The weights of the first visual encoder are frozen to keep the pre-trained results unchanged;

[0064] A second visual encoder, the input of the second visual encoder being an image resized to a preset high resolution, for extracting local detail information of the input image and outputting second image features;

[0065] A feature fusion layer for concatenating the first image features and the second image features in the channel dimension to form fused visual features;

[0066] A linear layer for converting the fused visual features into the input dimension required by the large language model;

[0067] A large language model for generating natural language answers based on the fused visual features and text inputs after dimension conversion.

[0068] Specifically, the first visual encoder adopting the Clip-ViT structure is used to extract the global visual features of natural general images and generate first image features. The size of the input image of the first visual encoder is fixed (such as 224×224), and the image is processed through the Transformer structure, divided into multiple Patches, and mapped into a high-dimensional Token structure. ViT outputs a set of visual features in the form of Tokens, and each Token represents the embedding information of an image Patch. Since the first visual encoder has been pre-trained on a large-scale image-text pair data, its weights are frozen in this model to ensure that the general visual features are not lost and to avoid catastrophic forgetting caused by new task training. In this way, the first visual encoder can continue to provide stable global visual features without being affected by the training data.

[0069] Exemplarily, the first visual encoder can adopt the Clip-ViT-L variant of the Clip-ViT structure, where "L" represents the Large version. ViT-L has more parameters, a larger amount of computation, and higher visual feature extraction capabilities.

[0070] The second visual encoder is used to process high-resolution image inputs, focusing on the extraction of local detail information and generating second image features. The input image size is adjusted to a preset high resolution (such as 1024×1024) to ensure sufficient detail information is retained in tasks such as dense text and chart parsing. Since images from different sources may have different resolutions, to maintain the consistency of the input format, all input images are adjusted to the preset high resolution.

[0071] The first visual encoder is good at processing natural images and provides global visual information; the second visual encoder is good at high-resolution image parsing and supplements local detail information. The feature fusion layer is used to concatenate the first image feature output by the first visual encoder and the second image feature output by the second visual encoder in the channel dimension to form a fused visual feature. This concatenation method ensures that the global feature and local detail information are complementarily fused. The linear layer converts the fused visual feature into the input dimension required by the large language model, ensuring that the visual information can be fully utilized by the large language model. The large language model performs natural language understanding and generation based on the transformed fused visual feature and text input, and finally outputs a semantically reasonable answer.

[0072] The multimodal visual understanding model provided in this embodiment combines two different visual encoders to extract global and local visual features and fuses them in the channel dimension to enhance the model's understanding ability of high-resolution images, dense text, and chart data. The first visual encoder maintains the pre-trained general visual features to ensure that the model has the cognitive ability of natural images, while the second visual encoder focuses on the detail extraction of high-resolution inputs to make up for the limitations of a single-path encoder in specific scenarios. The fused visual feature is linearly mapped to adapt to the input requirements of the large language model, thereby optimizing the utilization efficiency of visual information in multimodal tasks. Finally, the model can reason based on the input images and texts and generate natural language answers that conform to semantic understanding. Therefore, this application improves the image parsing ability of the multimodal visual understanding model by adopting a dual-path visual coding structure, optimizes the alignment method of visual features and the large language model, and enhances the application performance of the model in scenarios of high resolution, dense text, and chart data.

[0073] In this embodiment, the number of tokens of the first image feature and the second image feature is the same to ensure matching with the input dimension of the feature fusion layer;

[0074] The first image feature is composed of a first preset number of first tokens, and each first token corresponds to a first feature dimension;

[0075] After being flattened, the second image feature is composed of the same number of second tokens as the first preset number to ensure that the number of tokens output by the first visual encoder and the second visual encoder is consistent, and each second token corresponds to a second feature dimension;

[0076] The first image feature and the second image feature are concatenated in the channel dimension to form a fused visual feature obtained by multiplying the first preset number by the target feature dimension, where the target feature dimension is the sum of the first feature dimension and the second feature dimension.

[0077] Specifically, the first vision encoder is based on the ViT structure, receives the input image, and converts it into a fixed number of first Tokens as output. Suppose the first image features output by the first vision encoder, whose Token structure is the first preset number × the first feature dimension, which contains the first preset number of first Tokens, and the feature dimension corresponding to each first Token is the first feature dimension.

[0078] The second vision encoder is used to process high-resolution images and convert them into a fixed number of second Tokens as output. The final output second image features have a Token structure of the first preset number × the second feature dimension. That is, the second vision encoder generates the same number of second Tokens as the first vision encoder, and the feature dimension corresponding to each second Token is the second feature dimension. The first preset number remains unchanged throughout the inference process to ensure that the number of Tokens output by the two vision encoders is the same, so that their dimensions match during feature fusion.

[0079] By designing the vision encoder in this way, it is ensured that the feature fusion layer can smoothly splice the data and avoid the problem of image shape mismatch. The first vision encoder retains global visual information, and the second vision encoder enhances local detail information. After the data is spliced, the fused Token structure adapts to the input of the large language model, improving the visual understanding ability of multi-modal tasks.

[0080] In this embodiment, the second vision encoder includes a convolutional neural network and a feature pyramid network;

[0081] The convolutional neural network adopts a deep convolutional neural network structure to extract multi-scale image features and provide a basic feature map for the feature pyramid network;

[0082] The feature pyramid network adopts a feature fusion method of bottom-up, lateral connection, and top-down to enhance the multi-scale feature expression ability and generate multiple feature levels including low-level, intermediate-level, and high-level;

[0083] The output of the high-level feature level is obtained by downsampling the feature map output by the previous feature level to extract global semantic information and enhance the parsing ability for high-resolution images.

[0084] It should be noted that the convolutional neural network (CNN) is responsible for extracting the basic feature maps from the input images and providing multi-level feature information for the subsequent Feature Pyramid Network (FPN) processing. Exemplarily, the convolutional neural network adopts a deep convolutional neural network structure (such as ResNet-50), and through multiple convolutional layers and pooling layers, gradually extracts low-level to high-level visual features. For example, first, the initial convolutional layer extracts low-level visual features such as edges and textures, then multiple residual blocks are used for feature transformation to obtain richer mid-scale information. Next, the extracted features pay more attention to semantic information to provide a deeper context representation. Finally, the CNN outputs a multi-scale feature map for further processing by the FPN.

[0085] The FPN adopts a feature fusion method of bottom-up, lateral connection, and top-down to ensure the effective integration of low-level and high-level features and improve the adaptability of the model to targets of different scales. The bottom-up path extracts the original features and downsamples layer by layer. The top-down path upsamples the features through deconvolution and combines them with the low-level features to form a multi-scale feature pyramid. The lateral connection ensures that the high-level features can retain the low-level detail information through skip connections, improving the information flow ability. In the FPN network structure, it is generally named by levels, such as P2, P3, P4, P5, P6, etc. Among them, the P2 layer is the low-level layer (with a higher resolution and rich edge and texture detail information), the P3, P4, and P5 layers are the intermediate layers (gradually extracting higher-level semantic information), and the P6 layer is the high-level layer (the feature map has undergone a larger downsampling and contains more abstract global information). Compared with the intermediate layers, the P6 layer contains higher-level semantic information and is suitable for multi-modal task recognition and parsing.

[0086] As an example, the input of the second visual encoder is an image resized to a resolution of 1024×1024. The downsampling ratio of the high-level feature level is 64, and the shape of the feature map output by the high-level feature level is 16×16×1024. The feature map output by the high-level feature level is flattened through a flattening operation, flattening the 16×16 dimension to 256 dimensions to form a feature representation of 256×1024.

[0087] Specifically, the original image is adjusted to a resolution of 1024×1024 to retain more detailed information and adapt to high-resolution image parsing tasks. The second visual encoder adopts a CNN structure, specifically ResNet-50 including an FPN network. The FPN network has the ability to extract multi-scale features, and the P6 layer has a downsampling rate of 64 times. For the input image with a resolution of 1024×1024, after CNN feature extraction, the size of the feature map output by the P6 layer becomes 16×16×1024 (16×16 represents the spatial dimension, and 1024 represents the number of feature channels). The 0th and 1st dimensions of the flattened feature map are flattened, and the 16×16 dimension is flattened into 256, that is, 256 Tokens. Each second Token corresponds to a second feature dimension with 1024-dimensional feature information, so as to form an output of image features with a Token structure of 256×1024.

[0088] As an example, the first visual encoder adopts a vision coding network based on the Transformer structure, and converts the input image into the first image features through the following steps:

[0089] Receive an input image with a size of 224×224, and perform a fixed-size segmentation process on the image, dividing the image into 16×16-sized Patches to form 196 Patch units;

[0090] Perform a linear mapping on the Patch units, and transform each Patch into a 768-dimensional feature vector to generate an initial visual feature representation of 196×768 dimensions;

[0091] On the basis of the initial visual feature representation, add additional Tokens as global feature identifiers to form a feature sequence of 197×768;

[0092] Through the Transformer encoding module, perform multi-layer attention calculations on the 197×768 feature sequence to extract global visual features;

[0093] Adopt an interpolation expansion method to adjust the 197×768 visual features after Transformer encoding calculation to a 256×768 Token representation to match the Token structure output by the second visual encoder;

[0094] The first image features of 256×768 are used to be concatenated with the second image features of 256×1024 output by the second visual encoder in the channel dimension to form fused visual features.

[0095] Specifically, the first visual encoder processes the input image in a fixed block manner. The input size of the image is 224×224 (usually with three RGB channels). Before entering the Transformer, the image is divided into small patches, with the patch size being 16×16. ViT uses fixed-size patches and divides the 224×224 image into (224 / 16)×(224 / 16) = 14×14 = 196, that is, 196 patches. Each patch has 16×16×3 = 768-dimensional features. Each 16×16×3 patch is linearly projected and converted into a 768-dimensional vector. In this way, the input of ViT becomes a 196×768 token sequence, that is, the first feature dimension corresponding to each first token has 768-dimensional feature information. An additional [CLS] (classification token) is added to capture global information, and finally a 197×768 token sequence is formed. In multimodal applications, the model will additionally expand the number of tokens, usually padding to 256 first tokens to ensure consistency with the number of output tokens of the second visual encoder, thereby supporting the channel splicing of the feature fusion layer.

[0096] The 256×768 feature map output by the first visual encoder and the 256×1024 feature map output by the second visual encoder are spliced and fused in the channel dimension. Among them, the first preset number = 256, the target feature dimension = the first feature dimension + the second feature dimension = 768 + 1024 = 1792, and the final shape of the fused visual feature is 256×1792. The feature fusion layer is used to splice the visual features output by the first visual encoder and the second visual encoder in the channel dimension, ensuring information complementarity of cross-modal features and improving the integrity of visual information.

[0097] As an example, the linear layer includes a multi-layer perceptron network, which is used to transform the dimension of the fused visual feature from 1792 to 4096 to match the input format of the large language model.

[0098] Specifically, in multimodal tasks, visual features need to be fused with text features of the large language model in the same dimensional space to ensure semantic alignment and computational compatibility. The input of the linear layer is the fused visual feature, which is spliced from the outputs of the first visual encoder and the second visual encoder. A multi-layer perceptron (MLP) is used for linear transformation. For example, at least one fully connected layer is used for preliminary feature mapping, and the ReLU (or GELU) is used as the activation function of the fully connected layer to enhance the feature expression ability; then LayerNorm or BatchNorm is used for normalization to ensure feature stability and improve the generalization ability of the large language model.

[0099] Exemplarily, the large language model can be the LLaMA-7B model. Compared with larger-scale LLMs, LLaMA-7B is more suitable for deployment in environments with limited computing resources (such as edge computing, private servers) and can perform local inference. LLaMA-7B uses 4096-dimensional embedding vectors, that is, each input Token is mapped to a 4096-dimensional feature vector. LLaMA-7B requires 4096-dimensional input Tokens, while the image feature Tokens output by the dual-path visual encoder have a dimension of 1792. Therefore, a linear transformation layer is needed to complete the dimension conversion. The input dimension of the linear layer is 1792, and the output dimension is 4096. The linear layer uses a fully connected neural network for feature transformation and finally outputs visual features of 256×4096 to meet the input format requirements of the large language model. This feature will be combined with text Tokens and input into the large language model for cross-modal inference.

[0100] In summary, the multi-modal visual understanding model based on dual-path visual coding provided in this embodiment combines two different visual encoders to extract global and local visual features and fuse them in the channel dimension to enhance the model's understanding ability of high-resolution images, dense text, and chart data. The first visual encoder maintains pre-trained general visual features to ensure that the model has the cognitive ability of natural images, while the second visual encoder focuses on extracting details of high-resolution inputs to make up for the limitations of single-path encoders in specific scenarios. The fused visual features are linearly mapped to adapt to the input requirements of the large language model, thereby optimizing the utilization efficiency of visual information in multi-modal tasks. Finally, the model can perform inference based on the input images and texts and generate natural language answers that conform to semantic understanding. This application adopts a dual-path visual coding structure, improves the image parsing ability of the multi-modal visual understanding model, optimizes the alignment method of visual features with the large language model, and enhances the application performance of the model in scenarios of high resolution, dense text, and chart data.

[0101] Refer to Figure 2 As shown, the embodiment of the present application also provides a training method for a multi-modal visual understanding model based on dual-path visual coding. The training method includes:

[0102] Step A1: Initialize the second visual encoder with the weights pre-trained on the ImageNet dataset, and connect the second visual encoder to a small language model for warm-up training;

[0103] Step A2: After completing the warm-up training, connect the second visual encoder to the complete multi-modal visual understanding model, and use the image-text pair dataset to perform end-to-end pre-training on the multi-modal visual understanding model. In the end-to-end pre-training stage, only optimize the parameters of the linear layer and the second visual encoder, and freeze the weights of the first visual encoder and the large language model;

[0104] Step A3: Based on the task-specific supervised fine-tuning dataset, only fine-tune the linear layer and the large language model.

[0105] Specifically, the training method includes three stages: warm-up training, end-to-end pre-training, and supervised fine-tuning, to optimize the model's visual feature extraction ability and improve the inference performance of multi-modal tasks.

[0106] Since the second visual encoder uses the pre-trained weights recorded in the ImageNet dataset during pre-training, it is mainly applicable to natural image classification tasks, but has insufficient generalization ability for tasks such as dense text and chart data. And directly connecting to the LLaMA-7B model for training has a large computational cost. Therefore, first connect the second visual encoder to a small language model (such as OPT-125M) for autoregressive warm-up training, aiming to let the second visual encoder adapt to the task on a small model with lower computational cost first, and then connect to a large model for complete training. This can reduce the computational burden and enable it to extract stable visual features on dense text and chart data.

[0107] Specifically, for the input image in the autoregressive warm-up training, the second visual encoder extracts visual features and generates a token sequence. OPT-125M processes the visual tokens and text tokens in an autoregressive manner and predicts the next token. The model continuously uses its own prediction results to generate subsequent outputs, forming an autoregressive decoding mechanism. This ensures that the visual encoder can learn autoregressive characteristics and improve its adaptability in the LLaMA-7B model training.

[0108] After completing the warm-up training of the second visual encoder, connect it to the complete multi-modal visual understanding model, and use it together with the first visual encoder as the input processing module. Use the image-text pair dataset to perform end-to-end pre-training on the overall model, so that the model can learn the matching relationship between visual features and text descriptions. Since the first visual encoder has been pre-trained on a large-scale image-text pair data and already has general visual understanding ability, there is no need to adjust the parameters. As a large language model, LLaMA-7B is not fine-tuned in this stage to ensure that its existing natural language generation ability is not affected. Therefore, during the training process, only optimize the parameters of the linear layer (MLP) and the second visual encoder, and keep the weights of the first visual encoder and the large language model frozen to ensure that the pre-trained visual knowledge and language ability are not damaged. This training method can accelerate convergence and avoid catastrophic forgetting.

[0109] After end-to-end pre-training, the model has basic multi-modal vision-language understanding capabilities, but the performance in specific tasks (such as visual question answering, table parsing) still needs to be optimized. A supervised fine-tuning dataset is used to fine-tune the linear layer and the large language model, enabling the model to have stronger task adaptation capabilities. Select a dataset suitable for a specific task, such as datasets for visual question answering, table parsing, document understanding, etc. This dataset provides image inputs, task instructions, and target text outputs to enhance the multi-modal reasoning capabilities of the model. During the training process, the weights of the first vision encoder and the second vision encoder are kept frozen, and only the parameters of the linear layer and the large language model are adjusted. For the linear layer, it further optimizes the mapping method of the linear layer's visual features. For the large language model, it improves its understanding of visual inputs and enables it to generate more semantically consistent text answers.

[0110] Therefore, on the premise of reducing the computational overhead, the vision encoder, linear layer, and large language model are gradually optimized to enable the model to adapt to multi-modal tasks such as high-resolution images, dense text, and chart data, enhancing the generalization ability of the model's reasoning tasks.

[0111] Refer to Figure 3 As shown, the embodiment of the present application also provides a reasoning method for a multi-modal vision understanding model based on dual-path vision coding. The trained multi-modal vision understanding model is deployed to an inference device. The reasoning method includes:

[0112] Step B1: Receive multi-modal data input by the user, where the multi-modal data includes image data and text data related to the image;

[0113] Step B2: Input the multi-modal data into the multi-modal vision understanding model for inference calculation, and obtain the inference output of the multi-modal vision understanding model. The inference output includes a natural language answer related to the multi-modal data;

[0114] Step B3: Transmit the inference output to the user interaction interface to support natural language interaction, and perform multi-round inference processing based on the user's subsequent input.

[0115] Specifically, the multi-modal data can be a single image combined with relevant text, or continuous input interactive data. Step B1 is used to receive the image data and text data input by the user and convert them into a format that can be processed by the model. For example, adjust to a standardized size of 1024×1024 resolution and convert it into a format adapted to the model.

[0116] The first visual encoder in the dual-channel visual coding architecture extracts global visual features, and the second visual encoder processes high-resolution images to extract local detailed information. The image data is converted into a Token sequence to adapt to subsequent fusion processing. The model processes the text data, performs word segmentation and encoding, and converts it into a Token sequence. The visual Token sequence and the text Token sequence after feature extraction are concatenated in the feature fusion layer to form a complete fusion feature. The linear layer converts the fusion feature into a dimension that can be processed by the large language model. The large language model performs autoregressive inference, gradually generating answer Tokens in an autoregressive manner, and the generated text answer contains semantic information related to the input image and text.

[0117] The natural language answer generated by the model is transmitted to the user interaction interface (such as the ChatGPT dialog box, the APP application interface, or generates an API return value for other software to call through the API interface). The answer content can also be played by means of speech synthesis. The user can intuitively obtain the answer generated by the model to achieve an AI interaction experience. The model supports subsequent multi-round question and answer to enhance the interaction experience. If the user inputs a new question, the system will return to step B1, receive new input again, and perform a new round of reasoning. This process relies on the long context modeling ability of LLaMA-7B to ensure the coherence of the conversation.

[0118] For example, the user uploads a chart and enters the text question "Please describe the content of this chart". The model answers: "This chart shows the sales of different products, among which the sales of product A are the highest, accounting for 45%..." The user can ask further questions based on the first-round answer of the model. For example, "Please provide the detailed data of product A", and the model answers: "The sales of product A account for 30% of the total sales, 15% lower than that of product B..."

[0119] Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0120] In addition, some embodiments of the present application also provide an electronic device. The electronic device can be various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and so on. The electronic device can also be various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices.

[0121] The electronic device includes: one or more processors; and a memory storing computer program instructions that, when executed, cause the processors to execute the training method or the inference method of the multi-modal visual understanding model based on dual-channel visual coding provided in any one or more of the above embodiments. Figure 4 An exemplary structural diagram of the electronic device is disclosed. As Figure 4 shown, the electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component is interconnected using different buses and can be mounted on a common motherboard or otherwise installed as required. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory for displaying graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used in conjunction with multiple memories and multiple memories if needed. Similarly, multiple electronic devices can be connected, with each device providing some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Among them, the components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described herein and / or claimed.

[0122] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103, and the output device 1104 can be connected via a bus or other means, Figure 4 and the example of connection via a bus is taken herein.

[0123] The input device 1103 can receive input digital or character information and generate key signal inputs related to user settings and function controls of the electronic device, such as input devices like a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 1104 can include a display device, an auxiliary lighting device (such as an LED), and a haptic feedback device (such as a vibration motor), etc. The display device can include, but is not limited to, a liquid crystal display (LCD), a light-emitting diode (LED) display, and a plasma display. In some embodiments, the display device can be a touch screen.

[0124] To provide interaction with a user, the electronic device may be a computer. The computer has: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or an LCD monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0125] In the embodiments of the present application, a computer program / instructions is stored on a computer-readable medium. When the computer program / instructions are executed by a processor, the training method or the inference method of the multi-modal visual understanding model based on dual-channel visual coding provided by any one or more of the above embodiments is implemented. The computer-readable medium may be included in the electronic device described in the above embodiments; or may exist separately without being assembled into the device. The above computer-readable medium carries one or more computer-readable instructions.

[0126] The memory 1102 may be used as a non-transitory computer-readable storage medium for storing non-transitory software programs, non-transitory computer-executable programs, and modules. By running the non-transitory software programs, instructions, and modules stored in the memory 1102, the processor 1101 executes various functional applications and data processing of the server to implement the program instructions / modules corresponding to the methods provided by any one or more of the above embodiments in the embodiments of the present application.

[0127] The memory 1102 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 1102 may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 1102 may optionally include a memory remotely provided with respect to the processor 1101, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0128] Note that more specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0129] Computer-readable storage media include both permanent and non-permanent, removable and non-removable media, and can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0130] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., connected through the Internet using an Internet service provider).

[0131] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. For example, an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device can be used. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. In addition, some steps or functions of this application can be implemented by hardware, for example, as a circuit that cooperates with a processor to execute each step or function.

[0132] The computer program product provided by the embodiments of the present application includes one or more computer programs / instructions. When the computer programs / instructions are executed by a processor, they wholly or partly generate the processes or functions described in the embodiments of the present application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that integrates one or more available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive (SSD)).

[0133] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0134] As described above, the foregoing are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily make changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A multi-modal visual understanding model based on dual-channel visual coding, characterized in that, Including: A first visual encoder for extracting global visual features of a natural general image and outputting a first image feature, where the weights of the first visual encoder are frozen to keep the pre-trained result unchanged; A second visual encoder, where the input of the second visual encoder is an image resized to a preset high resolution, for extracting local detail information of the input image and outputting a second image feature; A feature fusion layer for concatenating the first image feature and the second image feature in the channel dimension to form a fused visual feature; A linear layer for converting the fused visual feature into the input dimension required by the large language model; A large language model for generating natural language answers based on the fused visual feature after dimension conversion and text input.

2. The multimodal visual understanding model based on dual-channel visual coding according to claim 1, wherein The number of tokens of the first image feature and the second image feature is the same to ensure matching with the input dimension of the feature fusion layer; The first image feature consists of a first preset number of first tokens, and each first token corresponds to a first feature dimension; After being flattened, the second image feature consists of the same number of second tokens as the first preset number to ensure that the number of tokens output by the first visual encoder and the second visual encoder is consistent, and each second token corresponds to a second feature dimension; The first image feature and the second image feature are concatenated in the channel dimension to form the fused visual feature obtained by multiplying the first preset number by the target feature dimension, where the target feature dimension is the sum of the first feature dimension and the second feature dimension.

3. The multimodal visual understanding model based on dual-channel visual coding according to claim 1, wherein The second visual encoder includes a convolutional neural network and a feature pyramid network; The convolutional neural network adopts a deep convolutional neural network structure to extract multi-scale image features and provide a basic feature map for the feature pyramid network; The feature pyramid network adopts a feature fusion method of bottom-up, lateral connection, and top-down to enhance the multi-scale feature expression ability and generate multiple feature levels including low-level, intermediate-level, and high-level; The output of the high-level feature level is obtained by downsampling the feature map output by the previous feature level to extract global semantic information and enhance the parsing ability for high-resolution images.

4. The multi-modal visual understanding model based on dual-channel visual coding according to claim 3, wherein The input of the second visual encoder is an image resized to a resolution of 1024×1024, the downsampling ratio of the high-level feature level is 64, the shape of the feature map output by the high-level feature level is 16×16×1024, and the feature map output by the high-level feature level is flattened through a flattening operation to flatten the 16×16 dimension to 256 dimensions, forming a feature representation of 256×1024.

5. The multimodal visual understanding model based on dual-channel visual coding according to claim 4, wherein The first visual encoder adopts a visual coding network based on the Transformer structure and converts the input image into a first image feature through the following steps: Receiving an input image with a size of 224×224, and performing fixed-size segmentation processing on the image to divide the image into patches of 16×16 size, forming 196 patch units; Perform a linear mapping on the Patch unit, transform each Patch into a 768-dimensional feature vector to generate an initial visual feature representation of 196×768 dimensions; Based on the initial visual feature representation, add additional Tokens as global feature identifiers to form a feature sequence of 197×768; Through the Transformer encoding module, perform multi-layer attention calculations on the 197×768 feature sequence to extract global visual features; Adopt an interpolation expansion method to adjust the 197×768 visual features after Transformer encoding calculations to a 256×768 Token representation to match the Token structure output by the second visual encoder; The 256×768 first image features are used to be concatenated with the 256×1024 second image features output by the second visual encoder in the channel dimension to form fused visual features.

6. The multimodal visual understanding model based on dual-channel visual coding according to claim 5, characterized in that, The linear layer includes a multi-layer perceptron network for transforming the dimension of the fused visual features from 1792 to 4096 to match the input format of the large language model.

7. A training method for a multi-modal visual understanding model based on dual-channel visual coding, characterized in that The training method includes: Initialize the second visual encoder with the weights pre-trained on the ImageNet dataset, and connect the second visual encoder to a small language model for warm-up training; After completing the warm-up training, connect the second visual encoder to a complete multi-modal visual understanding model, and use the image-text pair dataset to perform end-to-end pre-training on the multi-modal visual understanding model. In the end-to-end pre-training stage, only optimize the parameters of the linear layer and the second visual encoder, and freeze the weights of the first visual encoder and the large language model; Based on a task-specific supervised fine-tuning dataset, only perform fine-tuning training on the linear layer and the large language model.

8. An inference method for a multimodal visual understanding model based on dual-channel visual coding, characterized in that, Deploy the trained multi-modal visual understanding model to an inference device. The inference method includes: Receive multi-modal data input by the user, where the multi-modal data includes image data and text data related to the image; Input the multi-modal data into the multi-modal visual understanding model for inference calculations, and obtain the inference output of the multi-modal visual understanding model. The inference output includes a natural language answer related to the multi-modal data; Transmit the inference output to the user interaction interface to support natural language interaction, and perform multi-round inference processing based on the user's subsequent input.

9. An electronic device, characterized in that, The electronic device includes: One or more processors; and a memory storing computer program instructions, which when executed cause the processors to execute the training method of the multi-modal visual understanding model based on dual-path visual encoding as claimed in claim 7, or execute the inference method of the multi-modal visual understanding model based on dual-path visual encoding as claimed in claim 8.

10. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the training method of the multi-modal visual understanding model based on dual-path visual encoding as claimed in claim 7, or implement the inference method of the multi-modal visual understanding model based on dual-path visual encoding as claimed in claim 8.

11. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, it implements the training method of the multi-modal visual understanding model based on dual-channel visual coding as described in claim 7, or implements the inference method of the multi-modal visual understanding model based on dual-channel visual coding as described in claim 8.

Citation Information

Cited By

  • Chart analysis method and device, electronic equipment and storage medium

    CN120745603A

  • Complex spaceflight product man-machine cooperation assembly-oriented body-equipped agent packaging method and device

    CN120791810A

  • Chain thinking enhanced multi-modal spatial reasoning method for highway video data

    CN120822627A

  • A chain thought enhancement multi-modal spatial reasoning method for highway video data

    CN120822627B

  • Multi-modal large model deployment method and system for space governance

    CN121010861A