Multi-modal large model construction method and large model online updating method

By disassembling and combining submodules from pre-trained models, building a new multimodal large model, and performing partial parameter freezing and hybrid precision quantization, the efficiency and resource problems of multimodal large models in the existing technology are solved when deploying the end-side, and more efficient development and deployment are achieved.

CN120106172APending Publication Date: 2025-06-06INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510148683.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-10-08
Filing Date
2025-02-11
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing multimodal large models have problems such as high time overhead, high demand for computing resources and storage space, and poor performance on specific tasks when deploying on the end side.

Method used

By disassembling out submodules from different pre-trained models, building a new multimodal large model in combination, and freezing some parameters during the training process, only non-freezing parameters are updated. The trained model is then subjected to mixed precision quantization to reduce the amount of parameters.

Benefits of technology

Improves development efficiency, reduces training overhead, and makes multimodal large models easier to deploy on resource-constrained devices, improving performance capabilities in specific environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106172A_ABST
    Figure CN120106172A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal large model construction method and a large model online updating method, according to the technical scheme of the invention, sub-modules are disassembled from some existing pre-training modules, and can be combined as required to obtain a new multi-modal large model, so that the development efficiency is improved; then, when the new multi-modal large model is trained, part of parameters are frozen, and only a small number of other parameters are trained, so that the original knowledge from the sub-module of the pre-trained model is utilized, the training overhead is reduced, and the development efficiency is further improved; and finally, performing mixing precision quantification on the trained multi-modal large model so as to reduce the parameter quantity, thereby better deploying the multi-modal large model to resource-limited equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, specifically to the field of large model edge side deployment technology, and more specifically to a multimodal large model construction method and a large model online update method. Background Art

[0002] The existing multimodal large models have some shortcomings in the deployment on the end. First, it takes a lot of time to create a completely new multimodal large model. Second, multimodal large models are usually large in scale, requiring a lot of computing resources and storage space, and are difficult to run efficiently on end devices with limited computing power and storage resources. Secondly, the existing open source large models are mostly general-purpose designs and perform poorly on specific tasks. In order to meet the needs of downstream tasks on the end, it is usually necessary to fine-tune the pre-trained model.

[0003] It should be noted that this background technology is only used to introduce the relevant information of the present invention to help understand the technical solution of the present invention, but it does not mean that the relevant information is necessarily the prior art. The relevant information is submitted and disclosed together with the present invention solution. If there is no evidence that the relevant information has been disclosed before the application date of the present invention, the relevant information shall not be regarded as the prior art. Summary of the invention

[0004] Therefore, the purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a multimodal large model construction method and a large model online updating method.

[0005] The objective of the present invention is achieved through the following technical solutions:

[0006] According to a first aspect of the present invention, a method for constructing a multimodal large model based on a pre-trained model is provided, comprising the steps of: S1, disassembling a plurality of sub-modules from at least two different pre-trained models; S2, selecting some sub-modules from the plurality of sub-modules for combination to construct a new multimodal large model; S3, training the new multimodal large model to obtain a trained multimodal large model, wherein during training, some parameters in the new multimodal model are frozen, and only non-frozen parameters are updated; S4, performing mixed precision quantization on the trained multimodal large model to obtain a quantized multimodal large model, wherein during mixed precision quantization, multiple integer quantity precisons are used to quantize different parameters. This solution can at least achieve the following beneficial technical effects: this solution disassembles sub-modules from some existing pre-trained modules, which can be combined as needed to obtain a new multimodal large model to improve development efficiency; then, when training the new multimodal large model, some parameters are frozen and only a small number of remaining parameters are trained, thereby utilizing the original knowledge of the sub-modules from the pre-trained model, reducing training overhead and further improving development efficiency; finally, the trained multimodal large model is quantized with mixed precision to reduce the number of parameters so that it can be better deployed on resource-constrained devices.

[0007] Optionally, at least two pre-trained models include a MobileVLM-v2 model and a CLIP ViT-L / 14 model, wherein the new multimodal large model includes: a word segmenter, a downsampling mapper, and a language module from the MobileVLM-v2 model; and a visual encoder from the CLIP ViT-L / 14 model. This solution can at least achieve the following beneficial technical effects: This solution disassembles submodules from two multimodal large language models respectively, and then reconstructs them into a new multimodal large model, which can reduce the design cost of large models suitable for resource-constrained devices (such as edge devices, mobile devices), and improve product development efficiency.

[0008] Optionally, the new multimodal large model is configured as follows: using a word segmenter to process the input text into a text tag sequence; using a visual encoder to extract the visual features of the input image; using a connection layer to adjust the dimension of the visual features to couple the visual encoder and the downsampling mapper to obtain an adjusted visual feature; using the downsampling mapper to perform feature deformation, pooling and position information enhancement on the adjusted visual features to obtain a visual tag sequence; using a language module to determine a text response based on the text tag sequence and the visual tag sequence. This solution can at least achieve the following beneficial technical effects: This solution uses a connection layer to efficiently couple submodules from different pre-trained models, thereby efficiently completing the construction of a new model.

[0009] Optionally, in step S3, the parameters of the visual encoder and the language module are frozen, and the parameters of the connection layer and the downsampling mapper are updated. This solution can at least achieve the following beneficial technical effects: the solution retains the original knowledge of the visual encoder and the language module from the pre-trained model, and only uses the parameters of the connection layer and the downsampling mapper with relatively few parameters as non-frozen parameters for training and updating, which can reduce training overhead and further improve development efficiency.

[0010] Optionally, in step S4, only the trainable parameters of the model are quantized, and the excitation tensor is not quantized. This solution can at least achieve the following beneficial technical effects: This solution takes into account that the excitation tensor (intermediate result of reasoning) generated in matrix multiplication of a large Transformer-based model usually contains a large number of outliers. If it is truncated, these elements with large absolute values ​​may have a significant impact on the results during the model reasoning process, thereby causing the model effect to decline. Therefore, in the present invention, only the weights and / or biases are quantized, and the excitation tensor is not quantized.

[0011] According to the second aspect of the present invention, there is provided an online update method for a multimodal large model, comprising: obtaining a quantized multimodal large model obtained according to the method described in the first aspect, and deploying it on an edge device; the edge device uses the quantized multimodal large model to perform reasoning based on the input image and text, and collects the image and text question and answer corpus generated when serving users; and uses the Lora algorithm and the image and text question and answer corpus to update the quantized multimodal large model online. This scheme can at least achieve the following beneficial technical effects: the scheme deploys a quantized multimodal large model on the edge device, which can improve the reasoning efficiency on the edge side, and can combine the image and text question and answer corpus collected by the edge device for users in the service area, and use the Lora algorithm to fine-tune the quantized multimodal large model. Through a personalized fine-tuning method, the performance of the model in a specific environment is enhanced with a small amount of training overhead, ensuring excellent performance in end-side applications.

[0012] Optionally, the image-text question-and-answer corpus includes multiple samples, and the samples include images collected by edge devices and user prompt texts for the images, wherein the user prompt texts include questions and answers.

[0013] According to a third aspect of the present invention, there is provided a computer program product, comprising a computer program / instruction, which implements the steps of the method described in the first aspect or the second aspect when executed by a processor.

[0014] According to a fourth aspect of the present invention, an electronic device is provided, comprising: one or more processors; and a memory, wherein the memory is used to store executable instructions; the one or more processors are configured to implement the steps of the method described in the first aspect or the second aspect by executing the executable instructions.

[0015] Compared with the prior art, the advantages of the present invention are:

[0016] The method of the present invention disassembles sub-modules from some existing pre-trained modules, which can be combined as needed to obtain a new multimodal large model to improve development efficiency; then, when training the new multimodal large model, some parameters are frozen and only a small number of remaining parameters are trained, thereby utilizing the original knowledge of the sub-modules from the pre-trained model, reducing training overhead and further improving development efficiency; finally, the trained multimodal large model is quantized with mixed precision to reduce the number of parameters so that it can be better deployed on resource-constrained devices. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The embodiments of the present invention are further described below with reference to the accompanying drawings, in which:

[0018] Figure 1 A schematic diagram of a process of constructing a multimodal large model based on a pre-trained model according to an embodiment of the present invention;

[0019] Figure 2 A schematic diagram of the principle of an online updating method according to an embodiment of the present invention;

[0020] Figure 3 A schematic diagram of a system for deploying a large model on a terminal side and related implementation flow according to an embodiment of the present invention;

[0021] Figure 4 The figure is a schematic diagram of the overall process of model construction, model reasoning and model updating according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0023] As mentioned in the background technology section, existing multimodal large models have some shortcomings in terminal deployment. First, it takes a lot of time to create a completely new multimodal large model; second, multimodal large models are usually large in scale, requiring a lot of computing resources and storage space, and are difficult to run efficiently on terminal devices with limited computing power and storage resources. In this regard, the method of the present invention disassembles sub-modules from some existing pre-trained modules, which can be combined as needed to obtain a new multimodal large model to improve development efficiency; then, when training the new multimodal large model, some parameters are frozen, and only a small number of remaining parameters are trained, thereby utilizing the original knowledge of the sub-modules from the pre-trained model, reducing training overhead and further improving development efficiency; finally, the trained multimodal large model is quantized with mixed precision to reduce the number of parameters, so as to better deploy it on resource-constrained devices.

[0024] Before specifically introducing the embodiments of the present invention, some of the terms used therein are explained as follows:

[0025] A large model, also known as a large language model, is an artificial neural network model that contains a large number of parameters. Currently, a large model usually refers to a model that contains hundreds of millions or even more parameters. It should be understood that as technology develops, the minimum number of parameters required in a large model may increase.

[0026] A multimodal large model refers to a large model that can process at least two modal information (such as images and text).

[0027] According to one embodiment of the present invention, see Figure 1 , a method for constructing a multimodal large model based on a pre-trained model is provided, comprising steps S1, S2, S3 and / or S4. Each step is schematically described below.

[0028] Step S1: disassemble multiple sub-modules from at least two different pre-trained models.

[0029] According to one embodiment of the present invention, a word segmenter, a downsampling mapper and a language module are disassembled from a first pre-trained model for graph-text reasoning (such as a MobileVLM-v2 model), and a visual encoder is disassembled from a second pre-trained model for graph-text reasoning (such as a CLIP ViT-L / 14 model). Schematically, it is assumed that the first pre-trained model is a MobileVLM-v2 model and the second pre-trained model is a CLIP ViT-L / 14 model. The visual encoder (Vision Encoder module), the word segmenter (i.e., Tokenizer), the downsampling mapper (i.e., LDPv2) and the language module (i.e., Llama module) can be disassembled from the MobileVLM-v2 model. A complete visual encoder (Vision Encoder module) is disassembled from the CLIP ViT-L / 14 model. Alternatively, when disassembling sub-modules from a pre-trained module, some redundant processing layers in the original module can be deleted. For example, after disassembling the complete visual encoder from the CLIP ViT-L / 14 model, remove its last layer, and use the visual encoder with the last layer removed as the disassembled visual encoder. In this way, the number of parameters of the sub-modules can be reduced, making it more suitable for combining a multimodal large model for deployment on the end side (such as a mobile terminal, edge device (edge ​​server), etc.). Of course, implementers can also disassemble relevant sub-modules from other publicly available multimodal large models (such as: VisualGLM-6B model and / or mPLUG-Owl model, etc.), and the present invention does not impose any restrictions on this.

[0030] Step S2: Select some sub-modules from the multiple sub-modules to combine and construct a new multi-modal large model.

[0031] According to one embodiment of the present invention, the new multimodal large model includes a connection layer, a word segmenter, a downsampling mapper and a language module disassembled from the first pre-trained model, and a visual encoder disassembled from the second pre-trained model, wherein the new multimodal large model is configured in the following manner: using the word segmenter to process the input text into a text tag sequence; using the visual encoder to extract the visual features of the input image; using the connection layer to adjust the dimension of the visual feature to couple the visual encoder and the downsampling mapper to obtain the adjusted visual feature; using the downsampling mapper to perform feature deformation, pooling and position information enhancement on the adjusted visual feature to obtain a visual tag sequence; using the language module to determine the text response according to the text tag sequence (text encoding) and the visual tag sequence (image encoding). For example: the word segmenter (i.e., Tokenizer), downsampling mapper (i.e., LDPv2) and language module (i.e., Llama module) disassembled from the MobileVLM-v2 model (such as MobileVLM-v2 1.7B), and the visual encoder disassembled from the CLIP ViT-L / 14 model are combined into a new multimodal large model. Among them, since the output feature dimension of the visual encoder disassembled from the CLIP ViT-L / 14 model and the input feature dimension of the downsampling mapper (i.e., LDPv2) disassembled from the MobileVLM-v2 model do not match, a connection layer is used to connect the visual encoder and the downsampling mapper. The connection layer can be implemented as a single-layer or multi-layer fully connected layer. This solution can at least achieve the following beneficial technical effects: This solution allows users to design a multimodal large model from scratch without having to use a combination of sub-modules disassembled from an existing large model to build a required new multimodal large model, thereby reducing the development cost of deploying multimodal large models on electronic devices such as edge terminals and mobile terminals, thereby efficiently completing the construction of new models.

[0032] According to an optional embodiment of the present invention, when constructing a new multimodal large model, the precision of each parameter may not be unified, and the existing precision of the disassembled sub-modules may be retained.

[0033] According to an optional embodiment of the present invention, for a new multimodal large model, since its different parameters may use floating point numbers or fixed point numbers, different precisions may also exist. In order to reduce the difficulty of subsequent quantization, when constructing a new multimodal large model, the precision of each parameter can be unified first. In step S2, the precision of all trainable parameters of the new multimodal large model is unified to floating point values ​​of preset precision; for example, floating point numbers in F64 format or F32 format, so as to facilitate subsequent quantization.

[0034] It should be noted that when actually implementing this plan, especially when disassembling pre-trained models and building new multimodal large models, relevant laws and regulations should be observed, the intellectual property status of the relevant models should be understood, and implementation should be carried out after obtaining the permission of the relevant intellectual property rights holders.

[0035] Step S3: training a new multimodal large model to obtain a trained multimodal large model, wherein some parameters in the new multimodal model are frozen during training, and only non-frozen parameters are updated.

[0036] According to one embodiment of the present invention, the parameters of the visual encoder and the language module are frozen, and the parameters of the connection layer and the downsampling mapper are updated. Of course, there are other alternatives, such as only updating the parameters of the connection layer or the downsampling mapper.

[0037] Step S4: Perform mixed precision quantization on the trained multimodal large model to obtain a quantized multimodal large model, wherein the mixed precision quantization uses multiple integer quantity precisons to quantize different parameters.

[0038] According to one embodiment of the present invention, the operation of performing mixed precision quantization includes: setting integer quantization precisions of multiple bit widths and multiple floating point precisions for selection; based on the integer quantization precisions and floating point precisions of the multiple bit widths, generating a mixed precision quantization combination for quantizing each layer of the adjusted multimodal large model, wherein each mixed precision quantization combination indicates the parameter precision used by each layer of the multimodal large model. The user can set the precision range of quantization, for example: setting the integer quantization precision for selection includes Q2_K, Q4_0, Q4_K, Q5_0, Q6_K, Q8_0 and other quantization precisions supported by the Ollama platform, and can also set integer quantization precisions such as INT4 and INT8. In addition, the floating point quantization precision for selection can also be set, such as F16 or F8. Then, set the parameters that need to be quantized and / or selectively limit the quantization of the specified parameters; then, the user can select from a variety of integer quantization accuracies of bit width and a variety of floating-point accuracies, or a random algorithm can randomly select the parameter accuracy used by the layer where each parameter that needs to be quantized is located; thereby generating a mixed precision quantization combination for quantizing each layer of the trained multimodal large model. It should be understood that multiple mixed precision quantization combinations can be generated at one time to verify the performance of multiple quantized multimodal large models obtained by quantizing multiple mixed precision quantization combinations at a time in the later stage, and deploy them to edge devices or mobile terminals based on the best performance.

[0039] According to one embodiment of the present invention, in step S4, 6-bit integer quantization is used for the weight matrix and feedforward layer in the multi-head self-attention mechanism, and 4-bit integer quantization is used for the remaining layers. Thus, the capabilities of the relevant operators in the self-attention mechanism are better guaranteed, and the impact of quantization on model performance is reduced.

[0040] According to one embodiment of the present invention, in step S4, both the trainable parameters and the excitation tensor may be quantized, thereby improving the compression effect of the model. However, considering that the excitation tensor (intermediate result of reasoning) generated in the matrix multiplication of a large Transformer-based model usually contains a large number of outliers, if it is truncated, these elements with large absolute values ​​may have a significant impact on the results during the model reasoning process, thereby causing the model effect to decline. Therefore, preferably, in step S4, only the trainable parameters (weights and / or biases) of the model are quantized, and the excitation tensor is not quantized.

[0041] According to an embodiment of the present invention, there is provided a method for online updating of a multi-modal large model, comprising steps: K1, K2, and K3. Each step is schematically described below.

[0042] Step K1: Obtain a quantized multimodal large model obtained according to a multimodal large model construction method based on a pre-trained model, and deploy it on an edge device.

[0043] According to one embodiment of the present invention, the operating environment (inference environment) required by the multimodal large model can be constructed on the edge device, and then the quantized multimodal large model can be deployed to the edge device (end-side device). During actual deployment, the model can be imported into the edge device through the network, and the model can be deployed in the constructed inference environment to implement the inference operation.

[0044] Step K2: The edge device uses a quantized multimodal large model to perform reasoning based on the input image and text, and collects the image and text question and answer corpus generated when serving users.

[0045] According to one embodiment of the present invention, step K2 includes: obtaining input images and texts through the network for reasoning to obtain reasoning results. The text may be a question text posed by a user to an image through a user terminal. The image data may be an RGB image collected by a ground camera, or a remote sensing image collected by an on-orbit satellite. After the multimodal large model is inferred, the reasoning result is transmitted back to the user terminal through the network. During the reasoning period, the picture-text question-and-answer corpus generated when serving users may be collected to construct a database. The picture-text question-and-answer corpus includes multiple samples, the samples including images collected by edge devices and user prompt texts of the users for the images, wherein the user prompt texts include questions and answers. After the model gives the reasoning result, the user may correct the problem in the reasoning result. When collecting the picture-text question-and-answer corpus, if the user has made corrections, the corrected reasoning result and the related image are stored as samples. In addition, the samples in the picture-text question-and-answer corpus may be hierarchically managed, multiple quality levels may be set, and the multimodal large model may be used to rate the quality of each sample when collecting the sample, and then the sample may be selected based on the quality rating to perform online update on the multimodal large model. Alternatively, samples with higher quality levels are given a higher probability of being selected as training samples to better improve model performance.

[0046] Step K3: Use the Lora algorithm and the picture-text question-answering corpus to update the quantized multimodal large model online.

[0047] According to an embodiment of the present invention, based on the text-image question-answering corpus collected by the database of the edge device, the LoRA algorithm is used to fine-tune the deployment model to achieve online update of the model parameters. Figure 2 , the online update process includes:

[0048] Step K31: A bypass is added next to the parameters of the original deployed model to simulate the intrinsic rank structure of the model by first reducing the dimension and then increasing the dimension. The bypass includes the parameters of the LoRA adapter, including the trainable downsampling matrix A and upsampling matrix B;

[0049] Step K32: Initialize the A matrix to a Gaussian distribution matrix and the B matrix to a zero matrix; this initialization setting can make the neural network training more stable and help to quickly converge to a better solution;

[0050] Step K33: Use the image-text mapping corpus to fine-tune the model, keep the weight W of the original matrix unchanged, and only train the downsampling matrix A and the upsampling matrix B. During the fine-tuning process, keep the model dimensions and parameters unchanged. When outputting, add the parameters of the original deployed model and the parameters of the LoRA adapter, as shown in the following formula:

[0051]

[0052] in, Represents the original deployment model parameters, Indicates input, represents the downsampling matrix A, represents the upsampling matrix B, BA represents the matrix product of the upsampling matrix B and the downsampling matrix A; W∈R d×k express The dimension is d×k, A∈R r×k The dimension of A is r×k, B∈R d×r Indicates that the dimension of B is d×r;

[0053] Step K34: When fine-tuning stops, merge the original deployment model parameters and the matrix product BA to obtain a fine-tuned multimodal large model; that is, replace the original deployment model parameters W with the new parameters W merged , where W merged for:

[0054]

[0055] Step K35: Import the fine-tuned multimodal large model to the edge device through the network to update the parameters of the multimodal large model.

[0056] See also Figure 3 , providing a system for deploying a large model on the end side and related implementation processes. The device layer of the system includes: image acquisition devices, end-side devices (such as edge devices), cloud devices (such as cloud servers) and computer input devices. The devices can be connected by cables or wirelessly. The information layer of the system is configured as follows: the cloud device obtains the picture and text question and answer corpus collected by the end-side device, and uses the Lora algorithm to fine-tune the multimodal large model to obtain a fine-tuned multimodal large model.

[0057] See also Figure 4 , provides an overall schematic process of model construction, model reasoning and model updating based on the method of the present invention, wherein:

[0058] The model building method includes: pre-trained model disassembly, module modification, module combination, model training, model quantization and model deployment process. Among them, schematically:

[0059] Step A1: Disassembly of pre-trained model and module modification;

[0060] Taking the disassembly of the pre-trained MobileVLM v2 1.7B multimodal large model and CLIP ViT-L / 14 graphic multimodal model as an example, step A1 includes:

[0061] Step A11: Load and parse the MobileVLM-v2 model, remove the layers that make up the Vision Tower to get the Tokenizer, Vision Encoder module, Lightweight Downsampling Mapper (LDP v2 module) module, and Language Module (Llama module);

[0062] Step A12: Load and parse the CLIP ViT-L / 14 model, extract the layers that constitute the vision encoder from CLIP ViT-L / 14, delete its last layer (i.e., module modification), and output it as the Vision Encoder module.

[0063] Step A2: Module combination: combine the modules extracted from MobileVLM v2 1.7B and CLIP ViT-L / 14 to build a multimodal large model suitable for end-side deployment;

[0064] Step A2 includes:

[0065] Step A21: Combine the models, integrate the word segmenter and visual encoder extracted from CLIP ViT-L / 14 with the LDP v2 module extracted from MobileVLM v2 1.7B to align the weight dimension;

[0066] Step A22: uniformly convert the integrated model parameters into FP32 format;

[0067] Step A3: Train the model to achieve semantic transfer between modules to ensure that different modules can accurately understand and transfer semantics during the reasoning process;

[0068] Step A3 includes:

[0069] Step A31: Freeze the visual encoder and Llama modules, and only allow updating the LDP v2 modules and / or connection layer weights;

[0070] Step A32: Implement training to ensure that the image semantics output by the visual encoder can be accurately passed to the Llama module;

[0071] Step A4: Quantize the model and apply mixed precision quantization technology to improve the computing efficiency and resource utilization of the model while maintaining the performance and accuracy of the model.

[0072] Among them, different precision quantization is applied to each module in turn, and mixed precision quantization is applied within the module. 6-bit integer quantization is adopted for the weight matrix and feedforward layer in the multi-head self-attention mechanism, and 4-bit integer quantization is used for the remaining layers to further reduce the end-side operation pressure while ensuring the model performance. In step A4, multiple integer quantization precisions from 4 to 8 bits are supported. Considering that the excitation tensor generated in the matrix multiplication of the large Transformer-based model usually contains a large number of outliers, if it is truncated, these elements with large absolute values ​​may have a significant impact on the results during the model inference process, resulting in a decrease in the model effect. Therefore, in this example, only weight quantization is performed, and the excitation is not quantized.

[0073] Step A5: Import the model to the end device through the network, deploy the model in the established inference environment, and implement inference operations.

[0074] The model inference method includes: taking the user prompt and the collected image as the model input, the multimodal large model performs inference based on the model input, and obtains the inference result as the model output. That is, the model inference is implemented on the terminal side based on the user prompt and the image, including the following steps:

[0075] Step B1: The user inputs text to the client device through the network as input of the model;

[0076] Step B2: Collect images as input of the model;

[0077] Step B3: After the multimodal large model is inferred, the inference result is transmitted back to the user end through the network.

[0078] The model update method includes: using a database (text corpus) to collect a picture-text corresponding data set (picture-text question-answer corpus) to fine-tune the model, and then updating the parameters after the fine-tuning is completed. Among them, building a picture-text corresponding database based on terminal data, using the LoRA method to update the deployed model online, and implementing the model weight update, including the following steps:

[0079] Step C1: Collect the end-side picture and text question and answer corpus and build a database;

[0080] Step C1 includes:

[0081] Step C11: Transmit the image captured by the local camera to the server through the network.

[0082] Step C12: constructing a picture-text correspondence database based on user prompts and camera images.

[0083] Step C2: Use the LoRA method to fine-tune the multimodal large model;

[0084] Step C2 includes:

[0085] Step C21: A bypass containing the A matrix and the B matrix is ​​added next to the parameters of the multimodal large model to be trained, and the intrinsic rank structure of the model is simulated by the process of first reducing the dimension and then increasing the dimension;

[0086] Step C22: Initialize the A matrix to a Gaussian distribution matrix, and initialize the B matrix to a 0 matrix;

[0087] Step C23: fine-tune the model using the image-text mapping corpus, keep the parameter W of the multimodal large model to be trained unchanged, and only train the downsampled A matrix and the upsampled B matrix;

[0088] Step C24: When the training reaches convergence, new parameters are calculated based on the parameters W, A matrix and B matrix of the multimodal large model to be trained, and the parameters of the multimodal large model to be trained are updated using the new parameters.

[0089] In general, compared with the prior art, the implementation method of the present invention has significant beneficial effects in at least one of the following aspects:

[0090] Material savings: By reducing the consumption of training resources, the computing and storage requirements during model building are reduced;

[0091] Improved efficiency: Rapidly build models suitable for end-side deployment, significantly shortening the development and deployment cycle;

[0092] Performance optimization: During online updates, the model's performance in specific environments is enhanced through personalized fine-tuning methods, ensuring excellent performance in end-side applications.

[0093] The present invention is more competitive in resource-constrained scenarios (such as end-side applications) and demonstrates significant technical advantages. For example, in the image-text matching task, the reasoning speed of the method proposed by the present invention is more than 4 times faster than the current SOTA method CLIP, and the energy efficiency is more than 2 times higher.

[0094] It should be noted that although the above describes the various steps in a specific order, it does not mean that the various steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order as long as the required functions can be achieved.

[0095] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0096] A computer-readable storage medium may be a tangible device that holds and stores instructions used by an instruction execution device. Computer-readable storage media may include, for example, but are not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a protruding structure in a groove on which instructions are stored, and any suitable combination thereof.

[0097] The embodiments of the present invention have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A method for constructing a multimodal large model based on a pre-trained model, characterized in that: Includes steps: S1. Decompose multiple sub-modules from at least two different pre-trained models; S2, selecting some sub-modules from the multiple sub-modules to combine and construct a new multi-modal large model; S3. Training a new multimodal large model to obtain a trained multimodal large model, wherein during training, some parameters in the new multimodal model are frozen and only non-frozen parameters are updated; S4. Perform mixed precision quantization on the trained multimodal large model to obtain a quantized multimodal large model, wherein the mixed precision quantization uses multiple integer quantity precisons to quantize different parameters.

2. The multimodal large model construction method according to claim 1, characterized in that: The at least two pre-trained models include a MobileVLM-v2 model and a CLIP ViT-L / 14 model, wherein the new multimodal large model includes: The tokenizer, downsampling mapper, and language module from the MobileVLM-v2 model; and Vision encoder from the CLIP ViT-L / 14 model.

3. The multimodal large model construction method according to claim 2, characterized in that: The new multimodal large model is configured as follows: Use a tokenizer to process the input text into a sequence of text tokens; Using a visual encoder to extract visual features of an input image; The dimension of the visual feature is adjusted by using a connection layer to couple the visual encoder and the downsampling mapper to obtain an adjusted visual feature; Using a downsampling mapper to perform feature deformation, pooling, and position information enhancement on the adjusted visual features to obtain a visual marker sequence; A language module is used to determine a text response based on a sequence of textual tokens and a sequence of visual tokens.

4. The multimodal large model construction method according to claim 3, characterized in that: In step S3, the parameters of the visual encoder and language module are frozen, and the parameters of the connection layer and the downsampling mapper are updated.

5. The multimodal large model construction method according to claim 4, characterized in that: In step S4, only the trainable parameters of the model are quantized, and the activation tensor is not quantized.

6. A method for online updating of a multimodal large model, comprising: Obtain a quantized multimodal large model obtained according to the method of any one of claims 1 to 5, and deploy it on an edge device; The edge device uses a quantized multimodal large model to infer the input image and text, and collects the image and text question and answer corpus generated when serving users; The quantized multimodal large model is updated online using the Lora algorithm and the picture-text question-answering corpus.

7. The fine-tuning training method according to claim 6, characterized in that: The image-text question-and-answer corpus includes multiple samples, and the samples include images collected by edge devices and user prompt texts of users for the images, wherein the user prompt texts include questions and answers.

8. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: one or more processors; as well as A memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method of any one of claims 1-7 by executing the executable instructions.