Large language model modal expansion method and device based on parameter fusion and decoupling

Through task vector merging and modal-exclusive binary masking technology, the multimodal large language model is achieved without training, which solves the problems of high training costs and modal interference in the existing technology, maintains the performance of the original task and improves the stability and flexibility of the model.

CN120105352AActive Publication Date: 2025-06-06HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Patent Information

Application Number
CN202510582920.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-06
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

When the existing multimodal large language model expands the new modal capabilities, the training cost is high, the scalability is poor, and it is prone to modal interference problems, resulting in performance degradation.

Method used

The modal expansion method of large language model based on parameter fusion and decoupling is adopted. Through task vector merging and modal exclusive binary masking technology, multimodal capabilities are achieved without training and expansion, and parameter conflicts between modals are avoided.

Benefits of technology

Without retraining the model, it effectively expands multimodal capabilities, maintains the performance of the original task, solves the "catastrophic forgetting" problem, and improves the stability and flexibility of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105352A_ABST
    Figure CN120105352A_ABST
Patent Text Reader

Abstract

The invention provides a large language model modal expansion method and device based on parameter fusion and decoupling, and relates to the technical field of large language models. The method comprises the following steps: performing fine adjustment on a pre-training language model to obtain a plurality of multi-modal large language models; task vector extraction is carried out on each multi-modal large language model; sparsification is carried out on the original task vector by adopting a sparsification strategy to obtain sparse vectors, and fusion is carried out on the sparse vectors to obtain a fusion task vector; constructing model parameters according to the fusion task vector; constructing a mode-exclusive binary mask for each multi-mode large language model according to the fusion task vector; and constructing a fusion model according to the model parameters and the binary mask. The invention provides a multi-modal language model extension method with non-training fusion, modal decoupling, performance retention and continuous extension capabilities, which is suitable for efficiently integrating a plurality of MLLMs, reconstructing an original model structure, coping with application scenes such as continuous integration of new tasks and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models, and in particular to a large language model modal expansion method and device based on parameter fusion and decoupling. Background Art

[0002] Generative LLMs (Large Language Models) are an important breakthrough in the field of artificial intelligence in recent years. Based on deep neural networks and through large-scale text pre-training, they have powerful language understanding and generation capabilities, and are widely used in natural language processing tasks such as question-answering, translation, and dialogue, and have a profound impact in both academia and industry. However, language is not the only way for humans to perceive the world. Human cognition relies on the fusion of multimodal information, such as the collaborative understanding of signals such as images, voice, and video. Therefore, in order to give language models stronger general intelligence capabilities, researchers proposed MLLMs (Multimodal Large Language Models). By introducing visual, auditory and other modal encoders, the model has the ability to process multiple types of information, promoting the evolution of artificial intelligence systems to higher levels of perception and interaction.

[0003] The current mainstream multimodal modeling methods are usually based on the existing pre-trained language model architecture, introducing dedicated encoders for modalities such as images, audio, video, and point clouds. These encoders are responsible for feature extraction of raw data in non-language modalities, and by designing multimodal alignment modules or projection networks, the extracted modal features are mapped to a unified language semantic space. Subsequently, the entire multimodal system performs joint training of multimodal-text instruction data while keeping the main structure of the language model unchanged, thereby achieving the fusion of different modal input information and language-driven response generation.

[0004] In this process, two technical links are crucial. The first is the construction of high-quality training data. Since multimodal modeling requires learning the semantic mapping relationship between modalities, a large-scale and well-structured multimodal-language pairing dataset is required as a basic support. These data must not only cover multiple task types such as image description, audio question answering, and video understanding, but must also maintain a high degree of consistency at the semantic level to ensure that the model can fully grasp the intrinsic relationship between modalities during training, thereby improving its instruction understanding and task generalization capabilities. The second is the strong dependence on computing resources. Compared with pure text models, multimodal models have significantly improved input dimensions and structural complexity, especially the scalability of modalities such as video and audio in the time dimension, which requires the model to process long sequence inputs, multi-layer attention mechanisms, and cross-fusion structures between modalities, which significantly increases memory consumption and computational burden. In addition, to ensure generalization capabilities, the training process often requires the introduction of millions of samples and long-term optimization in ultra-large-scale parameter spaces.

[0005] Therefore, building a large multimodal language model is not only a complex model structure design task, but also a system engineering project that is highly dependent on data and computing resources. In practical applications, how to expand the multimodal capabilities of the model without repeated training has become the core challenge and key breakthrough direction in current multimodal artificial intelligence research.

[0006] Currently, there are two main technical paths in the construction of multimodal large language models: one is the end-to-end multimodal alignment method based on joint training, and the other is the method based on model parameter merging without training. The following is an analysis of each:

[0007] The first type of method is represented by OneLLM, MACAW-LLM, etc., which introduces multiple modal encoders (such as images, audio, video, etc.) or an encoder that can handle multiple modalities on the basis of a large language model, and performs joint training or instruction fine-tuning through multimodal-text pairing data to enable the model to have multimodal perception and language generation capabilities. The core structure of this type of method includes: the modal encoder is responsible for converting non-language modal (image, audio, etc.) input into semantic vectors; the alignment module projects the modal features to the input space of the language model; and then completes end-to-end training by constructing a large-scale instruction data set. This method has successfully achieved complex tasks such as visual question answering and audio and video understanding. However, since each modality needs to match a large amount of high-quality data, and the model needs to deal with a huge parameter space and modal cross-fusion mechanism during training, once you want to expand new modal capabilities, you often need to retrain the entire model, including rebuilding the modal encoder and reconstructing the training data set. This limits the scalability of the system and makes it difficult to quickly respond to the access needs of new modalities.

[0008] The second type of method is represented by the NaiveMC and DAMC frameworks, which explore the integration and expansion of multimodal capabilities of large language models by merging multiple trained MLLMs without training or lightweight training. The core idea of ​​the NaiveMC method is to synthesize the "task vector" (i.e., the difference between the fine-tuning parameters and the pre-training parameters) corresponding to the LLM parameters of each of the multiple MLLMs into a unified fused LLM through simple weighted averaging or linear superposition, while retaining all modal encoders. This method has low implementation cost and does not require additional training process. It is suitable for rapid prototyping verification and capability integration of existing models. However, experiments have shown that NaiveMC will cause serious modal interference problems after the fusion of multiple modal parameters, resulting in a decrease in the performance of the fused model on the original task.

[0009] To address the above issues, the DAMC framework introduces a parameter separation strategy, which explicitly distinguishes "language parameters" and "modality parameters" in the early stage of model training, and uses parameter-efficient training paradigms (such as Adapter and LoRA) to complete fine-tuning. Finally, when the models are merged, only the language parameters are fused, while the modality-related parts remain isolated, thereby reducing conflicts between different modalities. This method significantly improves the stability and task performance of the merged model, but its scope of application is limited: on the one hand, this method requires pre-setting the parameter structure and separation method in the fine-tuning stage, and cannot directly reuse the full-parameter fine-tuning model; on the other hand, its effect depends on the careful design and manual intervention during the training process, which does not meet the needs of rapid integration of large-scale open source models.

[0010] In summary, although the current existing technologies have their own advantages in multimodal fusion, they still face the following problems:

[0011] The joint training-based solution has extremely high training costs and lacks flexibility when expanding new modalities. Although NaiveMC can directly merge multiple modality models, there are parameter conflicts and performance degradation. DAMC can improve fusion quality, but it has strict requirements on training conditions and parameter formats and cannot be widely adapted.

[0012] None of the above methods can simultaneously achieve the technical goals of "no training", "scalability" and "original performance preservation". Therefore, there is an urgent need for a new solution that can achieve multimodal capability expansion without retraining, effectively solve the modal interference problem and retain the performance of the original model.

[0013] In the existing technology, most mainstream multimodal modeling methods rely on high-quality modality-text pairing data, and integrate multiple modal capabilities into the language model structure through end-to-end training. Although this type of method performs well in terms of performance, it has obvious expansion bottlenecks: each additional modality usually requires reconstruction of the modality encoder and retraining of the entire model, which not only consumes huge training resources, but also lacks flexibility, which is not conducive to the rapid evolution and application deployment of the model.

[0014] In order to meet the above challenges, some studies have proposed to merge multiple trained MLLMs at the parameter level by merging model parameters, so as to build a new model with multimodal capabilities. As representative frameworks, NaiveMC and DAMC respectively proposed two paths: "no training merging" and "parameter separation merging". Although such methods have alleviated the problem of heavy training to a certain extent, they still face key technical difficulties: although NaiveMC does not require training, there is serious interference between the merged modalities, and the model performance is significantly reduced; DAMC can reduce interference, but it relies on lightweight training strategies and is not compatible with full parameter fine-tuning models, so its scope of use is limited. Summary of the invention

[0015] In order to solve the technical problem in the prior art that it is difficult to balance between expanding the multimodal capabilities of a large multimodal language model without training and maintaining the original performance, the embodiment of the present invention provides a large language model modality expansion method and device based on parameter fusion and decoupling. The technical solution is as follows:

[0016] On the one hand, a large language model modal expansion method based on parameter fusion and decoupling is provided, the method is implemented by a large language model modal expansion device, and the method includes:

[0017] S1. Multiple multimodal large language models are obtained by fine-tuning the pre-trained language model.

[0018] S2. Extract a task vector for each of the multiple multimodal large language models to obtain an original task vector for each multimodal large language model.

[0019] S3. Use a sparsification strategy to sparse the original task vectors of multiple multimodal large language models to obtain sparse vectors, and fuse the sparse vectors to obtain a fused task vector.

[0020] S4. Construct model parameters based on the fusion task vector.

[0021] S5. Construct a modality-specific binary mask for each multimodal large language model based on the fusion task vector.

[0022] S6. Construct a fusion model based on model parameters, binary masks, and encoders of multiple multimodal large language models.

[0023] S7. Obtain the multimodal input to be processed, input it into the fusion model, and obtain the task processing result.

[0024] Optionally, the original task vector in S2 is the difference between the fine-tuning parameters of the multimodal large language model and the parameters of the pre-trained language model.

[0025] The task vector is shown in the following formula (1):

[0026] (1)

[0027] In the formula, Indicates The original task vector of the multimodal large language model with modalities, Indicates Fine-tuning parameters of a large multimodal language model with multiple modalities, Represents the parameters of the pre-trained language model.

[0028] Optionally, the fusion task vector in S3 is as shown in the following formula (2):

[0029] (2)

[0030] In the formula, represents the fusion task vector, Indicates The original task vector of the multimodal large language model with modalities, Indicates the number of multimodal large language models.

[0031] The model parameters in S4 are shown in equation (3):

[0032] (3)

[0033] In the formula, represents the model parameters, represents the parameters of the pre-trained language model, Represents an adjustable scaling factor.

[0034] Optionally, constructing a modality-specific binary mask for each multimodal large language model in S5 includes:

[0035] Select any parameter dimension of any multimodal large language model, obtain the sign of the selected parameter dimension in the original task vector and the sign in the fused task vector, and determine whether the sign in the original task vector is opposite to the sign in the fused task vector.

[0036] If the signs are opposite, the mask value of the selected parameter dimension is 0.

[0037] If the signs are consistent, it is determined whether the amplitude of the selected parameter dimension in the original task vector is greater than a preset threshold.

[0038] If it is greater than the preset threshold, the mask value of the selected parameter dimension is 1.

[0039] If it is not greater than the preset threshold, the mask value of the selected parameter dimension is 0.

[0040] Optionally, the calculation formula of the binary mask is as shown in the following formula (4): (4)

[0041] In the formula, Indicates The first parameters, Indicates the mode, represents the parameters in the mask matrix, Indicates The task vector of the modality parameters, represents the first parameters.

[0042] Optionally, the step of obtaining the multimodal input to be processed in S7 and inputting it into the fusion model to obtain the task processing result includes:

[0043] Get the multimodal input to be processed, and obtain the corresponding modal vector according to the modal type of the multimodal input.

[0044] The corresponding binary mask is dynamically selected according to the modality type of the multimodal input.

[0045] According to the modal vector and mask, modality-specific weighted calculations are performed on the Query, Key, and Value parameters of each layer of the Transformer of the fusion model to obtain the task processing result.

[0046] Optionally, the process of building the fusion model also includes:

[0047] Get a new multimodal large language model obtained by fine-tuning the pre-trained language model.

[0048] A new fusion model is constructed according to the new multimodal large language model and multiple multimodal large language models in the fusion model.

[0049] The sparsification strategy is used to thin the original task vector of the new fusion model to obtain a new sparse vector, and the new sparse vector is fused to obtain a new fusion task vector.

[0050] Construct new model parameters based on the new fusion task vector.

[0051] According to the new fusion task vector, a new modality-specific binary mask is constructed for each multimodal large language model.

[0052] A new fusion model is obtained according to the new model parameters and the new binary mask.

[0053] On the other hand, a large language model modal expansion device based on parameter fusion and decoupling is provided, and the device is applied to a large language model modal expansion method based on parameter fusion and decoupling, and the device includes:

[0054] The fine-tuning module is used to obtain multiple multimodal large language models by fine-tuning the pre-trained language model.

[0055] The extraction module is used to extract the task vector of each multimodal large language model in the multiple multimodal large language models to obtain the original task vector of each multimodal large language model.

[0056] The fusion module is used to use a sparse strategy to sparse the original task vectors of multiple multimodal large language models to obtain sparse vectors, and to fuse the sparse vectors to obtain a fused task vector.

[0057] The parameter construction module is used to construct model parameters according to the fusion task vector.

[0058] The binary mask construction module is used to construct a modality-specific binary mask for each multimodal large language model based on the fusion task vector.

[0059] The fusion model building module is used to build a fusion model based on model parameters, binary masks, and encoders of multiple multimodal large language models.

[0060] The output module is used to obtain the multimodal input to be processed, input it into the fusion model, and obtain the task processing result.

[0061] Optionally, the original task vector is a difference between fine-tuning parameters of the multimodal large language model and parameters of the pre-trained language model.

[0062] The task vector is shown in the following formula (1):

[0063] (1)

[0064] In the formula, Indicates The original task vector of the multimodal large language model with modalities, Indicates Fine-tuning parameters of a large multimodal language model with multiple modalities, Represents the parameters of the pre-trained language model.

[0065] Optionally, the task vector is fused as shown in the following formula (2):

[0066] (2)

[0067] In the formula, represents the fusion task vector, Indicates The original task vector of the multimodal large language model with modalities, Indicates the number of multimodal large language models.

[0068] The model parameters are shown in the following formula (3):

[0069] (3)

[0070] In the formula, represents the model parameters, represents the parameters of the pre-trained language model, Represents an adjustable scaling factor.

[0071] Optionally, the binary mask building module is further used to:

[0072] Select any parameter dimension of any multimodal large language model, obtain the sign of the selected parameter dimension in the original task vector and the sign in the fused task vector, and determine whether the sign in the original task vector is opposite to the sign in the fused task vector.

[0073] If the signs are opposite, the mask value of the selected parameter dimension is 0.

[0074] If the signs are consistent, it is determined whether the amplitude of the selected parameter dimension in the original task vector is greater than a preset threshold.

[0075] If it is greater than the preset threshold, the mask value of the selected parameter dimension is 1.

[0076] If it is not greater than the preset threshold, the mask value of the selected parameter dimension is 0.

[0077] Optionally, the calculation formula of the binary mask is as shown in the following formula (4): (4)

[0078] In the formula, Indicates The first parameters, Indicates the mode, represents the parameters in the mask matrix, Indicates The task vector of the modality parameters, represents the first parameters.

[0079] Optionally, the output module is further configured to:

[0080] Get the multimodal input to be processed, and obtain the corresponding modal vector according to the modal type of the multimodal input.

[0081] The corresponding binary mask is dynamically selected according to the modality type of the multimodal input.

[0082] According to the modal vector and mask, modality-specific weighted calculations are performed on the Query, Key, and Value parameters of each layer of the Transformer of the fusion model to obtain the task processing result.

[0083] Optionally, the process of building the fusion model also includes:

[0084] Get a new multimodal large language model obtained by fine-tuning the pre-trained language model.

[0085] A new fusion model is constructed according to the new multimodal large language model and multiple multimodal large language models in the fusion model.

[0086] The sparsification strategy is used to thin the original task vector of the new fusion model to obtain a new sparse vector, and the new sparse vector is fused to obtain a new fusion task vector.

[0087] Construct new model parameters based on the new fusion task vector.

[0088] According to the new fusion task vector, a new modality-specific binary mask is constructed for each multimodal large language model.

[0089] A new fusion model is obtained according to the new model parameters and the new binary mask.

[0090] On the other hand, a large language model modal expansion device is provided, comprising: a processor; a memory, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, any one of the large language model modal expansion methods based on parameter fusion and decoupling is implemented.

[0091] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the above-mentioned large language model modal expansion methods based on parameter fusion and decoupling.

[0092] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0093] In the present invention, a method for modal expansion and original performance retention of a large language model based on parameter fusion and mask decoupling is proposed. The core lies in that the multimodal large language model can be expanded without retraining through task vector merging and modal decoupling technology, and the original task performance can be effectively retained. Specifically, the present invention first merges the task vectors of multiple fine-tuned MLLMs through task vector merging technology, and uses scaling factors to adjust to generate a fused task vector, thereby effectively integrating multimodal information into the pre-trained language model. In order to avoid parameter conflicts between different modalities, the present invention introduces a modal decoupling mechanism, and accurately extracts a specific parameter subset of each modality by constructing a modality-specific binary mask. The method uses the principles of directional consistency and dominant saliency to ensure that the parameters of each modality can be processed independently to avoid interference. Finally, through task vector merging and modal masking technology, the present invention can not only expand multimodal capabilities without retraining a large amount of data, but also maintain the performance of the original task, solving the problem of "catastrophic forgetting". This method provides an efficient and flexible multimodal expansion method that can achieve seamless integration of multiple existing models and effectively support the continuous expansion of new tasks and multimodal capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0095] Figure 1 It is a flow chart of a large language model modality expansion method based on parameter fusion and decoupling provided by an embodiment of the present invention;

[0096] Figure 2 is an overall flow chart of MMER provided by an embodiment of the present invention;

[0097] Figure 3 It is a block diagram of a large language model modal expansion device based on parameter fusion and decoupling provided by an embodiment of the present invention;

[0098] Figure 4 It is a structural diagram of a large language model modality expansion device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0099] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0100] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0101] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.

[0102] In the embodiments of the present invention, sometimes the subscripts such as W 1 It may be written in non-subscript form such as W1. When the difference is not emphasized, the meaning is the same.

[0103] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0104] The embodiment of the present invention provides a large language model modality expansion method based on parameter fusion and decoupling, which can be implemented by a large language model modality expansion device, which can be a terminal or a server. Figure 1 The flowchart of the large language model modal expansion method based on parameter fusion and decoupling is shown. The processing flow of the method may include the following steps:

[0105] S1. Multiple multimodal large language models are obtained by fine-tuning the pre-trained language model.

[0106] In a feasible implementation, there is MLLMs (Multimodal Large Language Models), all of which are based on the same pre-trained language model Fine-tuned, recorded as parameter sets .

[0107] S2. Extract a task vector for each of the multiple multimodal large language models to obtain an original task vector for each multimodal large language model.

[0108] Optionally, the original task vector in S2 is the difference between the fine-tuning parameters of the multimodal large language model and the parameters of the pre-trained language model.

[0109] Extract the task vector from each fine-tuned model:

[0110] (1)

[0111] In the formula, Indicates The original task vector of the multimodal large language model with modalities, Indicates Fine-tuning parameters of a large multimodal language model with multiple modalities, Represents the parameters of the pre-trained language model.

[0112] S3. Use a sparsification strategy to sparse the original task vectors of multiple multimodal large language models to obtain sparse vectors, and fuse the sparse vectors to obtain a fused task vector.

[0113] In one feasible implementation, a sparse strategy (retaining the parameters with the top K% of absolute values) is used to sparse the task vector to obtain a sparse vector, and then multiple task vectors are fused into a unified task vector based on the consistency of the symbols in the parameter dimension. :

[0114] (2)

[0115] In the formula, represents the fusion task vector, Indicates The original task vector of the multimodal large language model with modalities, Indicates the number of multimodal large language models.

[0116] S4. Construct the fusion model parameters according to the fusion task vector, as shown in the following formula (3):

[0117] (3)

[0118] In the formula, represents the model parameters, represents the parameters of the pre-trained language model, Represents an adjustable scaling factor used to control the degree of fusion, which is usually determined by the average performance of the validation set or each modality task.

[0119] S5. Construct a modality-specific binary mask for each multimodal large language model based on the fusion task vector.

[0120] In a feasible implementation, in order to restore each original model from the fusion model or support modality-specific decoupling, the present invention constructs a modality-specific binary mask , used to identify the modality in the fusion vector The construction of the mask follows two principles:

[0121] 1. Directional consistency: If a parameter dimension of the task vector has opposite signs in the original and fused vectors, the mask value of this dimension is 0. Because there is a strong parameter conflict, masking is selected;

[0122] 2. Dominant saliency: If the magnitude of this dimension in the original task vector is large enough (such as greater than half of the fusion value) and the sign is consistent, it is set to 1. Therefore, this parameter is an important parameter and is therefore retained.

[0123] The final calculation formula of the mask is as follows: (4)

[0124] In the formula, Indicates The first parameters, Indicates the mode, Indicates The task vector of the modality parameters, represents the first parameters. Refers to a specific value in the matrix. By applying Hadamard multiplication to the fusion task vector through masking, the static decoupling and restoration mode can be achieved. The corresponding model:

[0125] (5)

[0126] The above mask can be used not only for model recovery, but also for dynamic modal decoupling in the inference phase. When the model performs multimodal input processing, the representations of different modalities (image, audio, text, etc.) are input into their respective encoders to obtain the modal vector After entering the language model calculation graph, the system dynamically selects the corresponding mask according to the modality type and performs modality-specific weighted calculations on the Query, Key, and Value parameters of each layer of Transformer. For example, the Query weight of the lth layer is processed as follows: (6)

[0127] in refers to the query parameter of the lth layer of the mask matrix corresponding to the first modality, and Refers to the parameters of the fused LLM The query parameters of the lth layer in the are the parameters of the original pre-trained LLM, Refers to the average value of all modality mask matrices, which is used as the mask matrix for the text modality. Subsequent calculations are based on this , complete the attention operation:

[0128] (7)

[0129] The same method is used for other linear layers. Through this structural design, the present invention realizes independent processing of parameter space between modalities, effectively avoids interference between different modalities, and improves the stability of the fusion model.

[0130] S6. Construct a fusion model based on model parameters, binary masks, and encoders of multiple multimodal large language models.

[0131] like Figure 2 As shown, the present invention provides a large language model modality expansion and original performance retention method based on parameter fusion and mask decoupling, named MMER.

[0132] Optionally, the process of building the fusion model also includes:

[0133] Get a new multimodal large language model obtained by fine-tuning the pre-trained language model.

[0134] A new fusion model is constructed according to the new multimodal large language model and multiple multimodal large language models in the fusion model.

[0135] The sparsification strategy is used to thin the original task vector of the new fusion model to obtain a new sparse vector, and the new sparse vector is fused to obtain a new fusion task vector.

[0136] Construct new model parameters based on the new fusion task vector.

[0137] According to the new fusion task vector, a new modality-specific binary mask is constructed for each multimodal large language model.

[0138] A new fusion model is obtained according to the new model parameters and the new binary mask.

[0139] In a feasible implementation, the present invention can also achieve "catastrophic forgetting suppression": when a new task of a certain modality needs to be fine-tuned, the new task model can still be integrated with the original model through the above-mentioned parameter merging and mask generation process, thereby improving the performance of the new task while retaining the performance of the old task, solving the problem of performance degradation caused by parameter coverage in traditional fine-tuning methods.

[0140] The MMER method provided by the present invention significantly improves the flexibility and application scope of the multimodal large language model through innovative multimodal expansion and retention technology. First, the method can expand the multimodal capabilities of the model without retraining, which means that new modalities can be quickly added to the existing system, avoiding the high cost and long cycle of retraining the entire model each time a new modality is added in the traditional method. This feature is particularly suitable for rapid iteration and multimodal system deployment, greatly shortening the development cycle and reducing the consumption of computing resources. In addition, the present invention ensures the independent processing of multimodal information through a precise modal decoupling mechanism, effectively avoiding mutual interference between different modalities. This not only improves the stability of the model, but also ensures the efficient processing of each modality after fusion, avoiding the common performance degradation problem in traditional multimodal fusion. In this way, the model can understand and respond to multimodal input more accurately, broadening its application scenarios in complex tasks such as visual question answering, audio and video analysis. The present invention also successfully solves the "catastrophic forgetting" problem, so that when expanding new tasks, the performance of the original tasks will not be affected. This advantage enables the multimodal large language model to maintain efficient and stable performance in long-term use, especially in applications that require the continuous addition of new tasks or new modalities. For systems that require long-term deployment and continuous optimization (such as smart assistants, autonomous driving, etc.), the present invention provides a more sustainable solution.

[0141] S7. Obtain the multimodal input to be processed, input it into the fusion model, and obtain the task processing result.

[0142] In order to verify the effectiveness of the scheme, three experiments were conducted on multiple tasks to test the method of the present invention. In the multimodal expansion experiment, the MMER method showed significant advantages, significantly surpassing the NaiveMC framework, and proved its effectiveness in expanding multimodal capabilities and improving the performance of merged MLLM. Compared with other baseline methods, MMER achieved the best performance in multiple multimodal tasks, showing that it effectively achieved the expansion of multimodal capabilities without retraining. In the multimodal performance retention experiment, the MMER method can effectively retain the performance of the original task. Finally, in the experiment of reducing catastrophic forgetting, MMER showed excellent robustness and was able to maintain the performance of the original task after fine-tuning the new task, avoiding the performance loss problem caused by task fusion in traditional methods.

[0143] The present invention proposes a fusion method that can achieve multimodal capability expansion of a large language model without training and has the ability to maintain the original performance, aiming to solve the following core technical problems: First, how to accurately identify and separate modality-specific parameters when merging multiple MLLMs, thereby avoiding performance loss caused by parameter conflicts between modalities; second, how to effectively retain the ability performance of the original task while fusing new modalities or new tasks, and avoid the occurrence of "catastrophic forgetting"; third, how to provide a universal parameter processing framework so that regardless of whether the original MLLM adopts a full parameter fine-tuning or lightweight fine-tuning strategy, it can be included in the fusion system for unified processing.

[0144] The design goal of the present invention is to achieve lossless modal expansion and original performance preservation of multimodal large language models through structured parameter fusion and decoupling mechanisms without relying on high-cost retraining. This invention can be widely used in the construction of multimodal systems, the integration and upgrading of existing models, the capability reorganization of open source models, and the development of cross-modal tasks under resource-constrained conditions, and has important research value and engineering promotion prospects.

[0145] In an embodiment of the present invention, a method for modal expansion and original performance retention of a large language model based on parameter fusion and mask decoupling is proposed. The core lies in that the multimodal large language model is expanded without retraining through task vector merging and modal decoupling technology, and the original task performance is effectively retained. Specifically, the present invention first merges the task vectors of multiple fine-tuned MLLMs through task vector merging technology, and uses scaling factors to adjust to generate a fused task vector, thereby effectively integrating multimodal information into the pre-trained language model. In order to avoid parameter conflicts between different modalities, the present invention introduces a modal decoupling mechanism, and accurately extracts a specific parameter subset of each modality by constructing a modality-specific binary mask. The method uses the principles of directional consistency and dominant saliency to ensure that the parameters of each modality can be processed independently to avoid interference. Finally, through task vector merging and modal masking technology, the present invention can not only expand multimodal capabilities without retraining a large amount of data, but also maintain the performance of the original task, solving the problem of "catastrophic forgetting". This method provides an efficient and flexible multimodal expansion method that can achieve seamless integration of multiple existing models and effectively support the continuous expansion of new tasks and multimodal capabilities.

[0146] Figure 3 The block diagram of a large language model modal expansion device based on parameter fusion and decoupling according to an exemplary embodiment is shown. The device is used in a large language model modal expansion method based on parameter fusion and decoupling. Figure 3The device includes a fine-tuning module 310, an extraction module 320, a fusion module 330, a parameter construction module 340, a binary mask construction module 350, a fusion model construction module 360 ​​and an output module 370. Among them:

[0147] The fine-tuning module 310 is used to obtain multiple multimodal large language models by fine-tuning the pre-trained language model.

[0148] The extraction module 320 is used to extract the task vector of each of the multiple multimodal large language models to obtain the original task vector of each multimodal large language model.

[0149] The fusion module 330 is used to use a sparse strategy to sparse the original task vectors of multiple multimodal large language models to obtain sparse vectors, and to fuse the sparse vectors to obtain a fused task vector.

[0150] The parameter construction module 340 is used to construct model parameters according to the fusion task vector.

[0151] The binary mask construction module 350 is used to construct a modality-specific binary mask for each multimodal large language model according to the fusion task vector.

[0152] The fusion model construction module 360 ​​is used to construct a fusion model according to model parameters, binary masks and encoders of multiple multimodal large language models.

[0153] The output module 370 is used to obtain the multimodal input to be processed, input it into the fusion model, and obtain the task processing result.

[0154] In an embodiment of the present invention, a method for modal expansion and original performance retention of a large language model based on parameter fusion and mask decoupling is proposed. The core lies in that the multimodal large language model is expanded without retraining through task vector merging and modal decoupling technology, and the original task performance is effectively retained. Specifically, the present invention first merges the task vectors of multiple fine-tuned MLLMs through task vector merging technology, and uses scaling factors to adjust to generate a fused task vector, thereby effectively integrating multimodal information into the pre-trained language model. In order to avoid parameter conflicts between different modalities, the present invention introduces a modal decoupling mechanism, and accurately extracts a specific parameter subset of each modality by constructing a modality-specific binary mask. The method uses the principles of directional consistency and dominant saliency to ensure that the parameters of each modality can be processed independently to avoid interference. Finally, through task vector merging and modal masking technology, the present invention can not only expand multimodal capabilities without retraining a large amount of data, but also maintain the performance of the original task, solving the problem of "catastrophic forgetting". This method provides an efficient and flexible multimodal expansion method that can achieve seamless integration of multiple existing models and effectively support the continuous expansion of new tasks and multimodal capabilities.

[0155] Figure 4 is a structural diagram of a large language model modality expansion device provided by an embodiment of the present invention, such as Figure 4 As shown, the large language model modality expansion device may include the above Figure 3 The large language model modality expansion device based on parameter fusion and decoupling is shown. Optionally, the large language model modality expansion device 410 may include a first processor 2001.

[0156] Optionally, the large language model modality expansion device 410 may further include a memory 2002 and a transceiver 2003 .

[0157] The first processor 2001, the memory 2002 and the transceiver 2003 may be connected via a communication bus.

[0158] Combine the following Figure 4 The components of the large language model modality expansion device 410 are specifically introduced as follows:

[0159] The first processor 2001 is the control center of the large language model modality expansion device 410, which can be a processor or a general term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement an embodiment of the present invention, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (field programmable gate arrays, FPGAs).

[0160] Optionally, the first processor 2001 can perform various functions of the large language model modality expansion device 410 by running or executing a software program stored in the memory 2002 and calling data stored in the memory 2002.

[0161] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 4 CPU0 and CPU1 are shown in FIG.

[0162] In a specific implementation, as an embodiment, the large language model modality expansion device 410 may also include multiple processors, such as Figure 4 The first processor 2001 and the second processor 2004 are shown in FIG. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions).

[0163] The memory 2002 is used to store the software program for executing the solution of the present invention, and is controlled to be executed by the first processor 2001. The specific implementation method can refer to the above method embodiment, which will not be repeated here.

[0164] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001, or may exist independently, and may be accessed through the interface circuit ( Figure 4 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0165] The transceiver 2003 is used to communicate with a network device or a terminal device.

[0166] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 4 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.

[0167] Optionally, the transceiver 2003 may be integrated with the first processor 2001, or may exist independently and communicate with the first processor 2001 through the interface circuit ( Figure 4 (not shown) is coupled to the first processor 2001, which is not specifically limited in this embodiment of the present invention.

[0168] It should be noted that Figure 4 The structure of the large language model modality expansion device 410 shown in the figure does not constitute a limitation on the router, and the actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0169] In addition, the technical effects of the large language model modal expansion device 410 can refer to the technical effects of the large language model modal expansion method based on parameter fusion and decoupling described in the above method embodiment, which will not be repeated here.

[0170] It should be understood that the first processor 2001 in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0171] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0172] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented by software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0173] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0174] In the present invention, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0175] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0176] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0177] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0178] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0179] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0180] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0181] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.

[0182] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A large language model modal expansion method based on parameter fusion and decoupling, characterized in that: The method comprises: S1. Obtain multiple multimodal large language models by fine-tuning the pre-trained language model; S2. Extracting a task vector from each of the multiple multimodal large language models to obtain an original task vector of each multimodal large language model; S3, using a sparse strategy to sparse the original task vectors of multiple multimodal large language models to obtain sparse vectors, and fusing the sparse vectors to obtain a fused task vector; S4, constructing model parameters according to the fusion task vector; S5. Constructing a modality-specific binary mask for each multimodal large language model according to the fusion task vector; S6. Building a fusion model according to the model parameters, the binary mask, and the encoders of the plurality of multimodal large language models; S7. Obtain the multimodal input to be processed, input it into the fusion model, and obtain the task processing result.

2. The large language model modality expansion method based on parameter fusion and decoupling according to claim 1 is characterized in that: The original task vector in S2 is the difference between the fine-tuning parameters of the multimodal large language model and the parameters of the pre-trained language model; The task vector is shown in the following formula (1): (1) In the formula, Indicates The original task vector of the multimodal large language model with modalities, Indicates Fine-tuning parameters of a large multimodal language model with multiple modalities, Represents the parameters of the pre-trained language model.

3. The large language model modality expansion method based on parameter fusion and decoupling according to claim 1 is characterized in that: The fusion task vector in S3 is shown in the following formula (2): (2) In the formula, represents the fusion task vector, Indicates The original task vector of the multimodal large language model with modalities, Indicates the number of multimodal large language models; The model parameters in S4 are shown in the following formula (3): (3) In the formula, represents the model parameters, represents the parameters of the pre-trained language model, Represents an adjustable scaling factor.

4. The large language model modality expansion method based on parameter fusion and decoupling according to claim 1 is characterized in that: The step of constructing a modality-specific binary mask for each multimodal large language model in S5 includes: Select any parameter dimension of any multimodal large language model, obtain the sign of the selected parameter dimension in the original task vector and the sign in the fused task vector, and determine whether the sign in the original task vector is opposite to the sign in the fused task vector; If the signs are opposite, the mask value of the selected parameter dimension is 0; If the signs are consistent, it is determined whether the amplitude of the selected parameter dimension in the original task vector is greater than a preset threshold; If it is greater than the preset threshold, the mask value of the selected parameter dimension is 1; If it is not greater than the preset threshold, the mask value of the selected parameter dimension is 0.

5. The large language model modality expansion method based on parameter fusion and decoupling according to claim 4 is characterized in that: The calculation formula of the binary mask is shown in the following formula (4): (4) In the formula, Indicates The first parameters, Indicates the mode, represents the parameters in the mask matrix, Indicates The task vector of the modality parameters, represents the first parameters.

6. The large language model modality expansion method based on parameter fusion and decoupling according to claim 1 is characterized in that: The step of obtaining the multimodal input to be processed in S7, inputting it into the fusion model, and obtaining the task processing result includes: Obtaining the multimodal input to be processed, and obtaining the corresponding modal vector according to the modal type of the multimodal input; Dynamically select the corresponding binary mask according to the modality type of the multimodal input; According to the modal vector and mask, modality-specific weighted calculations are performed on the Query, Key, and Value parameters of each layer of the Transformer of the fusion model to obtain the task processing result.

7. The large language model modality expansion method based on parameter fusion and decoupling according to claim 1 is characterized in that: The construction process of the fusion model also includes: Obtain a new multimodal large language model by fine-tuning the pre-trained language model; Constructing a new fusion model according to the new multimodal large language model and multiple multimodal large language models in the fusion model; Using a sparse strategy to sparse the original task vector of the new fusion model to obtain a new sparse vector, and fusing the new sparse vector to obtain a new fusion task vector; Constructing new model parameters according to the new fusion task vector; Constructing a new binary mask specific to each modality for each multimodal large language model according to the new fusion task vector; A new fusion model is obtained according to the new model parameters and the new binary mask.

8. A large language model modal expansion device based on parameter fusion and decoupling, the large language model modal expansion device based on parameter fusion and decoupling is used to implement the large language model modal expansion method based on parameter fusion and decoupling as claimed in any one of claims 1 to 7, characterized in that: The device comprises: A fine-tuning module is used to obtain multiple multimodal large language models by fine-tuning the pre-trained language model; An extraction module, configured to extract a task vector from each of the plurality of multimodal large language models to obtain an original task vector of each multimodal large language model; A fusion module, used to use a sparse strategy to sparse the original task vectors of multiple multimodal large language models to obtain sparse vectors, and to fuse the sparse vectors to obtain a fused task vector; A parameter construction module, used to construct model parameters according to the fusion task vector; A binary mask construction module, used for constructing a modality-specific binary mask for each multimodal large language model according to the fusion task vector; A fusion model construction module, used to construct a fusion model according to the model parameters, the binary mask and the encoders of the plurality of multimodal large language models; The output module is used to obtain the multimodal input to be processed, input it into the fusion model, and obtain the task processing result.

9. A large language model modal expansion device, characterized in that: The large language model modality expansion device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, which can be called by a processor to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-modal data processing method for enhancing large language model

    CN118070227A

  • Multi-modal reasoning method and device based on large language model and knowledge graph

    CN118193684A

  • Social network false message detection method based on large model and multi-modal fusion

    CN119475066A

Cited By

  • Financial condition analysis method and device based on large model, equipment and medium

    CN121052949A