Method and apparatus for modality expansion of large language models based on parameter fusion and decoupling
Through task vector merging and modal decoupling technology, the high training cost and performance degradation of multimodal large language models when expanding new modes are solved, seamless fusion and performance maintenance are achieved, and suitable for rapid iteration and multimodal system deployment.
Patent Information
- Application Number
- CN202510582920.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-05-07
AI Technical Summary
Existing multimodal large language models need to be retrained when expanding new modal capabilities, resulting in high training costs, lack of flexibility, and reduced performance of the combined model or limited scope of application.
Through task vector merging and modal decoupling technology, a sparse strategy and modal exclusive binary mask are used to build a fusion model to avoid conflicts between modals and maintain the original performance.
It realizes the expansion of multimodal capabilities without retraining, avoiding intermodal interference, maintaining the original task performance, and improving model stability and flexibility.
Smart Images

Figure CN120105352B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language models, and particularly to a method and device for expanding the modality of a large language model based on parameter fusion and decoupling. Background Art
[0002] Generative LLMs (Large Language Models) are important breakthroughs in the field of artificial intelligence in recent years. Based on deep neural networks and pre-trained with large-scale texts, they possess powerful language understanding and generation capabilities and are widely applied to natural language processing tasks such as question answering, translation, and dialogue, having a profound impact in both academic and industrial fields. However, language is not the only way for humans to perceive the world. Human cognition relies on multimodal information fusion, such as the collaborative understanding of signals like images, speech, and videos. Therefore, to endow language models with stronger general intelligent capabilities, researchers have proposed MLLMs (Multimodal Large Language Models). By introducing visual, auditory, and other modality encoders, the models are enabled to process various types of information, driving the evolution of artificial intelligence systems towards higher-level perception and interaction.
[0003] Current mainstream multimodal modeling methods generally build on existing pre-trained language model architectures and introduce dedicated encoders for modalities such as images, audio, videos, and point clouds. These encoders are responsible for extracting features from the raw data of non-verbal modalities and, by designing multimodal alignment modules or projection networks, mapping the extracted modality features to a unified language semantic space. Subsequently, on the basis of keeping the main structure of the language model unchanged, the entire multimodal system conducts joint training on multimodal-text instruction data, thereby achieving the fusion of input information of different modalities and the generation of language-driven responses.
[0004] In this process, two technical aspects are crucial. One is the construction of high-quality training data. Since multimodal modeling needs to learn the semantic mapping relationships between modalities, a large-scale and structurally clear multimodal-language paired dataset is required as the basic support. These data should not only cover various task types such as image description, audio question answering, and video understanding, but also maintain a high degree of consistency at the semantic level to ensure that the model can fully master the internal relationships between modalities during the training process, thereby enhancing its instruction understanding and task generalization capabilities. The other is the strong dependence on computing resources. Compared with pure text models, multimodal models have a significant increase in input dimension and structural complexity. Especially the extensibility of modalities such as videos and audio in the time dimension makes the model need to process long sequence inputs, multi-layer attention mechanisms, and cross-modal fusion structures, significantly increasing memory consumption and computational burden. In addition, to ensure generalization capabilities, the training process often needs to introduce millions of samples and conduct long-term optimization in an ultra-large-scale parameter space.
[0005] Therefore, building a multi-modal large language model is not only a complex model structure design task, but also a systematic project highly dependent on data and computing resources. In practical applications, how to expand the multi-modal capabilities of the model without repeated training has become the core challenge and key breakthrough direction in current multi-modal artificial intelligence research.
[0006] Currently, there are mainly two types of technical paths in the construction of multi-modal large language models: one is the end-to-end multi-modal alignment method based on joint training, and the other is the training-free method based on model parameter merging. The following is an analysis of each:
[0007] The first type of method is represented by OneLLM, MACAW-LLM, etc. It introduces multiple modal encoders (such as images, audio, video, etc.) or an encoder that can handle multiple modalities on the basis of a large language model, and performs joint training or instruction fine-tuning through multi-modal-text paired data to enable the model to have multi-modal perception and language generation capabilities. The core structure of this type of method includes: the modal encoder is responsible for converting the input of non-verbal modalities (images, audio, etc.) into semantic vectors; the alignment module projects the modal features into the input space of the language model; and then end-to-end training is completed by constructing a large-scale instruction-based dataset. This method has successfully achieved complex tasks such as visual question answering and audio-visual understanding. However, since each modality needs to match a large amount of high-quality data, and the model needs to process a huge parameter space and modality cross-fusion mechanism during training, once new modal capabilities are desired to be expanded, it often requires retraining the entire model from scratch, including reconstructing the modal encoder and reconstructing the training dataset. This limits the system scalability and makes it difficult to quickly respond to the access requirements of new modalities.
[0008] The second type of method is represented by the NaiveMC and DAMC frameworks, which explore integrating and expanding the multi-modal capabilities of large language models by merging multiple already trained MLLMs in a training-free or lightweight training manner. The core idea of the NaiveMC method is to synthesize a unified fused LLM by simply weighted averaging or linearly superimposing the "task vectors" corresponding to the LLM parameters in multiple MLLMs (i.e., the difference between the fine-tuning parameters and the pre-trained parameters), while retaining all modal encoders. This method has a low implementation cost and does not require an additional training process, and is suitable for rapid prototype verification and the integration of capabilities of existing models. However, experiments show that NaiveMC will cause serious modal interference problems after fusing multiple modal parameters, resulting in a decline in the performance of the fused model on the original tasks.
[0009] To address the above issues, the DAMC framework introduces a parameter separation strategy, that is, explicitly distinguishing "language parameters" and "modal parameters" in the initial stage of model training, and adopting parameter-efficient training paradigms (such as Adapter, LoRA) to complete fine-tuning. Finally, when merging the models, only the language parameters are fused, while the modal-related parts remain isolated, thereby reducing conflicts between different modalities. This method significantly improves the stability and task performance of the merged model, but its scope of application is limited: on the one hand, this method requires presetting the parameter structure and separation method in the fine-tuning stage and cannot directly reuse the fully parameter fine-tuned model; on the other hand, its effect depends on the elaborate design and manual intervention during the training process, which does not meet the requirements of rapid integration of large-scale open-source models.
[0010] In summary, although the current existing technologies have their own advantages in multimodal fusion, they still face the following problems:
[0011] The training cost is extremely high and lacks flexibility when expanding new modalities based on the joint training scheme; although NaiveMC can directly merge multiple modal models, there are parameter conflicts and performance degradation; DAMC can improve the fusion quality, but has strict requirements on training conditions and parameter formats and cannot be widely adapted;
[0012] None of the above methods can simultaneously achieve the technical goals of "no training", "scalability" and "original performance preservation". Therefore, there is an urgent need for a new scheme that can expand multimodal capabilities without retraining, effectively solve the modal interference problem and retain the performance of the original model.
[0013] In the existing technology, most mainstream multimodal modeling methods rely on high-quality modality-text paired data and fuse various modal capabilities into the language model structure through end-to-end training. Although such methods perform well in terms of performance, there are obvious expansion bottlenecks: for each additional modality, it is usually necessary to reconstruct the modal encoder and retrain the entire model, which not only consumes huge training resources but also lacks flexibility and is not conducive to the rapid evolution and application deployment of the model.
[0014] To address the above challenges, some studies propose to fuse multiple trained MLLMs at the parameter level by merging model parameters to construct a new model with multimodal capabilities. As representative frameworks, NaiveMC and DAMC respectively propose two paths of "no-training merging" and "parameter-separation merging". Although such methods alleviate the retraining problem to a certain extent, they still face key technical problems: although NaiveMC does not require training, there is serious interference between modalities after merging, and the model performance drops significantly; DAMC can reduce interference, but it relies on lightweight training strategies and is not compatible with fully parameter fine-tuned models, with a limited scope of use. Summary of the Invention
[0015] To solve the technical problem that existing multimodal large language models in the prior art are difficult to balance between expanding multimodal capabilities without training and maintaining the original performance, an embodiment of the present invention provides a method and device for expanding the modality of a large language model based on parameter fusion and decoupling. The technical solution is as follows:
[0016] On the one hand, a method for expanding the modality of a large language model based on parameter fusion and decoupling is provided. This method is implemented by a large language model modality expansion device, and the method includes:
[0017] S1. Obtain multiple multimodal large language models by fine-tuning a pre-trained language model.
[0018] S2. Extract task vectors for each multimodal large language model among the multiple multimodal large language models to obtain the original task vectors of each multimodal large language model.
[0019] S3. Sparsify the original task vectors of the multiple multimodal large language models using a sparsification strategy to obtain sparse vectors, and fuse the sparse vectors to obtain a fused task vector.
[0020] S4. Construct model parameters according to the fused task vector.
[0021] S5. Construct a modality-specific binary mask for each multimodal large language model according to the fused task vector.
[0022] S6. Construct a fused model according to the model parameters, the binary mask, and the encoders of the multiple multimodal large language models.
[0023] S7. Obtain the multimodal input to be processed, input it into the fused model, and obtain the task processing result.
[0024] Optionally, the original task vector in S2 is the difference between the fine-tuning parameters of the multimodal large language model and the parameters of the pre-trained language model.
[0025] The task vector is shown in the following formula (1):
[0026] (1)
[0027] In the formula, represents the original task vector of the -th modality of the multimodal large language model, represents the fine-tuning parameters of the -th modality of the multimodal large language model, represents the parameters of the pre-trained language model.
[0028] Optionally, the fused task vector in S3 is shown in the following formula (2):
[0029] (2)
[0030] In the formula, represents the fused task vector, represents the original task vector of the multimodal large language model of the th modality, represents the number of multimodal large language models.
[0031] The model parameters in S4 are as shown in the following formula (3):
[0032] (3)
[0033] In the formula, represents the model parameters, represents the parameters of the pre-trained language model, represents an adjustable scaling factor.
[0034] Optionally, a modality-specific binary mask is constructed for each multimodal large language model in S5, including:
[0035] Select any parameter dimension of any multimodal large language model, obtain the signs of the selected parameter dimension in the original task vector and in the fused task vector, and determine whether the signs in the original task vector and in the fused task vector are opposite.
[0036] If the signs are opposite, the mask value of the selected parameter dimension is 0.
[0037] If the signs are the same, determine whether the magnitude of the selected parameter dimension in the original task vector is greater than a preset threshold.
[0038] If it is greater than the preset threshold, the mask value of the selected parameter dimension is 1.
[0039] If it is not greater than the preset threshold, the mask value of the selected parameter dimension is 0.
[0040] Optionally, the calculation formula of the binary mask is as shown in the following formula (4):
[0041] (4)
[0042] In the formula, represents the th parameter in the th binary mask matrix, represents the modality, represents the parameter in the mask matrix, represents the th task vector of the parameters, indicating the th parameter of the fusion task vector.
[0043] Optionally, obtain the multi-modal input to be processed in S7, input it into the fusion model, and obtain the task processing result, including:
[0044] Obtain the multi-modal input to be processed, and obtain the corresponding modal vector according to the modal type of the multi-modal input.
[0045] Dynamically select the corresponding binary mask according to the modal type of the multi-modal input.
[0046] Perform modal-specific weighted calculations on the Query, Key, and Value parameters of each layer of Transformer in the fusion model according to the modal vector and the mask, and then obtain the task processing result.
[0047] Optionally, the construction process of the fusion model further includes:
[0048] Obtain a new multi-modal large language model obtained by fine-tuning the pre-trained language model.
[0049] Construct a new fusion model according to the new multi-modal large language model and multiple multi-modal large language models in the fusion model.
[0050] Sparsify the original task vector of the new fusion model using a sparsification strategy to obtain a new sparse vector, and fuse the new sparse vector to obtain a new fusion task vector.
[0051] Construct new model parameters according to the new fusion task vector.
[0052] Construct a new modal-specific binary mask for each multi-modal large language model according to the new fusion task vector.
[0053] Obtain a new fusion model according to the new model parameters and the new binary mask.
[0054] On the other hand, a large language model modal expansion device based on parameter fusion and decoupling is provided. This device is applied to the large language model modal expansion method based on parameter fusion and decoupling. The device includes:
[0055] A fine-tuning module for obtaining multiple multi-modal large language models by fine-tuning the pre-trained language model.
[0056] An extraction module for extracting the task vector of each multi-modal large language model in multiple multi-modal large language models to obtain the original task vector of each multi-modal large language model.
[0057] A fusion module, which is used to sparsify the original task vectors of multiple multimodal large language models by using a sparsification strategy to obtain sparse vectors, and fuse the sparse vectors to obtain a fused task vector.
[0058] A parameter construction module, which is used to construct model parameters according to the fused task vector.
[0059] A binary mask construction module, which is used to construct a modality-specific binary mask for each multimodal large language model according to the fused task vector.
[0060] A fusion model construction module, which is used to construct a fusion model according to the model parameters, the binary mask, and the encoders of multiple multimodal large language models.
[0061] An output module, which is used to obtain a multimodal input to be processed, input it into the fusion model, and obtain a task processing result.
[0062] Optionally, the original task vector is the difference between the fine-tuning parameters of the multimodal large language model and the parameters of the pre-trained language model.
[0063] The task vector is shown in the following formula (1):
[0064] (1)
[0065] In the formula, represents the original task vector of the th modality of the multimodal large language model, represents the fine-tuning parameters of the th modality of the multimodal large language model, represents the parameters of the pre-trained language model.
[0066] Optionally, the fused task vector is shown in the following formula (2):
[0067] (2)
[0068] In the formula, represents the fused task vector, represents the original task vector of the th modality of the multimodal large language model, represents the number of multimodal large language models.
[0069] The model parameters are shown in the following formula (3):
[0070] (3)
[0071] In the formula, represents the model parameters, represents the parameters of the pre-trained language model, Represents an adjustable scaling factor.
[0072] Optionally, the binary mask construction module is further configured to:
[0073] Select any parameter dimension of any multimodal large language model, obtain the signs of the selected parameter dimension in the original task vector and in the fused task vector, and determine whether the signs in the original task vector and in the fused task vector are opposite.
[0074] If the signs are opposite, the mask value of the selected parameter dimension is 0.
[0075] If the signs are the same, determine whether the amplitude of the selected parameter dimension in the original task vector is greater than a preset threshold.
[0076] If it is greater than the preset threshold, the mask value of the selected parameter dimension is 1.
[0077] If it is not greater than the preset threshold, the mask value of the selected parameter dimension is 0.
[0078] Optionally, the calculation formula of the binary mask is as shown in the following formula (4):
[0079] (4)
[0080] In the formula, represents the th parameter in the th binary mask matrix, represents the modality, represents the parameter in the mask matrix, represents the th parameter of the task vector of the th modality, represents the th parameter of the fused task vector.
[0081] Optionally, the output module is further configured to:
[0082] Obtain the multimodal input to be processed, and obtain the corresponding modality vector according to the modality type of the multimodal input.
[0083] Dynamically select the corresponding binary mask according to the modality type of the multimodal input.
[0084] Perform modality-specific weighted calculations on the Query, Key, and Value parameters of each layer of the Transformer of the fusion model according to the modality vector and the mask, and then obtain the task processing result.
[0085] Optionally, the construction process of the fusion model further includes:
[0086] Obtain a new multi-modal large language model obtained by fine-tuning a pre-trained language model.
[0087] Construct a new fusion model based on the new multi-modal large language model and multiple multi-modal large language models in the fusion model.
[0088] Adopt a sparsification strategy to sparsify the original task vector of the new fusion model to obtain a new sparse vector, and fuse the new sparse vector to obtain a new fused task vector.
[0089] Construct new model parameters according to the new fused task vector.
[0090] Construct a new binary mask exclusive to the modality for each multi-modal large language model according to the new fused task vector.
[0091] Obtain a new fusion model according to the new model parameters and the new binary mask.
[0092] On the other hand, a large language model modality expansion device is provided, and the large language model modality expansion device includes: a processor; a memory, and computer-readable instructions are stored on the memory. When the computer-readable instructions are executed by the processor, any one of the methods in the above-mentioned large language model modality expansion method based on parameter fusion and decoupling is implemented.
[0093] On the other hand, a computer-readable storage medium is provided, and at least one instruction is stored in the storage medium. The at least one instruction is loaded and executed by a processor to implement any one of the methods in the above-mentioned large language model modality expansion method based on parameter fusion and decoupling.
[0094] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0095] In the present invention, a method for large language model modality expansion and original performance retention based on parameter fusion and mask decoupling is proposed. The core lies in realizing the expansion of multi-modal large language models without retraining through task vector merging and modality decoupling techniques, and effectively retaining the original task performance. Specifically, the present invention first merges the task vectors of multiple fine-tuned MLLMs through the task vector merging technique, and adjusts them using a scaling factor to generate a fused task vector, thereby effectively integrating multi-modal information into the pre-trained language model. To avoid parameter conflicts between different modalities, the present invention introduces a modality decoupling mechanism, and precisely extracts a specific parameter subset for each modality by constructing a modality-specific binary mask. This method utilizes the principles of direction consistency and dominant saliency to ensure that the parameters of each modality can be processed independently and avoid interference. Finally, through the task vector merging and modality mask techniques, the present invention can not only expand multi-modal capabilities without a large amount of data retraining, but also maintain the performance of the original task, solving the problem of "catastrophic forgetting". This method provides an efficient and flexible multi-modal expansion method, which can realize the seamless fusion of multiple existing models and effectively support the continuous expansion of new tasks and multi-modal capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0097] Figure 1 is a flowchart of a method for large language model modality expansion based on parameter fusion and decoupling provided by an embodiment of the present invention;
[0098] Figure 2 is the overall flowchart of MMER provided by an embodiment of the present invention;
[0099] Figure 3 is a block diagram of a device for large language model modality expansion based on parameter fusion and decoupling provided by an embodiment of the present invention;
[0100] Figure 4 is a schematic structural diagram of a device for large language model modality expansion provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0101] The following describes the technical solutions in the present invention in conjunction with the drawings.
[0102] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0103] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same. "of", "corresponding, relevant" and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference between them is not emphasized, the meanings they intend to express are the same.
[0104] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.
[0105] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.
[0106] The embodiment of the present invention provides a large language model modality expansion method based on parameter fusion and decoupling, which can be implemented by a large language model modality expansion device, which can be a terminal or a server. Figure 1 The flowchart of the large language model modal expansion method based on parameter fusion and decoupling is shown. The processing flow of the method may include the following steps:
[0107] S1. Multiple multimodal large language models are obtained by fine-tuning the pre-trained language model.
[0108] In a feasible implementation, there is MLLMs (Multimodal Large Language Models), all of which are based on the same pre-trained language model Fine-tuned, recorded as parameter sets .
[0109] S2. Extract a task vector for each of the multiple multimodal large language models to obtain an original task vector for each multimodal large language model.
[0110] Optionally, the original task vector in S2 is the difference between the fine-tuning parameters of the multimodal large language model and the parameters of the pre-trained language model.
[0111] Extract task vectors from each fine-tuned model:
[0112] (1)
[0113] Wherein, represents the original task vector of the multi-modal large language model of the th modality, represents the fine-tuning parameters of the multi-modal large language model of the th modality, represents the parameters of the pre-trained language model.
[0114] S3. Sparsify the original task vectors of multiple multi-modal large language models using a sparsification strategy to obtain sparse vectors, and fuse the sparse vectors to obtain a fused task vector.
[0115] In a feasible implementation, a sparsification strategy (retaining the top K% of the parameters in absolute value) is used to sparsify the task vectors to obtain sparse vectors, and then multiple task vectors are fused into a unified task vector based on the consistency of the signs in the parameter dimension :
[0116] (2)
[0117] Wherein, represents the fused task vector, represents the original task vector of the multi-modal large language model of the th modality, represents the number of multi-modal large language models.
[0118] S4. Construct fused model parameters according to the fused task vector, as shown in the following formula (3):
[0119] (3)
[0120] Wherein, represents the model parameters, represents the parameters of the pre-trained language model, represents an adjustable scaling factor for controlling the degree of fusion, usually determined by the validation set or the average performance of each modal task.
[0121] S5. Construct a modality-specific binary mask for each multi-modal large language model according to the fused task vector.
[0122] In a feasible implementation, in order to restore each original model from the fused model or support modality-specific decoupling, the present invention constructs a modality-specific binary mask for identifying the part of the fused vector belonging to modality Valid parameters. The construction of this mask follows two principles:
[0123] 1. Direction consistency: If the sign of a certain parameter dimension in the task vector is opposite in the original and fused vectors, the mask value for this dimension is 0 because there is a strong parameter conflict and it is selected to be masked;
[0124] 2. Dominant significance: If the magnitude of this dimension in the original task vector is large enough (e.g., greater than half of the fused value) and the signs are consistent, it is set to 1. Therefore, this parameter is an important parameter and is selected to be retained.
[0125] The final calculation formula of the mask is as follows:
[0126] (4)
[0127] In the formula, represents the th parameter in the th binary mask matrix, represents the modality, represents the th parameter of the task vector of the th modality, represents the th parameter of the fused task vector. refers to a specific value in the matrix. By applying the Hadamard multiplication to the fused task vector through the mask, the modality can be statically decoupled and restored corresponding model:
[0128] (5)
[0129] The above mask can not only be used for model restoration but also for dynamic modality decoupling in the inference stage. When the model performs multi-modal input processing, the representations of different modalities (images, audio, text, etc.) are respectively input into their respective encoders to obtain modality vectors . After entering the computational graph of the language model, the system dynamically selects the corresponding mask according to the modality type and performs modality-specific weighted calculations on the Query, Key, and Value parameters of each layer of the Transformer. For example, the Query weights of the l-th layer are processed as follows:
[0130] (6)
[0131] where refers to the Query parameter of the l-th layer of the mask matrix corresponding to the first modality, and refers to the parameter of the fused LLM in the Query parameter of the l-th layer, is the parameter of the originally pre-trained LLM, It refers to the average value of all modal mask matrices, which is used as the mask matrix for the text modality. Subsequent calculations are based on this. , to complete the attention operation:
[0132] (7)
[0133] The same method is also used for other linear layers. Through this structural design, the present invention realizes the independent processing of the parameter space between modalities, effectively avoids interference between different modalities, and improves the stability of the fusion model.
[0134] S6. Construct a fusion model according to the model parameters, binary masks, and encoders of multiple multimodal large language models.
[0135] As Figure 2 shown, the present invention provides a method for expanding the modalities of a large language model and retaining the original performance based on parameter fusion and mask decoupling, named MMER.
[0136] Optionally, the construction process of the fusion model further includes:
[0137] Obtain a new multimodal large language model obtained by fine-tuning a pre-trained language model.
[0138] Construct a new fusion model according to the new multimodal large language model and the multiple multimodal large language models in the fusion model.
[0139] Adopt a sparsification strategy to sparsify the original task vector of the new fusion model to obtain a new sparse vector, and fuse the new sparse vector to obtain a new fused task vector.
[0140] Construct new model parameters according to the new fused task vector.
[0141] Construct a new binary mask exclusive to each modality for each multimodal large language model according to the new fused task vector.
[0142] Obtain a new fusion model according to the new model parameters and the new binary masks.
[0143] In a feasible implementation manner, the present invention can also achieve "catastrophic forgetting suppression": when a new task needs to be fine-tuned for a certain modality, the new task model can still be integrated with the original model through the above parameter merging and mask generation process, so as to retain the performance of the old tasks while improving the performance of the new tasks, and solve the problem of performance degradation caused by parameter coverage in traditional fine-tuning methods.
[0144] The MMER method provided by the present invention significantly improves the flexibility and application scope of multimodal large language models through innovative multimodal extension and retention techniques. First of all, this method can expand the multimodal capabilities of the model without retraining, which means that new modalities can be quickly added to the existing system, avoiding the high cost and long cycle of retraining the entire model every time a new modality is added in traditional methods. This feature is particularly suitable for rapid iteration and multimodal system deployment, greatly shortening the development cycle and reducing the consumption of computing resources. In addition, the present invention ensures the independent processing of multimodal information through an accurate modality decoupling mechanism, effectively avoiding mutual interference between different modalities. This not only improves the stability of the model but also ensures the efficient processing of each modality after fusion, avoiding the common performance degradation problems in traditional multimodal fusion. In this way, the model can more accurately understand and respond to multimodal inputs, broadening its application scenarios in complex tasks such as visual question answering and audio-video analysis. The present invention has also successfully solved the "catastrophic forgetting" problem, so that the performance of the original tasks will not be affected when new tasks are expanded. This advantage enables multimodal large language models to maintain high-efficiency and stable performance during long-term use, especially in applications that require continuous addition of new tasks or new modalities, showing significant advantages. For systems that need to be deployed for a long time and continuously optimized (such as intelligent assistants, autonomous driving, etc.), the present invention provides a more sustainable solution.
[0145] S7. Obtain the multimodal input to be processed and input it into the fusion model to obtain the task processing result.
[0146] To verify the effectiveness of this solution, the present invention conducted three experiments on multiple tasks to test the method of the present invention. In the multimodal extension experiment, the MMER method demonstrated significant advantages, significantly exceeding the NaiveMC framework, proving its effectiveness in expanding multimodal capabilities and improving the performance of merged MLLMs. Compared with other baseline methods, MMER achieved the best performance in multiple multimodal tasks, showing that it effectively realized the expansion of multimodal capabilities without retraining. In the multimodal performance retention experiment, the MMER method could effectively retain the performance of the original tasks. Finally, in the experiment of alleviating catastrophic forgetting, MMER demonstrated excellent robustness, being able to maintain the performance of the original tasks after fine-tuning on new tasks, avoiding the performance loss problems caused by task fusion in traditional methods.
[0147] The present invention proposes a fusion method that can expand the multimodal capabilities of large language models without training and maintain the original performance, aiming to solve the following core technical problems: First, how to accurately identify and separate modality-specific parameters when merging multiple MLLMs to avoid performance loss caused by parameter conflicts between modalities; Second, how to effectively retain the ability performance of the original task while integrating new modalities or new tasks and avoid the occurrence of the "catastrophic forgetting" phenomenon; Third, how to provide a general parameter processing framework so that whether the original MLLM adopts a full-parameter fine-tuning or lightweight fine-tuning strategy, it can be incorporated into the fusion system for unified processing.
[0148] The design goal of the present invention is to achieve lossless modality expansion and original performance retention of multimodal large language models through a structured parameter fusion and decoupling mechanism without relying on costly retraining. This invention can be widely applied to the construction of multimodal systems, the integration and upgrade of existing models, the ability recombination of open-source models, and cross-modal task development under resource-constrained conditions, and has important research value and engineering promotion prospects.
[0149] In an embodiment of the present invention, a method for modality expansion and original performance retention of a large language model based on parameter fusion and mask decoupling is proposed. The core lies in realizing the expansion of a multimodal large language model without retraining through task vector merging and modality decoupling techniques, and effectively retaining the performance of the original task. Specifically, the present invention first merges the task vectors of multiple fine-tuned MLLMs through the task vector merging technique, and adjusts them using a scaling factor to generate a fused task vector, thereby effectively integrating multimodal information into the pre-trained language model. To avoid parameter conflicts between different modalities, the present invention introduces a modality decoupling mechanism, and precisely extracts each modality-specific parameter subset by constructing a modality-specific binary mask. This method uses the principles of direction consistency and dominant significance to ensure that the parameters of each modality can be processed independently and avoid interference. Finally, through the task vector merging and modality mask techniques, the present invention can not only expand multimodal capabilities without a large amount of data retraining, but also maintain the performance of the original task, solving the "catastrophic forgetting" problem. This method provides an efficient and flexible multimodal expansion method, which can achieve seamless fusion of multiple existing models and effectively support the continuous expansion of new tasks and multimodal capabilities.
[0150] Figure 3 is a block diagram of a modality expansion device for a large language model based on parameter fusion and decoupling shown according to an exemplary embodiment. This device is used for the modality expansion method of a large language model based on parameter fusion and decoupling. Refer to Figure 3, the device includes a fine-tuning module 310, an extraction module 320, a fusion module 330, a parameter construction module 340, a binary mask construction module 350, a fusion model construction module 360, and an output module 370. Among them:
[0151] The fine-tuning module 310 is used to obtain multiple multimodal large language models by fine-tuning a pre-trained language model.
[0152] The extraction module 320 is used to extract task vectors for each multimodal large language model among the multiple multimodal large language models to obtain the original task vectors of each multimodal large language model.
[0153] The fusion module 330 is used to sparsify the original task vectors of the multiple multimodal large language models using a sparsification strategy to obtain sparse vectors, and fuse the sparse vectors to obtain a fused task vector.
[0154] The parameter construction module 340 is used to construct model parameters according to the fused task vector.
[0155] The binary mask construction module 350 is used to construct a modality-specific binary mask for each multimodal large language model according to the fused task vector.
[0156] The fusion model construction module 360 is used to construct a fusion model according to the model parameters, the binary mask, and the encoders of the multiple multimodal large language models.
[0157] The output module 370 is used to obtain the multimodal input to be processed, input it into the fusion model, and obtain the task processing result.
[0158] In an embodiment of the present invention, a method for modality extension and original performance retention of a large language model based on parameter fusion and mask decoupling is proposed. The core lies in realizing the extension of a multi-modal large language model without retraining through task vector merging and modality decoupling techniques, and effectively retaining the original task performance. Specifically, the present invention first merges the task vectors of multiple fine-tuned MLLMs through the task vector merging technique, and adjusts them using a scaling factor to generate a fused task vector, thereby effectively integrating multi-modal information into the pre-trained language model. To avoid parameter conflicts between different modalities, the present invention introduces a modality decoupling mechanism, and precisely extracts the parameter subsets specific to each modality by constructing modality-specific binary masks. This method utilizes the principles of direction consistency and dominant salience to ensure that the parameters of each modality can be processed independently and avoid interference. Finally, through the task vector merging and modality masking techniques, the present invention can not only extend the multi-modal capabilities without retraining a large amount of data, but also maintain the performance of the original tasks, solving the problem of "catastrophic forgetting". This method provides an efficient and flexible multi-modal extension method, which can achieve seamless integration of multiple existing models and effectively support the continuous expansion of new tasks and multi-modal capabilities.
[0159] Figure 4 It is a schematic structural diagram of a large language model modality extension device provided by an embodiment of the present invention. As Figure 4 shown, the large language model modality extension device may include the above Figure 3 shown large language model modality extension device based on parameter fusion and decoupling. Optionally, the large language model modality extension device 410 may include a first processor 2001.
[0160] Optionally, the large language model modality extension device 410 may further include a memory 2002 and a transceiver 2003.
[0161] Among them, the first processor 2001, the memory 2002, and the transceiver 2003 may be connected through a communication bus, for example.
[0162] Next, in conjunction with Figure 4 specific introduction will be made to each component of the large language model modality extension device 410:
[0163] Among them, the first processor 2001 is the control center of the large language model modality expansion device 410, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0164] Optionally, the first processor 2001 can execute various functions of the large language model modality expansion device 410 by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0165] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 4 CPU0 and CPU1 shown in
[0166] In a specific implementation, as an embodiment, the large language model modality expansion device 410 may also include multiple processors, such as Figure 4 the first processor 2001 and the second processor 2004 shown in
[0167] Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0168] Optionally, the memory 2002 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but not limited thereto. The memory 2002 can be integrated with the first processor 2001 or exist independently and be coupled to the first processor 2001 through the interface circuit ( Figure 4 not shown) of the large language model modality extension device 410. The embodiments of the present invention do not make specific limitations in this regard.
[0169] The transceiver 2003 is used to communicate with network devices or terminal devices.
[0170] Optionally, the transceiver 2003 can include a receiver and a transmitter ( Figure 4 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.
[0171] Optionally, the transceiver 2003 can be integrated with the first processor 2001 or exist independently and be coupled to the first processor 2001 through the interface circuit ( Figure 4 not shown) of the large language model modality extension device 410. The embodiments of the present invention do not make specific limitations in this regard.
[0172] It should be noted that Figure 4 the structure of the large language model modality extension device 410 shown in does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0173] In addition, the technical effects of the large language model modality extension device 410 can refer to the technical effects of the large language model modality extension method based on parameter fusion and decoupling described in the above method embodiments, which will not be elaborated here.
[0174] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0175] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM) and direct rambus RAM (DR RAM).
[0176] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0177] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context.
[0178] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0179] It should be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0180] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0181] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0182] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0183] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0184] In addition, the functional units in each embodiment of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0185] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0186] As described above, the above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for expanding the modality of a large language model based on parameter fusion and decoupling, characterized in that The method includes: S1. Obtain multiple multimodal large language models by fine-tuning a pre-trained language model; S2. Extract task vectors for each of the multiple multimodal large language models to obtain the original task vectors of each multimodal large language model; S3. Sparsify the original task vectors of the multiple multimodal large language models using a sparsification strategy to obtain sparse vectors, and fuse the sparse vectors to obtain a fused task vector; S4. Construct model parameters according to the fused task vector; S5. Construct a modality-specific binary mask for each multimodal large language model according to the fused task vector; S6. Construct a fusion model according to the model parameters, the binary mask, and the encoders of the multiple multimodal large language models; S7. Obtain a multimodal input to be processed, input it into the fusion model, and obtain a task processing result, where the multimodal representation includes images, audio, and text.
2. The method for expanding the modality of a large language model based on parameter fusion and decoupling according to claim 1, wherein The original task vector in S2 is the difference between the fine-tuning parameters of the multimodal large language model and the parameters of the pre-trained language model; The task vector is as shown in the following formula (1): (1) In the formula, represents the original task vector of the th modality of the multimodal large language model, represents the fine-tuning parameter of the th modality of the multimodal large language model, represents the parameter of the pre-trained language model.
3. The method for expanding the modality of a large language model based on parameter fusion and decoupling according to claim 1, characterized in that The fused task vector in S3 is as shown in the following formula (2): (2) In the formula, represents the fused task vector, represents the original task vector of the multi-modal large language model of the th modality, and represents the number of multi-modal large language models; The model parameters in S4 are as shown in the following formula (3): (3) In the formula, represents the model parameters, represents the parameters of the pre-trained language model, represents an adjustable scaling factor.
4. The method for expanding the modality of a large language model based on parameter fusion and decoupling according to claim 1, wherein Constructing a modality-specific binary mask for each multimodal large language model in S5 includes: Select any parameter dimension of any multimodal large language model, obtain the sign of the selected parameter dimension in the original task vector and the sign in the fused task vector, and determine whether the sign in the original task vector is opposite to the sign in the fused task vector; If the signs are opposite, the mask value of the selected parameter dimension is 0; If the signs are the same, determine whether the magnitude of the selected parameter dimension in the original task vector is greater than a preset threshold; If it is greater than the preset threshold, the mask value of the selected parameter dimension is 1; If it is not greater than the preset threshold, the mask value of the selected parameter dimension is 0.
5. The method for expanding the modality of a large language model based on parameter fusion and decoupling according to claim 4, wherein The calculation formula of the binary mask is as shown in the following formula (4): (4) In the formula, represents the th parameter in the th binary mask matrix, represents the modality, represents the parameter in the mask matrix, represents the th parameter of the task vector of the th modality, represents the th parameter of the fused task vector.
6. The method for expanding the modality of a large language model based on parameter fusion and decoupling according to claim 1, wherein, Obtaining a multimodal input to be processed, inputting it into the fusion model, and obtaining a task processing result in S7 includes: Obtain a multimodal input to be processed, and obtain a corresponding modality vector according to the modality type of the multimodal input; Dynamically select a corresponding binary mask according to the modality type of the multimodal input; Perform modality-specific weighted calculations on the Query, Key, and Value parameters of each layer of Transformer in the fusion model according to the modality vector and the mask, and then obtain a task processing result.
7. The method for expanding the modality of a large language model based on parameter fusion and decoupling according to claim 1, wherein The construction process of the fusion model further includes: Obtain a new multimodal large language model obtained by fine-tuning the pre-trained language model; Construct a new fusion model according to the new multimodal large language model and the multiple multimodal large language models in the fusion model; Sparsify the original task vector of the new fusion model using a sparsification strategy to obtain a new sparse vector, and fuse the new sparse vector to obtain a new fused task vector; Construct new model parameters according to the new fused task vector; Construct a new modality-specific binary mask for each multimodal large language model according to the new fused task vector; Obtain a new fused model according to the new model parameters and the new binary mask.
8. An apparatus for expanding the modality of a large language model based on parameter fusion and decoupling, the apparatus for expanding the modality of a large language model based on parameter fusion and decoupling is used to implement the method for expanding the modality of a large language model based on parameter fusion and decoupling as described in any one of claims 1-7, characterized in that, The device includes: A fine-tuning module for obtaining multiple multimodal large language models by fine-tuning a pre-trained language model; An extraction module for extracting task vectors from each of the multiple multimodal large language models to obtain the original task vectors of each multimodal large language model; A fusion module for sparsifying the original task vectors of multiple multimodal large language models using a sparsification strategy to obtain sparse vectors, and fusing the sparse vectors to obtain a fused task vector; A parameter construction module for constructing model parameters according to the fused task vector; A binary mask construction module for constructing a modality-specific binary mask for each multimodal large language model according to the fused task vector; A fused model construction module for constructing a fused model according to the model parameters, the binary mask, and the encoders of multiple multimodal large language models; An output module for obtaining a multimodal input to be processed, inputting it into the fused model, and obtaining a task processing result, where the multimodal representation includes images, audio, and text.
9. A large language model modality expansion device, characterized in that, The large language model modality expansion device includes: A processor; A memory having computer-readable instructions stored thereon, and when the computer-readable instructions are executed by the processor, implementing the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-modal data processing method for enhancing large language model
CN118070227A
Social network false message detection method based on large model and multi-modal fusion
CN119475066A