Data-free model fusion method based on orthogonal and projection double-space optimization

By adopting orthogonal and projection dual-space optimization technology in data-free model fusion, parameter conflict problem is solved, efficient model merging is achieved, performance and applicability are significantly improved, and the current technology is leading.

CN120234765AActive Publication Date: 2025-07-01CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510715352.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

Existing data-free fusion methods are prone to parameter conflicts, resulting in a decline in overall performance of the fusion model.

Method used

The method based on orthogonal and projection dual-space optimization is adopted, and the orthogonal subspace and dual-space constraint merging under calculation constraints is realized, the direct merge of multiple models is maximized, task sharing information is minimized, parameter conflicts are minimized.

Benefits of technology

It significantly reduces computing cost and privacy risks, improves the overall performance and applicability of data-free model fusion, avoids information loss, and reaches the leading level of current technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234765A_ABST
    Figure CN120234765A_ABST
Patent Text Reader

Abstract

The invention discloses a data-free model fusion method based on orthogonal and projection double-space optimization, which belongs to the technical field of data model fusion, is used for data-free model fusion, and comprises the following steps of: performing singular value decomposition on a shared orthogonal subspace, removing redundant vectors to obtain a redundancy-free subspace, and performing data fusion on the redundancy-free subspace; key parameters are projected into an orthogonal subspace parameter module through orthogonal space projection, an orthogonal subspace optimizer carries out gradient information updating and then feeds back to the orthogonal subspace parameter module, and a projection subspace optimizer carries out gradient updating and feeds back to a double-space constraint device. And carrying out model fusion on the pre-training model, the output of the projection subspace optimizer and the reformed vector. The method is suitable for merging a plurality of expert models in a multi-task scene, does not need to depend on additional data or retraining, and remarkably reduces the calculation cost and privacy risk; and task sharing information is maximized and parameter conflicts are minimized simultaneously in the subspace, so that the overall performance and applicability of data-free model fusion are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a data-free model fusion method based on the optimization of orthogonal and projection dual spaces, belonging to the technical field of data model fusion. Background Art

[0002] With the widespread application of the pre-training - fine-tuning paradigm in various artificial intelligence tasks, the number of expert models for different downstream tasks has increased sharply. Although each fine-tuned model has achieved significant performance improvement in its corresponding task, deploying multiple fine-tuned models separately will lead to a substantial increase in storage resources and operation and maintenance costs. To solve the above problems, model fusion technology has emerged. Model fusion forms a unified model that can take into account the requirements of multiple tasks by merging multiple fine-tuned models at the parameter level, thus realizing the unified management and deployment of multiple expert models. Compared with integrating or re-ranking the results of each model in the inference stage, the model fusion scheme directly completes the merger in the parameter space, which can not only significantly reduce the total storage volume of the model, but also avoid the computational overhead caused by the multi-model call during inference.

[0003] Currently, the mainstream model fusion methods can be divided into two categories: test-time adaptation based and data-free model fusion. The former usually needs to access the original or approximate datasets of each task to compensate or calibrate the merged results, but due to the need for additional data access, its applicability in scenarios where data is unavailable or security and privacy are concerned is limited. In response to this limitation, data-free model fusion methods have emerged, which can directly merge multiple models in the parameter space by only using the parameters of the pre-training and each fine-tuned model, without any additional data or re-training. Currently, the main data-free model fusion methods are as follows:

[0004] 1) Linear interpolation method. This type of method quickly retains the key information of each model by weighted averaging the parameter values of each model according to preset weights. However, simple interpolation cannot effectively solve the parameter conflict problem, often resulting in a decline in the overall performance of the fusion model; when the parameter differences between the models participating in the fusion are large, the mean operation is likely to introduce conflicts, and the effect of this method highly depends on the diversity and quality of the selected fine-tuned models, limiting the upper limit of its performance improvement.

[0005] 2) Weighted task arithmetic method. This type of method first obtains each task vector by subtracting the pre-trained model weight from the weight of the fine-tuned model, and then controls the behavior of the resulting model by performing arithmetic operations such as weighted addition or subtraction on the task vectors; simple addition and subtraction operations are difficult to fundamentally alleviate the conflict problem between task vectors, and the allocation of weights for each task vector lacks a systematic design, easily destroying the original beneficial parameters, thereby affecting the performance of the fused model on certain tasks.

[0006] 3) Task arithmetic methods based on subspaces. Such methods map each model parameter into an orthogonal or projection subspace for fusion, and reduce interference between tasks through subspace coordinate transformation. However, they focus on the characterization of a single subspace, ignore effective features in other potential subspaces, and it is difficult to balance information sharing and conflict elimination in multiple dimensions. Summary of the Invention

[0007] The purpose of the present invention is to provide a data-free model fusion method optimized based on orthogonal and projection dual spaces to solve the problem of parameter conflicts that easily occur in existing data-free model fusion methods.

[0008] A data-free model fusion method optimized based on orthogonal and projection dual spaces includes calculating an orthogonal subspace under constraints, merging dual-space constraints, and fusing the final model.

[0009] The calculation of the orthogonal subspace under constraints includes performing subtraction on the fine-tuned model set and the pre-trained model to obtain task vectors, performing singular value decomposition on the task vectors, splicing the decomposition results into a shared orthogonal subspace, performing singular value decomposition on the shared orthogonal subspace, removing redundant vectors to obtain a non-redundant subspace, and projecting key parameters into the orthogonal subspace parameter module through orthogonal space projection; the fine-tuned model set has branches for adaptive scaling, adjusting according to the layer scaling coefficient to form a reorganized vector, and then inputting it into the orthogonal subspace optimizer. The orthogonal subspace parameter module outputs gradient information, which is fused with the task vector through task vector preprocessing and then input into the orthogonal subspace optimizer. After the orthogonal subspace optimizer updates the gradient information, it feeds back to the orthogonal subspace parameter module.

[0010] Decompose the fine-tuned model set into task parameters, element parameters, and layer parameters. The merging of dual-space constraints includes inputting the reorganized vector, orthogonal subspace parameters, task parameters, element parameters, and layer parameters into the dual-space constraint device together, and then inputting it into the projection subspace optimizer. The projection subspace optimizer updates the gradient and feeds back to the dual-space constraint device.

[0011] The fusion of the final model includes fusing the pre-trained model, the output of the projection subspace optimizer, and the reorganized vector to finally obtain a fusion model, completing the data-free model fusion.

[0012] The subtraction operation includes:

[0013] ;

[0014] where is the th task vector, is the pre-trained model, which is the initial model parameters pre-trained based on a large-scale general dataset, is the th model parameter fine-tuned for different downstream tasks.

[0015] The singular value decomposition of the task vector includes:

[0016] ;

[0017] In the formula, is the left singular matrix after the decomposition of the task vector, is the singular value matrix after the decomposition of the task vector, is the right singular matrix after the decomposition of the task vector, represents matrix transpose.

[0018] The process of stitching the decomposition results into the shared orthogonal subspace includes retaining 's first columns to extract the main spatial features in the task vector , and stitching all together to construct the shared orthogonal subspace :

[0019] ;

[0020] In the formula, represents the stitching operation, represents 's maximum value.

[0021] The singular value decomposition of the shared orthogonal subspace includes:

[0022] ;

[0023] In the formula, is the left singular matrix after the decomposition of the shared orthogonal subspace, is the singular value matrix after the decomposition of the shared orthogonal subspace, is the right singular matrix after the decomposition of the shared orthogonal subspace. Store the key parameters into the main spatial features in the shared orthogonal subspace to form a non-redundant subspace.

[0024] Orthogonal space projection includes:

[0025] ;

[0026] In the formula, is a matrix of any dimension, is 's projection onto ;

[0027] The orthogonal subspace optimizer updates the gradient information including:

[0028] ;

[0029] Wherein, is the gradient of each layer, is the .

[0030] Orthogonal subspace parameter is:

[0031] ;

[0032] Wherein, is the after the update of the gradient information, , is the layer-by-layer scaling coefficient;

[0033] In the update of the gradient information, the loss function is:

[0034] .

[0035] The vector after rearrangement is:

[0036] ;

[0037] ;

[0038] Wherein, is the amplification coefficient for controlling the scaling coefficient.

[0039] The orthogonal subspace optimizer performs layer-by-layer iterative optimization until convergence, and calculates the gradient loss of each layer:

[0040] ;

[0041] Wherein, is the adjustable element parameter that can be learned for all elements within the layer, is the constraint parameter, is the of the layer, is the of the update adjustment number , dynamically balancing the contribution weights of and , is the of the .

[0042] Fusion Model is:

[0043] ;

[0044] In the formula, is the maximum value of

[0045] Compared with the prior art, the present invention has the following beneficial effects: The present invention is applicable to the combination of multiple expert models in a multi-task scenario, without relying on additional data or retraining, significantly reducing the computational cost and privacy risk; by simultaneously maximizing the task-sharing information and minimizing the parameter conflict in the subspace, to further improve the overall performance and applicability of the data-free model fusion; through the dual-space joint optimization and adaptive parameter fusion mechanism, efficient model combination is achieved without additional data or training; on the one hand, during the multi-task fusion process, it maximally suppresses the mutual interference between different tasks, significantly improving the overall performance of the fusion model in multi-modal tasks such as vision and language, reaching the leading level of the current technology, and at the same time effectively overcoming the information loss problem caused by the traditional single-subspace fusion strategy. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 is the overall flowchart of the present invention.

[0047] Figure 2 is the average accuracy of different fusion methods on ViT-B / 32;

[0048] Figure 3 is the influence of the constraint parameter P on different tasks. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0050] A data-free model fusion method based on orthogonal and projection dual-space optimization, including calculating the orthogonal subspace under constraints, dual-space constraint combination, and fusing the final model;

[0051] The orthogonal subspace under the described computational constraints includes performing a subtraction operation on the fine-tuning model set and the pre-trained model to obtain a task vector, performing singular value decomposition on the task vector, concatenating the decomposition results into the shared orthogonal subspace, performing singular value decomposition on the shared orthogonal subspace, removing redundant vectors to obtain a non-redundant subspace, and projecting the key parameters into the orthogonal subspace parameter module through orthogonal space projection; the fine-tuning model set has branches for adaptive scaling, adjusting according to the layer scaling coefficients to form a reorganized vector, and then inputting it into the orthogonal subspace optimizer. The orthogonal subspace parameter module outputs gradient information, which is fused with the task vector through task vector preprocessing and then input into the orthogonal subspace optimizer. After the orthogonal subspace optimizer updates the gradient information, it feeds back to the orthogonal subspace parameter module;

[0052] Decompose the fine-tuning model set into task parameters, element parameters, and layer parameters. The dual-space constraint merging includes inputting the reorganized vector, orthogonal subspace parameters, task parameters, element parameters, and layer parameters together into the dual-space constraint device, and then inputting it into the projection subspace optimizer. The projection subspace optimizer performs gradient update and feeds back to the dual-space constraint device;

[0053] The described fused final model includes performing model fusion on the pre-trained model, the output of the projection subspace optimizer, and the reorganized vector to finally obtain a fused model, completing model fusion without data.

[0054] The subtraction operation includes:

[0055] ;

[0056] In the formula, is the th task vector, is the pre-trained model, which is the initial model parameters pre-trained based on a large-scale general dataset, is the th model parameters after fine-tuning the fine-tuning model set for different downstream tasks.

[0057] Performing singular value decomposition on the task vector includes:

[0058] ;

[0059] In the formula, is the left singular matrix after the task vector decomposition, is the singular value matrix after the task vector decomposition, is the right singular matrix after the task vector decomposition, represents matrix transpose.

[0060] Concatenating the decomposition results into the shared orthogonal subspace includes retaining the first Column, extract the main spatial features in the task vector , and splice all together to construct a shared orthogonal subspace :

[0061] ;

[0062] In the formula, represents the splicing operation, represents the maximum value of

[0063] The singular value decomposition of the shared orthogonal subspace includes:

[0064] ;

[0065] In the formula, is the left singular matrix after the decomposition of the shared orthogonal subspace, is the singular value matrix after the decomposition of the shared orthogonal subspace, is the right singular matrix after the decomposition of the shared orthogonal subspace, and store the key parameters into the main spatial features in the shared orthogonal subspace to form a non-redundant subspace

[0066] The orthogonal space projection includes:

[0067] ;

[0068] In the formula, is a matrix of any dimension, is projected onto ;

[0069] The orthogonal subspace optimizer updates the gradient information, including:

[0070] ;

[0071] In the formula, is the gradient of each layer, is after the gradient information is updated

[0072] The orthogonal subspace parameter is:

[0073] ;

[0074] In the formula, is the th after the gradient information is updated, is the layer-by-layer scaling factor;

[0075] The loss function during gradient information update is:[[]]

[0076] .[[]]

[0077] The vector after reorganization is:[[]]

[0078] ;[[]]

[0079] ;[[]]

[0080] In the formula, is the amplification factor for controlling the scaling factor.[[]]

[0081] The orthogonal subspace optimizer performs layer-by-layer iterative optimization until convergence, and calculates the gradient loss of each layer :[[]]

[0082] ;[[]]

[0083] In the formula, is the adjustable element parameter that can be learned for all elements within the th layer, is the constraint parameter, is the of the th layer, is the of the update adjustment number of the th layer, dynamically balancing the contribution weights of and , is the of the th layer.[[]]

[0084] The fusion model is:[[]]

[0085] ;[[]]

[0086] In the formula, is the maximum value.[[]]

[0087] The technical process of the present invention is as shown in Figure 1As shown below. The present invention has been experimentally verified in two main fields: visual tasks and natural language processing tasks. In visual tasks, experiments were carried out based on two classic vision Transformer architectures (ViT-B / 32 and ViT-L / 14) of the CLIP framework, covering eight categories of image recognition scenarios, including fine-grained classification (StanfordCars), remote sensing image analysis (RESISC45 and EuroSAT), street view digit recognition (SVHN), traffic sign recognition (GTSRB), scene classification (SUN397), texture recognition (DTD), and handwritten digit benchmark (MNIST). The experimental results were evaluated using top-1 accuracy (the unit is percentage, and the unit in the tables of the present invention is percentage). In natural language processing tasks, the base and large versions of the Flan-T5 series were selected to verify the performance on eight core tasks (CoLA, MNLI, MRPC, QNLI, QQP, RTE, SST2, STSB) in the GLUE multi-task evaluation system.

[0088] The multi-task performance results when combining the ViT-B / 32 model on eight core task visual benchmarks are shown in Table 1.

[0089] Table 1 Multi-task performance results when combining the ViT-B / 32 model on eight core task visual benchmarks

[0090] ;

[0091] In Table 1, each English in the method column is the English representation of various methods in the prior art.

[0092] The multi-task performance results when combining the ViT-L / 14 model on eight core task visual benchmarks are shown in Table 2.

[0093] Table 2 Multi-task performance results when combining the ViT-L / 14 model on eight core task visual benchmarks

[0094] ;

[0095] The multi-task performance results when combining the Flan-T5-base (LoRA fine-tuning) model on all eight language tasks are shown in Table 3, and each English in the method column is the English representation of various methods in the prior art.

[0096] Table 3 Multi-task performance results when combining the Flan-T5-base (LoRA fine-tuning) model on all eight language tasks

[0097] ;

[0098] The multi-task performance results when merging the Flan-T5-large (LoRA fine-tuning) model on all eight language tasks are shown in Table 4.

[0099] Table 4 Multi-task performance results when merging the Flan-T5-large (LoRA fine-tuning) model on all eight language tasks

[0100] ;

[0101] The average accuracies of different fusion methods on ViT-B / 32 for a large number of tasks are shown in Table 5.

[0102] Table 5 Average accuracies of different fusion methods on ViT-B / 32

[0103] ;

[0104] The generalization results for two unseen tasks when merging the ViT-B / 32 model on six tasks are shown in Table 6.

[0105] Table 6 Generalization results for two unseen tasks when merging the ViT-B / 32 model on six tasks

[0106] ;

[0107] The influence of different components on task accuracy is shown in Table 7.

[0108] Table 7 Influence of different components on task accuracy

[0109] ;

[0110] The present invention also lists the average accuracies of different fusion methods on ViT-B / 32 for different numbers of tasks as Figure 2 shown, and the influence of the constraint parameter P on different tasks as Figure 3 shown. Figure 2 Among them, the vertical axis is the average accuracy rate in percentage, the horizontal axis is the task number without unit, ORION (Ours) represents the result of the present invention, and the rest are the results of other existing technologies. Figure 3 Among them, the left vertical axis is the average accuracy rate of visual tasks in percentage, the right vertical axis is the average accuracy rate of natural language processing tasks in percentage, and the horizontal axis is the different values of the constraint parameter P without unit.

[0111] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features, and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data-free model fusion method based on the optimization of orthogonal and projection dual spaces, characterized in that, Including the orthogonal subspace under computational constraints, the merging of dual-space constraints, and the fusion of the final model; The orthogonal subspace under computational constraints includes performing subtraction on the fine-tuned model set and the pre-trained model to obtain a task vector, performing singular value decomposition on the task vector, splicing the decomposition result into the shared orthogonal subspace, performing singular value decomposition on the shared orthogonal subspace, removing redundant vectors to obtain a non-redundant subspace, and projecting key parameters into the orthogonal subspace parameter module through orthogonal space projection; The fine-tuned model set has branches for adaptive scaling, adjusting according to the layer scaling coefficient to form a reorganized vector, and then inputting it into the orthogonal subspace optimizer. The orthogonal subspace parameter module outputs gradient information, which is fused with the task vector through task vector preprocessing and then input into the orthogonal subspace optimizer. After the orthogonal subspace optimizer updates the gradient information, it feeds back to the orthogonal subspace parameter module; Decompose the fine-tuned model set into task parameters, element parameters, and layer parameters. The merging of dual-space constraints includes inputting the reorganized vector, orthogonal subspace parameters, task parameters, element parameters, and layer parameters into the dual-space constraint device together, and then inputting it into the projection subspace optimizer. The projection subspace optimizer updates the gradient and feeds back to the dual-space constraint device; The fusion of the final model includes fusing the pre-trained model, the output of the projection subspace optimizer, and the reorganized vector to finally obtain a fusion model, completing the fusion of the model without data.

2. The method for fusing a data-free model based on orthogonal and projection dual-space optimization according to claim 1, wherein The subtraction operation includes: ; In the formula, is the th task vector, is the pre-trained model, which is the initial model parameters pre-trained based on a large-scale general dataset, is the th model parameters fine-tuned for different downstream tasks in the fine-tuned model set.

3. A data-free model fusion method based on orthogonal and projection dual-space optimization according to claim 2, wherein Performing singular value on the task vector includes: ; In the formula, is the left singular matrix after the decomposition of the task vector, is the singular value matrix after the decomposition of the task vector, is the right singular matrix after the decomposition of the task vector, represents the matrix transpose.

4. A data-free model fusion method based on orthogonal and projection dual space optimization according to claim 3, characterized in that, The process of splicing the decomposition result into the shared orthogonal subspace includes retaining the first columns, extracting the main spatial features in the task vector , and splicing all together to construct the shared orthogonal subspace : ; In the formula, represents the splicing operation, represents the maximum value of.

5. A data-free model fusion method based on orthogonal and projection dual space optimization according to claim 4, characterized in that Performing singular value decomposition on the shared orthogonal subspace includes: ; In the formula, is the left singular matrix after the shared orthogonal subspace decomposition, is the singular value matrix after the shared orthogonal subspace decomposition, is the right singular matrix after the shared orthogonal subspace decomposition. Store the key parameters into the main spatial feature in the shared orthogonal subspace to form a non-redundant subspace.

6. A data-free model fusion method based on orthogonal and projection dual-space optimization according to claim 5, wherein Orthogonal space projection includes: ; wherein, is a matrix of any dimension, is the projection of onto; The orthogonal subspace optimizer updates the gradient information includes: ; In the formula, is the gradient of each layer, is after the gradient information is updated .

7. A data-free model fusion method based on orthogonal and projection dual space optimization according to claim 6, characterized in that, Orthogonal subspace parameter is as follows: ; In the formula, is the after the update of the gradient information, and is the layer-by-layer scaling factor; Loss function during gradient information update is as follows: 。 8. A data-free model fusion method based on orthogonal and projection double-space optimization according to claim 7, characterized in that Reorganized vector is as follows: ; ; In the formula, is the magnification factor for controlling the scaling factor.

9. A data-free model fusion method based on orthogonal and projection dual-space optimization according to claim 8, wherein The orthogonal subspace optimizer performs layer-by-layer iterative optimization until convergence, and calculates the gradient loss of each layer : ; Wherein, is the adjustable element parameter that can be learned for all elements in the layer, is the constraint parameter, is the of the layer, is the of the update adjustment number , dynamically balancing the contribution weights of and , is the of the .

10. A data-free model fusion method based on orthogonal and projection dual space optimization according to claim 9, characterized in that, Fusion model is as follows: ; In the formula, is the maximum value.

Citation Information

Patent Citations

  • Image classification pre-training model continuous learning method based on low-rank adaptive combination

    CN117611913A

  • Model training method, voice processing method and corresponding device

    CN118865946A

  • Pre-training visual model parameter fine tuning method based on singular value

    CN119251621A

  • Large model parameter fine tuning method, device, equipment, medium and product

    CN119398127A

  • Parameter fine tuning method and device of pre-training model, equipment and medium

    CN119849576A