A model fusion method and device
Patent Information
- Application Number
- CN202610891955.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-25
AI Technical Summary
[0009]本发明针对现有技术中模型融合存在的特征强度失衡、弱特征保留不足、融合稳定性差等问题,提供一种模型融合方法及装置,通过特征尺度对齐、加权掩码生成、推理参数反缩放修正实现多任务参数统一尺度融合,结合掩码精准提取任务相关参数,在保持参数压缩效率的同时,显著提升弱特征任务性能与整体融合效果
1、缓解特征强弱不平衡:通过尺度对齐将不同任务参数映射至统一数值空间,避免强特征掩盖弱特征;
Smart Images

Figure CN122818218A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer and artificial intelligence technology, and specifically to a model fusion method and apparatus. Background Technology
[0002] Fine-tuning pre-trained large models for specific software engineering tasks has gradually become a paradigm because it not only leverages the rich knowledge learned in advance by the large model but also allows for customization for specific tasks. However, as the number of tasks requiring customization increases, the number of checkpoints obtained through fine-tuning also increases. This necessitates the simultaneous management, storage, and deployment of these checkpoints, ultimately posing challenges to the stable and sustainable development of large models.
[0003] Model fusion, as an emerging approach, operates on fine-tuned parameters (i.e., the difference between the fine-tuned parameters and the pre-trained model parameters) and merges fine-tuning checkpoints from multiple different tasks into one, thus enabling a single checkpoint to handle multiple tasks. The core function of model fusion lies in parameter compression and sharing: by fusing multiple models, it avoids repeatedly storing similar or redundant parameters, achieving efficient parameter compression while preserving the original performance of the fine-tuned model as much as possible. This significantly reduces the overall model size, improves resource utilization efficiency, significantly alleviates management pressure, and reduces GPU memory overhead during deployment. Furthermore, the fused model possesses multi-task processing capabilities, which is beneficial for deployment in resource-constrained environments (such as edge devices or mobile devices). However, current model fusion methods often face a challenge: strong features (dimensions with large parameter variations) typically overshadow weak features (dimensions with small parameter variations), leading to significant performance degradation in tasks with weak features (tasks with relatively small parameter variations). Inspired by feature normalization, we first normalize the fine-tuning parameters of different ranges to the same scale before fusion, and then perform subsequent operations. This allows the features of different tasks to be preserved, thereby greatly improving the performance of weak feature tasks after fusion without affecting the performance of other tasks.
[0004] With the development of deep learning technology, different tasks often require the training of independent models, leading to a continuous expansion of model parameter size and problems such as high storage costs, deployment difficulties, and low inference efficiency. To address these issues, model fusion technology has gradually gained attention. Model fusion refers to integrating the parameters of multiple trained models (usually models fine-tuned for different tasks or data distributions) to generate a unified model while maintaining the performance of each task as much as possible.
[0005] Existing model fusion methods typically rely on strategies such as parameter weighting, difference merging, or mask selection. For example, one type of method achieves fusion by linearly combining the parameters of multiple models; another type (such as mask-based fusion methods) attempts to select the most relevant parts from the parameters of different tasks for combination to improve fusion quality. However, these methods generally suffer from the following limitations: 1. Feature intensity imbalance: The magnitude of parameter changes (i.e., feature intensity) during fine-tuning varies significantly across different tasks. Strong features (large magnitude of change) tend to dominate the results during fusion, while weak features (small magnitude of change) are easily masked, leading to a significant performance degradation in some tasks.
[0006] 2. Insufficient ability to preserve weak features: Existing mask-based fusion methods (such as constructing selective masks for each task) have difficulty effectively preserving the important parameter dimensions corresponding to weak features when faced with large differences in feature strength, resulting in insufficient expressive ability of the fusion model for related tasks.
[0007] 3. Lack of a unified scale processing mechanism: Existing methods usually perform fusion directly in the original parameter space without uniformly processing the numerical range of parameters for different tasks. This results in a natural bias in the fusion of features at different scales, further exacerbating the performance instability problem.
[0008] Therefore, in various AI business systems that have high requirements for recognition and prediction accuracy and system robustness, such as computer vision, natural language processing, speech enhancement, financial risk control, time series prediction, autonomous driving, network security and multimodal intelligence, how to effectively alleviate the imbalance between strong and weak features, improve the expressive power of weak features, and achieve more stable and efficient parameter compression and multi-task fusion during model fusion has become a key technical problem that urgently needs to be solved. Summary of the Invention
[0009] This invention addresses the problems of feature intensity imbalance, insufficient retention of weak features, and poor fusion stability in existing model fusion technologies by providing a model fusion method and apparatus. It achieves unified scale fusion of multi-task parameters through feature scale alignment, weighted mask generation, and inference parameter inverse scaling correction. Combined with the mask, it accurately extracts task-related parameters, significantly improving the performance of weak feature tasks and the overall fusion effect while maintaining parameter compression efficiency.
[0010] To address the aforementioned technical problems, a first aspect of the present invention discloses a model fusion method, the method comprising: S1, obtain the parameter vectors for each task; S2, process the parameter vectors of each task to obtain the mask data of each task; S3, based on the task mask data, perform inference on the task to obtain task-specific inference parameters; S4. Load the task-specific inference parameters into the model, perform the corresponding task inference, and obtain and output the task result data.
[0011] As an optional implementation, in the first aspect of the present invention, the processing of the task parameter vectors to obtain task mask data includes: S21, Process the parameter vectors of each task to obtain the feature intensity of each task; S22, Based on the feature intensity of each task, the original parameters of each task are processed to obtain the scale alignment parameters of each task; S23, using the task contribution weight acquisition model, the scale alignment parameters of each task are processed to obtain the contribution weight of each dimension of each task and the global contribution weight of each task. S24. Based on the fusion parameter acquisition model, the contribution weights of each dimension of each task and the scale alignment parameters of each task are fused to obtain a fusion parameter set. S25, process the contribution weights of each dimension of each task, the scale alignment parameters of each task, and the fusion parameter set to obtain the mask data of each task.
[0012] As an optional implementation, in the first aspect of the present invention, the processing of the task parameter vectors to obtain the feature intensity of each task includes: The task feature intensity is calculated by processing the task parameter vectors using the task feature intensity calculation model. The expression for the task feature strength calculation model is as follows: , , in, For the first Task feature strength; Number the task; Total number of tasks; This represents the total dimension of the parameters in a single-task model. For the first Task No. Dimensional task parameter vector.
[0013] As an optional implementation, in the first aspect of the present invention, the step of processing the original parameters of each task based on the feature intensity of each task to obtain the scale alignment parameters of each task includes: S221, Based on the intensity of each task feature, determine the intensity of the pivot task feature; The expression for the pivot task characteristic strength is: , in, The characteristic strength of the pivot task; S222, Based on the pivot task feature intensity, process the feature intensity of each task to obtain the scaling factor of each task; The scaling factor expressions for each task are as follows: , in, For the first Scaling factor for the task; S223, using the scaling factors of each task, the original parameters of each task are scaled to obtain the scale alignment parameters of each task; The expressions for the task scale alignment parameters are as follows: , in, For the first The scale alignment parameters for each task.
[0014] As an optional implementation, in the first aspect of the present invention, the step of reasoning about the task based on the task mask data to obtain task-specific inference parameters includes: S31, using the inference stage parameter extraction model, process the mask data of each task and the fusion parameter set to obtain the relevant parameters of each task; The expression for the parameter extraction model in the inference stage is: , in, For element-wise product; For the first Task-related parameters for each task; For the first Task mask data for each task; For fusion parameters; S32, Perform inverse scaling on the parameters related to each task to obtain task-specific inference parameters; The expression for the task-specific inference parameters is: , in, For the first Task-specific inference parameters for each task. For the first Scaling factor for the task For the first Task-related parameters for each task.
[0015] As an optional implementation, in the first aspect of the present invention, the expression for the fusion parameter acquisition model is: , in, For the fusion parameter number Dimensional values; For the first Task No. Dimensional contribution weights; For numerical smoothing terms; For the first Task No. Dimensional scale alignment parameters.
[0016] As an optional implementation, in the first aspect of the present invention, the processing of the contribution weights of each dimension of each task, the scale alignment parameters of each task, and the fusion parameter set to obtain the mask data of each task includes: S251, using a binary mask to obtain a model, the scale alignment parameters of each task and the fusion parameter set are processed to obtain a binary mask for each task; S252, using the mask confidence weighted correction model, the contribution weights of each dimension of each task and the binary mask of each task are processed to obtain the mask data of each task; The expression for the mask confidence weighted correction model is: , in, For the first Task No. 3D mask data; Use the Sigmoid activation function; This is the scaling factor.
[0017] The second aspect of the present invention discloses a model fusion device, the device comprising: a task parameter vector acquisition module, a task mask calculation module, a task parameter inference module, and a task inference execution module; The task parameter vector acquisition module is used to acquire each task parameter vector; The task mask calculation module is used to process the task parameter vectors to obtain task mask data. The task parameter reasoning module is used to reason about the task based on the task mask data to obtain task-specific reasoning parameters. The task reasoning execution module is used to load the task-specific reasoning parameters into the model, execute the corresponding task reasoning, and obtain and output the task result data. The task parameter vector acquisition module, the task mask calculation module, the task parameter inference module, and the task inference execution module are connected in sequence.
[0018] A third aspect of the present invention discloses another model fusion apparatus, the apparatus comprising: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute some or all of the steps in the model fusion method disclosed in the first aspect of the present invention.
[0019] The fourth aspect of the present invention discloses a computer-readable storage medium, characterized in that the computer-readable storage medium stores computer instructions, which, when invoked, execute some or all of the steps in the model fusion method disclosed in the first aspect of the present invention.
[0020] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: 1. Alleviate feature imbalance: Map different task parameters to a unified numerical space through scale alignment to avoid strong features masking weak features; 2. Improve the performance of weak feature tasks: Weak feature parameters are effectively preserved after scale enhancement, which significantly improves the fusion performance of weak feature tasks; 3. Improve mask selection accuracy: The balanced parameter distribution enables the mask to accurately identify key dimensions of the task, reducing the accidental deletion of parameters; 4. Maintains efficient parameter compression: It does not increase the size of model parameters, has excellent parameter compression effect, and is suitable for resource-constrained deployment environments; 5. Improved fusion stability: The fused model can maintain almost 100% of the performance of the original models for each task; 6. Improve robustness of weak feature tasks: Further improve the robustness of weak feature tasks through dimensional weight normalization and soft mask correction. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram illustrating an application scenario of the model fusion device disclosed in an embodiment of the present invention; Figure 2 This is a schematic diagram of the model fusion method disclosed in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of a model fusion device disclosed in an embodiment of the present invention; Figure 4 This is a schematic diagram of another model fusion device disclosed in an embodiment of the present invention. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.
[0025] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0026] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0027] It should be noted that since the method in this application embodiment is executed in a computer device, the processing objects of each computer device exist in the form of data or information, such as time, which is essentially time information. It is understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, they are all corresponding data that exist so that the computer device can process them. Specific details will not be elaborated here.
[0028] It should be noted that a brief description of the artificial intelligence-related technologies that may be involved in this application is provided. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is the study of the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.
[0029] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0030] This application provides a model fusion method, apparatus, computer device, and computer-readable storage medium, which will be described in detail below.
[0031] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating a scenario in which the model fusion device provided in this application is used in a target detection system. The target detection system may include a computer device 100, which integrates a model fusion device, such as... Figure 1 Computer equipment in the country.
[0032] In this embodiment, the computer device 100 can be a standalone server, a server network, or a server cluster. For example, the computer device 100 described in this embodiment includes, but is not limited to, a computer, a network host, a single network server, a set of multiple network servers, or a cloud server composed of multiple servers. The cloud server is composed of a large number of computers or network servers based on cloud computing.
[0033] It is understood that the computer device 100 used in the embodiments of this application can be a device including receiving and transmitting hardware, that is, a device having receiving and transmitting hardware capable of performing bidirectional communication on a bidirectional communication link. Such a device may include: cellular or other communication devices having a single-line display or a multi-line display or a cellular or other communication device without a multi-line display. Specifically, the computer device 100 may be a desktop terminal or a mobile terminal, and may also be one of a mobile phone, tablet computer, laptop computer, etc.
[0034] Those skilled in the art will understand that Figure 1 The application environment given is merely one application scenario for the solution in this application and does not constitute a limitation on the application scenario of this application. Other application environments may include those that are more complex than those described above. Figure 1 The number of computer devices shown is more or less, for example Figure 1 Only one computer device is shown in the diagram. It is understood that the system may also include one or more other services, which are not limited here.
[0035] In addition, such as Figure 1 As shown, the target detection system may also include a memory 200 for storing model data, model parameter data, model training result data, and model training process data.
[0036] It should be noted that, Figure 1 The schematic diagram of the application scenario of the model fusion device shown is merely an example. The target detection system and scenario described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of target detection systems and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0037] This invention addresses the problems of feature intensity imbalance, insufficient retention of weak features, and poor fusion stability in the model fusion process by providing a model fusion method and apparatus. It achieves unified scale fusion of multi-task parameters through feature scale alignment and accurately extracts task-related parameters by combining masking. While maintaining parameter compression efficiency, it significantly improves the performance of weak feature tasks and the overall fusion effect.
[0038] Example 1 Please see Figure 2 , Figure 2 This is a flowchart illustrating a model fusion method disclosed in an embodiment of the present invention. Figure 2The described model fusion method is applied to target detection systems, such as local servers or cloud servers used in target detection systems; however, this embodiment of the invention is not limited to these applications. Figure 2 As shown, the model fusion method includes: It should be noted that, in this embodiment, based on a unified pre-trained model, six tasks, namely Java language program understanding (code search, code clone detection, vulnerability detection) and Python language program generation (code generation, code repair, code summarization), are independently fine-tuned. It should be noted that, in this embodiment, there is a significant difference in the parameter range between program understanding tasks and program generation tasks. The parameter variation range of the vulnerability detection task is significantly smaller, and it belongs to a weak feature task. It should be noted that in this embodiment, the performance of the fusion model is evaluated on six tasks, and the experimental conditions are as follows: The same pre-trained model was used as the base model; all tasks were trained and tested using standard datasets; the comparison method was a model fusion method without parameter scaling strategy. S1, obtain the parameter vectors for each task; It should be noted that obtaining the parameter vectors for each task refers to fine-tuning the pre-trained model for multiple different tasks to obtain the parameter vectors corresponding to each task. ,in (Task number) This represents the total dimension of the parameters in a single-task model. It should be noted that the parameter vector is the full incremental parameter vector after fine-tuning based on the same pre-trained model; It should be noted that in this embodiment, the fusion model is independently fine-tuned for the six types of tasks, resulting in six sets of task parameter vectors. ( ); S2, process the parameter vectors of each task to obtain the mask data of each task; S3, based on the task mask data, perform inference on the task to obtain task-specific inference parameters; S4, Load the task-specific inference parameters into the model, execute the corresponding task inference, and obtain and output the task result data; As can be seen, the model fusion method described in the embodiments of the present invention achieves unified scale fusion of multi-task parameters through feature scale alignment, and accurately extracts task-related parameters by combining masking. While maintaining parameter compression efficiency, it significantly improves the performance of weak feature tasks and the overall fusion effect, and solves the problems of feature intensity imbalance, insufficient retention of weak features, and poor fusion stability in the model fusion process.
[0039] Optionally, the processing of each task parameter vector to obtain each task mask data includes: S21, Process the parameter vectors of each task to obtain the feature intensity of each task; S22, Based on the feature intensity of each task, the original parameters of each task are processed to obtain the scale alignment parameters of each task; S23, using the task contribution weight acquisition model, the scale alignment parameters of each task are processed to obtain the contribution weight of each dimension of each task and the global contribution weight of each task. S24. Based on the fusion parameter acquisition model, the contribution weights of each dimension of each task and the scale alignment parameters of each task are fused to obtain a fusion parameter set. S25, process the contribution weights of each dimension of each task, the scale alignment parameters of each task and the fusion parameter set to obtain the mask data of each task; As can be seen, by implementing the model fusion method described in the embodiments of the present invention, the parameter vectors of each task are processed to obtain the mask data of each task. This provides data support for subsequent unified scale fusion of multi-task parameters through feature scale alignment and accurate extraction of task-related parameters by combining the mask. While maintaining parameter compression efficiency, it significantly improves the performance of weak feature tasks and the overall fusion effect, and solves the problems of feature intensity imbalance, insufficient retention of weak features, and poor fusion stability in the model fusion process.
[0040] Optionally, the processing of the task parameter vectors to obtain the feature intensity of each task includes: The task feature intensity is calculated by processing the task parameter vectors using the task feature intensity calculation model. The expression for the task feature strength calculation model is as follows: , , in, For the first Task feature strength; Number the task; Total number of tasks; This represents the total dimension of the parameters in a single-task model. For the first Task No. Dimensional parameters; It should be noted that the task feature intensity calculation model clearly distinguishes the feature strength of different tasks by quantifying the activation amplitude of the overall parameters of the task, providing a core benchmark for subsequent scale alignment and avoiding fusion deviation caused by differences in feature intensity. As can be seen, the model fusion method described in the embodiments of the present invention utilizes the task feature intensity calculation model to process the parameter vectors of each task and calculate the feature intensity of each task. This provides data support for subsequent unified scale fusion of multi-task parameters through feature scale alignment and accurate extraction of task-related parameters by combining masks. While maintaining parameter compression efficiency, it significantly improves the performance of weak feature tasks and the overall fusion effect, and solves problems such as feature intensity imbalance, insufficient retention of weak features, and poor fusion stability in the model fusion process.
[0041] Optionally, the step of processing the original parameters of each task based on the feature intensity of each task to obtain the scale alignment parameters of each task includes: S221, Based on the intensity of each task feature, determine the intensity of the pivot task feature; The expression for the pivot task characteristic strength is: , in, The characteristic strength of the pivot task; It should be noted that if multiple tasks have the same feature strength, the task that appears first will be selected as the pivot task. It should be noted that by using the feature strength of the pivot task, a unified benchmark for scale alignment is determined. With the strongest feature task as a reference, it is ensured that all task parameters can be mapped to the same numerical space, thereby alleviating the problem of imbalance between strong and weak features from the root. S222, Based on the pivot task feature intensity, process the feature intensity of each task to obtain the scaling factor of each task; The scaling factor expressions for each task are as follows: , in, For the first Scaling factor for the task; It should be noted that the scaling ratio is dynamically adjusted according to the difference in the strength of task features by using the scaling factor of each task. The weak feature task corresponds to a larger scaling factor to enhance the weak feature value, while the strong feature task has a moderate scaling ratio to avoid the strong feature from dominating too much. S223, using the scaling factors of each task, the original parameters of each task are scaled to obtain the scale alignment parameters of each task; The expressions for the task scale alignment parameters are as follows: , in, For the first Scale alignment parameters for each task; It should be noted that by aligning parameters at each task scale, all task parameters are mapped to a unified scale space, eliminating the natural differences in the numerical ranges of different task parameters, so that strong and weak features have equal influence in the fusion process, and completely solving the problem of weak features being masked. As can be seen, the model fusion method described in the embodiments of the present invention processes the original parameters of each task based on the feature strength of each task to obtain the scale alignment parameters of each task. This provides data support for subsequent unified scale fusion of multi-task parameters through feature scale alignment and accurate extraction of task-related parameters by combining masks. While maintaining parameter compression efficiency, it significantly improves the performance of weak feature tasks and the overall fusion effect, and solves the problems of feature strength imbalance, insufficient retention of weak features, and poor fusion stability in the model fusion process.
[0042] Optionally, the expression for the task contribution weight acquisition model is: , , in, For the first Global contribution weight of each task; For the first Task No. Dimensional contribution weights; For numerical smoothing terms; For the first Task No. Dimensional scale alignment parameters; It should be noted that, in this embodiment, the numerical smoothing term... ; It should be noted that the model obtains the contribution weights of the task dimensions by normalizing the contributions of each dimension to suppress the interference of extreme values on the fusion results; it quantifies the global importance of each task to provide a quantitative basis for subsequent mask correction; at the same time, it improves the dimensional survival rate of low parameter amplitude tasks, further protects weak features, and avoids the accidental deletion of key weak feature dimensions. As can be seen, the model fusion method described in the embodiments of the present invention obtains the model by utilizing the contribution weights of the task dimensions, normalizes the contributions of each dimension, and quantifies the global importance of each task. This provides data support for subsequent unified scale fusion of multi-task parameters through feature scale alignment and accurate extraction of task-related parameters by combining masks. While maintaining parameter compression efficiency, it significantly improves the performance of weak feature tasks and the overall fusion effect, and solves problems such as feature intensity imbalance, insufficient retention of weak features, and poor fusion stability in the model fusion process.
[0043] Optionally, the model expression for obtaining the fusion parameters is: , in, For the fusion parameter number Dimensional values; It should be noted that, in this embodiment, the numerical smoothing term... ; It should be noted that by using the fusion parameters to obtain the model and combining the dimensional contribution weights for weighted fusion, the problem of weak feature dimensions being ignored may be avoided by the fusion of a single maximum value, thereby further strengthening the parameter expression of weak feature tasks and improving the adaptability of the fusion model to weak feature tasks. As can be seen, implementing the model fusion method described in the embodiments of the present invention, using the fusion parameters to obtain the model, and combining the dimensional contribution weights for weighted fusion, provides data support for subsequent multi-task parameter unified scale fusion through feature scale alignment and accurate extraction of task-related parameters by combining masks. While maintaining parameter compression efficiency, it significantly improves the performance of weak feature tasks and the overall fusion effect, and solves the problems of feature intensity imbalance, insufficient retention of weak features, and poor fusion stability in the model fusion process.
[0044] Optionally, the process of processing the contribution weights of each dimension of each task, the scale alignment parameters of each task, and the fusion parameter set to obtain the mask data for each task includes: S251, using a binary mask to obtain a model, the scale alignment parameters of each task and the fusion parameter set are processed to obtain a binary mask for each task; The expression for the binary mask acquisition model is: , , in, For the first Task No. A 3D binary mask; For indicator functions; This is the relative threshold coefficient; It should be noted that the relative threshold coefficient takes values ranging from [value range missing]. The retention ratio is controlled between 10% and 50% to ensure the parameter compression effect; It should be noted that the first... Task No. A binary mask of dimension 1, where 1 indicates that dimension is retained and 0 indicates that it is discarded; its output range (0,1) values are determined to be retained if they are greater than 0.5, otherwise they are discarded. It should be noted that the binary mask acquisition model can accurately filter parameter dimensions related to the current task and eliminate redundant parameters that are irrelevant to the task. While achieving parameter compression, it retains the core features of the task and provides accurate parameter support for subsequent inference. S252, using the mask confidence weighted correction model, the contribution weights of each dimension of each task and the binary mask of each task are processed to obtain the mask data of each task; The expression for the mask confidence weighted correction model is: , in, For the first Task No. 3D mask data; Use the Sigmoid activation function; This is the scaling factor; It should be noted that the scaling factor... In this embodiment, the scaling factor is the default value. ; It should be noted that by using the mask confidence weighted correction model, the binary mask is upgraded to a soft mask. The importance of the parameter dimension is quantified by the confidence measure. Dimensions with a higher contribution than the global average are given a higher retention probability, which reduces the false deletion of key weak feature dimensions, further improves the mask accuracy of weak feature tasks, and enhances the stability of the fusion model. As can be seen, by implementing the model fusion method described in the embodiments of the present invention, the contribution weights of each dimension of each task, the scale alignment parameters of each task, and the fusion parameter set are processed to obtain the mask data of each task. This provides data support for the subsequent accurate extraction of task-related parameters by combining the mask. While maintaining the parameter compression efficiency, it significantly improves the performance of weak feature tasks and the overall fusion effect, and solves the problems of feature intensity imbalance, insufficient retention of weak features, and poor fusion stability in the model fusion process.
[0045] Optionally, the step of reasoning about the task based on the task mask data to obtain task-specific inference parameters includes: S31, using the inference stage parameter extraction model, process the mask data of each task and the fusion parameter set to obtain the relevant parameters of each task; The expression for the parameter extraction model in the inference stage is: , in, For element-wise product; For the first Task-related parameters for each task; For the first Task mask data for each task; For fusion parameters; It should be noted that the parameter extraction model in the inference stage can accurately extract parameters related to the current task, eliminate redundant parameters, reduce the amount of inference computation, and at the same time ensure that the extracted parameters can fully cover the core features of the task, thus guaranteeing the accuracy of inference. S32, Perform inverse scaling on the parameters related to each task to obtain task-specific inference parameters; The expression for the task-specific inference parameters is: , in, For the first Task-specific inference parameters for each task. For the first Scaling factor for the task For the first Task-related parameters for each task; It should be noted that the task-specific inference parameters restore the scale-aligned parameters to their original numerical scale, eliminating the impact of scaling operations on the inference results, ensuring the compliance of inference numerical values, and ensuring that the inference results of the fusion model are consistent with the inference accuracy of the independent fine-tuning model. In this embodiment, normalized accuracy is used as an evaluation index for model fusion technology. That is, for a task that needs to be fused, the ratio of the performance of the fused model to the performance of the original fine-tuned model before fusion indicates the degree to which the fused model retains the ability of the specific task. The larger the value, the better the performance of the model fusion method.
[0046] We tested our method on six different software engineering tasks, including three Java-based program comprehension tasks (code search, code clone detection, and vulnerability detection) and three Python-based program generation tasks (code generation, code repair, and code summarization). Experimental results show that, compared to other model fusion methods, our proposed method retains almost 100% of the performance of the original fine-tuned model in terms of normalization accuracy across all tasks. On the weak feature task (i.e., vulnerability detection), our method achieves a 63.49% improvement in normalization accuracy compared to other methods.
[0047] Table 1 shows the performance of the baseline method (Tall-Masks) and the model fusion method (Scaling-Masks) of this application. Experiments show that all six tasks maintained 100% consistent performance levels with their respective independently fine-tuned models, with no performance degradation. In the weak feature detection task, compared to the method without scaling strategy (Tall-Masks), a relative performance improvement of 63.49% was achieved. Introducing dimensional weights and soft masks further improved the stability of the weak feature task by more than 10%. As can be seen, by implementing the model fusion method described in the embodiments of the present invention, inference is performed on the tasks based on the mask data of each task to obtain task-specific inference parameters. While maintaining the parameter compression efficiency, the performance of weak feature tasks and the overall fusion effect are significantly improved, and the problems of feature intensity imbalance, insufficient retention of weak features and poor fusion stability in the model fusion process are solved.
[0048] Table 1 Performance of Scaling-Masks across all SE task fusions Example 2 Please see Figure 3 , Figure 3 This is a composition diagram of a model fusion device disclosed in an embodiment of the present invention. Wherein, Figure 3 The described model fusion device is applied in a target detection system, such as a local server or cloud server for the target detection system; however, this embodiment of the invention is not limited to this. Figure 3 As shown, the model fusion device includes: a task parameter vector acquisition module 101, a task mask calculation module 102, a task parameter inference module 103, and a task inference execution module 104. The task parameter vector acquisition module 101 is used to acquire each task parameter vector; The task mask calculation module 102 is used to process the task parameter vectors to obtain task mask data. The task parameter reasoning module 103 is used to reason about the task based on the task mask data to obtain task-specific reasoning parameters. The task reasoning execution module 104 is used to load the task-specific reasoning parameters into the model, execute the corresponding task reasoning, and obtain and output the task result data. The task parameter vector acquisition module 101, the task mask calculation module 102, the task parameter inference module 103, and the task inference execution module 104 are sequentially connected. As can be seen, the model fusion device described in this embodiment adopts the model fusion method described in Embodiment 1, which realizes unified scale fusion of multi-task parameters and accurate mask inference. It has the advantages of efficient parameter compression, excellent performance of weak feature tasks, and strong fusion stability. It can be widely used in scenarios such as deployment of multi-task models in artificial intelligence and lightweighting of edge device models.
[0049] Example 3 Please see Figure 4 , Figure 4 This is a schematic diagram of another model fusion device disclosed in an embodiment of the present invention. Figure 4The described apparatus can be applied to target detection systems, such as local servers or cloud servers used in target detection systems, and the embodiments of the present invention are not limited thereto. Figure 4 As shown, the device may include: Memory 201 storing executable program code; Processor 202 coupled to memory 201; The processor 202 calls the executable program code stored in the memory 201 to execute the steps in the model fusion method described in Embodiment 1.
[0050] Example 4 This invention discloses a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform the steps in the model fusion method described in Embodiment 1.
[0051] Example 5 This invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to perform the steps in the model fusion method described in Embodiment 1.
[0052] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0053] It should be noted that in all the calculation expressions or mathematical functions in the embodiments of the present invention, the variables involved have been dimensionlessized before calculation.
[0054] It should be noted that in all the calculation expressions or mathematical functions in the embodiments of the present invention, the values of the input independent variables all meet the reasonable requirements of the input value range of the calculation expression or mathematical function, and can ensure that the calculation expression or mathematical function can be calculated smoothly without violating physical laws or mathematical rules.
[0055] Through the detailed description of the above embodiments, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, including read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), once programmable read-only memory (OTPROM), electronically erasable rewritable read-only memory (EEPROM), read-only optical disc (CD-ROM) or other optical disc storage, disk storage, magnetic tape storage, or any other computer-readable medium that can be used to carry or store data.
[0056] Finally, it should be noted that the model fusion method and apparatus disclosed in the embodiments of the present invention are merely preferred embodiments of the present invention and are only used to illustrate the technical solutions of the present invention, not to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A model fusion method, characterized in that, The method includes: S1, obtain the parameter vectors for each task; S2, process the parameter vectors of each task to obtain the mask data of each task; S3, based on the task mask data, perform inference on the task to obtain task-specific inference parameters; S4. Load the task-specific inference parameters into the model, perform the corresponding task inference, and obtain and output the task result data.
2. The model fusion method according to claim 1, characterized in that, The process of processing the task parameter vectors to obtain task mask data includes: S21, Process the parameter vectors of each task to obtain the feature intensity of each task; S22, Based on the feature intensity of each task, the original parameters of each task are processed to obtain the scale alignment parameters of each task; S23, using the task contribution weight acquisition model, the scale alignment parameters of each task are processed to obtain the contribution weight of each dimension of each task and the global contribution weight of each task. S24. Based on the fusion parameter acquisition model, the contribution weights of each dimension of each task and the scale alignment parameters of each task are fused to obtain a fusion parameter set. S25, process the contribution weights of each dimension of each task, the scale alignment parameters of each task, and the fusion parameter set to obtain the mask data of each task.
3. The model fusion method according to claim 2, characterized in that, The process of processing the parameter vectors of each task to obtain the feature intensity of each task includes: The task feature intensity is calculated by processing the task parameter vectors using the task feature intensity calculation model. The expression for the task feature strength calculation model is as follows: , , in, For the first Task feature strength; Number the task; Total number of tasks; This represents the total dimension of the parameters in a single-task model. For the first Task No. Dimensional parameters.
4. The model fusion method according to claim 3, characterized in that, The process of processing the original parameters of each task based on the feature intensity of each task to obtain the scale alignment parameters of each task includes: S221, Based on the intensity of each task feature, determine the intensity of the pivot task feature; The expression for the pivot task characteristic strength is: , in, The characteristic strength of the pivot task; S222, Based on the pivot task feature intensity, process the feature intensity of each task to obtain the scaling factor of each task; The scaling factor expressions for each task are as follows: , in, For the first Scaling factor for the task; S223, using the scaling factors of each task, the original parameters of each task are scaled to obtain the scale alignment parameters of each task; The expressions for the task scale alignment parameters are as follows: , in, For the first The scale alignment parameters for each task.
5. The model fusion method according to claim 3, characterized in that, The expression for obtaining the fusion parameters is: , in, For the fusion parameter number Dimensional values; For the first Task No. Dimensional contribution weights; For numerical smoothing terms; For the first Task No. Dimensional scale alignment parameters.
6. The model fusion method according to claim 2, characterized in that, The process of processing the contribution weights of each dimension of each task, the scale alignment parameters of each task, and the fusion parameter set yields the mask data for each task, including: S251, using a binary mask to obtain a model, the scale alignment parameters of each task and the fusion parameter set are processed to obtain a binary mask for each task; S252, using the mask confidence weighted correction model, the contribution weights of each dimension of each task and the binary mask of each task are processed to obtain the mask data of each task; The expression for the mask confidence weighted correction model is: , in, For the first Task No. 3D mask data; Use the Sigmoid activation function; This is the scaling factor.
7. The model fusion method according to claim 1, characterized in that, The process of reasoning about tasks based on the task mask data to obtain task-specific inference parameters includes: S31, using the inference stage parameter extraction model, process the mask data of each task and the fusion parameter set to obtain the relevant parameters of each task; The expression for the parameter extraction model in the inference stage is: , in, For element-wise product; For the first Task-related parameters for each task; For the first Task mask data for each task; For fusion parameters; S32, Perform inverse scaling on the parameters related to each task to obtain task-specific inference parameters; The expression for the task-specific inference parameters is: , in, For the first Task-specific inference parameters for each task. For the first Scaling factor for the task For the first Task-related parameters for each task.
8. A model fusion device, characterized in that, The apparatus for implementing the model fusion method as described in any one of claims 1-7 includes: a task parameter vector acquisition module, a task mask calculation module, a task parameter inference module, and a task inference execution module; The task parameter vector acquisition module is used to acquire each task parameter vector; The task mask calculation module is used to process the task parameter vectors to obtain task mask data. The task parameter reasoning module is used to reason about the task based on the task mask data to obtain task-specific reasoning parameters. The task reasoning execution module is used to load the task-specific reasoning parameters into the model, execute the corresponding task reasoning, and obtain and output the task result data. The task parameter vector acquisition module, the task mask calculation module, the task parameter inference module, and the task inference execution module are connected in sequence.
9. A model fusion device, characterized in that, The device includes: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the model fusion method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, which, when invoked, are used to execute the model fusion method as described in any one of claims 1-7.