A model fusion method, device, and computer device based on topological structure
By constructing a ring topology in model fusion, local fusion of each device and up to two adjacent devices is solved, the problem of excessive resource peak demand in the prior art is solved, and efficient model fusion is achieved and suitable for resource-constrained devices.
Patent Information
- Application Number
- CN202510305666.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-03-14
AI Technical Summary
Existing model fusion techniques require loading all models onto a single device when processing a large number of models, resulting in a sharp increase in peak demand for storage and computing resources, limiting the application of model fusion on resource-constrained devices.
By building a ring topology, each device is locally fused with up to two adjacent devices, ensuring constant storage peaks, reducing storage requirements, and ensuring the efficiency and convergence of model fusion through limited rounds of updates.
It significantly reduces storage requirements and improves storage efficiency. It is suitable for resource-constrained devices, such as edge devices and GPUs with limited video memory, and promotes the wide deployment of model convergence technology in practical applications.
Smart Images

Figure CN119830225B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the technical field of machine learning, and in particular, to a model fusion method, apparatus, and computer device based on a topological structure. Background Art
[0002] Model fusion is a common method used in the fields of artificial intelligence and machine learning to enhance the generalization ability of models. Compared with traditional multi-task learning techniques, it does not require a large amount of training data and complex training processes. Model fusion obtains a new model with strong generalization ability by fusing multiple models fine-tuned on different tasks according to certain weights. Common model fusion methods include Ties-Merging, Task Arithmetic, Weight Averaging, and AdaMerging, etc. These methods enhance the effect of the fused model by constructing task vectors, eliminating conflicts in model fusion, and automatically determining model fusion parameters.
[0003] However, there are also defects in the prior art. For the method of manually setting model parameters, it usually takes a lot of time to determine a suitable parameter; while for the method of automatically determining fusion parameters, it usually requires a large amount of video memory, which limits their use in practical applications. In addition, the challenges faced by model fusion also include model selection, fusion strategy design, overfitting risk, etc.
[0004] For example, when dealing with a large number of models, layer-wise AdaMerging needs to load all models onto a single device at the same time. It is difficult for a single device to infinitely expand its storage and computing capabilities, which leads to a sharp increase in the peak demand for storage resources and computing resources. Especially when using GPU for model fusion, the consumption of video memory becomes a bottleneck, and it is difficult to implement on resource-constrained edge devices, which limits the application of model fusion technology in mobile devices or edge computing scenarios.
[0005] The Task Arithmetic and Weight Averaging methods do not require complex training processes and can quickly fuse multiple models into one model. However, these two methods usually require manual adjustment of parameters and a large number of experiments and professional knowledge to determine the best parameter combination, resulting in low efficiency, especially in cases where the parameter space is large or the model is complex; and these two methods still need to import all models into the same device and are not suitable for resource-constrained environments (such as edge devices); when the number of models to be fused increases, performance bottlenecks may be encountered. Summary of the Invention
[0006] To solve the problems in the prior art, this specification provides a model fusion method, apparatus, and computer device based on a topological structure. The method includes: constructing a topological structure, and determining the objects to be fused deployed in each node of the topological structure, where the topological structure includes multiple nodes, and each node corresponds to a device; according to a model fusion algorithm, merging and iterating the objects to be fused deployed in multiple nodes within a preset node range in the topological structure to obtain local fusion models generated after each merge iteration for all nodes in the topological structure; when the coefficients of the local fusion models generated after multiple merge iterations for all nodes are consistent, determining that the local fusion models corresponding to each node are completed in merging, and obtaining multiple multi-task models in the topological structure.
[0007] According to one aspect of the embodiments of this specification, merging and iterating the objects to be fused deployed in each node according to the model fusion algorithm includes: determining a certain node in the topological structure as the node to be fused; taking the adjacent nodes of the node to be fused in the topological structure as the local fusion nodes of the node to be fused; and merging and iterating the object to be fused of the node to be fused and the objects to be fused of the local fusion nodes.
[0008] According to one aspect of the embodiments of this specification, constructing a topological structure and determining the objects to be fused deployed in each node of the topological structure includes: configuring multiple devices, and taking each device as a node in the topological structure; storing an object to be fused in each node, where the object to be fused includes at least one of an initial fine-tuning model or a task vector.
[0009] According to one aspect of the embodiments of this specification, when the model fusion algorithm is a weighted average algorithm, the object to be fused for each node is an initial fine-tuning model. The step of merging and iterating the objects to be fused deployed in multiple nodes within a preset node range in the topological structure according to the model fusion algorithm to obtain local fusion models generated after each merge iteration for all nodes in the topological structure includes: based on a weighted algorithm, respectively assigning the same first initial weight to the initial fine-tuning models of each node; according to the first initial weight, merging and iterating the initial fine-tuning model of the node to be fused and the initial fine-tuning models of the local fusion nodes to obtain a merged first local fusion model.
[0010] According to one aspect of the embodiments of the present specification, when the fusion algorithm is a task algorithm, the object to be fused at each node is a task vector. According to the model fusion algorithm, the objects to be fused deployed at multiple nodes within a preset node range in the topological structure are merged and iterated to obtain the local fusion models generated after each merge iteration for all nodes in the topological structure, including: determining the preset number of merge iterations; determining whether the coefficients of the task vectors of all nodes after the preset merge iteration are basically the same; if so, determining the unified coefficient, and performing coefficient conversion according to the unified coefficient, the preset number of merge iterations, and the coefficients corresponding to the task vectors of each node in the naive task algorithm to determine the value of the unified coefficient. If not, updating the preset number of merge iterations, continuing the merge iteration, and repeating the above steps; determining the second local fusion model of the current node according to the unified coefficient.
[0011] According to one aspect of the embodiments of the present specification, after each node has experienced the preset number of rounds of merge iterations, the method includes:
[0012] Determining whether the weights corresponding to the first local fusion models after the merge iteration of each node are the same; if so, determining that the first local fusion models are merged, and obtaining multiple first local fusion models in the topological structure; determining whether the weights corresponding to the second local fusion models after the merge iteration of each node are the same; if so, determining that the second local fusion models are merged, and obtaining multiple second local fusion models in the topological structure.
[0013] According to one aspect of the embodiments of the present specification, when the model fusion algorithm is the inter-layer AdaMerging algorithm, the object to be fused at each node is a pre-trained model and a task vector. According to the model fusion algorithm, the objects to be fused deployed at each node are merged and iterated to obtain the local fusion models after each merge iteration for each node, including: based on the inter-layer AdaMerging algorithm, assigning a third initial weight to the task vector of each node; using the third initial weight, performing merge iteration on the pre-trained model, the task vector of the node to be fused, and the task vector of the local fusion node according to the information entropy loss function to obtain the merged third local fusion model.
[0014] An embodiment of this specification provides a model fusion device based on a topological structure. The device includes: a determination unit configured to construct a topological structure and determine objects to be fused deployed in each node in the topological structure, where the topological structure includes multiple nodes, and each node is connected to a device; a local fusion model determination unit configured to, according to a model fusion algorithm, merge and iterate the objects to be fused deployed in each node to obtain local fusion models generated after each merge and iteration of each node; and a fusion model determination unit configured to, when coefficients of the local fusion models generated after multiple merge and iterations of each node are consistent, determine that the local fusion models corresponding to each node are completed in merging, and obtain multiple multi-task models in the topological structure.
[0015] An embodiment of this specification further provides a computer device. The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method for model fusion based on a topological structure is implemented.
[0016] An embodiment of this specification further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the method for model fusion based on a topological structure is implemented.
[0017] This solution constructs a ring topological structure to enable each device to perform local fusion with at most two adjacent devices, thereby ensuring a constant storage peak value, significantly reducing storage requirements, not only improving storage efficiency, but also ensuring the efficiency of the model fusion process and the convergence of the final model through a limited number of rounds of updates. Description of the Drawings
[0018] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of this specification. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 Shown is a flowchart of a method for model fusion based on a topological structure according to an embodiment of this specification;
[0020] Figure 2 Shown is a flowchart of a method for merging and iterating the objects to be fused deployed in each node according to an embodiment of this specification;
[0021] Figure 3 Shown is a flowchart of a method for determining the objects to be fused in each node in the topological structure according to an embodiment of this specification;
[0022] Figure 4 The following is a flowchart of a method for merging iterations using a weighted average algorithm according to an embodiment of this specification;
[0023] Figure 5 The following is a flowchart of a method for merging iterations using a task algorithm according to an embodiment of this specification;
[0024] Figure 6 The following is a flowchart of a method for determining a first local fusion model according to an embodiment of this specification;
[0025] Figure 7 The following is a flowchart of a method for determining a third local fusion model according to an embodiment of this specification;
[0026] Figure 8a The following is a schematic diagram of a topological structure applied to a weighted average algorithm according to an embodiment of this specification;
[0027] Figure 8b The following is a schematic diagram of a topological structure applied to a task algorithm according to an embodiment of this specification;
[0028] Figure 8c The following is a schematic diagram of a topological structure applied to an inter-layer AdaMerging algorithm according to an embodiment of this specification;
[0029] Figure 9 The following is a schematic diagram of the structure of a model fusion device based on a topological structure according to an embodiment of this specification;
[0030] Figure 10 The following is a schematic diagram of a storage peak comparison according to an embodiment of this specification;
[0031] Figure 11 The following is a schematic diagram of the structure of a computer device according to an embodiment of this specification.
[0032] Explanation of the reference symbols in the drawings:
[0033] 901, Determination unit;
[0034] 902, Local fusion model determination unit;
[0035] 903, Multi-task model determination unit;
[0036] 1102, Computer device;
[0037] 1104, Processor;
[0038] 1106, Memory;
[0039] 1108, Driving mechanism;
[0040] 1110, Input / output module;
[0041] 1112. Input device;
[0042] 1114. Output device;
[0043] 1116. Presentation device;
[0044] 1118. Graphical user interface;
[0045] 1120. Network interface;
[0046] 1122. Communication link;
[0047] 1124. Communication bus. Detailed implementation
[0048] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this specification.
[0049] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of this specification are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this specification described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.
[0050] This specification provides method operation steps as described in the embodiments or flowcharts, but based on routine or non-creative labor, it may include more or fewer operation steps. The order of steps listed in the embodiments is only one way among the execution orders of many steps, and does not represent the only execution order. When the actual system or device product is executed, it can be executed in the order shown in the embodiments or the drawings or in parallel.
[0051] It should be noted that a model fusion method, device, and computer device based on a topological structure in this specification can be used in the field of machine learning technology and also in the field of artificial intelligence technology. This specification does not limit the application fields of a model fusion method, device, and computer device based on a topological structure.
[0052] In the fields of artificial intelligence and machine learning, model fusion (Model Fusion) or model ensemble (Model Ensemble) is a strategy for improving the overall prediction performance by combining the prediction results of multiple models. The basic idea of this strategy is to utilize the "collective wisdom" of multiple models to obtain better prediction results than any single model. Model fusion techniques can be divided into parameter-level fusion and prediction-level fusion, where parameter-level fusion refers to combining the parameters of multiple models, and prediction-level fusion refers to combining the prediction results of multiple models. With the rapid growth of data volume and the improvement of computing power, multi-model fusion techniques will play an increasingly important role in future artificial intelligence systems.
[0053] Figure 1 The following shows a flowchart of a model fusion method based on a topological structure according to an embodiment of this specification, which specifically includes the following steps:
[0054] Step 101: Construct a topological structure and determine the objects to be fused deployed in each node in the topological structure. Among them, the topological structure includes multiple nodes, and each node corresponds to a device.
[0055] In this step, construct a topological structure as shown in Figures 8a to 8c The topological structure shown in the figure can be applied to a computer network or a communication system. The topological structure shown in the figure is a ring topological structure, which includes multiple nodes, and each node deploys a device. Therefore, each node corresponds to each device. As shown in Figures 8a to 8c The topological structure has multiple circles. Each circle represents a node in the topological structure, and each node represents a device. Among them, the device can be a GPU or an edge device. The device in the node is used to receive, parse, and forward data. Further, each circle in the topological structure can have different colors, and nodes with different colors represent different devices. Each device stores the objects to be fused, and the objects to be fused will participate in model fusion in the subsequent process.
[0056] In this specification, devices in each node are deployed with corresponding objects to be fused. The objects to be fused in each node can be merged with the objects to be fused in the devices of adjacent nodes, realizing local merging or local fusion in the topological structure. In the embodiments of this specification, the objects to be fused are models or task vectors to participate in the merging. Among them, when the model to participate in the merging is stored in the node, it can be regarded as an initial fine-tuning model. Specifically, each initial fine-tuning model to participate in the merging is placed in a separate device, and these devices are linked through a ring topological structure, so that each device can only communicate with two adjacent devices. In each update, the models on each device are only locally fused with the models on adjacent devices, so as to ensure that the maximum number of models stored on each device does not exceed 3 (if the pre-trained model is included, it does not exceed 4). This local fusion process stops after a limited number of rounds, and finally the models on each device will converge and tend to be consistent, forming the final merged model.
[0057] In some embodiments of this specification, a topological structure as shown in Figures 8a to 8c can be deployed, or a star, mesh, graph structure topology, hierarchical topology or other forms of topological structures can be deployed. This application does not specifically limit the form of the topological structure composed of multiple nodes.
[0058] Step 102: According to the model fusion algorithm, perform merging iterations on the objects to be fused deployed in multiple nodes within a preset node range in the topological structure, and obtain the local fusion models generated after each merging iteration for all nodes in the topological structure.
[0059] In this step, the devices in the nodes of the topological structure can communicate with each other. In the embodiments of this specification, in order to reduce storage requirements and improve hardware efficiency and reduce dependence on hardware resources, a preset node range can be set, where the preset node range is used to determine the range for local fusion. Specifically, the preset node range refers to the node where the object to be fused is located and the nodes adjacent to this node. Therefore, in the Figures 8a to 8c shown topological structure, the devices on one node can communicate with the devices on two adjacent nodes, and the model parameters of the models deployed in their respective devices are transmitted to each other during the communication process.
[0060] Furthermore, after one merging of the objects to be fused deployed in the devices of each node, each node calculates the parameters of the objects to be fused in this node and the adjacent nodes, and obtains a merged model, which can be called the local fusion model after one merging iteration. This local fusion model is used as the new model deployed in this node and continues to participate in the next merging iteration until the merging iteration stop condition is met.
[0061] Step 103: When the coefficients of the local fusion models generated after multiple merging iterations by each node are consistent, it is determined that the local fusion models corresponding to each node are completed with merging, and multiple multi-task models in the topological structure are obtained.
[0062] In this step, local fusion needs multiple rounds of iteration to reach global consistency. After each node performs multiple merging iterations for local fusion, if it is determined that the local fusion models in all nodes in the topological structure converge and the coefficients before all local fusion models tend to be consistent, it is determined that the merging of the local fusion models corresponding to all nodes in the topological structure is completed, and the continuous merging iteration is stopped to form a final new fusion model. The fusion model finally stored in each node has strong generalization ability and can improve the overall prediction performance.
[0063] This specification reduces the demand for a large amount of storage (especially video memory), reduces the time cost of parameter adjustment, and improves the effect and generalization ability of model fusion to adapt to the growing data processing requirements and improve the performance of artificial intelligence systems.
[0064] Figure 2 The following is a flowchart of a method for merging and iterating the objects to be fused deployed by each node according to an embodiment of this specification, which specifically includes the following steps:
[0065] Step 201: Determine a certain node in the topological structure as the node to be fused.
[0066] Model fusion improves the overall prediction performance by combining the prediction results of multiple models. In this step, each device deployed in each node in the topological structure stores an initial model in the initial state, and each initial model needs to be merged. Therefore, each node can be called a node to be fused.
[0067] Step 202: Use the adjacent nodes of the node to be fused in the topological structure as the local fusion nodes of the node to be fused.
[0068] Figure 8a The following is a schematic diagram of a topological structure applied to the weighted average algorithm according to an embodiment of this specification. Taking the node as the node to be fused, its adjacent nodes in the topological structure are and , that is, and the nodes are used as the local fusion nodes of the node. If the node is used as the node to be fused, its adjacent nodes in the topological structure are and and , that is, The local fusion nodes of the nodes. During the communication between adjacent nodes, the initial fine-tuned models of each node are transmitted .
[0069] Figure 8b The figure shows a schematic diagram of a topological structure applied to a task algorithm in an embodiment of this specification. Taking the node as the node to be fused, its adjacent nodes in the topological structure are and , that is, taking and the nodes as the local fusion nodes of the node. If taking the node as the node to be fused, its adjacent nodes in the topological structure are and , then and the nodes can be taken as the local fusion nodes of the node. During the communication between adjacent nodes, the product of the task vector and the second weight of each node is transmitted .
[0070] Figure 8c The figure shows a schematic diagram of a topological structure applied to the inter-layer AdaMerging algorithm in an embodiment of this specification. Taking the node as the node to be fused, its adjacent nodes in the topological structure are and , that is, taking and the nodes as the local fusion nodes of the node. If taking the node as the node to be fused, its adjacent nodes in the topological structure are and , then and the nodes can be taken as the local fusion nodes of the node. During the communication between adjacent nodes, the product of the task vector and the third weight of each node is transmitted . In this specification, the inter-layer AdaMerging algorithm can also be called the AdaMerging algorithm.
[0071] Step 203, merge and iterate the objects to be fused of the node to be fused and the objects to be fused of the local fusion nodes.
[0072] In this specification, the Weight Averaging method, the task algorithm or Task Arithmetic method, and the inter-layer AdaMerging method can be used to perform merging iterations on the objects to be fused of the nodes to be fused and the objects to be fused of the local fusion nodes of the nodes to be fused. For different fusion methods, the specific objects to be fused of the nodes to be fused are also different. See Figures 3 to 5 for the description.
[0073] As methods for model fusion, Task Arithmetic and Weight Averaging are widely used to improve the generalization ability and performance of models, mainly relying on direct operations on model parameters. Among them, Weight Averaging realizes model fusion by calculating the average value of multiple model parameters. Specifically, the parameters of each model are assigned different weights according to their performance or confidence, and then the weighted average value is calculated to form the final fusion model.
[0074] Task Arithmetic adjusts the behavior of the pre-trained model through task vectors, where the task vectors represent the weight changes from the pre-trained model to the model after fine-tuning for a specific task. Specifically, these task vectors are combined through simple arithmetic operations (such as addition, subtraction, and negation) to achieve precise control of the model behavior.
[0075] Figure 3 The following is a flowchart of a method for determining the objects to be fused of each node in the topological structure according to an embodiment of this specification, which specifically includes the following steps:
[0076] Step 301: Configure multiple devices and deploy each device in a node in the topological structure.
[0077] Different from the traditional method in which all models that need to perform model fusion are loaded onto a single device at the same time, which will cause a sharp increase in the peak demand for storage resources, this specification deploys each device independently on a node, reducing the consumption of storage resources of a node and ensuring the storage and computing capabilities of the node.
[0078] Step 302: Store a to-be-fused object in each node, where the to-be-fused object includes at least one of an initial fine-tuning model or a task vector. When using the Weight Averaging method to perform iterative merging on each node, the to-be-fused object of the to-be-fused node is the initial fine-tuning model; when using the Task Arithmetic method to perform iterative merging on each node, the to-be-fused object of the to-be-fused node is the task vector; when using the inter-layer AdaMerging method to perform iterative merging, the to-be-fused object of the to-be-fused node is the initial fine-tuning model.
[0079] Figure 4 The following is a flowchart of a method for performing iterative merging using the weight averaging algorithm according to an embodiment of this specification, which specifically includes the following steps:
[0080] Step 401: Based on the weight algorithm, assign the same first initial weight to the initial fine-tuning model of each node.
[0081] In this specification, before the start of iterative merging, the weight averaging algorithm (Weight Averaging) assigns the same initial weight to the initial fine-tuning models in all nodes, and this initial weight is denoted as the first initial weight.
[0082] Step 402: According to the first initial weight, perform iterative merging on the initial fine-tuning model of the to-be-fused node and the initial fine-tuning model of the local fusion node to obtain a first locally fused model after merging.
[0083] The merging formula of naive Weight Averaging is: , represents the locally fused model. In this specification, k is designed to be 3. After the start of merging, the initial fine-tuning model on each device can only be merged with the initial fine-tuning models on adjacent devices.
[0084] For example, in Figure 8a In the shown topological structure, the formula for local merging is: , where represents the initial fine-tuning model in node i, represents the initial fine-tuning model in the adjacent node i - 1 of node i, represents the initial fine-tuning model in the adjacent node i + 1 of node i, represents the new locally fused model obtained by node i after the first round of local fusion, and the new locally fused model will replace the original initial fine-tuning model on the device of node i .
[0085] After repeating multiple rounds of iterative merging, the models on each device will converge to the final merged model, that is , where n is the number of merging iterations of the entire algorithm, represents the local fusion model on node i after the n-th round of local fusion, represents the multi-task model. At this time, the parameters of the models on all devices of all nodes are consistent and converge to the naive coefficients.
[0086] Figure 5 The following is a flowchart of a method for using task algorithm merging iterations according to an embodiment of the present specification. This method is applicable to the Task Arithmetic algorithm and specifically includes the following steps:
[0087] Step 501, determine the preset number of merging iterations.
[0088] In this step, an estimated number of merging iterations (assumed to be n) is first preset in advance to guide the task algorithm to perform local merging according to the preset number of merging iterations. In subsequent steps, the number of merging iterations needs to be further adjusted according to the results of local merging until after using the task algorithm for merging iterations, the coefficients corresponding to the models in the devices of each node converge and remain consistent.
[0089] In this specification, Task Arithmetic adjusts the behavior of the pre-trained model through task vectors. The task vector represents the weight change from the pre-trained model to the model after fine-tuning for a specific task. These task vectors are combined through simple arithmetic operations (such as addition, subtraction, and negation) to achieve precise control of the model behavior.
[0090] Step 502, determine whether the second weights corresponding to the task vectors of all nodes are basically the same after the preset merging iterations.
[0091] In this step, the condition for the end of the merging iteration using the task vector algorithm is that the coefficients corresponding to the models on all devices of each node are consistent and converge to the naive coefficients. Therefore, in step 501, after performing the preset number of merging iterations using the task vector algorithm, calculate whether the coefficients corresponding to the task vectors of each node are the same.
[0092] In this specification, the merging formula of the naive Task Arithmetic is as follows:
[0093] ; where represents the fine-tuned model, represents the pre-trained model, represents the fusion model, and K represents the number of nodes for local fusion. In order to simplify the operation in this specification, when using the task vector algorithm for local merging, the constant term is removed , only the merging of the task vector is performed. After the model converges after n merging iterations, the constant term is added back to the model.
[0094] In this step, first let
[0095] , represents the fine-tuned model minus the pre-trained model, which can be called the task vector.
[0096] In the method of this specification, the i-th node is merged iteratively.
[0097] For example, the task vector after the first merging iteration of the i-th node can be shown as follows:
[0098] ;
[0099] The task vector after the n-th merging iteration of the i-th node can be shown as follows:
[0100] .
[0101] In this specification, taking the topological structure with 5 nodes and corresponding 5 models as an example, the merging formula of the naive task algorithm is determined as follows:
[0102] ;
[0103] Among them, the task vector is defined as: ; where represents the initial fine-tuned model in node i, represents the pre-trained model, represents the model merging coefficient adjusted manually.
[0104] In this specification, through local fusion among 5 nodes, the initial fine-tuned models in each node are converged to , and the formula for the i-th corresponding node after n merging iterations is:
[0105] ; in the formula, is not the same as in the above formula. Taking the task vector corresponding to the first node as an example, its merging iteration process is as follows:
[0106] The first step: In the initial state, when no iterative merging is performed:
[0107] ; among them, taking Figure 8bTaking the topological structure in and as the task vectors to be fused.
[0108] Step 2: After the first iterative merge, the task vectors after local iterative merge of the first node are as follows:
[0109] ;
[0110] Step 3: Perform the second iterative merge. The task vectors after local iterative merge of the first node are as follows:
[0111]
[0112] = + ;
[0113] = + ;
[0114] + + + 2 ;
[0115] Step 30: The task vectors after local iterative merge of the first node are as follows:
[0116]
[0117] When n is large enough, at this time + . The current n is 30, that is, there is . According to the naive Task Arithmetic formula: ; Therefore, it is hoped that = , that is, make the coefficients of each task vector in consistent. Therefore, further calculate the maximum difference in the coefficients of the task vectors after 29 iterative merges, as follows:
[0118] It can be seen that the difference is extremely small and further shrinks with iteration. Therefore, it can be determined that after 29 iterations of the node in this step of the embodiment, the second weights corresponding to the task vectors are basically the same, and the situation of the merging formula of the naive Task Arithmetic can be achieved.
[0119] Step 503: If so, perform coefficient conversion based on the second weight, the preset merging iteration count, and the coefficients corresponding to the task vectors of each node in the naive task algorithm to determine the value of the second weight.
[0120] In this specification, the merging coefficient in the naive Task Arithmetic formula is different from the unified coefficient, but the two can be converted into each other. Taking the embodiment in Step 502 as an example, after performing the preset merging iteration count, all the coefficients corresponding to the task vectors in the obtained formula are unified as and used as the unified coefficient.
[0121] Further, make the unified coefficient the same as the merging coefficient in the naive Task Arithmetic formula, that is: ; when takes 0.3, calculate to obtain k Thus, the coefficient conversion is completed to determine the specific value of the unified coefficient.
[0122] Step 504: If not, update the preset merging iteration count, continue the merging iteration, and repeat the judgment step.
[0123] If after performing Step 502, it is determined that the coefficients corresponding to the task vectors of all nodes after the preset merging iteration count are inconsistent, it indicates that the current merging iteration count still cannot make the models of each node converge, and the models in each node have not reached the expected prediction effect. Therefore, it is necessary to increase the merging iteration count to control each node to continue the merging iteration.
[0124] Step 505: Determine the second local fusion model of the current node according to the second weight. In this step, after determining the second weight through n rounds of iterative model convergence, add the constant term back to the model, and finally obtain the fusion model, achieving the purpose of merging multiple fine-tuned models into one multi-task model.
[0125] Figure 6 The following shows the flowchart of a method for determining the third local fusion model according to an embodiment of this specification, which specifically includes the following steps:
[0126] Step 601: Based on the inter-layer AdaMerging algorithm, assign a third initial weight to the task vector of each node.
[0127] During the merging process of the inter-layer AdaMerging algorithm, unlabeled data is used to continuously update the model merging coefficient, thereby determining the optimal model merging coefficient. Since inferring the model is required to update the model merging coefficient using inter-layer AdaMerging, the entire merging process must use the GPU. To minimize the peak video memory of each device, each device in each node stores a different fine-tuned model in the initial state , and at the same time stores a copy of the pre-trained model .
[0128] The formula of the inter-layer AdaMerging algorithm is as follows:
[0129] ; where represents the set of all training data; represents the multi-task model, represents the initial fine-tuned model of node i. For the initial fine-tuned models on k nodes, each model is fine-tuned on the corresponding dataset , and a new model is constructed by merging N initial fine-tuned models through model merging, so that it performs well on all test sets .
[0130] Under the condition of using a ring topology, each device only communicates with the two adjacent devices and transfers their respective model parameters (without transferring ). In this way, even when a large number of models need to be merged in one merging iteration, the peak video memory of each device will not exceed 4, greatly reducing the hardware requirements during the merging process and increasing the usage scenarios of model merging.
[0131] Step 602: Using the third initial weight, perform a merging iteration on the pre-trained model, the task vector of the node to be fused, and the task vector of the local fusion node according to the information entropy loss function to obtain the merged third local fusion model.
[0132] In this specification, the information entropy function is used to perform a merging iteration on each node. Among them, the information entropy loss function in this specification replaces the traditional loss function, and the formula of the information entropy is as follows:
[0133] ; where represents the probability that the i-th sample predicted by the initial fine-tuned model belongs to category c. The information entropy is used to measure the uncertainty of the model prediction.
[0134] In this specification, during each round of the merging iteration process, the fusion coefficient α of the model is automatically determined and changes. The local merging formula for each round is as follows:
[0135] ; represents the pre-trained model, represents the initial fine-tuning model in node i, represents the task vector of node i, represents the task vector of node i + 1, represents the task vector of node i - 1.
[0136] In this specification, the optimization objective of inter-layer AdaMerging is to minimize the information entropy. In each iteration, the inter-layer AdaMerging algorithm calculates the gradient of the fusion coefficient α by inputting a small amount of unlabeled data and using stochastic gradient descent (SGD), and updates the fusion coefficient α of the model. Thus, the optimal model merging coefficient is determined, and then the merging operation is performed to replace the model with a new model. The merging iteration is repeated n times. When n is large enough, the models on all devices will converge and have unified parameters. At this time, the model on any one device can be used as the final model output, improving the generalization ability of the final model.
[0137] Figure 7 The following is a flowchart of a method for determining the first local fusion model according to an embodiment of this specification. In this specification, local fusion needs to be iterated multiple times to achieve global consistency. When globally consistent, the weights or coefficients of the current local fusion models on all devices need to be kept consistent. Therefore, when using different algorithms for local fusion, it is necessary to further determine whether the weights or parameters corresponding to the local models are consistent respectively. The specific steps of this method are as follows:
[0138] Step 701, determine whether the weights corresponding to the first local fusion models after the merging iteration of each node are consistent. As Figure 8a shown, after each merging iteration is completed, according to the first local fusion models currently formed by each node, determine whether the weights of the first local fusion models corresponding to all nodes in the topological structure are the same.
[0139] In this specification, when multiple merging iterations are completed and the first local fusion models corresponding to all nodes in the topological structure have converged, the weights of all the first local fusion models are basically consistent.
[0140] Step 702, if so, determine that the merging of the first local fusion models is completed, and obtain multiple first local fusion models in the topological structure. If they are consistent, it is determined that the model fusion of all nodes in the topological structure is completed using the weight averaging algorithm. At this time, the current local fusion model in each node has become a multi-task model and can be used to process complex tasks.
[0141] Step 703: Determine whether the weights corresponding to the third local fusion models after the merger iteration of each node are consistent. As Figure 8c shown, after each merger iteration is completed, based on the third local fusion models currently formed by each node, determine whether the weights corresponding to the third local fusion models of all nodes in the topological structure are the same.
[0142] In this specification, when multiple merger iterations are completed and the third local fusion models corresponding to all nodes in the topological structure converge, the weights of all the third local fusion models are basically consistent.
[0143] Step 704: If so, determine that the merger of the third local fusion models is completed, and obtain multiple third local fusion models in the topological structure. If they are consistent, it is determined that the model fusion of all nodes in the topological structure is completed using the weighted average algorithm. At this time, the current local fusion model in each node has become a multi-task model and can be used to process complex tasks.
[0144] To verify the effectiveness of this solution, the implementation of Tables 1 to 4 is adopted.
[0145] Table 1 is a multi-task performance table for merging 8 corresponding initial fine-tuning models on 8 image classification tasks. The right part of the first row in the table shows 8 different data sets, namely SUN397, Cars, RESISC45, EuroSAT, SVHN, GTSRB, MNIST, DTD, corresponding to 8 different initial fine-tuning models that are fused on these 8 different data sets respectively. The "reference method" row in the table shows the prediction accuracy of each initial pre-trained model on each data set without model fusion, which is 63.2, 59.8, 60.7, etc. It can be seen that the prediction accuracy is not high. The initial fine-tuning models can only perform well on their corresponding fine-tuning data sets and perform poorly on other data sets.
[0146] In the row of "Muti-Task Model Fusion Methods" in Table 1, the conventional model fusion method is adopted, and the prediction results of each initial fine-tuned model after local fusion on 8 image classification tasks are obtained respectively. The row of "Muti-Task Model Fusion Methodswith Memory Efficient Merging" in Table 1 represents the prediction accuracy after using the method of this specification and three algorithms for constant storage peak. The model performs the image classification prediction task, and the results in the table are the image classification prediction accuracy results. After each algorithm is iterated 20 times, the initial fine-tuned model on each node becomes a fusion model. It can be seen from Table 1 that the prediction results of Weight Averaging w / MEM and Task Arithmetic w / MEM on 8 image classification tasks are very close to the prediction values of the naive weight averaging method and the naive task arithmetic. If the number of iterations continues to increase, the gap between the prediction results of Weight Averaging w / MEM and Task Arithmetic w / MEM on 8 image classification tasks and the prediction results of the naive weight averaging method and the naive task arithmetic will be further reduced. Therefore, the model parameters of the multi-task model on the device initially storing the SUN397 fine-tuned model can be output as the final result.
[0147] In addition, the layer-wise AdaMerging w / MEM has a 0.9% improvement compared to the naive AdaMerging result.
[0148] Table 1
[0149]
[0150] Table 2 shows that after multiple merging iterations, the initial fine-tuned models stored in each of the 8 nodes (node 1, node 2... node 8) are the fine-tuned models corresponding to 8 image classification tasks. After multiple merging iterations, the prediction accuracies of the fusion models in the 8 nodes on 8 different datasets (SUN397, Cars, RESISC45, EuroSAT, SVHN, GTSRB, MNIST, DTD) are basically the same. It can be considered that the fusion models in each node are basically convergent, and any one of the fusion models can be used as the final output result.
[0151] Table 2
[0152]
[0153] Since the inter-layer AdaMerging w / MEM cannot pre-determine the final merging coefficients like the other two algorithms, more experiments are designed in this specification to prove the prediction effect of the inter-layer AdaMerging w / MEM. As shown in Table 3, Table 3 is a table of the generalization effect of the merged model using the inter-layer AdaMerging w / MEM algorithm on 6 image recognition tasks for performing two unseen tasks.
[0154] Table 3: Merging the CLIP-ViT-B / 32 model on 6 tasks
[0155]
[0156] In the first generalization task, the inter-layer AdaMerging w / MEM is higher than the inter-layer AdaMerging on both Seen Tasks and Unseen Tasks, with improvements of 1.7% and 2.8% respectively. In the second generalization task, the accuracy of the inter-layer AdaMerging w / MEM slightly decreases, but overall it still remains at the same level as the inter-layer AdaMerging.
[0157] In the robustness task test (see Table 4), 7 corrupted test datasets are created, namely: Motion Blur, Impulse Noise, Gaussian Noise, Pixelate, Spatter, Contrast, and JPEG Compression. The inter-layer AdaMerging w / MEM algorithm can maintain comparable accuracy with the inter-layer AdaMerging algorithm on the vast majority of tasks. At the same time, in the three tasks of Motion Blur, Contrast, and Compression, the inter-layer AdaMerging w / MEM leads the inter-layer AdaMerging by 0.7%, 0.7%, and 0.6% respectively.
[0158] The above results fully prove that the method in this specification can be applied to the three models of Weight Averaging, Task Arithmetic, and inter-layer AdaMerging.
[0159] Table 4: Robustness results of the merged CLIP-ViT-B / 32 model for four tasks
[0160]
[0161] During the model fusion process, the maximum number of models stored on each device does not exceed 4. In the task of merging 8 models, the original AdaMerging algorithm stores 9 models in the video memory. In contrast, this solution reduces the peak video memory by nearly 60% for inter-layer AdaMerging, greatly reducing the requirements for hardware. By reducing storage requirements and improving hardware efficiency, this specification reduces the dependence on hardware resources, is applicable to storage-constrained scenarios such as edge devices and GPUs with limited video memory, and promotes the wide deployment and technological innovation of model fusion technology in practical applications.
[0162] Figures 8a to 8c The following shows schematic diagrams of three topological structures according to embodiments of this specification.
[0163] In the figure, nodes of different colors represent different devices (which can be GPUs or end-side devices). , , \(T_i\) respectively represent the initial fine-tuned model, the pre-trained model, and the task vector, where . For the task algorithm and the inter-layer AdaMerging algorithm, it is also necessary to assign initial merging weights to each model .
[0164] Figure 8a In , taking the and nodes as the nodes to be fused, its adjacent nodes in the topological structure are and , that is, taking and nodes as the local fusion nodes of the node. If taking the node as the node to be fused, its adjacent nodes in the topological structure are and , that is, the node can take
[0165] Figure 8b In , taking the and nodes as the nodes to be fused, its adjacent nodes in the topological structure are and , that is, taking and nodes as the local fusion nodes of the node. If taking the node as the node to be fused, its adjacent nodes in the topological structure are and , that is, the The local fusion node of the node.
[0166] Figure 8c Among them, The node is used as the node to be fused, and its adjacent nodes in the topological structure are and , that is, and The nodes are used as The local fusion nodes of the node. If The node is used as the node to be fused, and its adjacent nodes in the topological structure are and , then and The nodes are used as The local fusion nodes of the node. During the communication process between adjacent nodes, the product of the task vectors and the third weight of each node is transmitted .
[0167] In this solution, by constructing a ring topological structure, each device can perform local fusion with at most two adjacent devices, thereby ensuring a constant storage peak value, significantly reducing the storage requirements, not only improving the storage efficiency, but also ensuring the efficiency of the model fusion process and the convergence of the final model through a limited number of rounds of updates.
[0168] Figure 9 The following figure shows the structural schematic diagram of a model fusion device based on a topological structure according to an embodiment of the present specification. The basic structure of the model fusion device based on the topological structure is described in this figure. Among them, the functional units and modules can be implemented in software, or a general-purpose chip or a specific chip can be used to implement the model fusion based on the topological structure. The device specifically includes:
[0169] A determination unit 901, configured to construct a topological structure and determine the objects to be fused deployed in each node in the topological structure, where the topological structure includes multiple nodes, and each node is connected to a device;
[0170] A local fusion model determination unit 902, configured to merge and iterate the objects to be fused deployed in multiple nodes within a preset node range in the topological structure according to a model fusion algorithm, and obtain the local fusion models generated after each merge iteration of all nodes in the topological structure;
[0171] A multi-task model determination unit 903, configured to determine that the merging of the local fusion models corresponding to each node is completed when the coefficients of the local fusion models generated after multiple merge iterations of all nodes are the same, and obtain multiple multi-task models in the topological structure.
[0172] Figure 10 The following figure shows a comparison schematic diagram of storage peak values according to an embodiment of the present specification.Figure 10 The storage peak comparison of two theoretically model fusion methods is shown. The advantage of the constant storage peak model fusion method proposed in this solution will continuously improve as the number of model fusions increases.
[0173] As Figure 11 shown, it is a schematic diagram of a computer device provided by an embodiment of this specification. The topology-based model fusion method described in this application can be applied to the computer device. The computer device 1102 may include one or more processors 1104, such as one or more central processing units (CPUs), and each processing unit may implement one or more hardware threads. The computer device 1102 may also include any memory 1106 for storing any kind of information such as code, settings, data, etc. Non-limiting examples include any type of RAM, any type of ROM, flash devices, hard disks, optical discs, etc. More generally, any memory may use any technology to store information. Further, any memory may provide volatile or non-volatile retention of information. Further, any memory may represent a fixed or removable component of the computer device 1102. In one case, when the processor 1104 executes the associated instructions stored in any memory or combination of memories, the computer device 1102 may perform any operation of the associated instructions. The computer device 1102 also includes one or more drive mechanisms 1108 for interacting with any memory, such as a hard disk drive mechanism, an optical disc drive mechanism, etc.
[0174] The computer device 1102 may also include an input / output module 1110 (I / O) for receiving various inputs (via the input device 1112) and for providing various outputs (via the output device 1114). A specific output mechanism may include a presentation device 1116 and an associated graphical user interface (GUI) 1118. In other embodiments, the input / output module 1110 (I / O), the input device 1112, and the output device 1114 may not be included, and it may only be a computer device in the network. The computer device 1102 may also include one or more network interfaces 1120 for exchanging data with other devices via one or more communication links 1122. One or more communication buses 1124 couple the components described above together.
[0175] The communication link 1122 may be implemented in any way, for example, through a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 1122 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc. governed by any protocol or combination of protocols.
[0176] Corresponding to Figures 1 to 7 In the method described above, an embodiment of this specification also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the above method are executed.
[0177] An embodiment of this specification also provides a computer-readable instruction. When a processor executes the instruction, the program therein causes the processor to execute the method as Figures 1 to 7 shown.
[0178] It should be understood that in various embodiments of this specification, the magnitudes of the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this specification.
[0179] It should also be understood that in the embodiments of this specification, the term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent three cases: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this specification generally represents an "or" relationship between the associated objects before and after.
[0180] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this specification can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this specification.
[0181] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0182] In several embodiments provided in this specification, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the displayed or discussed couplings or direct couplings or communication connections to each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can also be electrical, mechanical, or other forms of connection.
[0183] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of this specification.
[0184] In addition, each functional unit in the various embodiments of this specification can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0185] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this specification, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this specification. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical disks, and other media that can store program codes.
[0186] Specific embodiments are used in this specification to elaborate on the principles and implementation manners of this specification. The descriptions of the above embodiments are only used to help understand the method and its core idea of this specification. At the same time, for those of ordinary skill in the art, based on the idea of this specification, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this specification.
Claims
1. A model fusion method based on topological structure, characterized in that: The method comprises: Constructing a topological structure, and determining an object to be integrated deployed in each node in the topological structure, wherein the topological structure includes a plurality of nodes, and each node corresponds to a device; According to the model fusion algorithm, the objects to be fused deployed in multiple nodes within the preset node range in the topological structure are merged and iterated to obtain a local fusion model generated after each merging iteration of all nodes in the topological structure, including: when the model fusion algorithm is a weighted average algorithm, based on the weighted average algorithm, the same first initial weight is respectively assigned to the initial fine-tuning model of each node; When the fusion algorithm is a task algorithm, determine a preset number of merge iterations; determine whether the second weights corresponding to the task vectors of all nodes after the preset merge iterations are substantially consistent; if so, perform coefficient conversion according to the second weight, the preset number of merge iterations, and the coefficients corresponding to the task vectors of each node in the naive task algorithm to determine the value of the second weight; When the model fusion algorithm is an inter-layer AdaMerging algorithm, a third initial weight is assigned to the task vector of each node based on the inter-layer AdaMerging algorithm; When the coefficients of the local fusion models generated after multiple merging iterations of all nodes are consistent, the local fusion models corresponding to each node are determined to complete the merger, and multiple multi-task models in the topological structure are obtained.
2. The method according to claim 1, characterized in that According to the model fusion algorithm, merging and iterating the objects to be fused deployed in multiple nodes within a preset node range in the topological structure includes: Select a node in the topological structure as the node to be merged; Using the neighboring nodes of the node to be fused in the topological structure as local fusion nodes of the node to be fused in a preset node range; The objects to be fused of the nodes to be fused and the objects to be fused of the local fusion nodes are merged and iterated.
3. The method according to claim 2, characterized in that Constructing a topological structure, and determining the objects to be integrated deployed in each node in the topological structure includes: Configure multiple devices and deploy each device in a node in the topology; An object to be fused is stored in each node, and the object to be fused includes at least one of an initial fine-tuning model or a task vector.
4. The method according to claim 3, characterized in that When the model fusion algorithm is a weighted average algorithm, the object to be fused of each node is the initial fine-tuning model. According to the model fusion algorithm, the objects to be fused deployed in multiple nodes within a preset node range in the topological structure are merged and iterated to obtain a local fusion model generated after each merging iteration of all nodes in the topological structure, including: According to the first initial weight, the initial fine-tuning model of the node to be fused and the initial fine-tuning model of the local fusion node are merged and iterated to obtain a merged first local fusion model.
5. The method according to claim 3, characterized in that: When the fusion algorithm is a task algorithm, the object to be fused of each node is a task vector. According to the model fusion algorithm, the objects to be fused deployed in multiple nodes within a preset node range in the topological structure are merged and iterated to obtain a local fusion model generated after each merging iteration of all nodes in the topological structure, including: If not, update the preset number of merge iterations, continue to merge iterations, and repeat the judgment step; A second local fusion model of the current node is determined according to the second weight.
6. The method according to claim 5, characterized in that After each node undergoes a preset round of merging iterations, the method further includes: Determine whether the weights corresponding to the first local fusion model after merging and iterating each node are consistent; If so, it is determined that the first local fusion model is merged and completed, and a fusion model in the topological structure is obtained.
7. The method according to claim 3, characterized in that When the model fusion algorithm is the inter-layer AdaMerging algorithm, the objects to be fused of each node are the pre-trained model and the task vector. According to the model fusion algorithm, the objects to be fused deployed by each node are merged and iterated, and the local fusion model of each node after each merge iteration is obtained, including: Using the third initial weight, the pre-trained model, the task vector of the node to be fused, and the task vector of the local fusion node are merged and iterated according to the information entropy loss function to obtain a merged third local fusion model.
8. A model fusion device based on topological structure, characterized in that: The device comprises: A determination unit, configured to construct a topology structure and determine an object to be integrated deployed in each node in the topology structure, wherein the topology structure includes a plurality of nodes, and each node corresponds to a device; A local fusion model determination unit is used to merge and iterate the objects to be fused deployed in multiple nodes within a preset node range in the topological structure according to a model fusion algorithm to obtain a local fusion model generated after each merging iteration of all nodes in the topological structure, including: when the model fusion algorithm is a weighted average algorithm, based on the weighted average algorithm, the same first initial weight is respectively assigned to the initial fine-tuning model of each node; When the fusion algorithm is a task algorithm, determine a preset number of merge iterations; determine whether the second weights corresponding to the task vectors of all nodes after the preset merge iterations are substantially consistent; if so, perform coefficient conversion according to the second weight, the preset number of merge iterations, and the coefficients corresponding to the task vectors of each node in the naive task algorithm to determine the value of the second weight; When the model fusion algorithm is an inter-layer AdaMerging algorithm, a third initial weight is assigned to the task vector of each node based on the inter-layer AdaMerging algorithm; The fusion model determination unit is used to determine the local fusion models corresponding to each node to complete the merging when the coefficients of the local fusion models generated after multiple merging iterations of all nodes are consistent, so as to obtain multiple multi-task models in the topological structure.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Model training method and face recognition method based on adaptive split learning-federated learning
WO2023185485A1