Heterogeneous edge device-oriented adaptive large-model distributed fine tuning method

By designing a wireless federated learning framework with dynamic pruning and alignment mechanisms in a heterogeneous edge device environment, combining width and depth pruning technology, dynamically adjusting the model pruning ratio and narrowing the difference between sub-models and global models, the problem of insufficient resource utilization in heterogeneous device environment is solved, and efficient and reliable large-model federated learning is achieved.

CN120106169APending Publication Date: 2025-06-06BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510263311.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively adapt to the dynamic changes of equipment resources in heterogeneous edge equipment environments, resulting in uncertain overall performance, and pruning technology is insufficient in large-scale parameters and complex structures, affecting the aggregation effect of the global model.

Method used

A wireless federated learning framework based on dynamic pruning and alignment mechanism is designed, combining width pruning and depth pruning techniques, dynamically adjust the pruning ratio of the model, and reduce the gradient difference between the sub-model and the global model through hierarchical knowledge alignment and neuron-level alignment mechanism.

Benefits of technology

It significantly reduces communication and computing overhead, improves training efficiency, improves system learning performance and model accuracy, and solves the problem of insufficient resource utilization in heterogeneous equipment environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106169A_ABST
    Figure CN120106169A_ABST
Patent Text Reader

Abstract

A self-adaptive large model distributed fine tuning method for heterogeneous edge equipment belongs to the field of industrial intelligent large models, and comprises the following steps: a server constructs a sub-model through a two-dimensional pruning method, sequential neuron pruning is carried out on an FFN module through width pruning, and deep pruning is carried out on a transformer layer to remove part of a deep structure; sub-models are dynamically distributed according to the client capacity; training the sub-model by the client, keeping the trunk parameter of the pre-trained model frozen and only updating the LoRA; the client uploads the updated LoRA module and the sub-model parameters to the server; the server executes aggregation operation, updates a LoRA module of the global model, and keeps trunk parameters of the global model unchanged; in the global model alignment stage, the server executes hierarchical knowledge and neuron level alignment operation after the aggregation stage, and layer-by-layer distillation loss is minimized to obtain a heterogeneous agent sub-model. According to the method, the communication and calculation overhead is reduced, the training efficiency is improved, and the system learning performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence big models, and specifically relates to an adaptive big model distributed fine-tuning method for heterogeneous edge devices. Background Art

[0002] Large-model federated learning combines the strong generalization ability of large models with the privacy protection advantage of federated learning, and is an efficient distributed learning framework. By fine-tuning the pre-trained large model locally on the device and only transmitting the model update instead of the original data, the risk of privacy leakage is effectively reduced, while improving the model's adaptability to non-independent and identically distributed data. This framework has shown broad application potential in fields such as healthcare and law that have high requirements for data security and model performance, and provides strong support for large model training in a distributed environment.

[0003] The efficient parameter fine-tuning technology freezes the backbone parameters of the large model and only adjusts a small number of modules (such as LoRA, Adapter), thereby achieving efficient model optimization on resource-limited devices. This method can significantly reduce computing and memory requirements, while reducing communication overhead in model training, and is very suitable for scenarios with limited device resources. However, the efficient parameter fine-tuning technology achieves model optimization on resource-constrained devices by only adjusting a small number of key module parameters. When combined with pruning technology, changes in the model structure may cause the fine-tuning module to fail, thereby affecting the aggregation effect of the global model. In addition, existing technologies lack the ability to adapt to dynamic changes in device resources, and are difficult to operate stably in heterogeneous environments, resulting in uncertainty in overall performance.

[0004] Federated Parameter Efficient Tuning (FedPETuning) introduces efficient parameter fine-tuning technology into the federated learning framework. By freezing the weights of the pre-trained model, only the parameters of the local key modules are updated. This method effectively reduces the computing and communication costs in the distributed training process, so that resource-constrained devices can also efficiently participate in the fine-tuning and training of large models. Although federated parameter efficient tuning reduces communication and computing overhead in distributed training, it is usually based on the assumption of balanced device resources and ignores the actual situation of heterogeneous resources between devices. On resource-constrained devices, training delays increase significantly, slowing down the synchronous update of the global model, while the insufficient resource utilization of high-performance devices further limits the improvement of the overall efficiency of the system.

[0005] Heterogeneous-aware federated learning introduces model heterogeneity technology. Aiming at the differences in computing and communication capabilities of different clients, these methods design an adaptable federated learning framework. By using technologies such as width pruning and depth pruning to allocate heterogeneous sub-models to each device, they can effectively adapt to the resource constraints of different devices. At the same time, a single global inference model is generated through a stable aggregation strategy, which significantly improves the efficiency and flexibility of federated learning in heterogeneous environments. These methods effectively balance the personalized needs of resource-constrained devices with the performance of global models through heterogeneous pruning mechanisms, improving training efficiency and communication optimization capabilities. Federated learning methods for heterogeneous device environments customize adaptation models through pruning technology, improving system flexibility, but these methods are mainly applied to small models and are insufficient when faced with complex structures and large-scale parameters. At the same time, there are defects in the alignment mechanism between the pruned sub-models and the global model, which easily leads to a decrease in the aggregation effect of the global model and affects the overall performance stability.

[0006] The large model pruning federated learning strategy focuses on the differences between the pruned sub-model and the global model through strategies such as knowledge distillation and parameter alignment, ensuring the effective aggregation of the model and the stability of the global learning performance. Ensure the consistency of the structure and performance of the sub-model and the global model. For example, in the deep pruning method, the shallow model passes the information of the deep sub-model to the shallow layer by sharing shallow features and using self-distillation technology, thereby compensating for the training bias caused by insufficient data and limited resources, thereby ensuring the accuracy and generalization ability of the final global model. The large model pruning federated learning strategy combines width pruning and depth pruning techniques. However, the existing strategy pays too much attention to the adaptability of weak-performance devices and ignores the effective use of the resource potential of high-performance devices, resulting in insufficient optimization of the overall performance of the system. In a heterogeneous environment, this design limits the contribution of high-performance devices, and the participation delay of weak-performance devices also slows down the training process of the global model.

[0007] The synergy between large models and federated learning has important potential in terms of privacy protection and improving the generalization ability of artificial intelligence systems. Pre-trained large models are better adapted to specific tasks through fine-tuning, while federated learning allows mobile devices to fine-tune models without sharing original data, which not only protects data privacy but also improves generalization ability. However, the training of large models in federated learning requires a lot of computing and communication resources. Mobile devices with limited resources may face the problems of high time consumption and large amount of communication when fine-tuning. In addition, in practical applications, the communication bandwidth, computing power and memory between devices vary significantly. Since wireless federated learning needs to be performed synchronously on all devices, faster devices often reduce overall efficiency due to waiting for lagging devices with poorer resources. At the same time, the dynamic computing power and memory of the device are affected by background tasks, which limits the size of the sub-model that can be accommodated, and the difference between the sub-model and the global model will gradually increase, thus affecting the fine-tuning effect. Summary of the invention

[0008] In view of the defects of the prior art, the present invention proposes an adaptive large model distributed fine-tuning method for heterogeneous edge devices.

[0009] The present invention designs a wireless federated learning framework based on dynamic pruning and alignment mechanism. In the scenario of multiple edge base stations, it combines width pruning and depth pruning technology to dynamically adjust the pruning ratio of the model to adapt to the heterogeneous edge device environment, and reduces the gradient difference between the sub-model and the global model through hierarchical knowledge alignment and neuron-level alignment mechanism. In addition, the present invention also designs a resource allocation strategy that combines communication overhead and statistical heterogeneity, which significantly reduces energy consumption while improving system training performance and model accuracy, thereby providing an efficient and reliable solution for the practical application of large-model federated learning. The present invention significantly reduces communication and computing overhead, improves training efficiency, and improves system learning performance.

[0010] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0011] The present invention provides an adaptive large model distributed fine-tuning method for heterogeneous edge devices, which specifically includes the following steps:

[0012] Step 1: Sub-model construction and pruning;

[0013] The server decomposes the structure of the large model into multiple layers, and uses a two-dimensional large model pruning method that combines width and depth to build sub-models. Width pruning is to perform sequential neuron pruning on the FFN module of the large model to reduce computational complexity while retaining the MHA module and LoRA module. Depth pruning is to remove some deep structures from the entire transformer layer to reduce the overall depth, thereby building sub-models of different complexities based on device resources and improving the training efficiency of resource-constrained devices.

[0014] Step 2: Dynamic resource perception and sub-model allocation;

[0015] The server uses the client capability evaluation mechanism to calculate the communication, computation, and alignment capabilities of each client, and dynamically allocates sub-models of different complexities based on the capabilities of each client;

[0016] Step 3: Each client trains the assigned sub-model based on local private data. During the training process, the backbone parameters of the pre-trained model remain frozen, and only the LoRA module is updated.

[0017] Step 4: After local training is completed, each client uploads the updated LoRA module and local sub-model parameters to the server;

[0018] Step 5: The server performs aggregation operations based on the uploaded local sub-model parameters and LoRA modules, updates the LoRA modules of the global model, and keeps the backbone parameters of the global model unchanged;

[0019] Step 6: In the neuron-level alignment round, during the aggregation phase of the global model, the server performs hierarchical knowledge alignment and neuron-level alignment operations to obtain heterogeneous proxy sub-models by minimizing the layer-by-layer distillation loss to narrow the gap between the sub-models and the global model.

[0020] Step 7. Repeat steps 1 to 6 until convergence.

[0021] Furthermore, the large model includes: an embedding layer, a classifier based on a specific task and multiple transformer layers; each transformer layer contains two sublayers, namely an MHA module and an FFN module; in each transformer layer, sequential neuron pruning is used to perform width compression on the FFN module, while the MHA module remains unchanged.

[0022] Furthermore, the MHA module contains four model weight matrices: query matrix, key matrix, value matrix and output matrix, which are The model weight involved in the MHA module is W MHA = {W Q ,W K ,W V ,W O}, the sum of the parameters of the MHA module is The mathematical expression of the MHA module is:

[0023] MHA(x)=Concat(Attn 0 (x),...,Attn h (x))W O ,

[0024]

[0025] Among them, Attn 0 (x),…,Attn h (x) represents the output of each attention head in the multi-head attention mechanism, d model represents the dimension of the transformer layer input vector, and x represents the data input into the MHA module.

[0026] Furthermore, the FFN module includes two linear layers, and the weight matrices of the two linear layers are respectively and d ffRepresents the dimension of the FFN module, d model represents the dimension of the transformer layer input vector; the parameter quantity of the FFN module is The model weight involved in the FFN module is W FFN = {W 1 ,W 2}; The mathematical expression of the FFN module is:

[0027] FFN(x)=gelu(xW 1 +b 1 )W 2 +b 2 ,

[0028] Among them, x represents the data input to the FFN module, b 1 and b 2 Represent the bias vectors of the two layers of linear transformation respectively.

[0029] Furthermore, in the sequential neuron pruning process, the P-level pruning intensity is set, and for each level p∈{1,2,...,P}, there is a corresponding shrinkage ratio in both depth and width. The global model is denoted as W g = {W FFN ,W MHA ,W e}, the corresponding sub-model is expressed as W FFN Represents the model weights involved in the FFN module, W MHA Indicates the model weights involved in the MHA module, The corresponding shrinkage ratio is The model weights of the FFN module in the sub-model, Represents the model weight of the MHA module in the sub-model, Represents the model weight of the parameter efficient fine-tuning module in the sub-model; according to the shrinkage ratio The corresponding hierarchy of heterogeneous sub-models follows Each level of model is a subset of the previous level model; LoRA heterogeneous aggregation starts from the L P The pruning starts from the shallowest layer, with full participation of all clients, and gradually reduces the number of participating clients as the depth increases until only one level of clients participate in the aggregation in the deepest layer.

[0030] Furthermore, the calculation formula of the LoRA heterogeneous polymerization is:

[0031]

[0032] Among them, n i represents the amount of data owned by client i, nj represents the amount of data owned by client j, represents the model weight of the parameter efficient fine-tuning module corresponding to the lth layer, L represents the number of layers of the global model, and S(l) represents the set of clients participating in aggregation at the lth layer.

[0033] Furthermore, in step 2, the calculation formula for the communication, computing and alignment capabilities of each client is:

[0034]

[0035] in, represents the computation time of client i in round r, represents the transmission time of the parameter efficient fine-tuning module, represents the time spent on alignment in the alignment round, α represents the alignment interval in federated fine-tuning;

[0036] Submodel After normalization through normalized utility values, the client's capabilities are divided into P levels, each level corresponding to a sub-model of different complexity:

[0037]

[0038] in, Represents the overall normalized result.

[0039] Furthermore, in step 6, the hierarchical knowledge alignment operation is performed before the federated learning fine-tuning, and the heterogeneous proxy sub-model is obtained on the server side by minimizing the layer-by-layer distillation loss, which is defined as:

[0040]

[0041] Among them, L represents the number of layers of the global model, Indicates the shrinkage ratio of depth, M KD represents the size of the distillation dataset, and They represent the model weight W involved in the j-th sample in the l-th layer through the FFN module. FFN and its pruned version That is, the corresponding shrinkage ratio is The model weight of the FFN module in the sub-model, W 1 (l) and These are the model weights W involved in the FFN module FFN The corresponding weight matrix, and W 1 p(l) and The corresponding shrinkage ratio is The model weights of the FFN module in the sub-model The corresponding weight matrix, μ represents the regularization coefficient.

[0042] Furthermore, in step 6, the zero-value average percentage is used in the neuron-level alignment operation to select low-significance neurons for updating, helping the sub-model retain key local data knowledge; the APoZ score is used to determine which neurons need to be updated during the alignment process, which is defined as:

[0043]

[0044] Among them, M DT represents the size of the downstream dataset, S represents the sequence length of each sample; if the output of neuron n in the lth layer for the sth label of sample j is is zero, then the output will contribute to the APoZ score.

[0045] Furthermore, in the corresponding neuron-level alignment round, the APoZ score update process of the heterogeneous sub-model is:

[0046]

[0047] Among them, APoZ p The corresponding shrinkage ratio is The model weights of the FFN module in the sub-model APoZ score weight, APoZ p-1 The corresponding shrinkage ratio is The model weights of the FFN module in the sub-model The APoZ score weight of . The aggregated APoZ score is recorded as APoZ 1 , which is used to select the neurons in the FFN module that need to be updated in the current alignment round.

[0048] The beneficial effects of the present invention are:

[0049] a) The present invention realizes dynamic adaptation of device computing, communication and storage resources through adaptive pruning and dynamic sub-model allocation mechanism. High-performance devices are allocated complex sub-models to improve training performance; low-performance devices are allocated lightweight sub-models to ensure that they can participate in training smoothly, thus solving the problems of "performance bottleneck" and "resource waste" of heterogeneous devices in federated learning.

[0050] b) In the scenario where data distribution is not independent and identically distributed, the hierarchical knowledge alignment and neuron-level alignment mechanism proposed in this invention can effectively solve the alignment problem between heterogeneous sub-models and the global model. Through knowledge distillation technology, the features of local data from different devices are embedded into the sub-models, reducing the gradient differences between models and improving the convergence performance and stability of the global model under statistical heterogeneous conditions.

[0051] c) The pruning and dynamic allocation strategies proposed in this invention have wide applicability and are compatible with a variety of large model structures and efficient parameter fine-tuning methods (such as LoRA, Adapter, etc.). In addition, the dynamic sub-model allocation mechanism can adjust the model scale in real time according to the actual status of the device and adapt to resource fluctuations in a dynamic network environment, making it suitable for a variety of complex scenarios such as autonomous driving and intelligent edge computing.

[0052] d) The present invention effectively reduces the communication and computational overhead, and significantly improves the overall efficiency and convergence speed of federated learning. In the pruning design, the FFN module is selected as the main optimization target to reduce the number of parameters and computational complexity while maintaining the compatibility of the global model structure. In addition, the dynamic alignment mechanism reduces the information lost during the pruning process and maintains the stability and robustness of the model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 Optimize dynamic pruning and allocation strategies for adaptive large models on heterogeneous edge devices.

[0054] Figure 2 Pruning process for large models based on Transformer architecture.

[0055] Figure 3 Comparison of results of different pruning methods. DETAILED DESCRIPTION

[0056] The present invention is further described in detail below in conjunction with the accompanying drawings.

[0057] The present invention provides an adaptive large model distributed fine-tuning method for heterogeneous edge devices, and its specific implementation process is as follows:

[0058] Step S1: constructing a system model;

[0059] The present invention considers a wireless large-model federated learning system, such as Figure 1 As shown in the figure, the federated learning system includes a server and N clients, namely N users, and the basic weight of the pre-trained model is the global model weight W 0 Keep frozen in local training, the client only tunes and updates parameters Efficiently fine-tune module weights W e Each client has its own local private data set, denoted as {D 1 ,D 2 ,...,D N}, the client's local model set, i.e., the global model parameters, is represented as Θ = {W 1 ,W 2 ,...,W N}, where W i(i=1,2,…,N) is determined by the global model weight W 0 and parameters to efficiently fine-tune module weights W e Combination.

[0060] Step S2: Build a large-model heterogeneous perception federated learning model;

[0061] In each round of training, the server randomly selects K clients from N clients, and distributes the global model parameters Θ to the selected clients. The clients perform local training based on the local private dataset, and only tune and update the parameters to efficiently fine-tune the module weights W. e .

[0062] After all clients complete local training in round t, the client will use the updated parameters to efficiently fine-tune the module weights Upload to the server. The server performs weighted aggregation of updates from K clients using the following formula to generate new global parameters for efficient fine-tuning of weights

[0063]

[0064] Among them, |D k | represents the size of the local data set of client k, Indicates the global parameter efficient fine-tuning weight size updated by client k in the tth round of training; the new global parameter efficient fine-tuning weight generated Used for t+1 round model distribution and maintain the global model weight W 0 constant.

[0065] The following example uses the parameter efficient fine-tuning module (LoRA) to perform lightweight fine-tuning on a large model. The LoRA module introduces an additional parameter module without changing the core architecture of the pre-trained model. The LoRA module updates the matrix during training The server aggregates these update matrices to generate updated global parameters and efficiently fine-tune weights in Here r< <min(d g ,k g ), r represents the rank of the LoRA module, d g and k g Respectively represent the number of output channels and input channels of a certain layer of the global model. Through the above fine-tuning, the model only needs to tune 0.2% to 1.4% of the total parameters of the large model during uploading and updating, which significantly reduces the communication and computing overhead in federated learning. Among them, the forward propagation of the local model is defined as follows:

[0066]

[0067] in, represents the sub-model assigned to the client, α represents the scaling factor, x represents the input data of the client model, ΔW e It represents the parameter change learned by fine-tuning, that is, the update weight of the parameter efficient fine-tuning module.

[0068] Step S3: system target;

[0069] In order to efficiently fine-tune parameters of large models and reasonably implement device heterogeneity, improve global training efficiency, focus on the overall evaluation effect, and save communication costs to improve learning accuracy, the present invention uses system benefits to quantify the overall performance of federated learning.

[0070] The present invention defines the system benefits as follows: the system enables large models to acquire downstream task knowledge from resource-limited clients in heterogeneous and dynamic environments, and the fine-tuning parameters {W e ,W 0} is optimized to minimize the global objective function:

[0071]

[0072] Among them, ξ represents the local model loss function, y represents the target output, which is the true value that the model needs to predict. represents the expected operation, mathematically representing the local data distribution D from client i i Specifically, it takes a weighted average of all possible (x, y) pairs to reflect the average loss of the model on this data distribution. For the goal of global model generalization, the system accuracy can be defined as the overall test accuracy of all clients:

[0073]

[0074] Where G represents the global system accuracy, w i represents the weight of the i-th client, such as the number of test samples, g i Represents the accuracy of the global model on the test data of the i-th client.

[0075] Step S4: Adaptive large model distributed fine-tuning algorithm;

[0076] Since the existing efficient federated parameter tuning design for mobile devices is still insufficient in terms of architectural complexity and the adaptability of efficient parameter fine-tuning modules, the present invention proposes a two-dimensional large model pruning method that combines width and depth. By combining width pruning and depth pruning techniques, the pruning ratio of the model is dynamically adjusted to adapt to the heterogeneous edge device environment, and the model accuracy and training efficiency are balanced on the basis of ensuring that resource-constrained clients participate in training. The present invention combines efficient parameter fine-tuning with two-dimensional large model pruning strategies (width pruning and depth pruning) to cope with the heterogeneity of device computing and communication capabilities. Under the federated learning framework, the scale of the sub-model is dynamically adjusted to match the resource status of the client, while optimizing the training efficiency and performance of the global model.

[0077] In the present invention, the specific implementation process of the adaptive large model distributed fine-tuning algorithm is as follows:

[0078] S4.1: Sub-model construction and pruning strategy;

[0079] The server decomposes the structure of the large model into multiple layers (Layer 1 to Layer L), and uses a two-dimensional large model pruning method that combines width and depth to perform sequential neuron pruning for the Feed-Forward Network (FFN), while retaining the Multi-Head Attention (MHA) mechanism and the parameter efficient fine-tuning module (LoRA). The server constructs sub-models of different complexity, Sub-FFNA, Sub-FFN B, and Sub-FFN C, based on device resources.

[0080] The sub-model construction, i.e. the large model pruning process, is as follows: Figure 2 As shown. A large Transformer-based model usually consists of three parts: an embedding layer, a classifier based on a specific task, and multiple transformer layers. Each transformer layer contains two sublayers, namely a multi-head attention mechanism (MHA) module and a feedforward network (FFN) module. Among them, the multiple transformer layers are the parts that mainly affect the effect and memory usage. In the present invention, sequential neuron pruning is used to compress the width of the FFN module in each transformer layer, and the MHA module remains unchanged. This choice is based on the following two reasons: 1. Sequential neuron pruning can maintain the consistency of the model structure and reduce the differences in client models during the aggregation process; 2. The FFN module is focused on pruning because its parameter volume accounts for the highest proportion in the Transformer layer. Retaining the MHA module can ensure the structural compatibility between the sub-model and the global model, thereby minimizing the impact of pruning on model performance.

[0081] In the present invention, the MHA module can be expressed as:

[0082] MHA(x)=Concat(Attn 0 (x),...,Attn h (x))W O ,

[0083]

[0084] Among them, Attn 0 (x),…,Attn h (x) represents the output of each attention head in the multi-head attention mechanism, d model represents the dimension of the transformer layer input vector, and x represents the data input into the MHA module.

[0085] Among them, the four model weight matrices contained in the MHA module: query matrix, key matrix, value matrix and output matrix are The model weight involved in the MHA module is W MHA = {W Q ,W K ,W V ,W O}, the total number of parameters of the MHA module is approximately

[0086] In the present invention, the FFN module can be expressed as:

[0087] FFN(x)=gelu(xW 1 +b 1 )W 2 +b 2 ,

[0088] Among them, the weight matrices of the two linear layers in the FFN module are d ff Represents the dimension of the FFN module, usually set to 4×d model ; b 1 and b 2 Represent the bias vectors of the two-layer linear transformation respectively. Therefore, the number of parameters of the FFN module is approximately The model weight involved in the FFN module is W FFN = {W 1 ,W 2 Therefore, the MHA module is retained in width pruning, and sequential neuron pruning is performed on the FFN module.

[0089] Secondly, the present invention integrates the parameter efficient fine-tuning module (LoRA) in the MHA module to reduce the impact of pruning on model performance. Placing the LoRA module in the FFN module makes it more susceptible to the pruning of the FFN module during the aggregation process, while placing it in the MHA module can ensure the robustness of the fine-tuning process and reduce the interference of pruning on model performance.

[0090] Finally, to further reduce the memory usage and computational complexity of local sub-models (Sub-Foundation Models, Sub-FMs), the present invention performs depth pruning based on width pruning and removes the entire Transformer layer. By retaining the shallow structure and removing the deep structure, the present invention can retain the key low-order feature representation while maintaining the stability of the overall model performance.

[0091] In summary, for large model heterogeneous pruning, the present invention sets a total of P levels of pruning strength. For each level p∈{1,2,...,P}, there are corresponding shrinkage ratios in depth and width, which are expressed as The global model is denoted as W g = {W FFN ,W MHA ,W e}, the corresponding sub-model is expressed as Among them, W FFN Represents the model weights involved in the FFN module, W MHA Indicates the model weights involved in the MHA module, The corresponding shrinkage ratio is The model weights of the FFN module in the sub-model, Represents the model weight of the MHA module in the sub-model, Represents the model weight of the parameter efficient fine-tuning module in the sub-model; according to the corresponding shrinkage ratio, The corresponding hierarchy of heterogeneous sub-models follows Each level of model is a subset of the model at the previous level. The LoRA module will be affected by deep heterogeneous pruning, and the formula for LoRA heterogeneous aggregation is proposed, which is as follows:

[0092]

[0093] Among them, n i represents the amount of data owned by client i, n j represents the amount of data owned by client j, represents the model weight of the parameter efficient fine-tuning module corresponding to the lth layer, L represents the number of layers of the global model, and S(l) represents the set of clients participating in the aggregation at the lth layer. PThe pruning starts from the shallowest layer, with full participation of all clients, and gradually reduces the number of participating clients as the depth increases until only one level of clients participate in the aggregation in the deepest layer.

[0094] The present invention successfully solves the two major shortcomings of the existing methods by combining a two-dimensional pruning method of width pruning and depth pruning: on the one hand, width pruning fails to optimize the overall depth of the model, resulting in a high cumulative computing cost; on the other hand, depth pruning fails to effectively improve the computing efficiency within the layer. In addition, the present invention achieves compatibility between the pruned model and the LoRA module, thereby significantly reducing the communication and computing overhead of the model fine-tuning stage in federated learning. Figure 3 As shown in the figure, compared with the full model and the methods of using width pruning or depth pruning alone, the memory usage and computing time of 2D pruning in a single round of training are significantly reduced.

[0095] The present invention proposes an adaptive pruning mechanism that dynamically adjusts the size of the sub-model according to the computing and communication resources of the device by combining width pruning and depth pruning. In terms of width pruning, the feedforward network (FFN) module in the Transformer structure is pruned sequentially, while retaining the multi-head attention mechanism (MHA) structure to ensure that the pruned sub-model is consistent with the global large model structure and reduce the impact of pruning on model performance. In terms of depth pruning, the deep part of the large model is removed, the shallow structure is retained, and the complex deep features are replaced by low-order features, thereby reducing the computational complexity and storage consumption, allowing resource-constrained devices to participate in training.

[0096] S4.2: Submodel allocation and download;

[0097] The server dynamically allocates appropriate sub-models (including frozen backbone parameters and LoRA modules) to each client based on the client's computing power, communication bandwidth and other resource status, using dynamic resource perception and sub-model allocation strategy, and downloads them to the local device.

[0098] In order to reduce communication and computing delays and improve the training efficiency of large-model federated learning (FL) systems, the server needs to dynamically allocate appropriate sub-models based on the capabilities of participating clients. To this end, the present invention introduces a client capability evaluation mechanism that determines the actual load level of each client by calculating its communication, computing, and alignment capabilities, and dynamically allocates sub-models of different complexities. The communication, computing, and alignment capabilities of the client are derived from the available information using the following formula:

[0099]

[0100] in, represents the computation time of client i in round r, represents the transmission time of the parameter efficient fine-tuning module, represents the time spent on alignment in an alignment round, and α represents the alignment interval in federated fine-tuning.

[0101] Submodel After normalization through normalized utility values, the client's capabilities are divided into P levels, each level corresponds to a sub-model of different complexity, as follows:

[0102]

[0103] in, The client capability evaluation mechanism introduced in the present invention can dynamically adjust the scale of the heterogeneous sub-model according to the current capability of each client, thereby improving the delay efficiency and optimizing the resource allocation of the entire network.

[0104] In order to cope with the dynamic fluctuation of device resources, the present invention proposes a dynamic sub-model allocation strategy. The server comprehensively evaluates the overall capabilities of the device based on the client's computing time, communication time, and alignment time, and dynamically allocates sub-models of different complexities through the sub-model allocation formula. High-performance devices are allocated with more complex sub-models to give full play to their computing resources; weak-performance devices are allocated with lightweight sub-models to reduce the computing burden and ensure that all devices can participate in training efficiently.

[0105] S4.3: Local training;

[0106] Each client trains the assigned sub-model based on local private data. During the training process, the backbone parameters of the pre-trained model remain frozen and only the LoRA module is updated.

[0107] S4.4: LoRA module upload;

[0108] After local training is completed, each client uploads the updated LoRA module and local sub-model parameters to the server to reduce communication overhead.

[0109] S4.5: LoRA module aggregation;

[0110] The server performs aggregation operations based on the uploaded local sub-model parameters and LoRA modules, and updates the LoRA modules of the global model through weighted or other aggregation strategies, while keeping the backbone parameters of the global model unchanged.

[0111] S4.6: Neuron-level alignment;

[0112] Within the neuron-level alignment round, during the aggregation phase of the global model, the server performs hierarchical knowledge alignment and neuron-level alignment operations, narrowing the gap between the sub-model and the global model through methods such as knowledge distillation to ensure the performance and consistency of the global model.

[0113] The present invention designs an alignment mechanism between heterogeneous sub-models and the global model, including hierarchical knowledge alignment and neuron-level alignment. Hierarchical knowledge alignment aligns the pruned sub-models with the global model through a layer-by-layer distillation method, reducing the performance loss when the global model is aggregated. Neuron-level alignment selects low-significance neurons for update through the average percentage of zero values ​​(APoZ), retains the key knowledge of local data features, ensures that the training direction is consistent with the global model, and effectively solves the gradient offset problem.

[0114] S4.7: Repeat the above process until convergence.

[0115] In step S4.6, a sub-model and global model alignment strategy is adopted.

[0116] During the federated fine-tuning training process, the differences between the heterogeneous sub-models and the global model gradually accumulate, causing the training direction of the local heterogeneous sub-models to deviate from the ideal direction. The present invention takes the lead in studying the alignment mechanism between the heterogeneous sub-models and the global model. The traditional heterogeneous-aware pruning method ignores the gradient difference between the large model and the pruned heterogeneous sub-models. The present invention effectively narrows the gap between the heterogeneous sub-models and the global large model through the knowledge distillation process of hierarchical knowledge alignment and neuron-level alignment in heterogeneous scenarios.

[0117] Hierarchical knowledge alignment is performed before fine-tuning in federated learning. On the server side, the heterogeneous proxy sub-model is obtained by minimizing the layer-by-layer distillation loss. Its specific definition is as follows:

[0118]

[0119] Among them, L represents the number of layers of the global model, Indicates the shrinkage ratio of depth, M KD represents the size of the distillation dataset, and They represent the model weight W involved in the j-th sample in the l-th layer through the FFN module. FFN and its pruned version That is, the corresponding shrinkage ratio is The model weight of the FFN module in the sub-model, W 1 (l) and These are the model weights W involved in the FFN module FFN The corresponding weight matrix, and W 1 p(l) and The corresponding shrinkage ratio is The model weights of the FFN module in the sub-model The corresponding weight matrix, μ represents the regularization coefficient. At the beginning of each round of training, sub-models with different strengths are constructed through width and depth pruning, and the heterogeneous sub-model W is obtained on the server by minimizing the layer distillation loss p , distributed to the client for training.

[0120] During the federated fine-tuning process, the global model and the sub-model may gradually deviate, which requires a neuron-level alignment method to maintain consistency. However, over-alignment may limit the adaptability of the sub-model due to the domain difference between the distilled dataset and the local fine-tuning dataset. To reduce this risk, the present invention uses the Average Percentage of Zeros (APoZ) to select low-significance neurons for updating, thereby helping the sub-model retain key local data knowledge. The APoZ score is used to determine which neurons need to be updated during the alignment process, and its specific definition is as follows:

[0121]

[0122] Among them, M DT represents the size of the downstream dataset, and S represents the sequence length of each sample. If the output of neuron n in layer l for the sth token of sample j is If is zero, the output will contribute to the APoZ score. In the corresponding neuron alignment round, the APoZ score update process of the heterogeneous sub-model is as follows:

[0123]

[0124] Among them, APoZ p The corresponding shrinkage ratio is The model weights of the FFN module in the sub-model APoZ score weight, APoZ p-1 The corresponding shrinkage ratio is The model weights of the FFN module in the sub-model The APoZ score weight of . The aggregated APoZ score is recorded as APoZ 1 , which is used to select the neurons in the FFN module that need to be updated in the current alignment round.

[0125] In shallow layers, the APoZ score is usually smaller, while in deep layers it is larger, which indicates that shallow neurons are less involved in alignment, while deep neurons are more involved in the alignment process. This shows that shallow layers are mainly used to capture common features, and there are fewer neurons with high APoZ scores, so in the neuron-level alignment process, the proportion of shallow neurons is relatively small. The deep layers are responsible for extracting task-related high-level features and have stronger selectivity. Overall, hierarchical knowledge alignment and neuron-level alignment enhance the adaptability and performance of the model by ensuring consistency between sub-models and the global model in heterogeneous environments while retaining local knowledge.

[0126] The goal of the present invention is to achieve a trade-off between performance and training efficiency under the condition of limited heterogeneous device resources through dynamic pruning strategies and task-aware sub-model construction. The present invention proposes a strategy that combines width pruning and depth pruning, dynamically adjusts the pruning ratio according to the device computing power and memory limitations, and optimizes the hierarchy of sub-models for different tasks to ensure the retention of key features. In addition, the present invention introduces an adaptive alignment mechanism to reduce the impact of pruning on global model performance through hierarchical knowledge alignment and neuron-level alignment, while dynamically adjusting the client participation ratio and global aggregation weight to optimize resource allocation efficiency. Ultimately, the method of the present invention can maintain the robustness of model performance and the efficiency of global training while significantly reducing memory usage and computational overhead.

[0127] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. An adaptive large model distributed fine-tuning method for heterogeneous edge devices, characterized in that: The following steps are involved: Step 1: Sub-model construction and pruning; The server decomposes the structure of the large model into multiple layers and constructs sub-models using a two-dimensional large model pruning method that combines width and depth. Width pruning is to perform sequential neuron pruning on the FFN module of the large model to reduce computational complexity while retaining the MHA module and LoRA module. Depth pruning is to remove some deep structures from the entire transformer layer to reduce the overall depth, thereby constructing sub-models of different complexities based on device resource conditions. Step 2: Dynamic resource perception and sub-model allocation; The server uses the client capability evaluation mechanism to calculate the communication, computation, and alignment capabilities of each client, and dynamically allocates sub-models of different complexities based on the capabilities of each client; Step 3: Each client trains the assigned sub-model based on local private data; During the training process, the backbone parameters of the pre-trained model remain frozen and only the LoRA module is updated; Step 4: After local training is completed, each client uploads the updated LoRA module and local sub-model parameters to the server; Step 5: The server performs aggregation operations based on the uploaded local sub-model parameters and LoRA modules, updates the LoRA modules of the global model, and keeps the backbone parameters of the global model unchanged; Step 6: In the neuron-level alignment round, during the aggregation phase of the global model, the server performs hierarchical knowledge alignment and neuron-level alignment operations to obtain heterogeneous proxy sub-models by minimizing the layer-by-layer distillation loss to narrow the gap between the sub-models and the global model. Step 7. Repeat steps 1 to 6 until convergence.

2. According to claim 1, an adaptive large model distributed fine-tuning method for heterogeneous edge devices is characterized in that: The large model includes: an embedding layer, a classifier based on a specific task and multiple transformer layers; each transformer layer contains two sublayers, namely an MHA module and an FFN module; in each transformer layer, sequential neuron pruning is used to perform width compression on the FFN module, while the MHA module remains unchanged.

3. The adaptive large model distributed fine-tuning method for heterogeneous edge devices according to claim 2 is characterized in that: The MHA module contains four model weight matrices: query matrix, key matrix, value matrix and output matrix, which are The model weight involved in the MHA module is W MHA = {W Q ,W K ,W V ,W O }, the sum of the parameters of the MHA module is The mathematical expression of the MHA module is: MHA(x)=Concat(Attn0(x),...,Attn h (x))W O , Among them, Attn0(x),…,Attn h (x) represents the output of each attention head in the multi-head attention mechanism, d model represents the dimension of the transformer layer input vector, and x represents the data input into the MHA module.

4. The adaptive large model distributed fine-tuning method for heterogeneous edge devices according to claim 2 is characterized in that: The FFN module includes two linear layers, and the weight matrices of the two linear layers are and d ff Represents the dimension of the FFN module, d model represents the dimension of the transformer layer input vector; the parameter quantity of the FFN module is The model weight involved in the FFN module is W FFN ={W1, W2}; the mathematical expression of the FFN module is: FFN(x)=gelu(xW1+b1)W2+b2, Among them, x represents the data input to the FFN module, and b1 and b2 represent the bias vectors of the two-layer linear transformation respectively.

5. The adaptive large model distributed fine-tuning method for heterogeneous edge devices according to claim 1 is characterized in that: In the sequential neuron pruning process, the pruning intensity of level P is set, and for each level p∈{1,2,...,P}, there is a corresponding contraction ratio in depth and width. The global model is denoted as W g = {W FFN ,W MHA ,W e }, the corresponding sub-model is expressed as W FFN Represents the model weights involved in the FFN module, W MHA Indicates the model weights involved in the MHA module, The corresponding shrinkage ratio is The model weights of the FFN module in the sub-model, Represents the model weight of the MHA module in the sub-model, Represents the model weight of the parameter efficient fine-tuning module in the sub-model; according to the shrinkage ratio The corresponding hierarchy of heterogeneous sub-models follows Each level of model is a subset of the previous level model; LoRA heterogeneous aggregation starts from the L P The pruning starts from the shallowest layer, with full participation of all clients, and gradually reduces the number of participating clients as the depth increases until only one level of clients participate in the aggregation in the deepest layer.

6. The adaptive large model distributed fine-tuning method for heterogeneous edge devices according to claim 5 is characterized in that: The calculation formula of the LoRA heteropolymerization is: Among them, n i represents the amount of data owned by client i, n j represents the amount of data owned by client j, represents the model weight of the parameter efficient fine-tuning module corresponding to the lth layer, L represents the number of layers of the global model, and S(l) represents the set of clients participating in aggregation at the lth layer.

7. The adaptive large model distributed fine-tuning method for heterogeneous edge devices according to claim 1 is characterized in that: In step 2, the calculation formula for each client's communication, computing, and alignment capabilities is: in, represents the computation time of client i in round r, represents the transmission time of the parameter efficient fine-tuning module, represents the time spent on alignment in the alignment round, α represents the alignment interval in federated fine-tuning; Submodel After normalization through normalized utility values, the client's capabilities are divided into P levels, each level corresponding to a sub-model of different complexity: in, Represents the overall normalized result.

8. The adaptive large model distributed fine-tuning method for heterogeneous edge devices according to claim 1 is characterized in that: In step 6, the hierarchical knowledge alignment operation is performed before the federated learning fine-tuning, and the heterogeneous proxy sub-model is obtained on the server side by minimizing the layer-by-layer distillation loss, which is defined as: Among them, L represents the number of layers of the global model, Indicates the shrinkage ratio of depth, M KD represents the size of the distillation dataset, and They represent the model weight W involved in the j-th sample in the l-th layer through the FFN module. FFN And the corresponding shrinkage ratio is The model weights of the FFN module in the sub-model W1 (l) and These are the model weights W involved in the FFN module FFN The corresponding weight matrix, and W1 p(l) and The corresponding shrinkage ratio is The model weights of the FFN module in the sub-model The corresponding weight matrix, μ represents the regularization coefficient.

9. The adaptive large model distributed fine-tuning method for heterogeneous edge devices according to claim 1, characterized in that: In step 6, the zero-value average percentage is used in the neuron-level alignment operation to select low-significance neurons for updating, helping the sub-model retain key local data knowledge; the APoZ score is used to determine which neurons need to be updated during the alignment process, which is defined as: Among them, M DT represents the size of the downstream dataset, S represents the sequence length of each sample; if the output of neuron n in the lth layer for the sth label of sample j is is zero, then the output will contribute to the APoZ score.

10. The adaptive large model distributed fine-tuning method for heterogeneous edge devices according to claim 9, characterized in that: In the corresponding neuron-level alignment round, the APoZ score update process of the heterogeneous sub-model is: Among them, APoZ p The corresponding shrinkage ratio is The model weights of the FFN module in the sub-model APoZ score weight, APoZ p-1 The corresponding shrinkage ratio is The model weights of the FFN module in the sub-model APoZ score weight of The aggregated APoZ score is recorded as APoZ 1 , which is used to select the neurons in the FFN module that need to be updated in the current alignment round.

Citation Information

Cited By

  • Generalized enhanced sparse federated distillation and security defense fine tuning method for large model

    CN122293447A

  • Generalization-enhanced sparse federated distillation and security defense fine-tuning method for large models

    CN122293447B