Heterogeneous device sensing adaptive federated learning fine tuning method

By dynamically sensing the computing power of devices and network status, and adopting integer programming optimization and hierarchical weighted averaging strategies, the problems of idle resources and high computational complexity caused by static LoRA rank allocation are solved, and efficient federated learning fine-tuning of heterogeneous devices is achieved, thereby improving model accuracy and training efficiency.

CN120633890APending Publication Date: 2025-09-12TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510815398.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-12

Smart Images

  • Figure CN120633890A_ABST
    Figure CN120633890A_ABST
Patent Text Reader

Abstract

A self-adaptive federated learning fine tuning method for heterogeneous device perception comprises the steps of dynamically adjusting the LoRA rank of each device through integer programming optimization based on the real-time computing power and the network state of the device, constraining the time consumption of a single device and minimizing the global delay, realizing the rank improvement of a high-computing-power device to enhance contribution, and reducing the rank of a low-computing-power device to avoid the communication bottleneck. A local device performs fine tuning on a frozen pre-training model according to a distribution rank to generate a heterogeneous LoRA model, a server adopts a hierarchical aggregation strategy, aligns different rank parameters in an ascending order and only performs weighted average aggregation on common dimensions, the weight is jointly determined by a local data volume and rank square, high-dimensional matrix operation is avoided, and the calculation complexity of the server is remarkably reduced. And after aggregation, the global model adapts to each equipment target rank through matrix slice cutting, and redundant transmission is reduced. And iteratively executing until the model converges. And the problem of low efficiency caused by static resource allocation in federated learning is solved by dynamically allocating the low rank to adapt to the LoRA rank and an efficient aggregation mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to distributed machine learning technology, and in particular to a heterogeneous device-aware adaptive federated learning fine-tuning method. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, machine learning models have demonstrated powerful capabilities in tasks such as image recognition and natural language processing. However, centralized training requires the collection of massive amounts of user data, posing the risk of privacy breaches. Federated learning is a distributed machine learning paradigm that protects user privacy. Through a distributed training model, federated learning allows users to update models locally and upload only parameter updates, eliminating the need to share local private data. It is a leading solution for model training in privacy-sensitive scenarios. Federated learning fine-tuning is a pre-trained model fine-tuning technique based on the federated learning framework. It leverages data from multiple users to collaboratively fine-tune pre-trained models while preserving privacy.

[0003] In practical applications, the large number of pre-trained model parameters leads to problems such as insufficient memory on mobile devices, high communication overhead, and synchronization delays across heterogeneous devices, thus limiting the practicality of federated learning fine-tuning. To reduce communication and computational costs, efficient parameter fine-tuning techniques have been introduced into the federated learning fine-tuning framework. Among them, Low Rank Adaptation (LoRA) significantly reduces the amount of transmitted data by freezing the pre-trained model parameters and training only the low-rank incremental matrix. Traditional methods use a static LoRA rank assignment strategy, assigning the same fixed-rank parameter matrix to all devices. For example, some solutions directly set a uniform low-rank value for all users to reduce communication traffic.

[0004] FlexLoRA proposes a dynamic low-rank adaptation (LoRA) aggregation scheme for large language model (LLM) federated learning fine-tuning. By allowing clients to dynamically adjust the rank of local LoRA, it alleviates the "barrel effect" in traditional federated learning. Its core method is to integrate the resources of heterogeneous clients by synthesizing the full-size LoRA weight matrix contributed by each client and using singular value decomposition (SVD) to redistribute the weight matrix. Experiments show that FlexLoRA outperforms the existing federated learning method in downstream NLP tasks in various heterogeneous resource scenarios and is seamlessly compatible with the existing LoRA-based federated learning fine-tuning framework. Although FlexLoRA improves the utilization of heterogeneous resources through dynamic rank adjustment and SVD aggregation, it still has the following problems: (1) High computational complexity: It relies on SVD to decompose the full-size weight matrix, which has high time complexity (2) Insufficient static resource adaptation: It does not consider the real-time bandwidth fluctuation of the device and cannot dynamically respond to changes in network conditions. In contrast, the present invention proposes a hierarchical weighted average strategy to directly aggregate the shared dimension parameters of the heterogeneous LoRA matrix, thereby reducing the complexity of the aggregation process. Combined with historical communication data to predict bandwidth fluctuations, the LoRA rank of each device is adjusted in real time, simultaneously reducing the waiting time of high-resource devices and the communication overhead of low-resource devices.

[0005] FedPETuning proposes a Federated Parameter Efficient Fine-tuning (PETuning) benchmark framework for pre-trained language models (PLMs). Its core feature is a systematic evaluation of the comprehensive performance of four representative tuning methods, including LoRA, in federated learning fine-tuning. This method significantly reduces communication and computational overhead by freezing pre-trained parameters and fine-tuning only a small number of adapter modules (such as LoRA or Adapter). It also verifies its balance between defending against privacy attacks (such as gradient stealing) and maintaining model performance. However, FedPETuning has the following limitations: (1) Static parameter configuration: All clients use a fixed LoRA rank or adapter structure, resulting in high-resource devices being unable to enhance model contribution by increasing the LoRA rank, while low-resource devices face communication bottlenecks due to the fixed high rank; (2) It only supports the aggregation of homogeneous models and cannot adapt to scenarios where users use heterogeneous models.

[0006] Existing static LoRA rank allocation methods have significant flaws: First, differences in computing power, communication bandwidth, and data size between devices (spatial heterogeneity) are ignored. High-capability devices use low-rank matrices, resulting in idle resources, while low-capability devices suffer from slow uploads of high-rank matrices, hindering global synchronization efficiency. Second, dynamic fluctuations in network bandwidth (temporal heterogeneity) are not perceived. Fixed rank parameters cause a surge in communication time when bandwidth decreases, further exacerbating synchronization delays. Furthermore, existing model aggregation algorithms cannot efficiently handle heterogeneous model parameters. Some schemes achieve aggregation through singular value decomposition, but this has high computational complexity and is difficult to apply to large models. The core contradiction of the above issues lies in the inability of static LoRA rank allocation strategies to synergistically optimize the global efficiency and local resource utilization of federated learning fine-tuning.

[0007] 1. Static LoRA rank allocation leads to decreased training efficiency and resource utilization Existing methods use a fixed LoRA rank allocation strategy, ignoring the heterogeneity of computing power, bandwidth, and data size among devices, causing synchronization delays and resource waste. Specifically, (I) devices with strong computing and communication capabilities have idle computing resources because they use low-rank matrices; (II) devices with weak computing and communication capabilities have long upload times because they use high-rank matrices; (III) dynamic fluctuations in network bandwidth further exacerbate the differences in time consumption among different users. The above LoRA rank allocation strategy will significantly reduce the accuracy of the model. Figure 1 As shown in Figure 2, by increasing the LoRA rank of some users with strong performance, we can fully utilize user computing resources, reduce the time consumption differences between different users, and improve model accuracy.

[0008] 2. Low aggregation efficiency of heterogeneous LoRA models Existing aggregation methods cannot efficiently handle the aggregation problem of heterogeneous LoRA matrices. Existing solutions based on singular value decomposition require high-dimensional matrix multiplication operations on the server side, with a time complexity of (m and n are the maximum number of rows and columns in the LoRA matrix, respectively.) When applied to models with large parameter sizes, this approach generates extremely high computational load on the server side, leading to a sharp increase in server computing resource consumption. This increases time overhead and restricts the scalability of the federated learning fine-tuning system.

[0009] It should be noted that the information disclosed in the above background technology section is only used to understand the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0010] The main purpose of the present invention is to overcome the defects existing in the above-mentioned background technology and provide an adaptive federated learning fine-tuning method with heterogeneous device perception.

[0011] To achieve the above object, the present invention adopts the following technical solutions: A heterogeneous device-aware adaptive federated learning fine-tuning method includes the following steps: S1. Dynamic rank allocation: Based on the real-time computing capability of the device and the network status, the LoRA rank of each device is dynamically allocated through integer programming optimization to constrain the time consumption of a single device and minimize the maximum delay; S2. Local model update: Each device fine-tunes the local parameters of the frozen pre-trained model based on the assigned LoRA rank to generate a heterogeneous LoRA model; S3, hierarchical aggregation: The server aligns the LoRA model parameters of different ranks in ascending order, and only performs weighted average aggregation on the common dimensions. The weight is determined by the device data volume and the square of the LoRA rank. S4. Global model pruning: Slice and prune the aggregated global model according to the LoRA rank assigned to each device, and distribute the adapted local model parameters. S5. Iterative optimization: Repeat steps S1 to S4 until the model performance meets the requirements or the preset number of iterations is reached.

[0012] Furthermore, in step S1, the dynamic rank allocation specifically includes: Establish a global iteration time model to integrate device computing time and communication time; Predict real-time network bandwidth based on historical communication data and dynamically adjust the rank allocation scheme; The optimal LoRA rank combination is solved through an integer programming optimization problem with a low-rank penalty term.

[0013] Furthermore, in step S3, the hierarchical aggregation specifically includes: Align LoRA matrices of different ranks in ascending order and split them into multiple shared dimension intervals; Perform weighted averaging on the parameters within each shared dimension interval, with the weight calculated by normalizing the product of the device's local data volume and the square of the LoRA rank; High-dimensional matrix operations are completely avoided through layer-by-layer slicing and aggregation.

[0014] Furthermore, in step S4, the global model clipping specifically includes: According to the LoRA rank assigned to the device, the corresponding parameter subset of the global model is intercepted; Only slice parameters that match the target rank are distributed to each device to avoid redundant transmission.

[0015] Furthermore, the method further comprises: Collect device local computation time during initial communication and continuously update device network bandwidth estimates; The rank allocation constraint interval is dynamically adjusted through the time model to adapt to device resources and network fluctuations.

[0016] Furthermore, in the hierarchical aggregation, the weight of the weighted average is determined by: A normalized calculation is performed based on the local data volume of each device and the square value of its LoRA rank to enhance the contribution weight of high-rank devices and devices with large data volumes.

[0017] Furthermore, the objective function of the integer programming optimization problem is to minimize the maximum device delay, and the constraints include: The total time consumed by a single device is within the preset time target range; The global low-rank penalty term is inversely proportional to the sum of the LoRA ranks of each device.

[0018] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the heterogeneous device-aware adaptive federated learning fine-tuning method.

[0019] A computer program product includes a computer program, which, when executed by a processor, implements the heterogeneous device-aware adaptive federated learning fine-tuning method.

[0020] The present invention has the following beneficial effects: This paper proposes a heterogeneous device-aware adaptive federated learning fine-tuning method. This innovative approach addresses the low resource utilization and synchronization efficiency issues inherent in existing static low-rank adaptation (LoRA) rank allocation strategies in federated learning fine-tuning, as well as the high computational complexity of heterogeneous model aggregation. Its core technical advantage lies in dynamically sensing the real-time computing power and network status of devices. Using integer programming, it dynamically allocates LoRA ranks to each device. This allows high-computing devices to increase their rank within their time budget to enhance model contribution, while limiting the rank of low-computing devices to avoid communication latency, thereby synergistically optimizing global training efficiency and local resource utilization. Furthermore, a low-complexity hierarchical aggregation strategy is proposed. By aligning LoRA matrices of different ranks in ascending order and performing a weighted average aggregation only on shared dimensions, combined with a weighted design based on device data volume and LoRA rank square, the contribution of high-rank, data-intensive devices is enhanced while completely avoiding high-dimensional matrix operations, significantly reducing server-side aggregation time overhead. Furthermore, a global model pruning technique distributes adapted local parameter subsets through matrix slicing, further reducing redundant transmission. This method achieves 58% and 9% reduction in global training time in visual tasks (CIFAR-10) and natural language processing tasks (News Category), respectively, and the server aggregation time complexity is reduced from 10% to 10% of the traditional method. Reduce to It effectively supports efficient privacy-compliant fine-tuning of pre-trained large language models on mobile devices, adapts to differences in device computing power and dynamic network fluctuations, and provides a solution that balances accuracy, efficiency, and scalability for heterogeneous federated learning systems.

[0021] Other beneficial effects of the embodiments of the present invention will be further described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is an overall flow chart of the heterogeneous device-aware adaptive federated learning fine-tuning method of the present invention.

[0023] Figure 2 This is a diagram of the overall algorithm framework of the adaptive federated learning fine-tuning method according to an embodiment of the present invention.

[0024] Figure 3 Model accuracy curves of the federated learning fine-tuning algorithm with different LoRA ranks assigned to different users. DETAILED DESCRIPTION

[0025] The following is a detailed description of the embodiments of the present invention. It should be emphasized that the following description is only exemplary and is not intended to limit the scope of the present invention and its application.

[0026] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present invention, "plurality" means two or more, unless otherwise specifically defined.

[0027] To address the challenge that the static LoRA rank allocation strategy cannot collaboratively optimize the global efficiency and local resource utilization of federated learning fine-tuning, the present invention proposes a heterogeneous device-aware adaptive federated learning fine-tuning method. By real-time perception of device computing power and network status, the LoRA rank parameters of each device are dynamically adjusted, and an efficient hierarchical aggregation algorithm is designed for heterogeneous user models. This maximizes resource utilization while reducing communication overhead, ensuring the collaborative optimization of model accuracy and training efficiency.

[0028] See Figure 1 , an embodiment of the present invention provides a heterogeneous device-aware adaptive federated learning fine-tuning method, comprising the following steps: Step S1, dynamic rank allocation: Based on the real-time computing capability of the device and the network status, the LoRA rank of each device is dynamically allocated through integer programming optimization, which constrains the time consumption of a single device and minimizes the maximum delay; Step S2, local model update: each device fine-tunes the local parameters of the frozen pre-trained model based on the assigned LoRA rank to generate a heterogeneous LoRA model; Step S3, hierarchical aggregation: The server aligns the LoRA model parameters of different ranks in ascending order, and only performs weighted average aggregation on the common dimensions, with the weight determined by the device data volume and the square of the LoRA rank; Step S4, global model clipping: Slice and clip the aggregated global model according to the LoRA rank assigned to each device, and distribute the adapted local model parameters; Step S5, iterative optimization: Repeat the above steps of dynamic rank allocation, local update, hierarchical aggregation and pruning until the model performance meets the standard or reaches the preset number of iterations.

[0029] In some embodiments, in step S1, the dynamic rank allocation specifically includes: establishing a global iterative time model to integrate device computing time and communication time; predicting real-time network bandwidth based on historical communication data, and dynamically adjusting the rank allocation scheme; solving the optimal LoRA rank combination through an integer programming optimization problem containing a low-rank penalty term, ensuring that high-capability devices are assigned high ranks to improve their contributions, and low-capability devices are limited to low ranks to avoid timeouts.

[0030] In some embodiments, in step S3, the hierarchical aggregation specifically includes: aligning LoRA matrices of different ranks in ascending order and dividing them into multiple shared dimension intervals; performing weighted averaging on the parameters in each shared dimension interval, and the weight is calculated by normalizing the product of the local data volume of the device and the square of the LoRA rank; completely avoiding high-dimensional matrix operations through layer-by-layer slicing aggregation, thereby reducing the server computing complexity.

[0031] In some embodiments, in the hierarchical aggregation, the weight of the weighted average is determined by normalizing the local data volume of each device and the square value of its LoRA rank to enhance the contribution weight of high-rank devices and large data volume devices.

[0032] In some embodiments, in step S4, the global model clipping specifically includes: intercepting a corresponding parameter subset of the global model according to the LoRA rank assigned to the device; and only distributing slice parameters that match the target rank to each device to avoid redundant transmission.

[0033] In some embodiments, the heterogeneous device-aware adaptive federated learning fine-tuning method further includes: collecting the local computing time of the device during the first communication and continuously updating the device network bandwidth estimation value; dynamically adjusting the rank allocation constraint interval through the time model to adapt to device resources and network fluctuations.

[0034] In some embodiments, the objective function of the integer programming optimization problem is to minimize the maximum device delay, and the constraints include: the total time consumption of a single device (the sum of computing time and communication time) is within a preset time target interval; the global low-rank penalty term is inversely proportional to the sum of the LoRA ranks of each device.

[0035] The method of the embodiment of the present invention is applicable to scenarios including but not limited to: privacy-preserving fine-tuning of pre-trained large language models on mobile devices; collaborative training of distributed devices with heterogeneous computing power and network bandwidth; and federated learning tasks that require balancing model accuracy and synchronization efficiency in a dynamic network environment.

[0036] The following further describes specific embodiments of the present invention, its algorithm examples and experimental verification.

[0037] This paper proposes a heterogeneous device-aware adaptive federated learning fine-tuning method. This method addresses the imbalanced device resource utilization and inefficient synchronization caused by static LoRA rank allocation, as well as the high computational overhead of heterogeneous model aggregation. It achieves system-level optimization through a dynamic rank allocation mechanism and an efficient aggregation algorithm. The core solution is as follows: First, to address the issue of inefficient synchronization caused by differences in device computing power and network fluctuations, the present invention proposes a dynamic rank adjustment mechanism based on real-time resource awareness. By establishing a global iterative time model that integrates device computing time and communication time, an integer programming problem is designed to dynamically allocate the LoRA rank of each device. This time model aims to minimize maximum device latency, constraining the time consumption of individual devices to within a preset interval. The optimal LoRA rank allocation scheme is obtained by solving an optimization equation with a low-rank penalty term. This scheme allows high-computing-power, high-communication-capability devices to select a higher LoRA rank within their time budget to improve model contribution. At the same time, it limits the LoRA rank of low-computing-power, low-communication-capability devices to avoid exceeding the time budget. To address changes in user device network conditions, the dynamic bandwidth prediction module estimates the current network status based on historical communication delays, ensuring that the rank allocation scheme adapts to real-time fluctuations. After the aggregation step, a global model pruning technique uses matrix slicing to distribute only the subset of LoRA model parameters that matches each device's target rank, avoiding redundant model parameter transmission.

[0038] Secondly, to address the problem of low aggregation efficiency of heterogeneous LoRA models, traditional methods rely on high-dimensional matrix operations (such as singular value decomposition), which leads to a sharp increase in server computing load and makes it difficult to support dynamic LoRA rank scheduling requirements. The present invention proposes a low-complexity hierarchical aggregation strategy, the innovation of which is reflected in: aligning LoRA matrices of different ranks in ascending order, and only performing weighted aggregation on the common dimensional parts. The aggregation weight is determined by the amount of local data of the device and the square of its LoRA rank, which enhances the contribution of high-rank and large-data-volume devices. By directly calculating the weighted average value by slicing layer by layer, the matrix multiplication and decomposition operations are completely avoided. The above-mentioned efficient aggregation method reduces the time complexity to ,in is the maximum value of the LoRA rank of all users, satisfying .

[0039] Figure 2 The overall algorithm framework of the embodiment of the present invention is presented.

[0040] The federated learning fine-tuning system consists of a central server and several user devices. In each global iteration, the central server broadcasts the global model to the user devices. Each user device updates its local model in parallel and uploads the model to the central server to aggregate the new global model. The detailed steps are as follows: Step 1: The server broadcasts the pre-trained model and the randomly initialized LoRA model with the same rank to all user devices.

[0041] Step 2: All users update the LoRA model on the local dataset using the stochastic gradient descent algorithm in parallel while keeping the pre-trained model parameters frozen.

[0042] Step 3: The user uploads the updated LoRA model to the server, and then uploads the time cost of the LoRA model to the server. If it is the first communication, the local calculation time must also be uploaded.

[0043] Step 4: After collecting updates from all user devices, the server aggregates them using an efficient heterogeneous LoRA model aggregation method:

[0044]

[0045] in, is the global LoRA model, For users Local LoRA model. Represents the matrix arrive List. is the total number of users, Sort in ascending order There are different rank values, and the user numbers are arranged in ascending order of rank. The LoRA matrix contains part. And:

[0046] in , For users The amount of data used by users is compared to the total amount of data used by all users.

[0047] Step 5: The server assigns appropriate LoRA ranks to all users by solving an integer optimization problem:

[0048] Among them, is a non-negative hyperparameter used to adjust the weight of the low-rank penalty. is a non-negative hyperparameter that determines the target value for training time. is a non-negative hyperparameter that determines the maximum idle time allowed for a user device. By all User equipment in the A tensor consisting of the LoRA ranks of global iterations, where and denote the sizes of the output layer adapter model and the rank-1 LoRA model, respectively. For users The calculation time, For users In the The upload rate of a global iteration is estimated using the time consumption of the previous global iteration.

[0049] Step 6: Prune the global model according to the LoRA rank assigned to each user to obtain the local model:

[0050] The local model is broadcast to all users. Return to step 2.

[0051] Repeat the above cycle until the model performance meets the requirements or the number of iterations reaches the upper limit.

[0052] Let's take the example of fine-tuning a pre-trained large language model across mobile devices. The system consists of a central server and a large number of mobile device users (such as mobile phones, smartwatches, and tablets). Each user holds private local language data, such as keyboard input history and text messages. The goal is to fine-tune the large language model based on this private data, enabling it to learn the language style of the corresponding user group. The following are the specific implementation steps for this scenario: Step 1: The server broadcasts the pre-trained large language model and the randomly initialized LoRA model with the same rank to all user devices.

[0053] Step 2: All users update the LoRA model on the local language dataset using the stochastic gradient descent algorithm in parallel while keeping the pre-trained model parameters frozen.

[0054] Step 3: The user uploads the updated LoRA model to the server, and then uploads the time cost of the LoRA model to the server. If it is the first communication, the local calculation time must also be uploaded.

[0055] Step 4: After collecting updates from all user devices, the server aggregates them using an efficient heterogeneous LoRA model aggregation method.

[0056] Step 5: The server assigns appropriate LoRA ranks to all user devices by solving an integer optimization problem.

[0057] Step 6: Prune the global model according to the LoRA rank assigned to each user to obtain the local model.

[0058] Step 7: The server tests the accuracy of the user model. If the accuracy requirement is met, the federated fine-tuning is terminated; otherwise, the process returns to step 2.

[0059] The above steps achieve fine-tuning of the large language model while protecting the data privacy of mobile device users and avoiding high computational and communication overhead.

[0060] Table 1 shows the model accuracy of fine-tuning DistilBert on the News Category dataset. The results show that the algorithm of the present invention has advantages over the existing algorithms in terms of model performance at the same time and the same global iteration.

[0061] Table 1: Algorithm performance comparison

[0062] Figure 3The model accuracy curve of the federated learning fine-tuning algorithm with different LoRA ranks assigned to different users. The experiment shows the results when the rank of the poorly performing user is 1 and the ranks of the strong performing users are set to 2, 4, and 8. The experimental results show that by increasing the LoRA model rank of users with strong computing and communication capabilities, the model accuracy can be significantly improved under the same global iteration. Figure 3 As shown in Figure 2, by increasing the LoRA rank of some users with strong performance, we can fully utilize user computing resources, reduce the time consumption differences between different users, and improve model accuracy.

[0063] The present invention can be applied to scenarios such as fine-tuning of pre-trained large language models on mobile devices. Users around the world interact through mobile devices (such as smartphones and tablets), generating a large amount of language data that can be used to fine-tune large language models. However, user local data contains private information such as sensitive conversation records and input habits, and the hardware capabilities of different devices vary significantly, making it difficult to directly collect data for fine-tuning. In addition, some privacy regulations strictly prohibit the transmission of raw data, making traditional centralized fine-tuning methods difficult to implement. Through the present invention, mobile devices can efficiently fine-tune large language models based on local private data. The present invention enables various mobile devices to collaboratively optimize language models without uploading raw data: the resource usage of low-end devices is greatly reduced, supporting a wide range of models to participate in training; communication overhead and computing load are significantly reduced, meeting real-time response and privacy compliance requirements.

[0064] Compared with traditional technologies, the present invention has the following significant advantages: Significantly improved training efficiency: This paper uses a dynamic LoRA rank assignment mechanism to reduce global training time by 58% for vision tasks (CIFAR-10 dataset) and 9% for NLP tasks (News Category dataset), while maintaining the same model accuracy.

[0065] The computational overhead on the server side is significantly reduced. The efficient heterogeneous aggregation method says that the time complexity of server aggregation of heterogeneous user models is: Reduce to , significantly reducing the aggregation time overhead.

[0066] An embodiment of the present invention further provides a storage medium for storing a computer program, which at least performs the above method when executed.

[0067] An embodiment of the present invention further provides a control device, comprising a processor and a storage medium for storing a computer program; wherein the processor is configured to execute at least the method described above when executing the computer program.

[0068] An embodiment of the present invention further provides a processor, which executes a computer program and at least performs the method described above.

[0069] The storage medium can be implemented by any type of non-volatile storage device, or a combination thereof. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disc or a read-only optical disc (CD-ROM); the magnetic surface memory can be a magnetic disk memory or a magnetic tape memory. The storage medium described in the embodiments of the present invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0070] In the several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0071] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0072] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0073] Those skilled in the art will appreciate that all or part of the steps of the above-mentioned method embodiments may be implemented by hardware associated with program instructions, and the aforementioned program may be stored in a computer-readable storage medium. When the program is executed, the program executes the steps of the above-mentioned method embodiments. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0074] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0075] The methods disclosed in the several method embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments.

[0076] The features disclosed in several product embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new product embodiments.

[0077] The features disclosed in several method or device embodiments provided by the present invention can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0078] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. Those skilled in the art will recognize that, without departing from the scope of the present invention, several equivalent substitutions or obvious variations can be made, and the performance or use of the same should be considered to fall within the scope of protection of the present invention.

Claims

1. A heterogeneous device-aware adaptive federated learning fine-tuning method, characterized by: The following steps are involved: S1. Dynamic rank allocation: Based on the real-time computing capability of the device and the network status, the LoRA rank of each device is dynamically allocated through integer programming optimization to constrain the time consumption of a single device and minimize the maximum delay; S2. Local model update: Each device fine-tunes the local parameters of the frozen pre-trained model based on the assigned LoRA rank to generate a heterogeneous LoRA model; S3, hierarchical aggregation: The server aligns the LoRA model parameters of different ranks in ascending order, and only performs weighted average aggregation on the common dimensions. The weight is determined by the device data volume and the square of the LoRA rank. S4. Global model pruning: Slice and prune the aggregated global model according to the LoRA rank assigned to each device, and distribute the adapted local model parameters. S5. Iterative optimization: Repeat steps S1 to S4 until the model performance meets the requirements or the preset number of iterations is reached.

2. The method according to claim 1, characterized in that In step S1, the dynamic rank allocation specifically includes: Establish a global iteration time model to integrate device computing time and communication time; Predict real-time network bandwidth based on historical communication data and dynamically adjust the rank allocation scheme; The optimal LoRA rank combination is solved through an integer programming optimization problem with a low-rank penalty term.

3. The method according to claim 1, characterized in that In step S3, the hierarchical aggregation specifically includes: Align LoRA matrices of different ranks in ascending order and split them into multiple shared dimension intervals; Perform weighted averaging on the parameters within each shared dimension interval, with the weight calculated by normalizing the product of the device's local data volume and the square of the LoRA rank; Aggregation is performed by slicing layer by layer to avoid high-dimensional matrix operations.

4. The method according to claim 1, wherein In the hierarchical aggregation, the weight of the weighted average is determined by: A normalized calculation is performed based on the local data volume of each device and the square value of its LoRA rank to enhance the contribution weight of high-rank devices and devices with large data volumes.

5. The method according to claim 1, characterized in that In step S4, the global model clipping specifically includes: According to the LoRA rank assigned to the device, the corresponding parameter subset of the global model is intercepted; Only slice parameters that match the target rank are distributed to each device to avoid redundant transmission.

6. The method according to claim 1, characterized in that The method further comprises: Collect device local computation time during initial communication and continuously update device network bandwidth estimates; The rank allocation constraint interval is dynamically adjusted through the time model to adapt to device resources and network fluctuations.

7. The method according to claim 1, characterized in that The objective function of the integer programming optimization problem is to minimize the maximum device delay, and the constraints include: The total time consumed by a single device is within the preset time target range; The global low-rank penalty term is inversely proportional to the sum of the LoRA ranks of each device.

8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for adaptive federated learning fine-tuning based on heterogeneous device perception is implemented.

9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method for adaptive federated learning fine-tuning based on heterogeneous device perception is implemented.

Citation Information

Cited By

  • Multi-task energy consumption sensing rank self-adaptive federal fine tuning method for Internet of Vehicles

    CN121349661A