Distributed Transform training acceleration framework in federated learning

By constructing a distributed Transformer training acceleration framework, the problems of slow training speed, high communication overhead, and insufficient privacy protection in federated learning systems are solved, achieving efficient global model training and privacy protection.

CN120996140APending Publication Date: 2025-11-21XIANGTAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410759432.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-12
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing Transformer-based personalized federated learning systems suffer from slow training speeds, high communication overhead, and insufficient privacy protection.

Method used

A distributed Transformer training acceleration framework for federated learning was constructed, including a dynamic model segmentation module, an intelligent task scheduling module, a hierarchical communication optimization module, and an adaptive learning rate adjustment module. By dynamically segmenting the model, optimizing task allocation, asynchronous communication, and differential privacy technology, training efficiency is improved and communication overhead is reduced.

Benefits of technology

It significantly improves model training speed, optimizes the utilization of computing resources, reduces communication overhead, and effectively protects user data privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a distributed Transform training acceleration framework in federated learning. The training speed of a model based on Transform is relatively slow and is mainly limited by the aspects of calculation and communication. According to the scheme, firstly, dynamic model segmentation is introduced, and the Transform model is dynamically cut according to data distribution on local equipment, so that the calculation complexity of each piece of equipment is reduced; an intelligent task scheduling strategy ensures that each device can effectively participate in global model training, and computer resources of the device are fully utilized; the communication overhead is reduced through layered communication optimization, and the security and privacy are guaranteed by transmitting key information and adopting asynchronous communication and differential privacy at the same time; the system further comprises a model fusion acceleration unit which allows local equipment to finish partial model training locally, and then quickly fuses updating of model weights into a global model. The adaptive learning rate adjustment mechanism dynamically adjusts the learning rate according to the training progress and the model convergence condition of each device, and the training rate is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication technology, and specifically to personalized federated learning. Background Technology

[0002] In today's digital age, the decentralized nature of data and the need for privacy protection pose new challenges to the training of machine learning models. Federated learning, as a decentralized learning method, achieves global model training by training models on local devices without sharing raw data, thus avoiding the transmission of sensitive information to centralized servers. Against this backdrop, personalized federated learning systems based on Transformers have emerged, but their training speed still lags behind CNN-based models. To address this issue, we introduce a series of key technical modules to construct a more efficient distributed Transformer training acceleration framework for federated learning.

[0003] This invention relates to an innovative distributed Transformer training acceleration framework in federated learning. By introducing several key technical modules, it improves system training speed, reduces communication overhead, and provides an efficient global model while ensuring privacy. First, through a dynamic model segmentation module, we can analyze the distribution of local data and dynamically segment the Transformer model, allowing each local device to focus only on the model portion relevant to its data, thereby improving training efficiency. Second, through an intelligent task scheduling module, we evaluate the performance of each device, considering task priority, data characteristics, and heterogeneous device adaptability, and dynamically adjust task allocation weights to optimize the allocation of global training tasks. A hierarchical communication optimization module employs asynchronous communication and differential privacy technology to reduce data transmission volume between devices, ensuring communication security and privacy. A model fusion acceleration module allows each local device to complete part of the model training locally and quickly merges the updates of the local model into the global model through a fast fusion algorithm, improving training efficiency. Finally, an adaptive learning rate adjustment module dynamically adjusts the learning rate by monitoring the training progress and model convergence on each device in real time, ensuring the adaptability of the learning rate in the federated learning system. By integrating these key technology modules, we have constructed a distributed Transformer training acceleration framework in federated learning that has broad application prospects when dealing with large-scale, heterogeneous devices and privacy-sensitive decentralized data. Summary of the Invention

[0004] This invention relates to a distributed Transformer training acceleration framework in federated learning, aiming to overcome the problems of slow training speed, high communication overhead, and insufficient privacy protection in traditional systems. By introducing several key technical modules, we have successfully built an innovative system that provides an efficient global model training solution.

[0005] This invention proposes a distributed Transformer training acceleration framework in federated learning, aiming to overcome the problems of slow training speed, high communication overhead, and insufficient privacy protection in traditional systems. The method mainly includes: a dynamic model segmentation module, an intelligent task scheduling module, a hierarchical communication optimization module, a model fusion acceleration module, and an adaptive learning module. Dynamic segmentation of the Transformer model is achieved by analyzing local data; the intelligent task scheduling module dynamically adjusts and optimizes the allocation of global training tasks; the model fusion acceleration module accelerates the global model fusion process; and the adaptive learning rate adjustment module dynamically adjusts the learning rate.

[0006] The present invention has the following advantages:

[0007] 1. This invention can effectively improve the training speed of models;

[0008] 2. This invention can optimize the utilization of computing resources and improve the overall efficiency of the system;

[0009] 3. This invention effectively reduces communication overhead;

[0010] 4. This invention effectively protects the privacy of user data. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating the specific implementation method of the dynamic model segmentation module of the present invention;

[0012] Figure 2 This is a flowchart illustrating the specific implementation method of the intelligent task scheduling module of the present invention;

[0013] Figure 3 This is a flowchart illustrating a specific implementation method of the hierarchical communication optimization module of the present invention;

[0014] Figure 4 This is a flowchart illustrating the specific implementation method of the adaptive learning rate adjustment module of the present invention; Specific implementation methods

[0015]

[0016]

[0017] This invention designs a distributed Transformer training acceleration framework in federated learning, combining... Figure 1, Figure 2 , Figure 3 The specific implementation method is as follows:

[0018] Step 1: Specific Implementation Method of Dynamic Model Segmentation Module

[0019] 1) Utilize statistical methods and machine learning techniques to conduct a detailed analysis of the data on the local device, including data distribution (D), characteristics, and importance. Extract data features and use techniques such as principal component analysis (PCA) to obtain the feature matrix X;

[0020] 2) Based on the results of data analysis, formulate specific strategies for dynamic model segmentation; this may involve algorithms that dynamically adjust the model segmentation ratio, such as αi representing the model segmentation ratio on device i; ensure that the model segmentation strategy can be adaptively adjusted under different data distributions to optimize the training effect of each local device;

[0021] 3) Based on the established strategy, dynamically segment the Transformer model; this involves defining the segmentation ratio, selecting the segmentation level or model block, etc.; after segmentation, a local model Mi is generated that adapts to the data distribution of each local device, where i represents the index of the local device;

[0022] 4) The local device uses the segmented local model and trains the model only on the part that is related to its data; the training process can use the conventional backpropagation algorithm, but only updates the segmented model part;

[0023] 5) After local training is complete, the updated portions ΔMi of the local model from each local device are quickly fused into the global model; weighted fusion or other fast fusion algorithms can be used; the global model parameters are represented by M. global This indicates that the update rule is M. global ←M global +∑ i α i ΔM i ;

[0024] 6) Introduce a real-time monitoring mechanism to monitor the model's performance and training progress on various local devices; based on the monitoring results,

[0025] Dynamically adjust the model segmentation strategy to ensure the robustness and efficiency of the system under different data distribution conditions;

[0026] Step Two: Implementation Phase of the Intelligent Task Scheduling Module

[0027] 1) Evaluate the performance of each device participating in federated learning, including parameters such as computing power and storage capacity; obtain the task priority Pi, taking into account the device's performance evaluation, the importance of the task, and the data characteristics;

[0028] 2) Decompose the global training task into multiple subtasks, each corresponding to a local device; the task decomposition process considers task priority and data characteristics; define the task decomposition matrix T. ij , represents the weight assigned to subtask j on device i;

[0029] 3) Dynamically adjust task allocation weights based on real-time monitored equipment performance, task priorities, and data characteristics; the adjustment rules can adopt feedback control-based algorithms to ensure that task allocation can adapt to changes in equipment performance;

[0030] 4) Use optimization algorithms, such as genetic algorithms or particle swarm optimization, to optimize the task allocation strategy to maximize the overall utilization of computing resources; define the task allocation vector W. i , representing the weight of each subtask on device i;

[0031] 5) Based on the final task allocation weights, start the corresponding subtasks on each device; each device performs local model training according to the assigned subtasks;

[0032] 6) Consider the different characteristics of heterogeneous devices, such as the differences between mobile devices and servers; formulate corresponding task scheduling strategies to ensure that task allocation is adapted to various heterogeneous devices;

[0033] Step 3: Implementation Phase of the Layered Communication Optimization Module

[0034] 1) After the global model update, calculate the change (gradient or weight difference) of each model parameter to obtain the change matrix ΔM. i ;

[0035] 2) Divide the model parameters into multiple levels based on their importance, update frequency, or other characteristics; define a hierarchical matrix L. ii , indicating that parameter j belongs to the i-th layer;

[0036] 3) An asynchronous communication mechanism is used, transmitting only the changed portions of the model parameters, not the entire model; an asynchronous communication matrix A is defined. ii This indicates whether parameter j requires asynchronous communication on device i;

[0037] 4) Utilize communication compression technology to compress the changes in parameters that need to be transmitted; use encoding technology to encode the compressed data to improve communication efficiency;

[0038] 5) Dynamically adjust the communication frequency based on the importance of each parameter level and the computing power of the device; control the frequency matrix F. ij , representing the communication frequency of parameter j on i;

[0039] 6) Differential privacy techniques are employed during communication to add random noise to the model parameters in order to protect the privacy of the communication; a privacy matrix ∈ ij , representing the differential privacy protection parameter of parameter j on device i;

[0040] Step 4: Implementation Phase of the Model Fusion Acceleration Module

[0041] 1) On each local device, perform training of a portion of the model; this includes training the segmented local model for a certain number of rounds to generate locally updated weights;

[0042] 2) Introduce a locally updated model fusion algorithm, allowing each local device to complete part of the model training in its local environment; adopt a fast fusion algorithm to quickly fuse the locally updated weights into the global model;

[0043] 3) Use a distributed model fusion algorithm to ensure that local updates on multiple devices can be effectively merged into the global model; consider the asynchronous nature of communication to ensure the efficiency of distributed model fusion.

[0044] 4) Develop a fusion acceleration strategy to optimize the model fusion speed based on the performance of local devices, training progress, and model convergence.

[0045] 5) Adjust the parameter update weights of the global model to ensure that the contribution of each local device is fully reflected;

[0046] 6) Introduce a model fusion effect evaluation mechanism to monitor the performance improvement after each model fusion; adjust the model fusion strategy and parameters based on the evaluation results;

[0047] Step 5: Adaptive Learning Rate Adjustment Phase

[0048] 1) Monitor the training progress on each device in real time, including metrics such as the current training rounds, loss value, and model performance;

[0049] 2) Detect the model convergence on each device. By monitoring the loss value or other convergence metrics, determine whether the model has reached convergence; control the convergence state matrix C. ij , represents the convergence state of the model in the j-th round of training on device i;

[0050] 3) Evaluate the training performance on each device in real time, considering loss value, accuracy, or other metrics. Define the training performance evaluation matrix E. ij , representing the training effect of the j-th round of training on device i;

[0051] 4) Develop a learning rate adjustment strategy and dynamically adjust the learning rate based on the equipment training progress and model convergence; control the learning rate adjustment matrix;

[0052] 5) Use adaptive learning rate adjustment algorithms, such as Adagrad and Adam, to dynamically adjust the learning rate based on historical gradient information and training progress;

[0053] 6) Consider the different characteristics of heterogeneous devices, such as computing power and storage capacity, and adjust the learning rate to adapt to the training needs of different devices.

Claims

1. A distributed Transformer training acceleration framework for federated learning is proposed to address the issues of training speed and efficiency. The method includes at least the following steps: S1: Dynamic Model Segmentation Module Based on the data distribution on the local device, the dynamic model segmentation module enables each device to process only the Transformer model part related to its data, thereby reducing computational complexity; S2: Intelligent Task Scheduling Module This module evaluates the performance of each device by decomposing the global training task into subtasks, setting task priorities, and dynamically adjusting task allocation methods to optimize the utilization of computing resources. The intelligent task scheduling module comprehensively considers device performance evaluation, task priorities, data characteristics, and adaptability to heterogeneous devices to improve the overall system efficiency. S3: Layered Communication Optimization Module By employing asynchronous communication and differential privacy technology, only the changed parts of the model parameters or key information are transmitted, thereby reducing the amount of data transmitted between devices and ensuring the security and privacy of communication. S4: Model Fusion Acceleration Module This module allows local devices to quickly integrate local updates into the global model after completing partial model training, thereby reducing training time and improving model performance; S5: Adaptive Learning Rate Adjustment Module Based on the training progress and model convergence, ensure faster convergence speed and higher training efficiency during the training process.

2. The implementation of the dynamic model segmentation module in the distributed Transformer training acceleration framework in federated learning according to claim 1 further includes at least the following steps: 2.1: Determining the segmentation method for the dynamic model: First, the module obtains the distribution characteristics and importance of the local data through statistical analysis; second, it determines the key data characteristics based on the data distribution; finally, by comprehensively considering the data characteristics, importance, and distribution, the dynamic model segmentation module determines the dynamic segmentation method for each local device, thereby ensuring that each device only processes the Transformer model part related to its data. 2.2: Implementation method of dynamic model segmentation: Dynamic model segmentation is achieved by adjusting the weights of the Transformer model or removing model parts that are irrelevant to the local data; the formula is expressed as: TMi=f(Di) Among them, TM i It is the dynamically segmented Transformer model on each device i, D i This refers to the data distribution on the device; dynamic model segmentation is achieved by adjusting the TM. i Weighting or removing irrelevant parts ensures that each device focuses on the model relevant to its data.

3. The detailed implementation of the intelligent task scheduling module in the distributed Transformer training acceleration framework in federated learning according to claim 1, at least includes... Includes the following steps: 3.1: Performance evaluation and calculation of task allocation weights: 1) The module collects performance metrics for each device, such as computing speed, storage capacity, and bandwidth; 2) By comprehensively considering performance indicators, calculate the performance evaluation P for each device. i : P i =α·ComputeSpeed i +β·StorageCapacity i +Y·Bandwidth i Among them, ComputeSpeed i StorageCapacity represents the computing speed of the i-th device. i It refers to storage capacity, Bandwidth. i α is bandwidth, and β and γ are weighting parameters used to balance the importance of different performance metrics. 3) Through equipment performance evaluation P i and task priority i Calculate the task allocation weight W i : W i =g(P i ,Priority i ) Among them, W i is the task assignment weight for device i, and g(·) is a function for calculating the task assignment weight; 3.2: Considerations for adaptability to heterogeneous equipment: Heterogeneous device adaptability index, taking into account the heterogeneity between devices, such as different types of processors, storage devices and network connections; The impact of data characteristics is considered, and task allocation weights are adjusted according to the local data characteristics on each device to adapt to different data processing requirements.

4. The detailed implementation of the hierarchical communication optimization module in the distributed Transformer training acceleration framework in federated learning according to claim 1, at least... Includes the following steps: 4.1: Applications of asynchronous communication and differential privacy technologies: Asynchronous communication: Introducing an asynchronous communication mechanism allows devices to complete partial tasks at different times, thereby reducing the synchronization waiting time of communication; the formula is: Data i =h(AC) i Asynchronous communication characteristics (AC) i Used to determine the amount of data transmitted in communications; Differential privacy technology: To ensure the privacy of communications, differential privacy technology is introduced. By processing communication data, it reduces the risk of leakage of individual data privacy. 4.2: Minimizing data transfer volume: To reduce the amount of communication data transmitted, the module only transmits the changed parts of the model parameters or key information, rather than the entire model; the module analyzes the differences between the global model and the local model, and only transmits the changed parts of the model parameters to reduce the redundancy of communication data; through differential privacy technology, noise is processed on the transmitted model parameters to ensure the privacy of communication. 4.3: Security and Privacy Protection: To ensure the security and privacy of communication, the module employs encryption technology and access control policies to restrict access permissions for data transmission; communication data is encrypted using encryption algorithms to prevent unauthorized access; the module defines access control policies that only allow communication between specific devices, thus limiting the participants in the communication.

5. The distributed Transformer training acceleration framework in federated learning according to claim 1, wherein the model fusion acceleration module allows each local device to complete part of the model training in its local environment, further comprising at least the following steps: 5.1: Application of fast fusion algorithms: Define a fast fusion algorithm that quickly merges the updates of the local model into the global model based on the update amount of the local model and the current state of the global model. The formula is: Where GlobalModel is the global model, and LocalTrain is the local model. i It is the local training result of local device i; 5.2: Trade-off between fusion speed and model accuracy: The hyperparameters of the fusion algorithm are dynamically adjusted to adjust the fusion speed based on the accuracy of the global model and the training speed. The formula is: Speed-Accuracy_Tradeoff=φ·Speed+(1-φ)·Accuracy Here, φ is a trade-off parameter used to balance fusion speed and model accuracy.

6. The distributed Transformer training acceleration framework in federated learning according to claim 1, in order to achieve adaptive learning rate adjustment, the module dynamically monitors the training progress and model convergence on each device, and further includes at least the following steps: 6.1 Calculation of dynamically adjusted learning rate: Calculate the adaptive learning rate (LR) for each device. i The formula is: LR i =k(Progress i ) Progress i Let k represent the training speed on device i, and k(·) be the dynamic learning rate adjustment function. 6.2 Global model learning rate synchronization: After each training round, the adaptive learning rate on each device is applied to the global model; by coordinating the learning rate on the global model, the contributions of each device are effectively integrated globally.