Federal learning method based on computing power perception

By using a three-layer tree architecture and a computing power-aware model, the problems of heterogeneous computing power and communication bottlenecks in federated learning are solved, achieving balanced computing load and improved model convergence speed, thereby enhancing the efficiency and stability of federated learning.

CN120930830APending Publication Date: 2025-11-11INST OF WAR STUDIES ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511114051.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Traditional federated learning frameworks have problems such as heterogeneous computing power, communication bottlenecks, and inefficient synchronization mechanisms. Existing solutions such as model compression and asynchronous federated learning still have limitations and fail to fully consider the heterogeneity of device computing power and model oscillations.

Method used

A three-layer tree architecture is adopted to build a computing power awareness model to cluster clients and form clusters with different computing power. Devices are synchronously trained and aggregated within the same cluster, and asynchronously aggregated globally between clusters. The computing power awareness model is used to optimize computing load balancing and communication efficiency.

Benefits of technology

Optimize computational load balancing, reduce bandwidth bottlenecks, improve training efficiency and model convergence speed, enhance system robustness and global model stability, and adapt to complex computing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930830A_ABST
    Figure CN120930830A_ABST
Patent Text Reader

Abstract

The invention relates to the field of federated learning, and discloses a federated learning method based on computing power perception, which is used for solving the problems of heterogeneous computing power, communication bottleneck and low efficiency of a synchronization mechanism during federated learning, and comprises the following steps: constructing a three-layer architecture of computing power perception according to a three-layer tree architecture, and constructing a computing power perception model; clients are clustered according to the computing power perception model, clusters with different computing power are obtained, all devices synchronously execute local training tasks in the same cluster, aggregation is carried out after a certain round, aggregated model parameters are obtained, asynchronous global aggregation between the clusters is carried out according to the aggregated model parameters of each cluster, and the clustering efficiency is improved. And the classification accuracy, the loss curve and the convergence rate of the final global model are calculated, and performance comparison and analysis are performed, so that the heterogeneity of the equipment computing power is effectively reduced, and the problem of model oscillation caused by an asynchronous strategy is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of federated learning, and more specifically to a federated learning method based on computing power awareness. Background Technology

[0002] With the development of edge computing, more and more terminal devices are using federated learning for data processing. Federated learning is a decentralized machine learning method that can train models on multiple devices (such as edge computing devices, drone swarms, etc.) without centralized data storage, thereby improving privacy protection and communication efficiency.

[0003] However, traditional federated learning frameworks suffer from problems such as heterogeneous computing power, communication bottlenecks, and inefficient synchronization mechanisms. Existing solutions, such as model compression, hierarchical federated learning, and asynchronous federated learning, while alleviating these problems to some extent, still have limitations, such as reduced model accuracy, insufficient consideration of heterogeneous device computing power, and model oscillations caused by asynchronous strategies.

[0004] To address the above problems, this invention proposes a solution. Summary of the Invention

[0005] In order to overcome the above-mentioned defects of the prior art, the present invention provides a computing power-aware federated learning method to solve the problems existing in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A computationally-aware federated learning method includes the following steps:

[0008] Step 1: Based on the three-layer tree architecture, construct a three-layer architecture for computing power awareness, which includes a client, a cluster intermediate server, and a global server;

[0009] Step 2: Construct a computing power awareness model, and cluster clients based on the computing power awareness model to obtain clusters with different computing power.

[0010] Step 3: Within the same cluster, all devices synchronously execute local training tasks and aggregate them after a certain number of rounds to obtain the aggregated model parameters;

[0011] Step 4: Perform asynchronous global aggregation between clusters based on the aggregated model parameters of each cluster;

[0012] Step 5: Calculate the classification accuracy, loss curve, and convergence speed of the final global model, and conduct performance comparison and analysis.

[0013] Preferably, the step of constructing the computing power awareness model is as follows:

[0014] Collect the device computing power parameters of each client, including device computing frequency, device memory bandwidth, and device cache size, and upload the computing power parameters to the cluster intermediate server for recording;

[0015] The computing power is defined as the time required for the device to perform one local iteration, and the specific method for obtaining it is as follows:

[0016]

[0017] In the formula, C i f is the computing power of the i-th device. i Calculate the frequency for the i-th device;

[0018] The size of the transmitted data packets and the communication bandwidth between devices are obtained. Based on the data packet size and the communication bandwidth between devices, the communication latency between devices is estimated. The specific method for obtaining these parameters is as follows:

[0019]

[0020] In the formula, d i,j Let S be the communication delay between device i and device j, S be the data packet size, and B be the distance between the two devices. i,j The communication bandwidth between device i and device j;

[0021] Set computing power weight parameters and communication latency weight parameters, and calculate the computing power distance between devices based on device computing power and communication latency. The specific method for obtaining this information is as follows:

[0022] D i,j =α|C i -C j |+βd i,j ;

[0023] In the formula, D i,j Let be the computing power distance between device i and device j, α represent the computing power weight parameter, β represent the communication delay weight parameter, and |C i -C j | represents the absolute difference in computing power between device i and device j, d i,j The communication delay between device i and device j.

[0024] Preferably, the step of clustering clients based on the computing power awareness model is as follows:

[0025] The computing power distance between all devices is calculated based on the computing power perception model, and a weighted adjacency matrix is ​​constructed based on the computing power distance between all devices.

[0026] A bottom-up approach is adopted to gradually merge devices with similar computing power to form a cluster;

[0027] The clustering objective function is defined to minimize the difference in computing power distance within the cluster. The specific expression is as follows:

[0028]

[0029] In the formula, C k D represents the set of devices in cluster k, where K is the total number of clusters. i,j The computing distance between device i and device j;

[0030] Regularly update the computing power and communication status of the devices, and readjust the cluster partitioning based on the latest computing power distance;

[0031] Output the cluster number to which each device belongs, and record the computing power balance within the cluster.

[0032] Preferably, the step of constructing the weighted adjacency matrix based on the computing power distance between all devices is as follows:

[0033] Obtain the number of devices, and construct an adjacency matrix based on the computing power distance between devices and the number of devices, specifically:

[0034]

[0035] In the formula, A is the adjacency matrix, N is the number of devices, and D is the computing distance between devices.

[0036] Preferably, the step of all devices synchronously executing local training tasks and then aggregating them after a certain number of rounds is as follows:

[0037] Initialize the model parameters, assigning the same initial model parameters to each device in the cluster;

[0038] Each device trains a model based on local data, calculates the loss function, and updates the model parameters using gradient descent;

[0039] Set a termination round and a threshold. When the training reaches a fixed number of rounds or the rate of change of the objective function is lower than the threshold, stop training, output the model parameters at this time as the latest parameters, and perform synchronous aggregation between devices.

[0040] Preferably, the step of performing synchronous aggregation between devices is as follows:

[0041] Obtain the latest parameters and upload them to the cluster server;

[0042] The weighting factor for the device is calculated based on its computing power. The specific steps are as follows:

[0043]

[0044] In the formula, σ i C is the weighting factor for device i. iThe computing power of device i;

[0045] The aggregated model parameters are calculated by weighting the model parameters of each device with the weighting factor. The cluster server updates the aggregated model parameters and performs the next round of training.

[0046] Preferably, the step of performing asynchronous global aggregation between clusters based on the aggregated model parameters of each cluster is as follows:

[0047] Obtain the model parameters, training rounds, and data volume ratio for each cluster after aggregation. The data volume ratio is the proportion of the total number of non-participating devices in the cluster to the number of samples.

[0048] An asynchronous upload mechanism is used for submission, and the global server records the training rounds for each cluster.

[0049] The model weights are adjusted using a time decay factor, the expression of which is:

[0050]

[0051] In the formula, λ k Let γ represent the model weights of cluster k, γ be the time decay coefficient, and T represent the current global server training epoch. k Indicates the most recent training time of cluster k;

[0052] The final aggregate weights are obtained based on the model weights and the number of data samples in the cluster.

[0053] The weighted aggregation model of all clusters is calculated in the global server and used as the new global model. The new global model is then sent back to all clusters for the next round of training.

[0054] If the rate of change of the global model's loss falls below a threshold or the maximum number of training epochs is reached, training stops, the server stores the final global model, and notifies all clusters to stop training.

[0055] Preferably, the step of obtaining the final aggregate weight based on the model weight and the number of data samples in the cluster is as follows:

[0056] To obtain the number of data samples used during cluster training and the total number of all clusters, calculate the data volume weight for each cluster. The specific steps are as follows:

[0057]

[0058] In the formula, μ k N represents the weight of the data volume in cluster k. k Here, k represents the number of data samples used during training of cluster k, and K is the total number of clusters.

[0059] The final aggregate weight is obtained based on the model weight and the data volume weight. The specific steps are as follows:

[0060] JH k =λ k μ k ;

[0061] In the formula, JH k Let λ be the final aggregation weight for cluster k. k Let μ be the model weight for cluster k. k The data volume weight for cluster k.

[0062] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:

[0063] 1. Based on a three-layer tree architecture, a computationally-aware three-layer architecture is constructed to optimize computational load balancing, reduce bandwidth bottlenecks, and improve the training efficiency of federated learning. By dividing the architecture into a client layer, a cluster intermediate server layer, and a global server layer, direct communication between devices and the global server can be reduced, lowering data transmission pressure. Simultaneously, the cluster server enables local aggregation of devices with similar computing power, accelerating model updates. This approach fully utilizes the computing power of high-performance devices while preventing low-performance devices from slowing down the overall training progress, thereby improving system robustness and the convergence speed of the global model.

[0064] 2. Constructing a computing power-aware model allows for client clustering, resulting in clusters with varying computing power. This effectively improves the computational efficiency and communication stability of federated learning. By assigning devices with similar computing power to the same cluster, the imbalance of computing resources is reduced, preventing devices with lower computing power from slowing down the overall training progress. Simultaneously, devices within the cluster can perform local synchronous training, reducing bandwidth overhead from cross-cluster communication and improving training stability. This not only optimizes computational load distribution but also improves the convergence speed of the global model, making federated learning more efficient and adaptable to complex computing environments.

[0065] 3. Within the same cluster, all devices synchronously execute local training tasks and aggregate them after a certain number of rounds to obtain aggregated model parameters. This effectively improves training efficiency and reduces the impact of uneven computational load. Through synchronous training, devices within the cluster complete local model updates within the same number of rounds and aggregate them after a set number of training rounds, making the cluster-level model more stable and reducing the impact of individual device computing power differences on the overall training progress. Furthermore, cluster aggregation reduces the data transmission burden on the global server, lowers communication overhead, and improves model training consistency, providing higher-quality model parameters for global aggregation, thereby accelerating the overall convergence of federated learning.

[0066] 4. Asynchronous global aggregation of model parameters between clusters, based on the aggregated model parameters of each cluster, can effectively alleviate the problem of asynchronous training progress caused by differences in computing power and communication capabilities among different clusters, and improve the update efficiency of the global model. Through the asynchronous mechanism, each cluster does not need to wait for all clusters to complete training synchronously before uploading; instead, it independently submits the aggregated model parameters according to its own progress, reducing training blockage caused by low-computing-power clusters or high communication latency. Simultaneously, the global server can dynamically adjust the impact of each cluster on the global model based on the time decay factor and data contribution weight, ensuring that the cluster with the latest data and the largest amount of data contributes more, thereby improving the convergence stability of the model and optimizing the overall performance of federated learning. Attached Figure Description

[0067] Figure 1 This is the overall flowchart of the present invention. Detailed Implementation

[0068] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. In addition, the forms of the various structures described in the following embodiments are merely illustrative. The computing power-aware federated learning method involved in the present invention is not limited to the structures described in the following embodiments. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0069] This invention provides a computationally conscious federated learning method, such as... Figure 1 As shown, it includes the following steps:

[0070] Step 1: Based on the three-layer tree architecture, construct a computing power-aware three-layer architecture to optimize computing load balancing and reduce bandwidth bottlenecks;

[0071] The three-tiered tree architecture consists of a client layer, a cluster middleware server layer, and a global server layer to optimize computational load balancing and reduce bandwidth bottlenecks. The device layer handles local data training, the cluster layer performs dynamic device clustering based on computing power awareness and executes model aggregation, and the global server layer receives model updates from each cluster and performs global aggregation. This architecture reduces bandwidth consumption by minimizing direct communication between devices and the global server, while simultaneously balancing computational load and improving model convergence speed and federated learning efficiency.

[0072] In this embodiment, it should be specifically explained that the steps for constructing the three-layer architecture of computing power awareness are as follows:

[0073] Build a client to perform local model training and provide computing resources;

[0074] Build a cluster intermediate server to receive model updates from clients within the cluster, perform synchronous aggregation, and reduce the impact of heterogeneity;

[0075] Build a global server to collect models from each cluster, perform asynchronous aggregation, and update the global model.

[0076] Step 2: Construct a computing power awareness model, and cluster clients based on the computing power awareness model to obtain clusters with different computing power, thereby reducing the heterogeneity of computing power within the cluster;

[0077] Building a computational power-aware model and clustering clients based on it can effectively improve the computational efficiency and communication stability of federated learning. By grouping devices with similar computing power into the same cluster, computational heterogeneity within the cluster can be reduced, preventing weaker devices from slowing down the overall training progress and optimizing the utilization of computing resources. Furthermore, this grouping method can reduce communication latency between devices with different computing power, improve model synchronization efficiency, and reduce the computational burden on the global server, thereby improving the convergence speed and overall performance of federated learning.

[0078] In this embodiment, it should be specifically explained that the steps for constructing the computing power awareness model are as follows:

[0079] Collect computing capability parameters for each client device, including device computing frequency, device memory bandwidth, and device cache size, and upload these parameters to the cluster intermediate server for recording.

[0080] The computing power is defined as the time required for the device to perform one local iteration, and the specific method for obtaining it is as follows:

[0081]

[0082] In the formula, C i C represents the computing power of the i-th device, which is the reciprocal of the device's computing capacity. The higher the computing frequency, the higher C becomes. i The smaller the value, the faster the device can complete computational tasks. i The frequency is calculated for the i-th device, representing the number of floating-point calculations performed per second. Devices with lower computing power have higher computing power weights, thus taking into account the balance of computing power when allocating tasks.

[0083] The data packet size and inter-device communication bandwidth are obtained. Based on the data packet size and inter-device communication bandwidth, the inter-device communication latency is estimated. The specific method for obtaining these parameters is as follows:

[0084]

[0085] In the formula, d i,j The communication delay between device i and device j reflects the time required for data to be transmitted from device i to device j. S is the data packet size, indicating the amount of data transmitted. i,jThe communication bandwidth between device i and device j is denoted as . The lower the communication latency between devices, the faster the data transmission speed and the smaller the parameter synchronization latency during training.

[0086] Set computing power weight parameters and communication latency weight parameters, and calculate the computing power distance between devices based on device computing power and communication latency. The specific method for obtaining this information is as follows:

[0087] D i,j =α|C i -C j |+βd i,j ;

[0088] In the formula, D i,j Let be the computing power distance between device i and device j, representing the combined difference in computing power and communication latency between device i and device j. Let α represent the computing power weight parameter, and β represent the communication latency weight parameter. |C i -C j | represents the absolute difference in computing power between device i and device j, d i,j The communication delay between device i and device j.

[0089] In this embodiment, it should be specifically explained that the steps for clustering clients based on the computing power awareness model are as follows:

[0090] The computing power distance between all devices is calculated based on the computing power perception model, and a weighted adjacency matrix is ​​constructed based on the computing power distance between all devices to represent the computing power and communication relationship between devices.

[0091] A bottom-up approach is adopted to gradually merge devices with similar computing power to form a cluster;

[0092] The clustering objective function is defined to minimize the difference in computing power distance within the cluster. The specific expression is as follows:

[0093]

[0094] In the formula, C k D represents the set of devices in cluster k, where K is the total number of clusters. i,j The computing distance between device i and device j;

[0095] Regularly update the computing power and communication status of the devices, readjust the cluster partitioning based on the latest computing power distance, and use the sliding window method to calculate the recent average computing power to prevent short-term fluctuations from affecting the stability of the cluster.

[0096] Output the cluster number to which each device belongs and record the computing power balance within the cluster to optimize the allocation of training tasks.

[0097] Bottom-up clustering is a hierarchical clustering strategy primarily used to gradually merge devices with similar computing power into clusters. In power-aware federated learning, this method first treats each device as a separate cluster, then gradually merges the two closest devices or sub-clusters based on the computing power distance between devices (a weighted distance of computing power and communication latency) until a preset number of clusters or a computing power balance standard is reached. This method can adaptively adjust device grouping, making the computing power of devices within the same cluster more similar, thereby optimizing the allocation of computing resources, reducing synchronization blocking problems during training, and improving the overall computational efficiency of federated learning.

[0098] The sliding window method is a dynamic computation strategy that smooths out data fluctuations. In power-aware cluster partitioning, it is used to calculate the average computing power of recent periods, preventing short-term computing power fluctuations from affecting system stability. This method records the computing power values ​​of devices over multiple recent training epochs and calculates their average value to reduce the interference of abnormal computing loads in individual epochs on cluster partitioning. By smoothing computing power data through the sliding window, the system can group devices more stably and avoid frequent cluster adjustments.

[0099] In this embodiment, it should be specifically explained that the step of constructing the weighted adjacency matrix based on the computing power distance between all devices is as follows:

[0100] Obtain the number of devices, and construct an adjacency matrix based on the computing power distance between devices and the number of devices, specifically:

[0101]

[0102] In the formula, A is the adjacency matrix and N is the number of devices.

[0103] Step 3: Within the same cluster, all devices synchronously execute local training tasks and aggregate them after a certain number of rounds to obtain the aggregated model parameters, in order to ensure the stability of model convergence.

[0104] In this embodiment, it should be specifically explained that the step of all devices synchronously executing local training tasks and then performing aggregation after a certain number of rounds is as follows:

[0105] Initialize the model parameters, assigning the same initial model parameters to each device in the cluster;

[0106] Each device trains a model based on local data, calculates the loss function, and updates the model parameters using gradient descent;

[0107] A loss function is a mathematical function used to measure the difference between a model's predictions and the true values. In federated learning, the device calculates the loss function based on local data to evaluate the current model's performance. A larger loss function value indicates a greater model error, requiring further optimization; a smaller value indicates that the model is closer to the true values.

[0108] Gradient descent is an optimization algorithm used to minimize a loss function, gradually converging the model parameters to their optimal values. In federated learning, each device calculates the gradient of the model parameters—the rate of change of the loss function with respect to the parameters—based on the calculated loss function, and updates the parameters in the opposite direction of the gradient to reduce error.

[0109] Set a termination round and a threshold. When training reaches a fixed number of rounds or the rate of change of the objective function is lower than the threshold, stop training, output the model parameters at this point as the latest parameters, and perform synchronous aggregation between devices, including the rate of change of the objective function and the rate of change of the loss function.

[0110] In this embodiment, it should be specifically explained that the steps for inter-device synchronization and aggregation are as follows:

[0111] Obtain the latest parameters and upload them to the cluster server;

[0112] The weighting factor for the device is calculated based on its computing power. The specific steps are as follows:

[0113]

[0114] In the formula, σ i C is the weighting factor for device i. i The computing power of device i;

[0115] The aggregated model parameters are calculated by weighting the model parameters of each device with the weighting factor. The cluster server updates the aggregated model parameters and performs the next round of training.

[0116] Step 4: Perform asynchronous global aggregation between clusters based on the aggregated model parameters of each cluster;

[0117] In this embodiment, it should be specifically explained that the asynchronous global aggregation between clusters based on the aggregated model parameters of each cluster is as follows:

[0118] Obtain the model parameters, training rounds, and quantity ratios for each cluster after aggregation. The data volume ratio is the proportion of the total number of non-participating devices in the cluster to the number of samples.

[0119] Because different clusters have different computing speeds and communication bandwidths, an asynchronous upload mechanism is adopted to ensure that model submissions are not affected by slow clusters. The global server records the training rounds of each cluster to ensure that the timeliness of the model is taken into account when performing aggregate calculations.

[0120] Because model parameters uploaded from different clusters may differ over time, models submitted earlier may become outdated. Therefore, a time decay factor is used to adjust model weights, the expression of which is:

[0121]

[0122] In the formula, λ k Let γ represent the model weights of cluster k, γ be the time decay coefficient used to control the degree of influence of time decay, and T represent the current global training epoch, i.e., the current training epoch of the global server. k Indicates the most recent training time of cluster k, and represents the timestamp of the last commit of the cluster. When TT k When it increases, λ k Approaching 0 reduces the influence of the old model;

[0123] The final aggregate weights are obtained based on the model weights and the number of data samples in the cluster.

[0124] The weighted aggregation model of all clusters is calculated in the global server and used as the new global model. The new global model is then sent back to all clusters for the next round of training.

[0125] If the rate of change of the global model's loss falls below a threshold or the maximum number of training epochs is reached, training stops, the server stores the final global model, and notifies all clusters to stop training.

[0126] Asynchronous upload is a method to optimize communication efficiency in federated learning. Addressing the differences in computational speed and communication bandwidth among different clusters, it allows each cluster to independently submit its model after local training, without waiting for all clusters to synchronize and complete training. This avoids low-computing-power or high-latency clusters slowing down the overall progress, enabling the global server to receive and aggregate new models more quickly, thus improving training efficiency. Furthermore, asynchronous upload reduces network congestion, improves bandwidth utilization, and allows high-computing-power clusters to update more promptly, contributing to faster global model convergence and enhancing the robustness and stability of the federated learning system.

[0127] The time decay factor is a weight adjustment strategy used to reduce the impact of outdated model updates. In federated learning, because models are uploaded at different times from different devices or clusters, earlier submitted models may be relatively old, and directly using them may affect the accuracy of the global model. Therefore, a time decay factor is introduced to attenuate the contribution of a model based on the difference between its submission time and the current time, giving higher weights to recently updated models and gradually decreasing the weights of earlier models. This method effectively reduces the interference of outdated models on global aggregation, improving the real-time performance and convergence stability of the global model.

[0128] In this embodiment, it should be specifically explained that the step of obtaining the final aggregate weight based on the model weight and the number of data samples in the cluster is as follows:

[0129] To obtain the number of data samples used during cluster training and the total number of all clusters, calculate the data volume weight for each cluster. The specific steps are as follows:

[0130]

[0131] In the formula, μ k N represents the weight of the data volume in cluster k. k The number of data samples used during training of cluster k is K, where K is the total number of clusters. The larger the data volume of a cluster, the higher the weight of the data volume.

[0132] The final aggregate weight is obtained based on the model weight and the data volume weight. The specific steps are as follows:

[0133] JH k =λ k ·μ k ;

[0134] In the formula, JH k Let λ be the final aggregation weight for cluster k. k Let μ be the model weight for cluster k. k The data volume weight for cluster k.

[0135] Step 5: Calculate the classification accuracy, loss curve, and convergence speed of the final global model, and perform performance comparison and analysis. The loss curve and classification accuracy are used to verify whether the federated learning algorithm can effectively guarantee the convergence of the model, and the convergence speed is used to verify whether the algorithm is resource-efficient.

[0136] Calculating and analyzing the classification accuracy, loss curve, and convergence speed of the final global model helps to comprehensively evaluate the performance of the federated learning system. Classification accuracy reflects the model's predictive ability on test data, the loss curve shows the trend of model error changes during training, and the convergence speed measures the number of training epochs required for the model to reach a stable state. These metrics allow for intuitive comparison of the effects of different training strategies, identification of optimization space, and improvement of the model's generalization ability. Furthermore, comparative analysis with other federated learning methods helps to verify the effectiveness of optimization strategies and provides a theoretical basis for further improvements.

[0137] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0138] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A federated learning method based on computing power awareness, characterized in that, Includes the following steps: Step 1: Based on the three-layer tree architecture, construct a three-layer architecture for computing power awareness, which includes a client, a cluster intermediate server, and a global server; Step 2: Construct a computing power awareness model, and cluster clients based on the computing power awareness model to obtain clusters with different computing power. Step 3: Within the same cluster, all devices synchronously execute local training tasks and aggregate them after a certain number of rounds to obtain the aggregated model parameters; Step 4: Perform asynchronous global aggregation between clusters based on the aggregated model parameters of each cluster; Step 5: Calculate the classification accuracy, loss curve, and convergence speed of the final global model, and conduct performance comparison and analysis.

2. The federated learning method based on computing power awareness according to claim 1, characterized in that: The steps for constructing the computing power awareness model are as follows: Collect the device computing power parameters of each client, including device computing frequency, device memory bandwidth, and device cache size, and upload the computing power parameters to the cluster intermediate server for recording; The computing power is defined as the time required for the device to perform one local iteration, and the specific method for obtaining it is as follows: In the formula, C i f is the computing power of the i-th device. i Calculate the frequency for the i-th device; The size of the transmitted data packets and the communication bandwidth between devices are obtained. Based on the data packet size and the communication bandwidth between devices, the communication latency between devices is estimated. The specific method for obtaining these parameters is as follows: In the formula, d i,j Let S be the communication delay between device i and device j, S be the data packet size, and B be the distance between the two devices. i,j The communication bandwidth between device i and device j; Set computing power weight parameters and communication latency weight parameters, and calculate the computing power distance between devices based on device computing power and communication latency. The specific method for obtaining this information is as follows: D i,j =α|C i -C j |+βd i,j ; In the formula, D i,j Let be the computing power distance between device i and device j, α represent the computing power weight parameter, β represent the communication delay weight parameter, and |C i -C j | represents the absolute difference in computing power between device i and device j, d i,j The communication delay between device i and device j.

3. The federated learning method based on computing power awareness according to claim 1, characterized in that: The steps for clustering clients based on the computing power awareness model are as follows: The computing power distance between all devices is calculated based on the computing power perception model, and a weighted adjacency matrix is ​​constructed based on the computing power distance between all devices. A bottom-up approach is adopted to gradually merge devices with similar computing power to form a cluster; The clustering objective function is defined to minimize the difference in computing power distance within the cluster. The specific expression is as follows: In the formula, C k D represents the set of devices in cluster k, where K is the total number of clusters. i,j The computing distance between device i and device j; Regularly update the computing power and communication status of the devices, and readjust the cluster partitioning based on the latest computing power distance; Output the cluster number to which each device belongs, and record the computing power balance within the cluster.

4. The federated learning method based on computing power awareness according to claim 3, characterized in that: The steps for constructing the weighted adjacency matrix based on the computing power distance between all devices are as follows: Obtain the number of devices, and construct an adjacency matrix based on the computing power distance between devices and the number of devices, specifically: In the formula, A is the adjacency matrix, N is the number of devices, and D is the computing distance between devices.

5. The federated learning method based on computing power awareness according to claim 1, characterized in that: The process of all devices synchronously executing local training tasks and then aggregating them after a certain number of rounds is as follows: Initialize the model parameters, assigning the same initial model parameters to each device in the cluster; Each device trains a model based on local data, calculates the loss function, and updates the model parameters using gradient descent; Set a termination round and a threshold. When the training reaches a fixed number of rounds or the rate of change of the objective function is lower than the threshold, stop training, output the model parameters at this time as the latest parameters, and perform synchronous aggregation between devices.

6. The federated learning method based on computing power awareness according to claim 5, characterized in that: The steps for performing synchronous aggregation between devices are as follows: Obtain the latest parameters and upload them to the cluster server; The weighting factor for the device is calculated based on its computing power. The specific steps are as follows: In the formula, σ i C is the weighting factor for device i. i The computing power of device i; The aggregated model parameters are calculated by weighting the model parameters of each device with the weighting factor. The cluster server updates the aggregated model parameters and performs the next round of training.

7. The federated learning method based on computing power awareness according to claim 1, characterized in that: The steps for asynchronous global aggregation between clusters based on the aggregated model parameters of each cluster are as follows: Obtain the model parameters, training rounds, and data volume ratio for each cluster after aggregation. The data volume ratio is the proportion of the total number of non-participating devices in the cluster to the number of samples. An asynchronous upload mechanism is used for submission, and the global server records the training rounds for each cluster. The model weights are adjusted using a time decay factor, the expression of which is: In the formula, λ k Let γ represent the model weights of cluster k, γ be the time decay coefficient, and T represent the current global server training epoch. k Indicates the most recent training time of cluster k; The final aggregate weights are obtained based on the model weights and the number of data samples in the cluster. The weighted aggregation model of all clusters is calculated in the global server and used as the new global model. The new global model is then sent back to all clusters for the next round of training. If the rate of change of the global model's loss falls below a threshold or the maximum number of training epochs is reached, training stops, the server stores the final global model, and notifies all clusters to stop training.

8. The federated learning method based on computing power awareness according to claim 7, characterized in that: The step of obtaining the final aggregate weight based on the model weight and the number of data samples in the cluster is as follows: To obtain the number of data samples used during cluster training and the total number of all clusters, calculate the data volume weight for each cluster. The specific steps are as follows: In the formula, μ k N represents the weight of the data volume in cluster k. k Here, k represents the number of data samples used during training of cluster k, and K is the total number of clusters. The final aggregate weight is obtained based on the model weight and the data volume weight. The specific steps are as follows: JH k =λ k ·m k ; In the formula, JH k Let λ be the final aggregation weight for cluster k. k Let μ be the model weight for cluster k. k The data volume weight for cluster k.