Decentralized Machine Learning Method Based on Model Pruning and Topology Construction

By adaptively allocating pruning ratios and dynamic construction topology in edge computing environments, the limited resource, heterogeneity and non-IID data problems in decentralized machine learning are solved, efficient model training and communication are achieved, and training speed and resource utilization are improved.

CN115766465BActive Publication Date: 2025-07-25SUZHOU INST FOR ADVANCED STUDY USTC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211412314.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-11
Publication Date
2025-07-25
Estimated Expiration
2042-11-11

AI Technical Summary

Technical Problem

In the edge computing environment, decentralized machine learning faces challenges of limited resources, heterogeneity of systems, network dynamics and non-IID data, and existing methods have failed to effectively solve the problems of computing efficiency and communication overhead.

Method used

Through the joint optimization of model pruning and topology construction, each node is adaptively assigned a pruning ratio and a dynamic construction network topology, allowing each node to train and transmit sub-models suitable for its capabilities, reducing computational and communication overhead, and utilizing high-speed links to alleviate gradient dispersion.

Benefits of technology

The training process of decentralized machine learning is accelerated, computing and communication efficiency is improved, network blocking and poor capabilities are avoided as bottlenecks, and model convergence speed is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115766465B_ABST
    Figure CN115766465B_ABST
Patent Text Reader

Abstract

The present invention discloses a decentralized machine learning method based on model pruning and topology construction. In each training round, the method includes: the coordinator obtains the network topology and the pruning ratio of each node according to the status information of each node and sends them to each node; each node performs structured model pruning on the local model according to the received pruning ratio, trains the pruned sub-model using the local dataset, and updates the corresponding model parameters; each node communicates with the corresponding neighbor nodes according to the network topology and restores the model structure, performs parameter aggregation on the restored model to obtain the latest local model, and then starts the next training round. The present invention fully solves the key challenges of limited resources, system heterogeneity, network dynamics, and non-IID data by using the joint optimization of model pruning and topology construction, thereby accelerating the training process of decentralized machine learning in the edge computing environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of Distributed Machine Learning, and in particular, to a decentralized machine learning method based on model pruning and topology construction. Background Art

[0002] With the rapid development of the Internet of Things and mobile computing, a vast amount of data is generated by intelligent devices at the network edge globally. The traditional centralized machine learning paradigm needs to transmit data to the central cloud for model training, which will incur a large amount of communication overhead and pose a risk of privacy leakage. The edge computing paradigm directly stores and processes data at the network edge close to the data source without uploading the data to the cloud, effectively addressing the above challenges. With the help of edge computing, distributed machine learning effectively utilizes data from edge devices to train models and develop intelligent applications. In the distributed machine learning under the PS architecture, all nodes need to communicate with the central server, which is likely to cause network congestion and increase the risk of single-point failure. On the contrary, in decentralized machine learning, each node collaborates to train the model in a P2P form and only communicates with its neighbor nodes, which can effectively avoid the bottleneck of the central server and greatly improve scalability.

[0003] However, deploying decentralized machine learning in the edge computing scenario still faces some key challenges: limited resources, system heterogeneity, network dynamics, and non-IID data. (1) Limited resources: Different from the resource-rich central cloud, the computing and communication resources of edge nodes are very limited, and it is difficult to train deep neural networks with a large number of parameters. (2) System heterogeneity: The scope of edge nodes may include general gateways and dedicated base stations, and there is a large gap in their computing and communication capabilities. Nodes with poor capabilities may become the bottleneck of model training. (3) Network dynamics: Edge nodes are usually distributed in different geographical regions. Due to factors such as background noise and bandwidth competition, the wireless communication conditions between nodes are dynamically changing. (4) Non-IID data: Edge nodes collect data according to the local environment and user preferences, and there may be a huge gap in the data distribution between different nodes, and none of them can represent the global data distribution, which will affect the convergence speed of model training.

[0004] Existing solutions aim to address the challenges of communication resource limitations, network dynamics, and non-IID data, but there are still two limitations. On the one hand, existing methods mainly reduce communication overhead but ignore computational efficiency. The computing power of edge nodes is very limited, and training a large deep neural network on resource-constrained edge nodes is very slow or even infeasible. On the other hand, existing methods assume that the local models of each node (even after compression) have the same structure and size, which ignores the heterogeneity of edge nodes. Nodes with poor capabilities will hinder model aggregation and slow down the training process. Summary of the Invention

[0005] To solve the existing technology, the present invention provides a decentralized machine learning method based on model pruning and topology construction, which fully addresses the key challenges of limited resources, system heterogeneity, network dynamics, and non-IID data by using the joint optimization of model pruning and topology construction, thereby accelerating the training process of decentralized machine learning in an edge computing environment.

[0006] The present invention provides a decentralized machine learning method based on model pruning and topology construction. In each training round, it includes:

[0007] S1. The coordinator obtains the network topology and the pruning ratio of each node according to the status information of each node and sends them to each node;

[0008] S2. Each node performs structured model pruning on the local model according to the received pruning ratio, and uses the local dataset to train the pruned sub-model and updates the corresponding model parameters;

[0009] S3. Each node communicates with the corresponding neighbor nodes according to the network topology and restores the model structure, performs parameter aggregation on the restored model to obtain the latest local model, and then starts the next training round.

[0010] Optionally, the status information of each node includes: the link speed and computing power of each node.

[0011] Optionally, the network topology in S1 is dynamically constructed according to time-varying network conditions and non-IID data distributions.

[0012] Optionally, when determining the pruning ratio of each node, by fixing the network topology, considering the time resource constraint and the pruning ratio range constraint, the optimization goal is to minimize the convergence bound of model training, and the optimization problem is solved by linear programming.

[0013] Optionally, when determining the network topology construction, by fixing the pruning ratio of each node, calculating the consensus speed of each communication link, and removing the links with a consensus speed lower than a certain threshold from the fully connected topology.

[0014] Optionally, the consensus speed of each link is equal to the consensus distance between the nodes at both ends of the link divided by the communication time.

[0015] Optionally, the S2 specifically includes:

[0016] Each node performs structured model pruning on the local model according to the received pruning ratio, removes a corresponding proportion of the model parameters, and obtains a sub-model that matches its own capabilities and the indexes of the retained parameters.

[0017] Each node trains the pruned sub-model on the local dataset using stochastic gradient descent and updates the corresponding model parameters.

[0018] Optionally, each node performs structured model pruning on the local model according to the received pruning ratio, removes a corresponding proportion of the model parameters, and obtains a sub-model corresponding to the original local model, including:

[0019] Use the L1 norm to evaluate the importance of each structure in the local model of the model, and remove the model structures with low importance scores lower than a certain threshold and the corresponding feature maps according to the pruning ratio to obtain a sub-model corresponding to the original local model.

[0020] Optionally, a binary mask is used to record the indexes of the retained parameters.

[0021] Optionally, the S3 includes:

[0022] Use a machine learning algorithm to calculate the weights of the models of each neighbor node;

[0023] Perform parameter aggregation on the restored model according to the weights of the models of each neighbor node to obtain the latest local model.

[0024] Based on the decentralized machine learning scenario, the present invention mainly aims to fully address the challenges of limited resources, system heterogeneity, network dynamics, and non-IID data in the edge environment through the joint optimization of model pruning and topology construction, and accelerate the training process. This method is different from previous methods, mainly reflected in: adaptive model pruning can reduce the computational and communication overhead of training at the same time, match the heterogeneous capabilities of nodes, and the dynamically constructed network topology can effectively cope with time-varying network conditions and non-IID data distributions. The joint optimization of the two realizes the balance between resource overhead and training performance.

[0025] Compared with the solutions in the prior art, the advantages of the present invention are:

[0026] 1. Based on the decentralized machine learning scenario, each node only communicates with its own neighbor nodes instead of a central server, avoiding network congestion and improving scalability;

[0027] 2. The method of the present invention utilizes model pruning technology to simultaneously reduce the computational and communication overheads of decentralized machine learning and accelerate the training process.

[0028] 3. In the method of the present invention, the pruning ratio of adaptive decision-making enables each node to train and transmit a pruned model suitable for heterogeneous capabilities, solves the straggler problem, and prevents nodes with poor capabilities from becoming the bottleneck of model training.

[0029] 4. The method of the present invention dynamically constructs the network topology through the consensus speed, can make full use of high-speed links, and reduces the weight dispersion caused by non-IID data, thereby improving the convergence speed.

[0030] The present invention discloses an acceleration method for decentralized machine learning in a heterogeneous edge computing scenario. The method determines the pruning ratio and network topology through a coordinator, and distributes the optimization scheme to each node. Then, each node performs structured model pruning according to different pruning ratios, and trains and transmits sub-models adapted to heterogeneous capabilities on the dynamically constructed network topology, thereby improving the computational and communication efficiency and accelerating the model convergence. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a flowchart of a decentralized machine learning method based on model pruning and topology construction provided by an embodiment of the present invention;

[0032] Figure 2 is an effect diagram of model pruning and topology construction provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that only parts related to the present invention are shown in the drawings for the sake of convenience of description, rather than all the structures. Embodiment

[0034] Figure 1 is a flowchart of a decentralized machine learning method based on model pruning and topology construction provided by an embodiment of the present invention, which specifically includes the following steps:

[0035] S1. The coordinator obtains the network topology and the pruning ratio of each node according to the status information of each node and sends them to each node.

[0036] The above S1 is the information collection and decision-making stage, which specifically includes:

[0037] S11. At the beginning of each training round, all nodes report their status information (link speed and computing power) to the coordinator;

[0038] S12. According to the status information of this round, the coordinator solves the optimization problem to obtain the network topology and the pruning ratio of each node, and sends the obtained strategy to each node.

[0039] Optionally, each node can use tools such as Dstat to collect the current link speed and computing power in each round, and report them to the coordinator for decisions on the pruning ratio and network topology.

[0040] In this embodiment, the pruning ratio of each node is different to adapt to the heterogeneous capabilities of each node. Nodes with stronger capabilities are assigned smaller pruning ratios, so as to train and transmit larger sub-models; nodes with weaker capabilities are assigned larger pruning ratios, so as to train and transmit smaller sub-models. The adaptive pruning ratio enables each node to train and transmit sub-models suitable for its capabilities, reducing the impact of stragglers on the training efficiency.

[0041] Among them, the network topology solved by the coordinator is dynamically constructed according to time-varying network conditions and non-IID (non-identically distributed) data distributions, which can make full use of high-speed communication links and reduce gradient dispersion.

[0042] In this embodiment, when determining the network topology and the pruning ratio of each node, considering the tight coupling between the pruning ratio and the network topology, a joint optimization algorithm is designed for pruning ratio decision-making and topology construction, fixing one decision and optimizing the other decision, and iteratively solving the pruning ratio and the network topology to achieve a balance between resource overhead and training performance.

[0043] Among them, when formulating the joint optimization problem in S12, time resource constraints are mainly considered: the total time of model training cannot exceed the time budget; network topology constraints: the network topology of each round should form a connected graph; pruning ratio constraints: the pruning ratio range of each node should satisfy being greater than or equal to 0 and less than 1; the goal of the joint optimization problem is to minimize the convergence bound of model training.

[0044] The preferred technical solution is: in the above S12, the coordinator decouples the long-term optimization problem into multiple sub-problems, and uses the joint optimization algorithm to solve the sub-problems of the current round in each training round to make online decisions.

[0045] Specifically, when determining the pruning ratio of each node, by fixing the network topology, considering time resource constraints and pruning ratio range constraints, the optimization goal is to minimize the convergence bound of model training, and the optimization problem is solved by linear programming.

[0046] When determining the network topology construction, by fixing the pruning ratio of each node, calculating the consensus speed of each communication link, removing the links with the consensus speed lower than a certain threshold from the fully connected topology, and updating the network topology.

[0047] The consensus speed of each of the above - mentioned links is equal to the consensus distance between the nodes at both ends of the link divided by the communication time. The consensus distance reflects the magnitude of the data distribution difference between two nodes. The greater the data distribution difference, the greater the consensus distance. The communication time reflects the communication efficiency of the link. The faster the link speed, the shorter the communication time. Retaining links with a large consensus distance and / or a short communication time is beneficial for reducing the weight dispersion between models and making full use of high - speed links to improve the training efficiency.

[0048] Furthermore, when calculating the consensus distance between nodes, if the two nodes did not communicate in the previous round and the latest model weights of the other party cannot be obtained, historical weights are used for calculation.

[0049] S2. Each node performs structured model pruning on the local model according to the received pruning ratio and trains the pruned sub - model using the local dataset to update the corresponding model parameters.

[0050] The above - mentioned S2 is the local training stage of the node, which specifically includes:

[0051] S21. After each node receives the pruning ratio solved in S12, it performs structured model pruning on the local model, removes the corresponding proportion of model parameters, and obtains a sub - model that matches its own capabilities and the indexes of the retained parameters.

[0052] S22. Each node trains the pruned sub - model on the local dataset using stochastic gradient descent to update the corresponding model parameters.

[0053] When performing structured pruning on the local model, L1 - norm is used to evaluate the importance of each structure in the model. According to the pruning ratio, model structures with lower importance scores (such as filters or neurons) and their related feature maps are removed to obtain a sub - model of the original model. Through this pruning method, it is possible to retain the important model structures in the original model to the greatest extent, and local training with the pruned model can effectively reduce the computational overhead compared with training the original model.

[0054] Furthermore, in this embodiment, a binary mask is used to record the indexes of the retained parameters, and this mask helps the node identify the corresponding positions of the removed and retained parameters.

[0055] Exemplarily, the local data trained in this embodiment can be image segmentation data, image recognition data, etc. During the training process, by selecting the target client and determining the corresponding compression ratio for the target client, the data processing efficiency can be greatly improved.

[0056] S3. Each node communicates with its corresponding neighbor nodes according to the network topology, restores the model structure, performs parameter aggregation on the restored model to obtain the latest local model, and then starts the next training round. For the effect diagrams of model pruning and topology construction during the training process, see Figure 2 .

[0057] The above S3 is the model exchange and aggregation stage, specifically including:

[0058] S31. Each node communicates with its neighbor nodes according to the network topology solved in S12, and exchanges the trained sub-models and the indexes of the reserved parameters in S22 with each other;

[0059] S32. After each node receives the models of all its neighbor nodes, it restores the model structure according to the corresponding indexes.

[0060] S33. Perform parameter aggregation on the restored model to obtain the latest local model, and then start the next training round.

[0061] Among them, in this embodiment, the neighbor nodes exchange the pruned models after training, which can effectively reduce the communication overhead compared with exchanging the original models.

[0062] When aggregating the models of each neighbor, a machine learning algorithm is used to calculate the weights of the models of each neighbor node, and parameter aggregation is performed on the restored model according to the weights of the models of each neighbor node to obtain the latest local model. Exemplarily, the above machine learning model can be the Metropolis-Hasting algorithm.

[0063] In addition, to restore the model structure according to the corresponding indexes, each node first uses the corresponding indexes to restore the model structure of the neighbor model, and then uses the parameters of the local model to fill in the removed parameters of the neighbor model to complete the model restoration.

[0064] The technical solution of this embodiment, by assigning different pruning ratios to heterogeneous edge nodes, each node can train and transmit sub-models suitable for its capabilities after performing model pruning, reducing the computational and communication overheads, and preventing nodes with poor capabilities from becoming training bottlenecks. In addition, this method dynamically constructs the network topology considering time-varying network conditions and non-IID data distributions, making full use of high-speed links and reducing gradient dispersion. Considering the tight coupling between the pruning ratio and the network topology, a joint optimization algorithm is designed for pruning ratio decision-making and topology construction to achieve a balance between resource overheads and training performance.

[0065] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, it can also include more other equivalent embodiments, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A decentralized machine learning method based on model pruning and topology construction, characterized in that In each training round, it includes: S1. The coordinator solves the optimization problem according to the status information of each node to obtain the network topology and the pruning ratio of each node, and sends them to each node; Formalize the time resource constraint of the optimization problem: The total time for model training cannot exceed the time budget; Network topology constraint: The network topology of each round should form a connected graph; Pruning ratio constraint: The pruning ratio range of each node should satisfy being greater than or equal to 0 and less than 1; The goal of the joint optimization problem is to minimize the convergence bound of model training; When determining the pruning ratio of each node, by fixing the network topology, considering the time resource constraint and the pruning ratio range constraint, the pruning ratio of each node is optimally solved through linear programming; When determining the network topology construction, by fixing the pruning ratio of each node, calculate the consensus speed of each communication link, remove the links with the consensus speed lower than a certain threshold from the fully connected topology, and update the network topology; S2. Each node performs structured model pruning on the local model according to the received pruning ratio, and uses the local dataset to train the pruned sub-model, and updates the corresponding model parameters; S3. Each node communicates with the corresponding neighbor nodes according to the network topology and restores the model structure, performs parameter aggregation on the restored model to obtain the latest local model, and then starts the next training round.

2. The method according to claim 1, wherein The status information of each node includes: the link speed and computing power of each node.

3. The method according to claim 1, wherein The network topology in S1 is dynamically constructed according to the time-varying network conditions and non-IID data distribution.

4. The method according to claim 1, wherein The consensus speed of each link is equal to the consensus distance between the two end nodes of the link divided by the communication time.

5. The method according to claim 1, wherein S2 specifically includes: Each node performs structured model pruning on the local model according to the received pruning ratio, removes the corresponding proportion of model parameters, and obtains a sub-model that matches its own capabilities and the index of the retained parameters; Each node uses stochastic gradient descent to train the pruned sub-model on the local dataset and updates the corresponding model parameters.

6. The method according to claim 5, wherein Each node performs structured model pruning on the local model according to the received pruning ratio, removes the corresponding proportion of model parameters, and obtains a sub-model corresponding to the original local model, including: Use the L1 norm to evaluate the importance of each structure in the local model of the model, and remove the model structures with low importance scores lower than a certain threshold and the associated feature maps according to the pruning ratio to obtain a sub-model corresponding to the original local model.

7. The method according to claim 5, wherein Use a binary mask to record the index of the retained parameters.

8. The method according to claim 5, characterized in that In S3, a machine learning algorithm is used to calculate the weights of the models of each neighbor node, and parameter aggregation is performed on the restored model according to the weights of the models of each neighbor node to obtain the latest local model.

Citation Information

Patent Citations

  • Industrial Internet of Things distributed federal learning method based on Markov chain consensus

    CN114139688A

  • Personalized collaborative learning method and device based on neural network model pruning

    CN114418085A