SYSTEM AND METHOD FOR TRAINING MACHINE LEARNING MODELS IN A DISTRIBUTED SYSTEM - Patent application
By compressing updates using dense arrays of parameter deltas in federated learning systems, the method addresses the challenge of efficient data transmission in distributed systems, enhancing training consistency and efficiency, especially in resource-constrained networks.
Patent Information
- Application Number
- JP2024027641
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-09-21
- Filing Date
- 2024-02-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-02-27
AI Technical Summary
Existing federated learning systems face challenges in efficiently transmitting large amounts of data between nodes, particularly in scenarios where IoT devices have limited network quality, leading to inconsistent training processes.
The proposed method compresses the size of updates shared between nodes in a distributed system by using dense arrays of parameter deltas, which are arranged based on the magnitude of corresponding parameters in the reference model, thereby reducing the overall data transmission during the training phase.
This approach allows for more frequent and efficient updates between nodes, improving the consistency and efficiency of the training process, even in environments with limited network resources.
Smart Images

Figure 0007675879000016 
Figure 0007675879000017 
Figure 0007675879000018
Abstract
Description
[Technical field]
[0001] The present disclosure relates to methods and systems for training machine learning models in a distributed system. In particular, but not by way of limitation, the present disclosure relates to methods for efficiently performing federated learning by reducing the size of updates shared among devices in the system. [Background technology]
[0002] Machine learning (ML) methods aim to train models based on observed data. In traditional ML approaches, raw data collected by edge devices (such as in an Internet of Things (IoT) network) is communicated back to a central server to train a global model.
[0003] Federated Learning (FL) and Distributed Learning (DL) are decentralized ML frameworks that aim to parallelize the training process by using multiple connected computing devices simultaneously to train a single model. An edge device is a piece of hardware in a network (e.g. in IoT) that provides an entry point to the network and can constantly collect raw data. In some scenarios, IoT devices may be limited in terms of network quality. In such cases, communication needs to be further restricted to allow for consistent training.
[0004] Deep learning is a subset of ML, where large datasets are used to train ML models in the form of neural networks (NNs). A neural network is a connected system of functions whose structure is inspired by the human brain. It has multiple interconnected nodes, and each connection can transmit data like a signal transmitted through a synapse. The connections between the nodes have weights, which are parameters that are optimized, resulting in training the model. Summary of the Invention
[0005] The configuration of the embodiments will be more fully understood and appreciated from the following detailed description, given by way of example only, and taken in conjunction with the drawings in which: [Brief description of the drawings]
[0006] [Figure 1] FIG. 1 illustrates a system architecture for implementing federated learning in one configuration. [Diagram 2] FIG. 1 shows a flowchart detailing a federated learning method with full local model updates from workers and full global updates from the server. [Diagram 3] FIG. 1 shows a flowchart detailing a method for federated learning with reduced update size, according to one implementation. [Figure 4] FIG. 4 shows a flow chart detailing the training update step of the method of FIG. 3. [Figure 5A] FIG. 13 shows graphs of different performance metrics for different parameter pruning methods. [Figure 5B] FIG. 13 shows graphs of different performance metrics for different parameter pruning methods. [Figure 6A] FIG. 13 shows graphs of different performance metrics for different parameter pruning methods. [Figure 6B] FIG. 13 shows graphs of different performance metrics for different parameter pruning methods. [Figure 7A] FIG. 13 shows graphs of different performance metrics for different parameter pruning methods. [Figure 7B] FIG. 13 shows graphs of different performance metrics for different parameter pruning methods. [Figure 8A] A graph showing different performance metrics for different federated machine learning methods. [Figure 8B] A graph showing different performance metrics for different federated machine learning methods. [Figure 9A]A graph showing different performance metrics for different federated machine learning methods. [Figure 9B] A graph showing different performance metrics for different federated machine learning methods. [Figure 10] FIG. 1 illustrates a computing device for practicing the methods described herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0007] Implementations described herein provide improvements to federated learning by compressing the size of each update sent between nodes in a distributed system. This can result in a reduction in the overall amount of data sent during the training phase of a distributed system. Particular implementations adjust the size of each update based on the quality of service of the communication link over which the updates are sent. This makes efficient use of network capacity, for example, by allowing larger updates to be sent in response to improved quality of service and reducing the size of updates in response to reduced quality of service.
[0008] According to one aspect of the present disclosure, a computer-implemented method for training a machine learning model in a distributed system is provided. The distributed system includes a plurality of nodes that exchange updates to jointly train the machine learning model. Each node of the plurality of nodes maintains a local version of the machine learning model. The local version of the machine learning model of each of the plurality of nodes is initialized with the same respective one or more parameter values. The method is performed on a node of the plurality of nodes, the method comprising: receiving a first update to a local model from at least one other node in the distributed system; the local model comprising a local version of the machine learning model, the first update comprising a dense array of one or more first parameter deltas, the one or more first parameter deltas arranged in the dense array in an order determined by a reference model, each first parameter delta representing a difference between a parameter of the local model and a corresponding parameter of an updated version of the machine learning model maintained by the at least other node; updating the local model based on the received first update and the reference model to determine an updated local model; determining one or more second parameter deltas, each second parameter delta representing a difference between a parameter of the updated local model and a corresponding parameter of a previous version of the local model; and transmitting the second update to the at least one other node in the distributed system, wherein the second update comprises a dense array of one or more second parameter deltas, the one or more second parameter deltas arranged in the dense array in an order determined by the reference model.
[0009] By receiving an update comprising a dense array of one or more first parameter deltas from at least one other node and transmitting an update comprising a dense array of one or more second parameter deltas to at least one other node, the one or more first parameter deltas and the one or more second parameter deltas are arranged in the respective dense arrays in an order determined by the reference model, and may not need to include one or more complete parameters and / or one or more corresponding parameter identifiers. This may allow the size of each update to be reduced. This may also allow the number of first parameter deltas and second parameter deltas included in each update to be increased without increasing the size of each update.
[0010] The one or more first parameter deltas may be arranged in a dense array according to the magnitude of the corresponding one or more parameters of the reference model. The one or more second parameter deltas may be arranged in a dense array according to the magnitude of the corresponding one or more parameters of the reference model.
[0011] The multiple nodes may include multiple workers and a server. Each of the multiple workers may be configured to train a respective local model, and the multiple workers report a first update to the local model back to the server. The server may be configured to aggregate the first updates from the workers to update the global model. The server may be configured to report the updates to the global model back to the workers. The server may be configured to aggregate the first updates from the workers based on a reference model.
[0012] The reference model may comprise a copy of a previous version of the global model, for example, the reference model may comprise a copy of a version of the global model that precedes an updated version of the global model at the server.
[0013] The node may be a worker, and the method may comprise receiving a corresponding one of the first updates to the local model from a server. The one or more first parameter deltas may indicate a current state of all updated parameters of the global model. Updating the local model may comprise applying the corresponding one of the first updates to the local model to make the local model compliant with the global model.
[0014] The node may be a worker, and the method may comprise determining a level of reduction in the number of parameters of the local model. The one or more parameters of the local model to be removed or pruned may be selected randomly. Thus, for example, pre-training of the local model and / or training data may not be required to select the one or more parameters to be removed or pruned.
[0015] The method may include applying the determined level of reduction in the number of parameters to the local model to generate a reduced local model. The method may include applying the determined level of reduction in the number of parameters of the local model to a reference model. This may enable the reference model to act as a mask, for example, when the local model is updated based on updates received at at least one other node and / or during aggregation of updates from workers.
[0016] The method may include training the reduced local model based on the training data to obtain an updated local model, The step of training the reduced local model may be part of or comprised in the step of updating the local model.
[0017] The method may comprise sending a corresponding one of the first updates to a server for use in updating the global model.
[0018] Determining a level of reduction in the number of parameters of the local model may comprise determining a quality of service of a communication link between the worker and the server. Determining a level of reduction in the number of parameters of the local model may comprise determining a level of reduction in the number of parameters of the local model based on the quality of service.
[0019] The method may further comprise adjusting the level of reduction in the number of parameters of the local model based on a quality of service. Adjusting the level of reduction in the number of parameters of the local model based on the quality of service may comprise decreasing the level of reduction in the number of parameters of the local model in response to an increase in the quality of service. Adjusting the level of reduction in the number of parameters of the local model based on the quality of service may comprise increasing the level of reduction in the number of parameters of the local model in response to a decrease in the quality of service.
[0020] For example, when the level of reduction in the number of parameters of the local model is decreased, applying the determined level of reduction in the number of parameters to the local model may comprise including one or more additional parameters in the local model. The one or more additional parameters to be included in the local model may be determined based on the reference model.
[0021] The local model may comprise multiple layers. Applying the determined level of reduction in the number of parameters to the local model may comprise, for example, distributing the reduced number of parameters across multiple layers, such that each of the multiple layers comprises the same or equal number of parameters. Thus, the methods disclosed herein may enable layer-wise pruning, which may be understood as a global pruning level, based, for example, on the determined level of reduction in the number of parameters of the local model.
[0022] Additionally, by distributing a reduced number of parameters across multiple layers, such that each of the multiple layers has the same or equal number of parameters, one or more smaller layers of the local model may be less affected by the global pruning level than one or more relatively larger layers of the local model.
[0023] Applying the determined level of reduction in the number of parameters to the local model when at least one layer of the multiple layers of the local model is full may further include excluding the at least one layer from further distributing the reduced number of parameters. Applying the determined level of reduction in the number of parameters to the local model when at least one layer of the multiple layers of the local model is full may further include, for example, distributing the reduced number of parameters across one or more remaining layers of the multiple layers of the local model such that each of the one or more remaining layers of the multiple layers of the local model comprises the same or equal number of parameters.
[0024] The reference model may comprise multiple layers. Applying the determined level of reduction in the number of parameters to the reference model may comprise, for example, distributing the reduced number of parameters across multiple layers of the reference model, such that each of the multiple layers of the reference model comprises the same or equal number of parameters. One or more parameters of each of the multiple layers of the reference model may correspond to one or more parameters of each of the multiple layers of the local model.
[0025] Applying the determined level of reduction in the number of parameters to the reference model when at least one layer of the multiple layers of the reference model is full may further include excluding the at least one layer from further distributing the reduced number of parameters. Applying the determined level of reduction in the number of parameters to the reference model when at least one layer of the multiple layers of the reference model is full may further include distributing the reduced number of parameters across one or more remaining layers of the multiple layers of the reference model, e.g., such that each of the one or more remaining layers of the multiple layers of the reference model comprises the same or equal number of parameters.
[0026] The method may further comprise, for example, sending by the worker a full update of the local model representing a current state of all parameters of the local model when the quality of service exceeds an upper threshold. The method may further comprise, for example, omitting sending a full update by the worker when the quality of service is below a lower threshold.
[0027] The node may be a server. The local model maintained by the node may be a global model maintained locally by the server. Receiving a first update to the local model may comprise receiving a plurality of updates from a plurality of workers. Each update may comprise a dense array of one or more second parameter deltas. Updating the local model may comprise aggregating the updates from the plurality of workers to update the global model based on the reference model. A corresponding one of the updates to the global model may be transmitted by the server to each of the workers for use in updating the worker's respective local model.
[0028] The one or more second parameter deltas may indicate the current state of all updated parameters of the global model.
[0029] The local version of the machine learning model of each of the multiple nodes may be randomly initialized with the same parameter value or values of each. For example, in a first iteration of the method, updates received from at least one other node may be used to initialize the local model. In the first iteration, the updates received from the at least one other node may comprise, for example, a random seed for randomly initializing the local model.
[0030] According to a further aspect of the present disclosure, there is provided a node for use in a distributed system. The distributed system comprises a plurality of nodes exchanging updates to jointly train a machine learning model. Each node of the plurality of nodes maintains a local version of the machine learning model. The local version of the machine learning model of each of the plurality of nodes is initialized with the same respective one or more parameter values. The node comprises a storage configured to store a local model comprising a local version of the machine learning model; and a processor, the processor configured to: receive a first update to the local model from at least one other node in the distributed system, the first update comprising a dense array of one or more first parameter deltas, the one or more first parameter deltas arranged in the dense array in an order determined by the reference model, each first parameter delta representing a difference between a parameter of the local model and a corresponding parameter of an updated version of the machine learning model maintained by the at least other node; update the local model based on the received first update and the reference model to determine an updated local model; determine one or more second parameter deltas, each second parameter delta representing a difference between a parameter of the updated local model and a corresponding parameter of a previous version of the local model; and send the second update to the at least one other node in the distributed system, wherein the second update comprises a dense array of one or more second parameter deltas, the one or more second parameter deltas arranged in the dense array in an order determined by the reference model.
[0031] According to a further aspect of the present disclosure, a non-transitory computer-readable medium is provided that comprises computer-executable instructions that, when executed by a computer, configure the computer to operate as a node in a distributed system. The distributed system comprises a plurality of nodes that exchange updates to jointly train a machine learning model. Each node of the plurality of nodes maintains a local version of the machine learning model. The local version of each of the plurality of nodes is initialized with the same respective one or more parameter values. The computer-executable instructions cause a computer to receive a first update to the local model from at least one other node in the distributed system, the local model comprising a local version of the machine learning model, the first update comprising a dense array of one or more first parameter deltas, the one or more first parameter deltas arranged in the dense array in an order determined by the reference model, each first parameter delta representing a difference between a parameter of the local model and a corresponding parameter of the updated version of the machine learning model maintained by the at least other node, updating the local model based on the received first update and the reference model to determine an updated local model, determining one or more second parameter deltas, each second parameter delta representing a difference between a parameter of the updated local model and a corresponding parameter of a previous version of the local model, and sending the second update to the at least one other node in the distributed system, wherein the second update comprises a dense array of one or more second parameter deltas, the one or more second parameter deltas arranged in the dense array in an order determined by the reference model.
[0032] 1 shows a system architecture for implementing federated learning (FL) in one configuration. The system architecture includes a parameter server (PS) node 10 and multiple worker nodes 20. A global model 12 with one or more parameters is stored on the server 10. The worker nodes 20 contribute to training the global model 12.
[0033] In the following description, the parameter server nodes are referred to as servers and the worker nodes are referred to as workers.
[0034] In this architecture, workers 20 submit model updates to the server 10. These updates are aggregated at the server 10 to update the global model 12. The parameters of the updated global model are then propagated back to each worker 20.
[0035] More specifically, each worker 20 stores a local model 22, which is a locally maintained version of the global model 12. Each local model 22 is periodically updated to match the global model 12 based on updates sent from the server 10 to the worker 20.
[0036] Each worker 20 trains a local model 22 based on locally available training data 24. This training data 24 may be specific to each worker node 20 and may be stored in memory (e.g., from previous measurements) or received in real time (e.g., from sensor readings).
[0037] Each worker 20 may train a local model 22 by one or more updates of the local model 22 based on the training data. In general, training involves adjusting parameters (weights) of the local model 22, which may be implemented as a neural network (NN), to optimize a function (e.g., to reduce an error). Gradient Descent (GD), such as Stochastic Gradient Descent, is one optimization technique for learning the weights of a neural network (NN), although other methods for parameter optimization are available.
[0038] After training the local models, the workers 20 communicate the parameter updates to the server 10 , which aggregates the parameter updates across the workers 20 and applies them to update the global model 12 .
[0039] The parameter updates reported back to the server 10 may include updated parameters of the local models 22, which may then be used by the server 10 to determine updates to the global model 12. Aggregation of the updates may be in the form of taking an average (e.g., mean or median) over the updates, or by any other form of aggregation.
[0040] After the global model 12 has been updated, the updated parameters of the global model 12 may be communicated back to the workers 20 so that the workers 20 can update their respective local models 22 .
[0041] Updates may be exchanged after each step of training of the local models, or multiple training steps may be performed before the global model is updated. Updating the global model after each step of local model training may ensure that the local models do not diverge from each other. Updating after several optimization iterations may reduce the number of updates that need to be transferred.
[0042] Each iteration of updating the global model 12 includes sending updates from the workers 20 to the server 10 and sending parameters of the global model 12 back to the workers 20. The communication of large amounts of data over the network can be a major bottleneck in large-scale federated learning model training. In possible applications, devices may be limited in terms of allowed energy usage, communication bandwidth (BW), and other network resources. This may prevent participation in federated learning model training because the quality of service of the network is insufficient to withstand such a volume / speed of data transfer. It is a concern to keep these federated learning workers / nodes / agents / servers connected and ensure that they can still engage in the training process.
[0043] Note that in this implementation there is a separate server for the workers, but this server may also be a worker and therefore perform local training either directly to the global model or as an emulated worker that updates a separate local version and aggregates its own updates with updates from other workers.
[0044] This application proposes a novel adaptive federated learning model parameter compression method to reduce the overall amount of data transmitted during the training phase of a distributed system.
[0045] In the parameter compaction method described herein, parameters of the global model 12 and each local model 22 are initialized with the same respective values. For example, the global model 12 and each local model 22 may be randomly initialized with the same parameter values. Each update sent by a worker 20 includes a dense array of one or more parameter deltas. Each parameter delta represents the difference between a parameter of the updated local model 22 and a corresponding parameter of a previous version of the local model 22. The parameter delta sent by a worker may also be referred to as a second parameter delta.
[0046] A dense array of parameter deltas may be understood as an array of parameter deltas in which most parameter delta values are non-zero or not equal to zero. A dense array of parameter deltas may also be referred to as a non-sparse array of parameter deltas. The parameter deltas are arranged in a dense array in an order determined by the reference model. By including a dense array of parameter deltas in each update sent by the worker 20, sending the complete parameters may be avoided. Because each update sent by the worker 20 includes a dense array of parameter deltas in an order determined by the reference model, it may not be necessary to send an identifier of the parameter delta (e.g., an index of the associated parameter weight) that may enable the server to determine which of the respective parameters of the local model the parameter delta is associated with. Instead, the reference model is used to reconstruct the local model 22 of each worker 20 during aggregation.
[0047] Similar to compressing updates from workers 20 to the server 10, the same compression strategy can be implemented when sending local model updates from the server 10 to workers 20. Each update to each worker 20 includes a dense array of one or more parameter deltas. Each parameter delta represents the difference between a parameter of the local model 22 and a corresponding parameter of an updated version of the global model 12. The parameter delta indicates the current state of all updated parameters of the global model 12. The parameter delta sent by the server may also be referred to as the first parameter delta.
[0048] The parameter deltas are arranged in a dense array in an order determined by the reference model. The local model 22 is updated based on the received updates and the reference model. The updates are applied to the local model 22 to make the local model 22 conform to the global model 12. For example, the parameter deltas may be added to the corresponding parameters of the local model 22.
[0049] The reference model comprises a copy of a global model of a version earlier than the updated version of the global model in the server. In this embodiment, the dense array parameter deltas are ordered according to the magnitude of each corresponding parameter of the reference model. For example, the dense array parameter deltas can be ordered in descending order according to the magnitude of each corresponding parameter of the reference model. The term "magnitude" as used herein may be understood as the absolute value of the parameter, or the L1 criterion or measure.
[0050] The parameter compression method proposed herein may allow the size of each update transmitted between a worker and a server to be reduced. In addition, the number of parameter deltas to be included in each update may be increased without increasing the size of the update. Thus, updates may be transmitted more frequently and / or more efficiently between workers and servers.
[0051] It should be understood that in other implementations, the parameter deltas in each of the above-mentioned updates may be replaced with a corresponding parameter or parameters. For example, each update may comprise a dense array of one or more changed or updated parameters arranged in an order determined by the reference model. Any of the features described herein relating to parameter deltas may also apply to the corresponding parameters.
[0052] Figure 2 shows a flowchart detailing the federated learning method with full local model updates from the workers and full global updates from the server. The operation of a parameter server and a single worker node are shown, with dotted lines representing the network interface between the two. Solid arrows represent communication within a node (e.g., each server or worker), while dashed arrows represent communication across the network to other nodes.
[0053] The method begins with the server waiting for update requests from the workers (30). Once an update request is sent by a worker (40) and received by the server (32), the server then sends the complete global model to each worker (34).
[0054] Once the worker receives the complete global model (42), it replaces the local model stored in the worker with the global model and performs a training update of the global model (44). This training update adjusts the parameters according to an optimization method. In this example, gradient descent is used to update the global model parameters based on the training data available to the worker. Thus, the global model is locally updated (46) to generate new model parameters from training.
[0055] The worker then sends the locally updated complete global model as an update to the server (48). The worker then loops back to step 40 to send requests for further updates from the server.
[0056] When the server receives an update (36), it aggregates (38) the locally updated global model with other locally updated global models received from other workers, and updates the global model based on this aggregation (39). The server then loops back to step 30 to wait for further requests from the workers.
[0057] This approach can place a heavy burden on the network, as the server sends the global model to each worker, and each worker sends back the complete, locally updated global model to the server, regardless of network conditions. Therefore, as mentioned above, the approach proposed herein instead involves sending a dense array of parameter deltas to reduce the size of updates sent between the server and workers.
[0058] 3 shows a flowchart detailing a method for associative learning with reduced update size according to one embodiment, which is similar to the one shown in FIG.
[0059] Similar to Figure 2, the operation of the parameter server and a single worker node are shown, with dotted lines representing the network interface between the two. Solid arrows represent communication within a node (e.g., each server or worker), while dashed arrows represent communication across the network to other nodes.
[0060] The method begins with the server waiting for an update request from a worker (50). An update request is sent by the worker (60) and received by the server (52), and the server then sends the update to the worker (54).
[0061] In the first iteration of the method, the update may include, for example, a random seed to randomly initialize the worker's local version of the global model. The server sends this update to each worker, so that the global model and each worker's local version of the global model are initialized with the same respective parameter values. The local versions of the global model may also be referred to as the local model.
[0062] In subsequent iterations of the method, the updates include the parameter deltas. The updates may be considered to include only the parameter deltas associated with the parameters to be updated. The parameter deltas are provided to the workers as a dense array of parameter deltas. The order of the parameter deltas in the dense array is determined by the reference model, as described above.
[0063] In the first iteration of the method, when a worker receives an update (62), the worker initiates a training update of the local version of the global model stored in the worker (64), as described above.
[0064] In a subsequent iteration of the method, when the worker receives an update (62), the worker applies the update to the local model to make the local model conform to the global model. For example, the worker is configured to determine one or more updated parameters of the local model based on the parameter delta and the reference model. The worker may be configured to use the reference model to determine which parameter of the local model is associated with the received parameter delta, e.g., based on the magnitude of the corresponding parameter of the reference model. The worker may be configured to add the parameter delta to the associated parameter of the local model to determine an updated parameter of the worker's local model. The worker is configured to replace the corresponding parameter of the local model with the determined updated parameter, thereby making the local model conform to the global model. In this embodiment, the reference model may be used as a mask to reconstruct the local model stored in the worker. The mask may also be referred to as a deterministic mask. The worker performs training updates of the local model (64).
[0065] This training update adjusts the parameters according to an optimization method. In this example, a gradient descent method, such as stochastic gradient descent, is used to update the parameters of the local model based on the training data available to the worker. It should be understood that in other implementations, other methods for parameter optimization may be used. Thus, the local model is updated (66) to generate new model parameters from the training. The new model parameters include the parameters that changed during the training update (64).
[0066] The worker is configured to determine new parameter deltas, each new parameter delta representing the difference between a new parameter in the updated local model and a corresponding parameter in the previous version of the local model. The worker is configured to send (68) the new parameter deltas as updates to the server. The new parameter deltas are provided to the server as a dense array of new parameter deltas, as described above. The order of the new parameter deltas in the dense array is determined by the reference model. The worker then loops back to step 60 to send a request for further updates from the server.
[0067] When the server receives the update (56), the server maintains a copy of the current version of the global model stored on the server. The server may be configured to determine new model parameters based on the received new parameter delta and a reference model corresponding to the server's current version of the global model. The server may be configured to use the reference model to determine which parameters of the global model are associated with the received new parameter delta. The server may be configured to add the received new parameter delta to the associated parameters of the global model to determine new parameters of the global model. The server is then configured to aggregate (58) the new model parameters with other new model parameters determined based on the received new parameter deltas from other workers and the reference model. The server is configured to update (59) the global model based on the aggregation.
[0068] The dense arrays of parameter deltas received by the server from the workers may have different sizes. Thus, two or more sets of new parameters determined based on the received dense arrays of parameter deltas and the reference model may have different sizes. The server may be configured to aggregate one or more overlapping new parameters and / or one or more non-overlapping new parameters based on the number of workers associated with the received dense arrays of parameter deltas.
[0069] The server reconstructs the local model associated with each worker based on the reference model, e.g. a copy of the version of the global model that the server maintained in step 56, and the new parameters determined for each worker during aggregation step 58. As described above, the server may be configured to use the reference model as a mask, e.g. a deterministic mask, to reconstruct each worker's local model. The server then loops back to step 50 to wait for further requests from the workers.
[0070] Figure 4 shows a flow chart detailing the training update steps 64a-64e of the method shown in Figure 3. The training update steps may be performed for a number of iterations I. At the start of the method, the number of iterations I is set to zero.
[0071] The quality of service of the communication link between the server and the worker is monitored (64a). Quality of service may also be referred to as network quality. Quality of service may be defined in terms of a number of different metrics, such as bandwidth, signal-to-noise ratio, channel quality, received signal strength, error rate, network availability, etc. In this embodiment, the bandwidth BW of the communication link between the server and the worker is monitored and / or determined.
[0072] Active probing or passive monitoring techniques may be used to determine the bandwidth of a communication link between a server and a worker. Active probing may include sending one or more data packets having a known size from a worker to a server and measuring the duration between when the data packets are received at the server and when an acknowledgment of receipt from the server is received by the worker. The worker is configured to determine the bandwidth based on the measured duration and the data packet size. For example, the worker may be configured to determine the bandwidth by dividing the data packet size by the duration.
[0073] In passive monitoring, a worker may be configured to monitor ongoing communication with a server, which may include sending and / or receiving updates as described above. For example, the worker may be configured to determine a bandwidth based on data rates observed during this communication.
[0074] If the number of iterations I is equal to zero (64b), a level of reduction in the number of parameters of the local model stored in the worker is determined (64c). The level of reduction in the number of parameters of the local model may also be referred to as a pruning level of the local model. Pruning may comprise removing one or more parameters from the local model. The level of reduction in the number of parameters of the local model is determined based on the determined bandwidth. The level of reduction in the number of parameters of the local model may be adjusted based on the determined bandwidth, as described in more detail below.
[0075] The level of reduction in the number of parameters of the local model may be understood as a global pruning level of the local model. The global pruning level is based on the total number of parameters of the local model. The global pruning level may be set by a global pruning parameter k=1-C, where C is the data compression ratio. The data compression ratio may be defined as the ratio of the uncompressed data rate to the compressed data rate. The data compression ratio C may be determined as follows:
[0076]
number
[0077] In this embodiment, the data compression ratio C is determined using active probing. However, it should be understood that in other embodiments, the data compression ratio may be determined using passive monitoring, for example, as described above. The data compression ratio C may be determined based on the maximum data rate R measured in bits per second, the time σ allocated to the worker to send updates to the server, measured in seconds, and the size (θ) of the complete local model, measured in bits. The duration σ may be understood as a hyperparameter of the local model. This hyperparameter may be determined differently. For example, in one embodiment, the duration σ is determined based on the expected data rate and the allocated transmission cost for sending updates between the server and the worker. In this embodiment, the duration may be determined as follows:
[0078]
number
[0079] where E[R] is the expected data rate and C k is the quota sending cost per worker, measured in currency, and C bit is the cost per transmitted bit.
[0080] In another embodiment, the duration is the maximum allowable duration σ max In some implementations, there may be a large difference between the bandwidth of a communication link between a server and a worker and another bandwidth of another communication link between the server and another worker. The duration σ may be determined based on the maximum duration σ assigned to a worker associated with a communication link with a server that has a smaller bandwidth than the bandwidth of another communication link with the server associated with the other worker. max can be based on.
[0081] The compression ratio may be limited. For example, an upper threshold for the compression ratio may be determined. When the compression ratio C exceeds the upper threshold, the update includes a full update of the global model representing the current state of all parameters of the global model, or a full update of the local model representing the current state of all parameters of the local model. For example, in an embodiment where C>1, the update may include a full update. In such an embodiment, the full update may include an uncompressed copy of the global model or the local model.
[0082] A lower threshold for the compression ratio may be determined. For example, in an implementation where C→0, sending updates may be inefficient. If the compression ratio C is below the lower threshold, the worker or server may not send updates until the maximum data rate R increases. Thus, the worker or server omits sending updates. The lower threshold for the compression ratio C may be selected to be between 0.03 and 0.1, such as 0.05.
[0083] The determined level of reduction in the number of parameters is applied to the local models stored in the workers (64d). For example, the total number of parameters of the local models is reduced. The determined level of reduction in the number of parameters may also be applied to a reference model such that the reference model may be used as a mask when reconstructing the local and / or global models during aggregation.
[0084] The local model comprises multiple layers, and the reduced number of parameters is distributed across the layers of the local model such that each of the layers comprises the same or an equal number of parameters.
[0085] The reduced number of parameters may be distributed based on an order determined by the number of parameters of each of the layers. For example, the reduced number of parameters may be distributed in ascending order of the number of parameters of each layer of the local model. An equal portion of the reduced number of parameters may be assigned to each layer of the local model. The portion of the reduced number of parameters may be assigned first to the layer with the fewest number of parameters.
[0086] When at least one of the layers becomes full, the full layer is excluded from further distribution of the reduced number of parameters. For example, no further parameters are assigned to the full layer. For example, if the smallest layer cannot accommodate the assigned number of reduced parameters, the layer is considered full. Then, the assignment of the remaining reduced number of parameters to the remaining layers is determined. The remaining reduced number of parameters can then be distributed, for example, throughout the remaining layer or layers of the local model, such that the remaining layers have the same or equal number of parameters. A layer of the local model can be considered full when the number of parameters assigned to this layer is equal to the number of parameters of this layer. It should be understood that when one layer of the local model is considered full, the number of parameters assigned to each of the remaining layers is increased compared to the number of parameters of the full layer, for example, by the number of parameters that could not be assigned to the full layer.
[0087] The distribution of the number of reduced parameters can be described as follows: A is an ordered set of the number of parameters per layer of the local models stored in the worker, and is defined as follows:
[0088]
number
[0089] where I={1,2,...,N} is an ordered set of N indices, and a nrepresents a parameter.
[0090] The number of reduced parameters, Θ, is defined as:
[0091]
number
[0092] where k=(1-C) as above, k∈[0,1]. The number of reduced parameters Θ is rounded to the nearest integer. It can be seen that the number of reduced parameters Θ depends on the compression ratio C mentioned above, and therefore on the bandwidth of the communication link between the worker and the server. Thus, the reduction of parameters per layer depends on the global pruning parameter k, which in turn depends on the compression ratio C.
[0093] Parameter a n The number of reduced parameters for layer n, including
[0094]
number
[0095] is defined as follows:
[0096]
number
[0097] Where:
[0098]
number
[0099] is the parameter to be distributed up to layer n.
[0100]
number
[0101] is the previous parameter
[0102]
number
[0103] Equation 6 defines the sum of the parameters a n The number of reduced parameters for each layer n, including
[0104]
number
[0105] Define the variance of .
[0106]
number
[0107] It should be understood that when A is equal to A for layer n, then layer n is considered to be a full layer.
[0108] An exemplary process for distribution of a reduced number of parameters is shown in Algorithm 1. Algorithm 1: Require: A, k∈[0,1], Φ=0
[0109]
number
[0110] Θ is a constant representing a reduced number of parameters to be spread out. for n in{1,2,…,|A|}do
[0111]
number
[0112] end for
[0113] Additionally or alternatively, the reduced number of parameters can be distributed across layers of the local model by creating a map, which can be of the following form:
[0114]
number
[0115] Equation 7 can be thought of as mapping the number of parameters to a reduced number of parameters per layer of the local model.
[0116] The distribution of the reduced number of parameters is also applied to the reference model. For example, the reference model comprises multiple layers, and the reduced number of parameters are distributed across the layers of the reference model in the same way as across the layers of the local model described above. After the distribution of the reduced number of parameters of the reference model is determined, one or more parameters of the reference model corresponding to the parameters removed from the local model may also be removed from the reference model. Thus, one or more parameters of each of the layers of the reference model correspond to one or more parameters of each of the layers of the local model. This allows the reference model to be used as a reference or mask for reconstructing the global model stored on the server and the local version of the global model stored on the worker.
[0117] Pruning methods can be used to reduce the overall amount of data sent during the learning phase of a distributed system. For example, a global pruning method can be used to remove a percentage of parameters throughout the machine learning model. This results in fewer layers of the machine learning model being more affected by pruning, particularly at high pruning levels, such as removing 99% or more of the machine learning model's parameters. For example, a convolutional neural network, such as LeNet-5, can include a convolutional kernel and a fully connected neural network, each having different dimensions. If a high pruning level, such as 99.9%, is to be achieved, the convolutional kernel and the fully connected neural network are affected differently by the pruning level. Depending on the dimensions of the convolutional kernel and the fully connected neural network, the probability of setting all parameter values, for example, in a column of a fully connected neural network to zero may be significantly lower than the probability of setting all parameter values, for example, in a column of a convolutional kernel to zero.
[0118] Structured pruning may allow different pruning levels to be defined for different layers of a machine learning model. However, it may be difficult to adjust the pruning of layers to achieve a desired global pruning level. In addition, when the global pruning level changes, it may be necessary to adjust the pruning levels of different layers individually.
[0119] Other pruning methods based on an assessment of the importance of parameters, such as, for example, the absolute value of a parameter or the Euclidean norm of a substructure (e.g., a channel, filter, layer, or other substructure of a machine learning model), require the machine learning model to be trained, or at least partially trained, prior to pruning, which may require sending a full model update of the local model in the first iteration of the federated learning method, for example, as described above in connection with FIG.
[0120] The present application proposes a novel pruning method, which is part of the parameter compaction method described herein. The pruning method disclosed herein does not require pre-training of the machine learning model or training data, e.g., selecting one or more parameters to be kept in the machine learning model. Instead, the above-mentioned step of determining the level of reduction in the number of parameters (64c) and / or step of applying the determined level of reduction in the number of parameters (64d) can be performed before training the local model or during training of the local model. The parameters of the local model to be removed or pruned are selected randomly, which may allow a higher sparsity level to be achieved, for example, compared to pruning methods based on evaluation of the magnitude, importance, and / or other attributes of the parameters.
[0121] For example, the number of reduced parameters for layer n, defined in Eq.
[0122]
number
[0123] may be adaptively adjusted based on the determined bandwidth of the communication link between the server and the worker. Thus, the pruning method disclosed herein can generate layered pruning levels for varying global pruning levels, e.g., for varying bandwidth.
[0124] As described above, the reduced number of parameters is distributed across the layers such that each of the layers has the same or equal number of parameters, which can prevent one or more small layers of the local version of the global model from being more affected by the global pruning level than one or more relatively large layers of the local model, while still allowing the global pruning level to be achieved.
[0125] The pruning methods described herein may be used with any machine learning model and do not require a particular architecture or structure of the machine learning model.
[0126] The reduced local model is trained (64e) based on the training data available to the worker. The training adjusts one or more parameters of the reduced local model according to an optimization method. In this example, a gradient descent method, such as stochastic gradient descent, is used to update the reduced local model parameters based on the training data available to the worker. The one or more parameters of the reduced local model to be adjusted during training may be selected based on an attribute or measure, e.g., magnitude, of each parameter of the reduced local model. For example, one or more parameters of each layer of the reduced local model may be selected for training based on the magnitude of each parameter of the reduced local model. For example, only the parameter with the largest magnitude per layer of the reduced local model may be adjusted during training. It should be understood that in other implementations, another attribute or measure, e.g., L2 norm, may be used to select one or more parameters of the reduced local model to be trained.
[0127] The number of iterations I is the given maximum number of iterations I MAX If it is less than the maximum iteration count I, the worker loops back to step 64a. MAX If so, the worker proceeds to step 66 shown in Figure 3 and described above. The maximum number of iterations may be selected based on the desired accuracy of the local model and / or stopping parameters, for example when the accuracy of the local model does not improve.
[0128] If for any subsequent iteration, e.g., I>0, the determined bandwidth BW(I) for the current iteration is less than or equal to the determined bandwidth BW(I-1) for the previous iteration, the worker proceeds to step 64e. However, if for any subsequent iteration, e.g., I>0, the determined bandwidth BW(I) for the current iteration is greater than the determined bandwidth BW(I-1) for the previous iteration, the worker proceeds to step 64c.
[0129] As mentioned above, the level of reduction in the number of parameters of the local model may be adjusted based on the determined bandwidth. For example, in response to an increase in the bandwidth, the level of reduction in the number of parameters of the local model may then be reduced in step 64c. In another expression, the global pruning level may be reduced. Then, one or more additional parameters may be included in the local model. The additional parameters to be included in the local model may be determined based on the reference model. For example, one or more positions of the additional parameters to be included may be selected based on the reference model. In response to a decrease in the bandwidth, the level of reduction in the number of parameters of the local model may be increased. In this way, the level of reduction in the number of parameters may be adjusted based on the determined bandwidth. This may also be referred to as adaptive pruning of the local model. It should be understood that in this embodiment, the level of reduction in the number of parameters of the local model is adjusted based on the determined bandwidth, but in other embodiments, the level of reduction in the number of parameters may be adjusted based on another metric, such as a signal-to-noise ratio, a channel quality, a received signal strength, an error rate, network availability, etc.
[0130] 5A-7B show graphs of different performance metrics for different parameter pruning methods. Each of the figures shows a graph of the accuracy of different parameter pruning methods depending on the percentage of the number of remaining parameters of the pruned model. The experimental data obtained using the pruning method disclosed herein is represented by squares in the figures.
[0131] Experimental data obtained using the pruning method described in "Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science" by Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong Nguyen, Madeleine Gibescu, Antonio Liotta, Nature Communications, Vol. 9, Article No. 2383 (2018) are represented by crosses in Figures 5A-7B. This method is also called ERK. This method does not require any training of parameters prior to pruning.
[0132] Experimental data obtained using the pruning method described in Lee, J., Park, S., Mo, S., Ahn, S., Shin, J., "Layer-adaptive sparsity for the Magnitude-based Pruning" (2020), arXiv:2010.07611, are represented by triangles in Figures 5A to 7B. This method is also called LAMP. In this method, an initially trained model is gradually pruned by alternating training and the progressive removal of more parameters. This method can be considered as an iterative pruning method.
[0133] Both ERK and LAMP methods can be considered as size-based pruning methods. In these methods, the number of parameters to be removed or kept is determined solely based on the size of each parameter. For example, the number of parameters with the smallest size can be set to zero and not trained.
[0134] The pruning method was applied to a convolutional neural network model for image recognition, namely VGG-16, using the CIFAR-10 dataset. Figures 5A, 6A, and 7A show graphs of performance metrics of different pruning methods applied to the VGG-16 model. The pruning method was also applied to another convolutional neural network, namely EfficientNet-B0, using the CIFAR-10 dataset. Figures 5B, 6B, and 7B show graphs of performance metrics of different pruning methods applied to the EfficientNet-B0 model.
[0135] The experimental data shown in Figures 5A and 5B was obtained by first training different models to achieve high accuracy. Then, the number of parameters of each model was iteratively reduced as described in Lee, J., Park, S., Mo, S., Ahn, S., Shin, J., "Layer-adaptive sparsity for the Magnitude-based Pruning" (2020), arXiv:2010.07611. From Figures 5A and 5B, it can be seen that the pruning method disclosed herein determined a higher accuracy than the ERK pruning method. The accuracy determined for the pruning method disclosed herein is comparable to that of the LAMP pruning method.
[0136] The experimental data shown in Figures 6A and 6B were obtained by randomly removing 90% of the parameters of each model in the first iteration after initialization. The parameters to be removed were selected based on the respective magnitude of each parameter. The magnitude of each parameter is also sometimes referred to as the L1 scale. In this example, the parameters with the smallest magnitude were removed. Following pruning, each of the models was iteratively trained as described in Lee, J., Park, S., Mo, S., Ahn, S., Shin, J., "Layer-adaptive sparsity for the Magnitude-based Pruning" (2020), arXiv:2010.07611. The number of parameters of each model was iteratively reduced as described above in connection with Figures 5A and 5B. This scenario can be considered equivalent to initially setting a global pruning level of 90% in the first iteration, and then gradually limiting the bandwidth of the communication link between the server and the workers. It can be seen from FIGS. 6A and 6B that the accuracy determined for the pruning method disclosed herein is higher than the accuracy determined for both the ERK and LAMP methods.
[0137] The experimental data shown in Figures 7A and 7B was obtained by randomly initializing the parameters of the model and resetting the parameters after each training iteration, which may also be considered equivalent to "initialization pruning". This was repeated for different levels of parameter reduction, e.g., different pruning levels. It can be seen from Figures 7A and 7B that the accuracy determined for the pruning method disclosed herein is higher than the accuracy determined for both the ERK method and the LAMP method.
[0138] Figures 8A-9B show graphs of different performance metrics for different federated machine learning methods. Figures 8A and 9A each show a graph of accuracy determined over a period of time for each of the methods. Figures 8B and 9B each show a graph of the amount of data communicated throughout the network as recorded by the parameter server over a period of time.
[0139] Experimental data obtained using the federated machine learning method without any compression is represented by solid lines in Figures 8A-9B and labeled "FL". Experimental data obtained using the federated machine learning method described in U.S. Patent Application Publication No. 2022 / 0156633A1, which is incorporated herein by reference in its entirety, is represented by crosses in Figures 8A-9B and labeled "US20220156633A1". Experimental data obtained using the method disclosed herein is represented by circles in Figures 8A-9B and labeled "PC". Experimental data obtained using the method described in Shaoxiong Ji, Wenqi Jiang, Anwar Walid, and Xue Li, "Dynamic Sampling and Selective Masking for Communication-Efficient Federated Learning" (2021), arXiv:2003.09603, is represented by dashed lines in Figures 8A-9B and labeled "SM".
[0140] The experimental data shown in Figures 8A and 8B was obtained using the Modified National Institute of Standards and Technology (MNIST) dataset. From Figure 8A, it can be seen that the accuracy determined for the different methods is similar. However, a low accuracy was determined for the method described in US Patent Application Publication No. 2022 / 0156633A1. The communication data amount determined for the method described herein is less than the communication data amount determined for any of the other methods mentioned above, as shown in Figure 8B.
[0141] The experimental data shown in Figures 9A and 9B was obtained using the CIFAR-10 dataset. From this figure, it can be seen that the accuracy determined for the method described herein is similar to that determined for the method described in US Patent Application Publication No. 2022 / 0156633A1. The communication data amount determined for the method described herein is less than the communication data amount determined for any of the other methods described above, as shown in Figure 9B.
[0142] 10 illustrates a computing device 100 for practicing the methods described herein. The computing device 100 may be a server 10 or one of the workers 20.
[0143] Computing device 100 includes a bus 110 , a processor 120 , a memory 130 , a persistent storage device 140 , an input / output (I / O) interface 150 , and a network interface 160 .
[0144] Bus 110 interconnects the components of computing device 100. The bus may be any circuit suitable for interconnecting the components of computing device 100. For example, if computing device 100 is a desktop or laptop computer, bus 110 may be an internal bus located on a computer motherboard of the computing device. As another example, if computing device 100 is a smartphone or tablet, bus 110 may be a global bus of a system-on-chip (SoC).
[0145] Processor 120 is a processing device configured to execute computer-executable instructions loaded from memory 130. Before and / or during execution of the computer-executable instructions, the processor may load the computer-executable instructions from memory 130 via a bus into one or more caches and / or one or more registers of the processor. Processor 120 may be a central processing unit including a suitable computer architecture, for example, x86-64 or ARM architecture. Processor 120 may include, or alternatively be, dedicated hardware adapted for application-specific operations.
[0146] Memory 130 is configured to store instructions and data for use by processor 120. Memory 130 may be a non-transitory volatile memory device, such as a random access memory (RAM) device. In response to one or more operations by the processor, instructions and / or data may be loaded into memory 130 from persistent storage device 140 via the bus in preparation for one or more operations by the processor that utilize those instructions and / or data.
[0147] The persistent storage device 140 is a non-transient, non-volatile storage device, such as a flash memory, a solid state disk (SSD), or a hard disk drive (HDD). A non-volatile storage device maintains data stored on the storage device even after a loss of power. The persistent storage device 140 may have significantly higher access latency and lower bandwidth than the memory 130, e.g., it may take significantly longer to read and write data to the persistent storage device 140 than to the memory 130. However, the persistent storage device 140 may have a significantly larger storage capacity than the memory 130.
[0148] I / O interface 150 facilitates connections between the computing device and external peripherals. I / O interface 150 can receive signals from a given external peripheral, e.g., a keyboard or mouse, convert the signals into a format understandable by processor 120, and relay them to the bus for processing by processor 120. I / O interface 150 can also receive signals from processor 120 and / or data from memory 130, convert them into a format understandable by a given external peripheral, e.g., a printer or display, and relay them to the given external peripheral.
[0149] Network interface 160 facilitates a connection between the computing device and one or more other computing devices over a network. For example, network interface 160 may be an Ethernet network interface, a Wi-Fi network interface, or a cellular network interface.
[0150] Implementations of the subject matter and operations described herein may be realized as digital electronic circuitry, or computer software, firmware, or hardware, including the structures disclosed herein and structural equivalents thereof, or as a combination of one or more of these. For example, hardware may include a processor, a microprocessor, an electronic circuit, an electronic component, an integrated circuit, and the like. Implementations of the subject matter described herein may be realized with one or more computer programs, i.e., one or more modules of computer program instructions, encoded on a computer storage medium for execution by or for controlling the operation of a data processing device. Alternatively, or in addition, the program instructions may be encoded in an artificially generated propagated signal (e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to an appropriate receiver device for execution by a data processing device). The computer storage medium may be or may be included in a computer readable storage device, a computer readable storage substrate, a random or serial access memory array or device, or a combination of one or more of these. Furthermore, although a computer storage medium is not a propagating signal, a computer storage medium may be a source or destination of computer program instructions encoded as an artificially generated propagating signal. A computer storage medium may also be, or be contained in, one or more separate physical components or media (eg, multiple CDs, disks, or other storage devices).
[0151] Although several configurations have been described, these configurations are presented merely as examples and do not limit the scope of protection. The inventive concept described herein may be implemented in various other forms. In addition, various omissions, substitutions and modifications to the specific embodiments described herein may be made without departing from the scope of protection defined in the appended claims.
Claims
1. 1. A computer-implemented method for training a machine learning model in a distributed system, the distributed system comprising a plurality of nodes exchanging updates to jointly train the machine learning model, each node of the plurality of nodes maintaining a local version of the machine learning model, the local version of the machine learning model of each of the plurality of nodes being initialized with a respective same one or more parameter values, the method being performed on a node of the plurality of nodes, the method comprising: receiving a first update of a local model from at least one other node in the distributed system; the local model comprising the local version of the machine learning model, the first update comprising an array of non-zero values of one or more first parameter deltas, the one or more first parameter deltas arranged in the array in an order determined by a reference model, each first parameter delta representing a difference between a parameter of the local model and a corresponding parameter of an updated version of the machine learning model maintained by at least the other node, the reference model being a model prior to the updated version; updating the local model by determining updated parameters where each first parameter delta is added to a corresponding parameter of the local model based on the received first updates and the reference model, and determining an updated local model; determining one or more second parameter deltas, each second parameter delta representing a difference between a parameter of the updated local model and a corresponding parameter of a previous version of the local model; sending a second update to the at least one other node in the distributed system, wherein the second update comprises an array of non-zero values of the one or more second parameter deltas, the one or more second parameter deltas being arranged in the array in an order determined by the reference model. A method comprising:
2. the one or more first parameter deltas are arranged in the array according to magnitudes of corresponding one or more parameters of the reference model; The method of claim 1 , wherein the one or more second parameter deltas are ordered in the array according to magnitudes of corresponding one or more parameters of the reference model.
3. The plurality of nodes comprises a plurality of workers and a server; Each of the plurality of workers is configured to train a respective local model, and the plurality of workers reports a plurality of first updates of the local model back to the server; 2. The method of claim 1 , wherein the server is configured to aggregate the first updates from the workers to update a global model and report the updates of the global model back to the workers, and the server is configured to aggregate the first updates from the workers based on the reference model.
4. The method of claim 3 , wherein the reference model comprises a copy of the global model of a version earlier than the updated version of the global model.
5. the node is a worker, and the method comprises: receiving a corresponding one of the plurality of first updates of the local model from the server, the one or more first parameter deltas indicating a current state of all updated parameters of the global model; The method of claim 3 , wherein updating the local model comprises applying a corresponding one of the first updates of the local model to make the local model conform to the global model.
6. the node is a worker, and the method comprises: determining a level of reduction in the number of parameters of the local model; applying the determined level of reduction in number of the plurality of parameters to the local model to generate a reduced local model; applying the determined level of reduction in number of the plurality of parameters of the local model to the reference model; training the reduced local model based on training data to obtain the updated local model; and transmitting a corresponding one of the first updates to the server for use in updating the global model.
7. determining the level of reduction in number of the plurality of parameters of the local model; determining a quality of service of a communication link between said worker and said server; and determining the level of reduction in the number of the plurality of parameters of the local model based on the quality of service.
8. The method of claim 7 , further comprising adjusting the level of reduction in the number of the plurality of parameters of the local model based on the quality of service.
9. adjusting the level of reduction in number of the plurality of parameters of the local model based on the quality of service; decreasing the level of the reduction in the number of the parameters of the local model in response to an increase in the quality of service; and increasing the level of the reduction in the number of the plurality of parameters of the local model in response to a decrease in the quality of service.
10. applying the determined level of reduction in the number of parameters to the local model when the level of reduction in the number of parameters of the local model is decreased; The method of claim 9 , comprising including one or more additional parameters in the local model, the one or more additional parameters to be included in the local model being determined based on the reference model.
11. the local model comprises a plurality of layers, and applying the determined level of reduction in number of the plurality of parameters to the local model; 7. The method of claim 6, comprising distributing a reduced number of parameters across multiple layers of the local model such that each of the multiple layers of the local model comprises the same number of parameters.
12. applying the determined level of reduction in the number of the plurality of parameters to the local model when at least one layer of the plurality of layers of the local model is full; excluding the at least one layer from further distributing the reduced number of parameters; 12. The method of claim 11, further comprising distributing the reduced number of parameters across one or more remaining layers of the multiple layers of the local model such that each of the one or more remaining layers of the multiple layers of the local model comprises the same number of parameters.
13. the reference model comprises a plurality of layers, and applying the determined level of reduction in number of the plurality of parameters to the reference model; 7. The method of claim 6, comprising distributing a reduced number of parameters across a plurality of layers of the reference model such that each of the layers of the reference model comprises the same number of parameters, and wherein one or more parameters of each of the layers of the reference model correspond to one or more parameters of each of the layers of the local model.
14. applying the determined level of reduction in the number of the plurality of parameters to the reference model when at least one layer of the plurality of layers of the reference model is full; excluding the at least one layer from further distributing the reduced number of parameters; 14. The method of claim 13, further comprising: distributing the reduced number of parameters across one or more remaining layers of the plurality of layers of the reference model such that each of the one or more remaining layers of the plurality of layers of the reference model comprises the same number of parameters.
15. sending, by the worker, a complete update of the local model representing the current state of all parameters of the local model when the quality of service exceeds an upper threshold; The method of claim 7 , further comprising: skipping sending the full update by the worker when the quality of service is below a lower threshold.
16. the node is the server, and the local model maintained by the node is the global model maintained locally by the server; receiving the first update of the local model comprises receiving a plurality of updates from the plurality of workers, each update comprising an array of non-zero values of one or more second parameter deltas; updating the local model comprises aggregating the updates from the workers to update the global model based on the reference model; 4. The method of claim 3, wherein a corresponding one of the updates to the global model is sent by the server to each of the workers for use in updating their respective local models.
17. The method of claim 16 , wherein the one or more second parameter deltas indicate a current state of all updated parameters of the global model.
18. 2. The method of claim 1, wherein the local version of the machine learning model of each of the plurality of nodes is randomly initialized with the same respective one or more parameter values.
19. 1. A node for use in a distributed system comprising a plurality of nodes exchanging updates to jointly train a machine learning model, each node of the plurality of nodes maintaining a local version of the machine learning model, the local version of the machine learning model of each of the plurality of nodes being initialized with the same respective one or more parameter values, the nodes comprising: storage configured to store a local model comprising the local version of the machine learning model; a processor, the processor comprising: receiving a first update of a local model from at least one other node in the distributed system, the first update comprising an array of non-zero values of one or more first parameter deltas, the one or more first parameter deltas being arranged in the array in an order determined by a reference model, each first parameter delta representing a difference between a parameter of the local model and a corresponding parameter of an updated version of the machine learning model maintained by at least the other node, the reference model being a model prior to the updated version; updating the local model by determining updated parameters where each first parameter delta is added to a corresponding parameter of the local model based on the received first updates and the reference model, and determining an updated local model; determining one or more second parameter deltas, each second parameter delta representing a difference between a parameter of the updated local model and a corresponding parameter of a previous version of the local model; sending a second update to the at least one other node in the distributed system, wherein the second update comprises an array of non-zero values of the one or more second parameter deltas, the one or more second parameter deltas being arranged in the array in an order determined by the reference model. A node that is configured to:
20. 1. A non-transitory computer-readable medium comprising a plurality of computer-executable instructions that, when executed by a computer, configures the computer to operate as a node in a distributed system, the distributed system comprising a plurality of nodes exchanging a plurality of updates to collectively train a machine learning model, each node of the plurality of nodes maintaining a local version of the machine learning model, the local version of the machine learning model of each of the plurality of nodes being initialized with a respective same one or more parameter values, the plurality of computer-executable instructions configuring the computer to: receiving a first update to a local model from at least one other node in the distributed system; the local model comprising the local version of the machine learning model, the first update comprising an array of non-zero values of one or more first parameter deltas, the one or more first parameter deltas arranged in the array in an order determined by a reference model, each first parameter delta representing a difference between a parameter of the local model and a corresponding parameter of an updated version of the machine learning model maintained by at least the other node, the reference model being an earlier model than the updated version; updating the local model by determining updated parameters, where each first parameter delta is added to a corresponding parameter of the local model based on the received first updates and the reference model, to determine an updated local model; determining one or more second parameter deltas, each second parameter delta representing a difference between a parameter of the updated local model and a corresponding parameter of a previous version of the local model; sending a second update to the at least one other node in the distributed system, wherein the second update comprises an array of non-zero values of the one or more second parameter deltas, the one or more second parameter deltas being arranged in the array in an order determined by the reference model. A non-transitory computer-readable medium for causing
Citation Information
Patent Citations
Federal learning-based stationary parameter freezing sparsification method and system, and medium
CN116562367A
Signaling of gradient vectors for joint learning in wireless communication systems
CN116711249A
Personalized federal learning method for realizing efficient communication on non-independent identically distributed data
CN116757276A
Method, federated learning system, and computer program (vertical federated learning with compressed embeddings)
JP2022181195A
Iterative learning process using over-the-air transmission and unicast digital transmission
WO2023160816A1