Development of machine learning models using distributed learning

By comparing the trainable parameter values of the ML model in distributed learning and selecting the appropriate training state, the communication volume between nodes is reduced, the problem of high communication costs is solved, and more efficient communication and energy use is achieved.

CN120380485APending Publication Date: 2025-07-25TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280102605.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-14
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

In distributed learning, communication costs and power consumption are high, especially in the communication process between client and server nodes, existing methods fail to effectively reduce communication overhead and power consumption.

Method used

By comparing the trainable parameter values of the current node version and the reference version of the ML model between the first and second nodes of the computer network, selecting the appropriate training state and assigning to components of the ML model, and transmitting only the necessary information to reduce traffic.

Benefits of technology

The communication cost and power consumption between the first node and the second node in distributed learning is reduced, and communication efficiency and energy utilization are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120380485A_ABST
    Figure CN120380485A_ABST
Patent Text Reader

Abstract

A computer-implemented method (100) for developing ML models using distributed learning is disclosed. The method is performed by a first node of a computer network, the method comprising: obtaining a reference version of an ML model (110); receiving a representation of an updated version of the ML model from a second node in the computer network (120); and generating a current node version of the ML model using at least the received representation (130). The method further comprises, for each component of the ML model, comparing (140) a trainable parameter value from a current node version of the ML model to a corresponding trainable parameter value from a reference version of the ML model, and determining (140) a trainable parameter value from the current node version of the ML model according to a measure of a difference between the compared trainable parameter values. A training state selected from the at least two candidate sets of training states is assigned to the component of the ML model (150). The method also includes informing a second node of the computer network of at least one of: the assigned training status of the component of the ML model, and / or trainable parameter values from a current node version of the ML model (160).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to methods for developing machine learning (ML) models using distributed learning. These methods can be performed, for example, by a first node and a second node of a communication network. The present disclosure also relates to a first node, a second node, and a computer program product configured to perform methods for developing an ML model using distributed learning when run on a computer. Background Art

[0002] For the purposes of this specification, distributed learning refers to a set of learning techniques in which at least some training is performed on multiple distributed nodes. Federated learning is an example of a distributed learning technique and is a method for collaboratively training a global machine learning (ML) model. Federated learning involves a server node and a number of clients (e.g., edge devices or core network nodes), and typically a large number of clients participate jointly. Learning begins with the server initializing a global model with a fixed architecture and sending the model to candidate clients. The version of the model is trained for a certain number of rounds using their local data in the clients. Updates to the global model are then sent to the server, where they are aggregated (e.g., by averaging) and then sent back to the clients. This results in a shared global training model that combines knowledge from the clients without the clients sharing their raw data.

[0003] In federated learning (FL), typically a large number of clients participate jointly. After downloading the model from the server, the clients perform training on their local data (e.g., multiple iterations of stochastic gradient descent). In a simple form of federation, each client sends a complete model update (e.g., gradients, hyperparameters, or weights) to the server. The process is repeated until the model training converges.

[0004] In some FL use cases, the agents communicate over the same radio channel or network path, and in this case, the communication channel can become a bottleneck. The communication cost can be very high, for example considering dozens to hundreds of federated rounds and the millions of model weight parameters that should be exchanged between the server and the clients in each federated round.

[0005] Therefore, FL requires the client and server nodes to send and receive a large amount of data in each round of federation. Although a larger bandwidth is expected in the server nodes, sending all model parameters to many clients still results in a large communication cost and power consumption. The communication cost is determined by the number of parameters exchanged between the server node and the client (i.e., converted to bits). Several proposals have been tried to reduce the communication overhead of FL, mainly by any of the following ways: 1) sending only the updated parameters; 2) reducing the model size; and / or 3) reducing the number of federation rounds by faster convergence. Examples of each of these methods are discussed below.

[0006] Sending only the updated parameters: In some existing proposals (e.g., "Federated Learning: Strategies to improve communication efficiency" by Jakub Konečný, H.B. McMahan, F.X. Yu, P. Richtarik, A.T. Suresh, and D. Bacon, NIPS Workshop on Private Multi-Party Machine Learning, 2016), only a part of the model is updated in the client and transmitted to the server node. This can be achieved, for example, by freezing some layers and forcing the update of other layers. In other proposals, the model parameters that have been updated by the client are identified and only the updated parameters are sent.

[0007] Reducing the model size: This includes gradient or model compression by applying data compression techniques. Other methods of reducing the model size are, for example, pruning, subsampling, encoding, and quantization. This method is also used by Konečný et al. in the above reference work.

[0008] Reducing the number of federation rounds: The communication cost can also be reduced by algorithms that result in faster convergence and thus reduce the number of federation rounds, as proposed, for example, in "Two-Stream Federated Learning: Reduce the Communication Costs" by X. Yao, C. Huang, and L. Sun, 2018 IEEE Visual Communications and Image Processing (VCIP), 2018, pp. 1-4, doi: 10.1109 / VCIP.2018.8698609.

[0009] In the heuristic method, the clients exchange parameters with each other in turn (i.e., ring topology) and perform pre-aggregation. Then the final result is sent to the server. This method reduces the communication volume between the client and the server node, but security is an issue due to the large amount of parameter exchange between the clients.

[0010] Since the bandwidth in the 3GPP radio uplink is limited compared to the downlink, the communication cost is typically discussed for communication from the client to the server node, as in the above-mentioned reference works. In these proposals, the locally trained models are sent to the server, and the server aggregates them to create a global training model. This global model is then sent to all candidate clients, and the process is repeated for a given number of federated rounds. However, in these proposals, the communication cost from the server node to the client is ignored, and typically the entire global model is transmitted to the client in each federated round.

[0011] In previous works, it has been proposed to use ensemble methods or to reduce the server-to-client communication by means of compression and joint dropout techniques. However, these practices only bring limited benefits in terms of communication cost. Summary of the Invention

[0012] An object of the present disclosure is to provide a method, a first node and a second node, and a computer program product that at least partially solve one or more of the above challenges. Another object of the present disclosure is to provide a method, a first node and a second node, and a computer program product for collaborating to facilitate reducing the communication overhead in distributed machine learning techniques and thus reducing power consumption.

[0013] According to a first aspect of the present disclosure, there is provided a computer-implemented method for developing a machine learning (ML) model using distributed learning, the ML model including a plurality of components, and each component including at least one trainable parameter. The method is performed by a first node of a computer network, and the method includes: obtaining a reference version of the ML model, receiving a representation of an updated version of the ML model from a second node in the computer network, and generating a current node version of the ML model using at least the received representation. The method further includes: for each component of the ML model, comparing the trainable parameter values from the current node version of the ML model with the corresponding trainable parameter values from the reference version of the ML model, and assigning a training state selected from a candidate set of at least two training states to the component of the ML model based on a measure of the difference between the compared trainable parameter values. The method further includes informing the second node of the computer network of at least one of: the assigned training state of the component of the ML model, and / or the trainable parameter values from the current node version of the ML model.

[0014] According to another aspect of the present disclosure, there is provided a computer-implemented method for developing a machine learning (ML) model using distributed learning, the ML model including a plurality of components, and each component including at least one trainable parameter. The method is performed by a second node of a computer network and includes: for each component of the ML model, receiving from a first node at least one of the following: an assigned training state of the component of the ML model, the assigned training state being selected from a candidate set of at least two training states; and / or a value of a trainable parameter of the component of the ML model. The method further includes: generating an updated version of the ML model at least using the received training state and trainable parameter value, and sending a representation of the updated version of the ML model to the first node.

[0015] According to another aspect of the present disclosure, there is provided a computer program product including a computer-readable non-transitory medium having computer-readable code implemented therein, the computer-readable code being configured to cause, when executed by a suitable computer or processor, the computer or processor to perform the method according to any one aspect or example of the present disclosure.

[0016] According to another aspect of the present disclosure, there is provided a first node for developing a machine learning (ML) model using distributed learning, the ML model including a plurality of components, and each component including at least one trainable parameter. The first node includes processing circuitry configured to cause the first node to: obtain a reference version of the ML model, receive a representation of an updated version of the ML model from a second node in a computer network, and generate a current node version of the ML model at least using the received representation. The processing circuitry is further configured to: for each component of the ML model, cause the first node to compare the value of the trainable parameter from the current node version of the ML model with the corresponding value of the trainable parameter from the reference version of the ML model, and assign a training state selected from a candidate set of at least two training states to the component of the ML model based on a measure of the difference between the compared values of the trainable parameters. The processing circuitry is further configured to cause the first node to inform the second node in the computer network of at least one of the following: the assigned training state of the component of the ML model, and / or the value of the trainable parameter from the current node version of the ML model.

[0017] According to another aspect of the present disclosure, there is provided a second node for developing a machine learning (ML) model using distributed learning, the ML model including a plurality of components, and each component including at least one trainable parameter. The second node includes a processing circuit configured to: for each component of the ML model, cause the second node to receive from a first node at least one of the following: an assigned training state of the component of the ML model, the assigned training state being selected from a candidate set of at least two training states; and / or a value of a trainable parameter of the component of the ML model. The processing circuit is further configured to: cause the second node to generate an updated version of the ML model using at least the received training state and trainable parameter value, and send a representation of the updated version of the ML model to the first node.

[0018] Thus, aspects of the present disclosure provide methods and nodes capable of reducing the communication cost between a first node and a second node in distributed learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To better understand the present disclosure and to more clearly show how the present disclosure may be implemented, reference will now be made, by way of example, to the accompanying drawings in which:

[0020] Figure 1 is a flowchart showing process steps in a computer-implemented method for developing an ML model using distributed learning;

[0021] Figure 2 is a flowchart showing process steps in another example of a computer-implemented method for developing an ML model using distributed learning;

[0022] Figures 3a to 3d is a flowchart showing process steps in another example of a computer-implemented method for developing an ML model using distributed learning;

[0023] Figure 4 is a flowchart showing process steps in another example of a computer-implemented method for developing an ML model using distributed learning;

[0024] Figure 5 is a block diagram showing functional modules in an example first node;

[0025] Figure 6 is a block diagram showing functional modules in another example first node;

[0026] Figure 7 is a block diagram showing functional modules in an example second node;

[0027] Figure 8 is a block diagram showing functional modules in another example second node;

[0028] Figure 9 shows an overview of a system capable of implementing the example methods of the present disclosure;

[0029] Figure 10 shows Figure 9 example operations at a server node of;

[0030] Figure 11 shows Figure 9 example operations at a client node of;

[0031] Figure 12 shows Figure 9 example operations at another example of a client node of;

[0032] Figure 13 shows the process flow of an example of a method for implementing Figures 3a to 3d ; and

[0033] Figure 14 shows an example of signaling between a server and multiple client nodes. Detailed Description

[0034] Examples of the present disclosure propose methods that can reduce the communication overhead from a first node to a second node in a distributed learning scenario. The first node can be a server node or a client node, and thus the methods of the present disclosure can reduce the two-way communication cost. Assigning training states to different components of an ML model, based on a comparison of the training parameter values in the current node version and the reference version of the ML model, allows for a flexible determination of which parts of the model will be trained and exchanged between the first node and the second node, thereby ensuring that only the most important information is transmitted, reducing the communication cost, and thus reducing power consumption.

[0035] Figure 1is a flowchart showing the process steps in a computer-implemented method 100 for developing or training a machine learning (ML) model using distributed learning, the ML model including multiple components, and each component including at least one trainable parameter. The method is performed by a first node of a computer network, which can be a server node or a client node in different examples. Whether operating as a server node or a client node, the first node can include a physical node or a virtual node and can be implemented in a computer system, a computing device, or a server apparatus, and / or in a virtualized environment (e.g., cloud, edge cloud, Open Radio Access Network (O-RAN), or fog deployment). Examples of virtual nodes can include a piece of software or a computer program, a code segment operable to implement the computer program, a virtualization function, or any other logical entity. The first node can be implemented, for example, in the core network of a communication network. In other examples, the computer network can be a commonly referred to communication network, such as an LTE network, a New Radio (NR) network, or any other existing or future communication network system, and the first node can be implemented in a radio access node, which itself can include a physical node and / or virtualized network functions operable to exchange wireless signals. In some examples, the radio access node can include a base station node, such as a NodeB, an eNodeB, a gNodeB, or any future implementation of the function. The first node can include multiple logical entities (discussed in more detail below) and can, for example, include a virtualized network function (VNF). In other examples, the first node can be implemented in a device, which can be a wireless device, a wired device, a constrained device, etc.

[0036] Reference Figure 1 , method 100 includes obtaining a reference version of the ML model in step 110. If the first node is a server node, in one example, the reference version can, for example, include the current global version of the ML model in a federated learning system. In another example where the first node is a server node, the reference version of the ML model can be specific to a particular second node communicating with the first node, and the first node can, for example, maintain multiple reference versions of the ML model, each reference version specific to a particular second node. If the first node is a client node, the reference version of the ML model can include the current node version of the ML model, which the client node maintains in local memory.

[0037] Method 100 further includes receiving, at step 120, a representation of an updated version of the ML model from a second node in a computer network. This representation provides information about the updated version of the ML model and can thus be understood as an instance of the updated version of the model. The representation or instance can include a tensor (such as a vector or matrix) that defines the complete model including all weights, or can include a tensor that defines only the trainable parameters of the model. In other examples, the representation or instance can include a tensor that indicates only how the parameters have changed relative to the parameters in a previous training round, or can include any other form that is capable of reconstructing or approximating the parameters using information from a previous training round.

[0038] Then, method 100 includes, at step 130, generating a current node version of the ML model using at least the received representation. As discussed further below, for example if the first node is a server node and is using a single reference model for all second nodes it communicates with, this can include aggregating the representation with other received representations to generate an updated global model version. In other examples, if the first node is a client node, step 130 can include replacing the parameter values in the previous client node version of the model with the parameter values received from the second (server) node in the representation, and then using a dataset to train the ML model using the local data of the client node to generate the current node version of the ML model.

[0039] Then, method 100 includes performing steps 140 and 150 for each component of the ML model, as shown at 140i. At step 140, method 100 includes comparing the trainable parameter values from the current node version of the ML model with the corresponding trainable parameter values from a reference version of the ML model. At step 150, method 100 includes assigning a training state selected from a candidate set of at least two training states to the component of the ML model based on a measure of the difference between the compared trainable parameter values. Finally, at step 160, method 100 includes informing a second node of the computer network of at least one of: the assigned training state of the component of the ML model, and / or the trainable parameter values from the current node version of the ML model.

[0040] It will be understood that, for the purposes of this disclosure, a component of an ML model can include any part or segment of the model architecture that includes at least one trainable model parameter. Examples of components of an ML model can include the individual layers of a neural network (such as an LSTM, RNN, CNN, GNN, etc.), or individual nodes or groups of nodes. Further examples of components of an ML model can include individual nodes or groups of nodes in a random forest, or a population / genome in a genetic algorithm.

[0041] Method 100 may be supplemented by method 200 executed by a second node. Figure 2 is a flowchart showing process steps in another computer-implemented method 200 for developing an ML model using distributed learning, the ML model including multiple components, and each component including at least one trainable parameter. The method is executed by a second node of a computer network, which may be a client node or a server node. For the first node described above and executing method 100, the second node may include a physical node or a virtual node and may be implemented in any of the ways described above for the first node executing method 100.

[0042] Reference Figure 2 , method 200 includes, in step 210, receiving from the first node at least one of the following: the assigned training state of the component of the ML model, the assigned training state being selected from a candidate set of at least two training states; and / or the trainable parameter value of the component of the ML model. As shown at 201i, the second node receives the assigned training state and / or parameter value of each component of the ML model, as shown at 210i. In step 220, method 200 includes generating an updated version of the ML model using at least the received training state and trainable parameter values. If the second node is a server node, this may include aggregation, or if the second node is a client node, this may include training. Then, method 200 includes, in step 230, sending a representation of the updated version of the ML model to the first node. The representation or instance may take many forms, as discussed above with reference to method 100.

[0043] It will be understood that in some examples, a single client or server node may execute both method 100 and method 200, thus operating as both the first node and the second node: assigning a state (method 100), and taking an action based on a state assigned by another node (method 200).

[0044] In some examples of the present disclosure, the training state may include an indication of at least one of the following:

[0045] (i) whether the value of the trainable parameter of the component should be updated by either or both of the first node and the second node;

[0046] (ii) whether the updated trainable parameter value should be provided by either or both of the first node and the second node to the other node; and / or

[0047] (iii) a representation of the hyperparameter value used by the first node or the second node in updating the value of the trainable parameter.

[0048] It will be understood that different combinations of the above features can be included in any one training state. In some examples, the candidate set of training states can include at least two of active, stopped, or frozen.

[0049] In some examples, the active training state can indicate that the values of the trainable parameters of the component should be updated by each of the second node and the first node, and each of the second node and the first node should provide the updated parameter values to the first node and the second node, respectively.

[0050] In additional examples, the stopped training state can indicate that the values of the trainable parameters of the component should be updated by each of the second node and the first node, and only the second node should provide the updated values to the first node.

[0051] In additional examples, the frozen training state can indicate that the values of the trainable parameters of the component should not be updated by the second node or the first node.

[0052] In additional examples, the active, stopped, and frozen states can be combined with additional indications. For example, the "stop and boost" state is a stopped state and thus indicates that the values of the trainable parameters of the component should be updated by each of the second node and the first node, and only the second node should provide the updated values to the first node. However, the "stop and boost" state can additionally indicate that the learning rate used by the second node should be increased. The state can indicate the increment or factor by which the learning rate should be increased, or can indicate the new learning rate used by the second node. The learning rate is an example of a hyperparameter value used by the first node or the second node when updating the values of the trainable parameters.

[0053] Other enhancements can be made to the active, stopped, and frozen states, including changes to the hyperparameters to be used by the second node. Another example of a combined state is "stopped with new learning rate". This is similar to what was discussed above, where a new learning rate is specified to be added to the stopped state. Another example of a combined state is "collective stop", according to which the parameters are sent to the first node only after a specified number of training epochs. Another example of a combined state is "whole network stop / active / frozen", according to which, instead of layer-by-layer state assignment, one state can be assigned to the whole network. The state of the whole network can be stopped, active, or frozen.

[0054] Methods 100 and 200 are used to develop an ML model. For the purposes of this disclosure, the term "ML model" encompasses the following concepts within its scope:

[0055] Machine learning algorithms, including processes or instructions capable of generating model artifacts for performing a given task or representing a real-world process or system through their use of data during the training process; and

[0056] Model artifacts, which are created through this training process and include a computational architecture that performs the task.

[0057] In some examples, the ML models developed according to methods 100, 200 can include classification models, regression models, prediction models, clustering models, control models, and / or other types of ML models. Example uses of classification models include: classifying different types of faults based on hardware or software log files of communication network base stations, classifying compliance or violation of service level agreements based on statistical data observed from infrastructure, etc. Example uses of regression models can include estimating the energy consumption of radio units. Example uses of prediction models can include predicting KPIs using PM / CM (Performance Management / Configuration Management) data and LSTM models, performing handover prediction using graph neural networks, etc. Example uses of clustering include unsupervised joint learning for detecting anomalies. Example uses of control models include predictive maintenance. It will be understood that a series of specific model architectures have been commercialized for the above example tasks, and this list is not exhaustive but only shows different types of ML models and associated tasks that can be developed according to the methods disclosed herein.

[0058] Figures 3a to 3d A flowchart of another example of a computer-implemented method for developing an ML model using distributed learning is shown. The ML model includes multiple components, and each component includes at least one trainable parameter. Like method 100 discussed above, method 300 is performed by a first node in a computer network, which can include a physical node or a virtual node. The first node can act as a server node or a client node, as discussed in further detail below. Method 300 shows various examples of how to implement and supplement the steps of method 100 to provide the functions discussed above and additional functions.

[0059] First refer to Figure 3a, in the first step 310, the first node obtains a reference version of the ML model. If the first node is a server node, in one example, the reference version of the ML model may include the current global version of the ML model maintained by the server node. In another example, the reference version of the ML model may include a version specific to a particular second node. Thus, for example, the server node may maintain a reference version of the ML model specific to each second node communicating with the server node. In yet another example, the server node may construct the reference version based on the global version maintained at the server node and the model parameter values received from the client node in step 320. The server node may receive a representation including a plurality of ML model parameter values from the client node. Then, the server node may construct the reference version of the model by using the values obtained from the global version maintained by the server node to fill in any missing parameter values in the received representation. Alternatively, if the first node is a client node, the reference version may include, for example, the current client node version of the ML model, which was trained using the client's local dataset in the previous learning round.

[0060] In step 320, the first node receives a representation of an updated version of the ML model from a second node in the computer network. This representation or instance can take many forms, as discussed above with reference to method 100. As shown in 320a, this may include receiving updated values of the trainable parameters of the components of the ML model with assigned active or stopped training status from the second node. As discussed above, these training statuses include instructions to provide the updated trainable parameter values to the first node. Such a status may include a combined or intermediate training status, where the actions associated with active or stopped (as discussed above) are combined with changes to hyperparameters. Generally, any existing model format (such as Open Neural Network Exchange, ONNX (https: / / onnx.ai)) or the formats used in common toolkits (such as Keras or PyTorch) can be used to send or transmit the ML model or a representation of the ML model between nodes.

[0061] In step 330, the first node generates a current node version of the ML model using at least the received representation. The operations involved in completing step 330 can vary depending on whether the first node is a server node or a client node. If the first node is a server node and the second node is a client node, as shown in 330a, generating a current node version of the ML model using at least the received representation can include: aggregating the received representation with representations received from other client nodes. This aggregation can be performed by taking a mean, weighted mean, any non-linear transformation of weights, etc. In such an example, the current node version of the ML model can include an updated global version of the model. In some examples, for components of the ML model with a frozen training state, generating a current node version of the ML model using at least the received representation can include copying the values from the reference version of the ML model into the current node version (as in some examples, the values of the frozen layers are not included in the representation from the second node).

[0062] Alternatively, if the first node is a client node and the second node is a server node, then in step 330, generating a current node version of the ML model using at least the received representation can include: at step 330b, initially generating a first client node version of the ML model by replacing the values of the trainable parameters in the reference version of the ML model with the values of the trainable parameters included in the received representation. Then, completing step 330 can include: at step 330c, generating a current client node version of the ML model by performing training on the first client node version of the ML model using the training dataset associated with the client node.

[0063] Now referring to Figure 3b , the first node performs steps 340 and 350 on the respective components of the ML model. In step 340, the first node compares the values of the trainable parameters from the current node version of the ML model with the corresponding values of the trainable parameters from the reference version of the ML model, where the reference version is as described above with reference to step 310. In step 350, the first node assigns a training state selected from a candidate set of at least two training states to the component of the ML model based on a measure of the difference between the compared values of the trainable parameters. As discussed above and as shown in 350i, the training state can include an indication of at least one of the following:

[0064] (i) whether the values of the trainable parameters of the component should be updated by either or both of the first node or the second node;

[0065] (ii) whether the updated trainable parameter values should be provided by either or both of the first node and the second node to the other node; and / or

[0066] (iii) a representation of hyperparameter values used by the first node or the second node when updating the values of the trainable parameters.

[0067] As shown in 350ii, the candidate set of training states can include at least two of active, stopped, or frozen. The active, stopped, and frozen training states can include the following indications:

[0068] The active training state indication: The values of the trainable parameters of this component should be updated by each of the second node and the first node, and each of the second node and the first node should provide the updated parameter values to the first node and the second node accordingly.

[0069] The stopped training state indication: The values of the trainable parameters of this component should be updated by each of the second node and the first node, and only the second node should provide the updated values to the first node.

[0070] The frozen training state indication: The values of the trainable parameters of this component should not be updated by the second node or the first node. As discussed above, any one of the active, stopped, and frozen states can be combined with additional indications, such as adjusting the hyperparameter values to be used by the second node.

[0071] As shown in 350iii, in some examples, assigning a training state to a component of the ML model based on a measure of the difference between the compared trainable parameter values includes generating a mask for the ML model that includes entries for each component of the ML model, each entry indicating the assigned training state of the corresponding component.

[0072] Figure 3d The steps that can be performed to execute the assigned training state at step 350 are shown.

[0073] Now referring to Figure 3d , the first node can initially check in step 350a whether the measure of the difference between the compared trainable parameter values is equal to or higher than a difference threshold. If the measure of the difference between the compared trainable parameter values is equal to or higher than the difference threshold, the first node can assign an active state to the component in step 350b. If the measure of the difference between the compared trainable parameter values is lower than the difference threshold, the first node can assign a stopped or frozen training state to the component in step 350c. The measure of the difference can be the L2 distance, cosine similarity, calculated discrete index, or any other suitable difference measure.

[0074] As shown in 350ai, in the first learning round, the difference threshold can be set to zero. This can ensure that: in the first learning round, all components of the ML model are assigned an active training state. As discussed below, in subsequent learning rounds in subsequent method steps, the difference threshold can be increased.

[0075] After it has been determined at step 350c that a component should be assigned a stopped or frozen training state (because the measure of the difference is below the difference threshold), the first node can then check in step 350d whether the number of learning rounds in which the component already has the assigned stopped training state is equal to or higher than the freeze threshold. If the number of learning rounds in which the component already has the assigned stopped training state is lower than the freeze threshold, the first node can assign a stopped training state to the component in step 350e.

[0076] If the number of learning rounds in which the component already has the assigned stopped training state is equal to or higher than the freeze threshold, the first node can assign a frozen training state to the component in step 350f. For a component of the ML model that already has the frozen training state assigned in step 350f, the first node can then check in step 350g whether a restart condition is met. If the restart condition is not met, the first node can maintain the assigned frozen training state. If the restart condition is met, the first node can assign an active or stopped training state to that component of the ML model in step 350h. Examples of the restart condition can include a threshold number of combined rounds in which the component has been frozen, and combined conditions related to the availability of communication bandwidth and power, the convergence of other components of the ML model, etc. The restart condition can take actions to ensure that the second node does not remain at a local minimum and to ensure that when suitable resources are available, the training possibilities are fully utilized.

[0077] When assigning any one of the active, frozen, or stopped states in steps 350b, 350e, and 350f, the first node can additionally determine whether it is appropriate to adjust the hyperparameters used by the second node. Thus, for example, when assigning a stopped state, the first node can determine whether it would be beneficial to increase the learning rate or change any other hyperparameters used by the second node based on available resources or other constraints.

[0078] Referring again to Figure 3b , after the training state has been assigned to the components of the ML model at step 350, the first node then checks in step 352 whether all components of the model have been assigned a training state. If not all components have been assigned a training state, the first node returns to step 340 to compare and assign a training state for the next component of the model.

[0079] When all components of the model have been assigned a training status, reference is now made to Figure 3c , whereupon the first node can then update the values of the trainable parameters in the reference version of the ML model to be the same as the values of the trainable parameters in the current node version of the ML model in step 354. It will be understood that the completion of this step can depend on the nature of the reference model and that, in some examples, if the first node is a server node and the server node constructs a reference version specific to each client node, the server node can update the global version of the model to include the values of the trainable parameters from the current node version of the model.

[0080] In step 360, the first node informs a second node of the computer network of at least one of the following: the assigned training status of each component of the ML model, and / or the values of the trainable parameters from the current node version of the ML model. As shown in 360a, this can include sending to the second node at least one of the following: the mask entries or the values of the trainable parameters of each component of the ML model from the current node version of the ML model.

[0081] As shown in 360b, informing the second node can include: informing the second node of the values of the trainable parameters from the current node version of the ML model only for the components of the ML model that have an assigned active training status. For components with a stopped or frozen training status, the first node can send only the mask entries. It will be understood that, depending on the number of possible state changes, the mask entries can be represented by a very limited number of bits and, thus, greatly reduce the total amount of information to be sent compared to sending the training parameter values of all components of the ML model.

[0082] In step 362, the first node can check whether an update condition is met and, when the update condition is met, the first node can increase the difference threshold. In some examples, the update condition can include at least one of the following items:

[0083] Completing a threshold number of learning rounds;

[0084] The available bandwidth for communicating with the second node drops below a communication threshold;

[0085] The available computing resources for performing the method drop below a computing threshold;

[0086] The available memory is below a memory threshold;

[0087] The power consumption at the first node or the second node exceeds a power threshold;

[0088] External intervention.

[0089] The effect of the update condition can be that whenever there is a need to reduce energy consumption, processing, or communication resource usage, the difference threshold is increased (meaning that more components of the ML model will enter a stopped or frozen mode, and thus there will be fewer parameter value updates and exchanges). It will be understood that the external intervention can be a manual intervention from an expert or an administrator, or any other external factor.

[0090] If the update condition is not met, the first node returns to step 310 to complete another iteration of the method within another round of learning rounds.

[0091] It will be understood that the examples of the present disclosure thus provide a method that can both reduce communication costs in a distributed learning scenario and provide considerable flexibility in its application. With respect to which entities update and exchange trainable parameter values, different states provide considerable variation and granular control, while reducing communication costs by: sending trainable parameter values only for components suitable for sending trainable parameter values, and for other components, sending training states that can be represented extremely efficiently with a small number of bits.

[0092] Figure 4 is a flowchart showing process steps of another example of another computer-implemented method 400 for developing an ML model using distributed learning, the ML model including a plurality of components, and each component including at least one trainable parameter. Similar to method 200 discussed above, method 400 is performed by a second node in a computer network, which can include a physical node or a virtual node. The second node can act as a client node or a server node, as discussed in further detail below. Method 400 shows various examples of how to implement and supplement the steps of method 200 to provide the functions discussed above and additional functions.

[0093] Refer to Figure 4 , the second node performs step 410 for each component of the ML model, as shown in 410a. In step 410, the second node receives from the first node at least one of the following: the assigned training state of the component of the ML model, the assigned training state being selected from a candidate set of at least two training states; and / or the trainable parameter value of the component of the ML model.

[0094] The characteristics and examples of the training state were discussed above with reference to method 300, and that discussion equally applies to method 400.

[0095] As shown at 410b, performing step 410 can include receiving at least one of a masked entry or a trainable parameter value of the component of the ML model, where the masked entry indicates an assigned training state of the component of the ML model. As shown at 410c, performing step 410 can additionally or alternatively include: receiving only the trainable parameter values of the components of the ML model having an assigned active training state from the first node.

[0096] In step 420, the second node generates an updated version of the ML model using at least the received training state and trainable parameter values. Step 420 can be performed differently depending on whether the second node is a client node or a server node. In one example, if the second node includes a client node and the first node includes a server node, performing step 420 can initially include: in step 420a, generating a first client node version of the ML model by replacing the values of the trainable parameters in the previous client version of the ML model with the trainable parameter values received from the first node. Performing step 420 can include: in step 420b, generating an updated version of the ML model by performing training on the first client node version of the ML model using a training dataset associated with the client node. In such an example, generating an updated version of the ML model by performing training on the first client node version of the ML model using a training dataset associated with the client node can include: during training, updating the values of the trainable parameters only for the components of the ML model having an active or stopped training state. Thus, while only the active component parameter values can be received from the first node (the server node in this example), the parameters of both the active components and the stopped components can be trained at the second node (the client node in this example).

[0097] In another example, if the second node includes a server node and the first node includes a client node, generating an updated version of the ML model using at least the received training state and trainable parameter values in step 420 can include aggregating the received trainable parameter values with the trainable parameter values received from other client nodes in step 420c.

[0098] In some examples of the present disclosure, comparison of the training parameter values and assignment of the training state can be performed at both the first node and the second node. In such an example, the second node can perform steps 422 and 424 of method 400 for each component of the ML model.

[0099] In step 422, the second node compares the trainable parameter values of the updated version of the ML model with the corresponding trainable parameter values of the previous node version of the ML model, and in step 424, the second node updates the assigned training status of the components of the ML model based on a measure of the difference between the compared trainable parameter values. It will be understood that all of the details and options discussed above regarding the comparison and assignment of training status, etc., performed by the first node as part of method 300 apply equally to steps 422 and 424 of method 400 performed by the second node.

[0100] In step 430, the second node sends a representation of the updated version of the ML model to the first node. This representation or instance can take many forms, as discussed above with reference to method 100. In an example where the second node has performed steps 422 and 424, the second node can also send the updated training status along with the representation of the updated version of the ML model to the first node. As shown in 430a, sending a representation of the updated version of the ML model to the first node can include: sending the trainable parameter values from the updated version of the ML model to the first node only for the components of the ML model with an assigned active or stopped training status.

[0101] As discussed above, methods 100 and 300 can be performed by the first node, and the present disclosure provides a first node adapted to perform any or all of the steps of the above methods. The first node can include a physical node such as a computing device, server, etc., or can include a virtual node. The virtual node can include any logical entity such as a virtualized network function (VNF), which itself can operate in a cloud, edge cloud, O-RAN, or fog deployment. The first node can be operable to be instantiated in a cloud-based deployment and / or can be a device. In one example, in a centralized or cloud-based deployment, the first node can be instantiated in a physical server or a virtual server.

[0102] Figure 5 is a block diagram showing an example first node 500 that can implement, for example, methods 100 and / or 300 of an example according to the present disclosure when receiving appropriate instructions from a computer program 550 as Figure 1 and Figures 3a to 3d shown. Referring to Figure 5 , the first node 500 includes a processor or processing circuit 502 and can include a memory 504 and an interface 506. The processing circuit 502 is operable to perform some or all of the steps of methods 100 and / or 300 as discussed above with reference to Figure 1 and Figures 3a to 3d . The memory 504 can contain instructions executable by the processing circuit 502 such that the first node 500 is operable to perform asFigure 1 and Figures 3a to 3d some or all of the steps of the methods 100 and / or 300 shown. The instructions may also include instructions for performing one or more telecommunication and / or data communication protocols. The instructions may be stored in the form of a computer program 550. The memory may also contain, for example, ML model parameters. In some examples, the processor or processing circuitry 502 may include one or more microprocessors or microcontrollers, as well as other digital hardware (which may include a digital signal processor (DSP), dedicated digital logic, etc.). The processor or processing circuitry 502 may be implemented by any type of integrated circuit (such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc.). The memory 504 may include one or several types of memory suitable for the processor, such as read only memory (ROM), random access memory, cache memory, flash memory devices, optical storage devices, solid state disks, hard disk drives, etc. The first node 500 may also include an interface that may be operable to facilitate communication with a second node and / or other nodes or modules via a suitable communication channel.

[0103] Figure 6 illustrates functional modules in another example of a first node 600, which may, for example, execute examples of the methods 100 and / or 300 of the present disclosure according to computer-readable instructions received from a computer program. It will be understood that Figure 6 the modules shown are functional modules and may be implemented in any suitable combination of hardware and / or software. The modules may include one or more processors and may be integrated to any extent.

[0104] Referring to Figure 6, the first node 600 is used to develop a machine learning (ML) model using distributed learning. The ML model includes multiple components, and each component includes at least one trainable parameter. The first node 600 includes a reference module 610 for obtaining a reference version of the ML model. The first node 600 also includes a transceiver module 620 for receiving a representation of an updated version of the ML model from a second node in a computer network. The first node 600 also includes a learning module 630 for generating a current node version of the ML model using at least the received representation. The first node 600 also includes a state module 640. The state module is configured to: for each component of the ML model, compare the trainable parameter values from the current node version of the ML model with the corresponding trainable parameter values from the reference version of the ML model, and assign a training state selected from a candidate set of at least two training states to the component of the ML model based on a measure of the difference between the compared trainable parameter values. The transceiver module 620 is also used to inform the second node of the computer network of at least one of the following: the assigned training state of the component of the ML model, and / or the trainable parameter values from the current node version of the ML model. The second node 600 may also include an interface 650, which may be operable to facilitate communication with the second node and / or other nodes or modules via a suitable communication channel.

[0105] As discussed above, methods 200 and 400 may be performed by a second node, and the present disclosure provides a second node suitable for performing any or all of the steps of the above methods. The second node may include a physical node such as a computing device, a server, etc., or may include a virtual node. The virtual node may include any logical entity such as a virtualized network function (VNF), which itself may operate in a cloud, edge cloud, O-RAN, or fog deployment. The second node may be operable to be instantiated in a cloud-based deployment and / or may be a device. In one example, in a centralized deployment or a cloud-based deployment, the second node may be instantiated in a physical server or a virtual server.

[0106] Figure 7 is a block diagram showing an example first node 700, which can implement methods 200 and / or 400 according to an example of the present disclosure when receiving appropriate instructions from a computer program 750 as shown in Figure 2 and Figure 4 shown. Referring to Figure 7 , the second node 700 includes a processor or processing circuit 702 and may include a memory 704 and an interface 706. The processing circuit 702 is operable to execute as described above with reference to Figure 2 and Figure 4Some or all of the steps of the methods 200 and / or 400 discussed. The memory 704 may contain instructions executable by the processing circuitry 702 such that the second node 700 is operable to perform some or all of the steps of the methods 200 and / or 400 as Figure 2 and Figure 4 shown. The instructions may also include instructions for performing one or more telecommunication and / or data communication protocols. The instructions may be stored in the form of a computer program 750. In some examples, the processor or processing circuitry 702 may include one or more microprocessors or microcontrollers, as well as other digital hardware (which may include a digital signal processor (DSP), dedicated digital logic, etc.). The processor or processing circuitry 702 may be implemented by any type of integrated circuit (such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), etc.). The memory 704 may include one or several types of memory suitable for the processor, such as read only memory (ROM), random access memory, cache memory, flash memory devices, optical storage devices, solid state disks, hard disk drives, etc. The second node 700 may also include an interface, which may be operable to facilitate communication with the first node and / or other nodes or modules via a suitable communication channel.

[0107] Figure 8 FIG. shows functional modules in another example of the second node 800, which may perform examples of the methods 200 and / or 400 of the present disclosure, for example, according to computer-readable instructions received from a computer program. It will be understood that Figure 8 the modules shown are functional modules and may be implemented in any suitable combination of hardware and / or software. The modules may include one or more processors and may be integrated to any extent.

[0108] Referring to Figure 8 , the second node 800 is used to develop a machine learning (ML) model using distributed learning. The ML model includes multiple components, and each component includes at least one trainable parameter. The second node 800 includes a transceiver module 810 for causing the second node to receive, from the first node, for each component of the ML model, at least one of the following: the assigned training state of the component of the ML model, the assigned training state being selected from a candidate set of at least two training states; and / or the value of the trainable parameter of the component of the ML model. The second node 800 also includes a learning module 820 for generating an updated version of the ML model using at least the received training state and the trainable parameter values. The transceiver module 810 is also used to send a representation of the updated version of the ML model to the first node. The second node 800 may also include an interface 830, which may be operable to facilitate communication with the first node and / or other nodes or modules via a suitable communication channel.

[0109] The above discussion Figures 1 to 4 provides an overview of methods that can be performed according to different examples of the present disclosure. These methods can be performed by a first node and a second node respectively, as Figures 5 to 8 shown. These methods can reduce the communication overhead between nodes during distributed learning. Specifically, the method can reduce the communication overhead from the server node to the client node during distributed learning. Now, a detailed discussion will be given on how the different process steps shown in Figures 1 to 4 and discussed above can be implemented. The functions and implementation details described below are with reference to the modules of examples that execute methods 100, 200, 300, and / or 400, Figures 5 to 8 substantially as described above.

[0110] Figure 9 shows an overview of a system capable of implementing an example method of the present disclosure. The system includes a server node that communicates with M client nodes. In one example, the server node can operate as the first node and execute examples of methods 100, 300, and the client nodes can operate as the second node and execute examples of methods 200, 400. In additional examples, these roles can be interchanged, where the server node operates as the second node and the client nodes operate as the first node. It will be understood that in some examples of method 400, the second node can also perform a comparison of trainable parameter values and assign a training status to the components of the ML model being trained. In such examples, the training status can be assigned at both the server node and the client node.

[0111] In Figure 9 the example implementation shown, the server node operates as the first node and the client nodes operate as the second node. The server node assigns a training status to the components of the ML model (in this case, to the layers of the neural network being trained), and the server node only sends the parameters of the active layers to the clients. In this implementation, the latest trained model is stored in the server node, and after model aggregation, parameters that have a significant similarity (i.e., below a difference threshold) to the reference model are not sent to the clients. In the shown implementation, a mask is created by labeling the layers as active, stopped, or frozen. This mask is sent to the candidate clients together with the parameters of the active layers.

[0112] Figure 10 shows Figure 9Example operations at a server node that operates as a first node and executes examples of method 100 and / or 300. The server node includes two logical components that are particularly relevant to method implementation: an aggregator and a comparator. In the example shown, the ML model being trained is a neural network, and the comparator is a layer-by-layer comparator. Comparators for other types of components and other types of ML models can be envisioned based on the nature of the ML model being trained. Refer to Figure 10 , aggregation is performed after all model updates have been received from candidate client nodes. In some examples, not all client nodes are considered by the server node to be candidate clients for aggregation and parameter sharing. The aggregated model is then compared to a reference model, which in this example is the current model at the server. The comparator assigns one of three training states (also referred to as modes) to each layer: active, stopped, or frozen. The parameters of stopped and frozen layers are masked and not sent to the client nodes. The comparator works as follows: if the similarity score of a layer is less than a threshold (i.e., if the difference score is equal to or higher than the threshold), the layer is marked as active, meaning it should be further trained in the client using the updated parameter values provided by the server. If the similarity score of a layer is greater than the threshold (i.e., the difference score is below the threshold) and the previous mask for that layer was active, the layer enters the stopped mode, meaning no significant changes have been identified in the trained model at the client. Therefore, there is limited benefit in sending the updated parameters back to the client, although the client node can continue to update the parameter values during its training. If it is detected that the similarity score is greater than the threshold (i.e., if the difference score is below the threshold), and at the same time, the previous state of the layer in the reference model was stopped, the layer can enter the frozen mode, meaning the parameters of the layer remain unchanged at the client node during the training phase. This state takes advantage of the fact that further training is unlikely to further change the parameters of the layer. In some examples, the method can ensure that components (such as NN layers) can maintain the stopped state for a given number of learning rounds before transitioning to the frozen state.

[0113] Figure 11 illustrates Figure 9Example operations at a client node that operates as a second node and performs examples of methods 200 and / or 400. The client node includes three logical components that are particularly relevant to method implementation: a Null replacer, a trainer, and a mask placer. In the example shown, the ML model being trained is a neural network, so the trainer is an NN trainer. After receiving the server node model, the client node replaces the null layers (with a stopped or frozen state) with the corresponding layers in the current client model. The new model is then sent to the NN trainer component to train the model. During the training phase, the frozen layers remain unchanged, but the parameters of the active and stopped layers can evolve freely. Finally, in the trained model, the frozen layers are set to null before the parameter values of the stopped and active layers are sent to the server node.

[0114] Figure 12 illustrates Figure 9 Example operations of another example of a client node. Figure 12 The client node operates as a second node and performs example method 400, where the second node also performs a comparison of trainable parameter values and assigns a training status. In Figure 12 the client node, the mask placer is enhanced to also perform the function of the layer-by-layer comparator discussed above. Referring to Figure 12 , after receiving the server node model, the client node replaces the null layers (with a stopped or frozen state) with the corresponding layers in the current client model. The new model is then sent to the NN trainer component to train the model. During the training phase, the frozen layers remain unchanged, but the parameters of the active and stopped layers can evolve freely. After training, the trained model is compared with the current client model (which acts as a reference model in this use case). For the neural network of this example, the comparison is layer by layer, as discussed above. As a result of the comparison, and based on the similarity or difference score calculated for the corresponding values of the trainable parameters in the trained and reference versions of the model, the client node assigns a training status to each layer, as fully discussed above with reference to the operations at the server node and Figure 10 as discussed in full.

[0115] In some examples, the models at different clients may vary slightly from a shared global model, allowing a degree of specialization to be achieved at each client based on its local data.

[0116] Figure 13 illustrates a process flow of an example implementing method 300 (when executed by a first node including a server node).

[0117] Referring to Figure 13 , the server node can perform the following steps:

[0118] 1 - Receive models from all candidate clients (step 320 of method 300)

[0119] 2 - Retrieve the current server node model as a reference model (step 310 of method 300)

[0120] 3 - Perform aggregation (steps 330, 330a of method 300)

[0121] 4 - Compare the aggregated model with the reference model layer by layer (steps 340, 350a of method 300)

[0122] 5 - If the similarity score of a layer is less than the threshold, set that layer to active in the mask (steps 350a, 350b of method 300)

[0123] 6 - If the similarity score is greater than the threshold and the previous state of that layer in the reference model was active, set that layer to stopped in the mask (steps 350a, 350d, 350e of method 300)

[0124] 7 - If the similarity score is greater than the threshold and the previous state of that layer in the reference model was stopped or frozen, set that layer to frozen in the mask (steps 350a, 350d, 350f of method 300)

[0125] 8 - If all layers have been compared, proceed to step 9 (step 352 of method 300), otherwise return to step 4 (steps 340, 350a of method 300)

[0126] 9 - Send the parameters of the active layers along with the mask to all candidate clients. Do not send the parameters of the stopped and frozen layers (steps 360, 360a, 360b of method 300)

[0127] 10 - Replace the aggregated model with the reference model and store it along with the mask (step 354 of method 300)

[0128] A corresponding process flow can be envisioned to implement an example of method 400 (when executed by a second node including a client node). Such a process flow may include the following steps:

[0129] 1 - Receive a model and mask entries from a server node (steps 410, 410a, 410b, 410c of method 400)

[0130] 2 - Retrieve the current client node model as a reference model

[0131] 3 - In the model received from the server node, replace the empty layers with their corresponding layers in the reference model (steps 420, 420a of method 400)

[0132] 4 - Perform NN training (step 420b of method 400)

[0133] 5 - Compare the trained model with the reference model layer by layer (step 422 of method 400). If the similarity score of a layer is less than the threshold, set that layer to be empty; otherwise, retain the parameters of that layer and select that layer (step 424 of method 400, also see steps 350a to 350h of method 300)

[0134] 6 - Send the parameters of the selected layers to the server node (step 430 of method 400)

[0135] 7 - Replace the trained model with the reference model.

[0136] By considering the above two example process flows, it will be understood that the method steps of methods 100, 200, 300, and 400 can be executed in an order different from the order shown in the figures. Thus, for example, before obtaining the reference version of the ML model by retrieving it from the memory, for example, the first node can receive a representation of the updated version of the ML model from the second node. Other examples of being able to execute the method steps in an order different from the order presented above can be envisioned.

[0137] Referring to the comparison of the trainable parameter values of different versions of the ML model, it will be understood that a series of different algorithms can be used to calculate the similarity or difference between the values of the components of the ML model, such as the similarity or difference between the layers of the aggregated version and the reference version of the neural network model. Example difference metrics can include the L2 distance, cosine similarity, or calculating a discrete index, which is a summary of statistical data calculated based on the ratio of variance to the mean.

[0138] As discussed above, in the first learning round (e.g., the first joint round in the case of federated learning), the difference threshold can be set to zero such that the parameter values of all components (e.g., all layers of the neural network) are sent. In this case, all layers are assigned an active training state. As the server model converges, the difference threshold can be gradually increased to achieve the communication cost savings discussed above. For example, if the server node acting as the first node stops or freezes half of the parameters of the model within half of the training time, the communication cost from the server to the client will achieve a 25% reduction.

[0139] To avoid getting permanently stuck in a local optimum, the first node can decide to unfreeze layers in some learning (or joint) rounds. The learning rounds continue until an end condition is met, such as until a certain learning round is reached or when the loss value drops below a predefined value.

[0140] Figure 14An example of signaling between a server and multiple clients is shown. Figure 14 Continuing with the following example: The server node acts as the first node and implements examples of methods 100, 300, and the client node acts as the second node and implements examples of methods 200, 400. Refer to Figure 14 , in steps (1) and (2), the server receives the locally trained models M_local_1 and M_local_2. In step (3), the server performs steps 2 to 8 described above with reference to Figure 13 . In steps (4) and (5), the server transmits the simplified model M_global_active to the client, which only includes the selected active layers. Then, the client node uses the active layers of the received global model to perform the next round of training.

[0141] Thus, the examples of the present disclosure provide methods and nodes capable of reducing the communication cost of distributed learning. This reduction is achieved by reducing the amount of information to be exchanged between the first node and the second node, and can be used in combination with or as a supplement to other techniques for reducing the amount of information (such as compression, etc.). The reduction of communication cost for both server-to-client communication and client-to-server communication can be achieved.

[0142] The reduction of communication cost is achieved by eliminating the transmission of parameter values that change little or not at all between learning rounds, and thus the reduction of power consumption is achieved. By freezing the components of the model that have been observed to have limited changes, faster convergence of the entire model can also be achieved. The use of intermediate states or stop states (where parameters can be updated but not provided from one node to another) provides an intermediate option between updating and exchanging parameters and completely freezing the update and exchange of parameters. Additionally, the use of updates to one or more hyperparameters is used to enhance any state described herein in order to further fine-tune the learning process and adapt to situations including available bandwidth, processing resources, memory resources, etc. It will be understood that the present disclosure provides a very low-overhead method. The size of the mask for transmitting the training state can be as small as 2 bits per component (i.e., each layer of the NN), which can be negligible compared to the amount of data sent.

[0143] As discussed above, the update rate of parameters during distributed learning can be controlled by dynamically changing the difference threshold. For example, a smaller threshold is used in the early learning rounds (i.e., most components are active), and a larger threshold is used in the later rounds (i.e., forcing more components into the stop or freeze mode). The threshold can also be set or changed depending on the accuracy required by the ML model, the nature of the computer network implementing the method (such as the current network load when the method is executed in a communication network), or the nature of the first node or the second node.

[0144] The methods of the present disclosure may be implemented in hardware or as software modules running on one or more processors. The methods may also be performed in accordance with the instructions of a computer program, and the present disclosure also provides a computer-readable medium having stored thereon a program for performing any of the methods described herein. The computer program embodying the present disclosure may be stored on a computer-readable medium, or it may, for example, be in the form of a signal (e.g., a downloadable data signal provided from an Internet website), or it may be in any other form.

[0145] It should be noted that the above examples illustrate rather than limit the present disclosure, and those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims or the numbered embodiments. The word "comprising" does not exclude the presence of elements or steps other than those listed in the claims or embodiments, "a" or "an" does not exclude a plurality, and a single processor or other unit may perform the functions of several units recited in the claims or numbered embodiments. Any reference signs in the claims or numbered embodiments should not be construed as limiting their scope.

Claims

1. A computer - implemented method (100) for developing a machine - learning (ML) model using distributed learning, the ML model comprising a plurality of components, and each component comprising at least one trainable parameter, the method being performed by a first node of a computer network, the method comprising: Obtaining a reference version (110) of the ML model; Receiving a representation of an updated version of the ML model from a second node in the computer network (120); Generating a current - node version of the ML model using at least the received representation (130); and For each component (140i) of the ML model: Comparing the trainable - parameter values from the current - node version of the ML model with the corresponding trainable - parameter values from the reference version of the ML model (140); Assigning a training state selected from a candidate set of at least two training states to the component of the ML model based on a measure of the difference between the compared trainable - parameter values (150); and Informing the second node of the computer network of at least one of the following: The assigned training state of the component of the ML model; The trainable - parameter values from the current - node version of the ML model (160).

2. The method according to claim 1, wherein, The training state includes an indication (350i) of at least one of the following: Whether the value of the trainable parameter of the component should be updated by either or both of the first node and the second node; Whether the updated trainable - parameter value should be provided by either or both of the first node and the second node to the other node; A representation of the hyper - parameter value used by the first node or the second node when updating the value of the trainable parameter.

3. The method according to claim 1 or 2, wherein, The candidate set of training states includes at least two of active, stopped, or frozen (350ii).

4. The method according to claim 3, wherein: The active training state indicates that the value of the trainable parameter of the component should be updated by each of the second node and the first node, and each of the second node and the first node should accordingly provide the updated parameter value to the first node and the second node; The stopped training state indicates that the value of the trainable parameter of the component should be updated by each of the second node and the first node, and only the second node should provide the updated value to the first node; And The frozen training state indicates that the value of the trainable parameter of the component should not be updated by the second node or the first node.

5. The method according to claim 3 or 4, wherein Assigning a training state to the component of the model based on a measure of the difference between the compared trainable - parameter values for each component of the ML model includes: If the measure of the difference between the compared trainable - parameter values is equal to or higher than a difference threshold (350a), then assigning an active state (350b) to the component.

6. The method according to any one of claims 3 to 5, wherein Assigning a training state to the component of the model based on a measure of the difference between the compared trainable - parameter values for each component of the ML model includes: If a measure of the difference between the comparable trainable parameter values is below a difference threshold (350a), assign a stopped or frozen training state (350c) to the component.

7. The method according to any one of claims 3 to 6, wherein Assigning a training state to a component of the ML model based on a measure of the difference between the comparable trainable parameter values for each component of the ML model includes: If a measure of the difference between the comparable trainable parameter values is below a difference threshold and if the number of learning rounds for which the component has been assigned a stopped training state is equal to or higher than a freeze threshold (350d), assign a frozen training state (350f) to the component.

8. The method according to any one of claims 3 to 7, wherein, Assigning a training state to a component of the ML model based on a measure of the difference between the comparable trainable parameter values for each component of the ML model includes: If a measure of the difference between the comparable trainable parameter values is below a difference threshold and if the number of learning rounds for which the component has been assigned a stopped training state is below a freeze threshold (350d), assign a stopped training state (350e) to the component.

9. The method according to any one of the preceding claims, wherein, Assigning a training state to a component of the ML model based on a measure of the difference between the comparable trainable parameter values includes generating a mask for the ML model, the mask including entries for each component of the ML model, each entry indicating the assigned training state of the corresponding component (350iii).

10. The method according to claim 9, wherein, Inform at least one of the following to a second node of the computer network: The assigned training state of the component of the ML model; Trainable parameter values from the current node version of the ML model, including: Send at least one of the following (360a) to the second node: a mask entry or a trainable parameter value of the component of the ML model from the current node version of the ML model.

11. The method according to any one of claims 3 to 10, wherein Inform at least one of the following to a second node of the computer network: The assigned training state of the component of the ML model; Trainable parameter values from the current node version of the ML model, including: Inform the trainable parameter values from the current node version of the ML model to the second node only for components of the ML model that have an assigned active training state (360b).

12. The method according to any one of claims 3 to 11, wherein, Receiving a representation of an updated version of the ML model from a second node in the computer network includes: Receiving updated values of the trainable parameters of components of the ML model that have an assigned active or stopped training state from the second node (320a).

13. The method according to any one of the preceding claims, further comprising: Updating the trainable parameter values in a reference version of the ML model to be the same as the trainable parameter values in the current node version of the ML model (354).

14. The method according to any one of the preceding claims, wherein, The first node includes a server node and the second node includes a client node, and wherein generating the current node version of the ML model using at least the received representation includes: Aggregate the received representation with representations received from other client nodes (330a).

15. The method according to any one of claims 1 to 13, wherein The first node includes a client node, and the second node includes a server node, and wherein generating the current node version of the ML model using at least the received representation includes: Generating a first client node version of the ML model (330b) by replacing the values of the trainable parameters in a reference version of the ML model with the values of the trainable parameters included in the received representation; and Generating the current node version of the ML model (330c) by performing training on the first client node version of the ML model using a training data set associated with the client node.

16. The method according to any one of claims 5 to 15, wherein In a first learning round, the difference threshold is set to zero (350ai).

17. The method according to any one of claims 5 to 16, further comprising: When an update condition is satisfied (362), increasing the difference threshold (364).

18. The method according to claim 17, wherein, The update condition includes at least one of the following items: Completing a threshold number of learning rounds; The available bandwidth for communicating with the second node drops below a communication threshold; The available computing resources for performing the method drop below a computing threshold; The available memory is below a memory threshold; The power consumption at the first node or the second node exceeds a power threshold; External intervention.

19. The method according to any one of claims 3 to 18, further comprising: For a component of the ML model having an assigned frozen training state: Checking whether a restart condition is satisfied (350g); And When the restart condition is satisfied, assigning an active or stopped training state to the component of the ML model (350h).

20. A computer-implemented method (200) for developing a machine learning ML model using distributed learning, the ML model including a plurality of components, and each component including at least one trainable parameter, the method being performed by a second node of a computer network, the method comprising: For each component of the ML model (210i), receiving (210) from a first node at least one of the following: The assigned training state of the component of the ML model, the assigned training state being selected from a candidate set of at least two training states; The values of the trainable parameters of the component of the ML model; Generating an updated version of the ML model (220) using at least the received training state and trainable parameter values; And Sending a representation of the updated version of the ML model to the first node (230).

21. The method according to claim 20, wherein The training state includes an indication of at least one of the following: Whether the value of the trainable parameter of the component should be updated by either or both of the first node and the second node; Whether the updated trainable parameter values should be provided to the other node by either or both of the first node and the second node; A representation of the hyperparameter values used by the first node or the second node when updating the value of the trainable parameter.

22. The method according to claim 20 or 21, wherein, The candidate set of training states includes at least two of active, stopped, or frozen.

23. The method according to claim 22, wherein: The active training state indicates that the values of the trainable parameters of the component should be updated by each of the second node and the first node, and each of the second node and the first node should correspondingly provide the updated parameter values to the first node and the second node; The stopped training state indicates that the values of the trainable parameters of the component should be updated by each of the second node and the first node, and only the second node should provide the updated values to the first node; And The frozen training state indicates that the values of the trainable parameters of the component should not be updated by the second node or the first node.

24. The method according to any one of claims 20 or 23, wherein, For each component of the ML model, receive from the first node at least one of the following: The assigned training state of the component of the ML model, the assigned training state being selected from a candidate set of at least two training states; The values of the trainable parameters of the component of the ML model, including: Receive at least one of the masked entries or the values of the trainable parameters of the component of the ML model, the masked entries indicating the assigned training state of the component of the ML model (410b).

25. The method according to any one of claims 22 to 24, wherein For each component of the ML model, receive from the first node at least one of the following: The assigned training state of the component of the ML model, the assigned training state being selected from a candidate set of at least two training states; The values of the trainable parameters of the component of the ML model, including: Receive only the values of the trainable parameters of the components of the ML model having the assigned active training state from the first node (410c).

26. The method according to any one of claims 22 to 25, wherein Sending a representation of the updated version of the ML model to the first node includes: sending only the values of the trainable parameters from the updated version of the ML model to the first node for the components of the ML model having the assigned active or stopped training states (430a).

27. The method according to any one of claims 20 to 26, wherein, The first node includes a server node, and the second node includes a client node, and wherein generating the updated version of the ML model using at least the received training state and the values of the trainable parameters includes: Generating a first client node version of the ML model by replacing the values of the trainable parameters in the previous client version of the ML model with the values of the trainable parameters received from the first node (420a); and Generating the updated version of the ML model by performing training on the first client node version of the ML model using the training dataset associated with the client node (420b).

28. The method according to claim 27 when claim 27 depends on claim 22, wherein Generating the updated version of the ML model by performing training on the first client node version of the ML model using the training dataset associated with the client node includes: During training, updating only the values of the trainable parameters for the components of the ML model having the active or stopped training states (420b).

29. The method according to any one of claims 20 to 26, wherein The first node includes a client node, and the second node includes a server node, and wherein generating an updated version of the ML model using at least the received training status and trainable parameter values includes: Aggregating (420c) the received trainable parameter values with trainable parameter values received from other client nodes.

30. The method according to any one of claims 20 to 29, further comprising: For each component of the ML model: Comparing (422) the trainable parameter values from the updated version of the ML model with the corresponding trainable parameter values from a previous node version of the ML model; Updating (424) the assigned training status of the component of the ML model based on a measure of the difference between the compared trainable parameter values; And Sending (430) the updated training status along with a representation of the updated version of the ML model to the first node.

31. A computer program product comprising a computer-readable medium having computer-readable code embodied therein, the computer-readable code being configured to, when executed by a suitable computer or processor, cause the computer or processor to perform the method according to any one of claims 1 to 30.

32. A first node (500) for developing a machine learning ML model using distributed learning, the ML model including a plurality of components, and each component including at least one trainable parameter, the first node including processing circuitry (502), the processing circuitry (502) being configured to cause the first node to: Obtain a reference version of the ML model; Receive a representation of an updated version of the ML model from a second node in the computer network; At least use the received representation to generate the current node version of the ML model; And For each component of the ML model: Comparing the trainable parameter values from the current node version of the ML model with the corresponding trainable parameter values from the reference version of the ML model; Assigning a training status selected from a candidate set of at least two training statuses to the component of the ML model based on a measure of the difference between the compared trainable parameter values; and Informing the second node of the computer network of at least one of the following: The assigned training status of the component of the ML model; The trainable parameter values from the current node version of the ML model.

33. The first node according to claim 32, wherein, The processing circuitry is further configured to cause the first node to perform the steps according to any one of claims 2 to 19.

34. A second node (700) for developing a machine learning ML model using distributed learning, the ML model including a plurality of components, and each component including at least one trainable parameter, the second node including processing circuitry (702), the processing circuitry (702) being configured to cause the second node to: For each component of the ML model, receive from a first node at least one of the following: The assigned training state of the component of the ML model, the assigned training state being selected from a candidate set of at least two training states; The trainable parameter values of the component of the ML model; Generate an updated version of the ML model using at least the received training state and trainable parameter values; And Send a representation of the updated version of the ML model to the first node.

35. The second node according to claim 34, wherein, The processing circuitry is further configured to cause the second node to perform the steps according to any one of claims 21 to 30.