Device and method for realizing asynchronous federated learning
Through the asynchronous process of asynchronous federated learning and appropriate client selection, the problems of training delay and inefficiency in traditional federated learning are solved, and more efficient model training and resource optimization are achieved.
Patent Information
- Application Number
- CN202380082150.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-30
- Filing Date
- 2023-05-03
- Publication Date
- 2025-07-08
AI Technical Summary
During the traditional federated learning process, synchronization is difficult due to changes in the timetable of the FL client, resulting in training delay and inefficiency, making it difficult to meet the performance and time requirements of model training.
Using an asynchronous process, the FL server generates global updates and distributes them after receiving local updates, allowing the client to train independently, selects the appropriate client through rating information and data statistics, splits the training tasks, and processes outdated updates to accelerate model training.
The federated learning process is accelerated, the efficiency and accuracy of model training is improved, the training time and performance requirements are met, and resource utilization is optimized.
Smart Images

Figure CN120283243A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to network analysis and machine learning in the context of mobile communications, and in particular to a technique for training a machine learning model according to a federated learning scheme. Background Art
[0002] Network analysis and AI / ML (Artificial Intelligence / Machine Learning) are deployed in a 5G (fifth-generation mobile communication) core network by introducing NWDAF (Network Data Analytics Function), which considers support for various types of analysis, such as UE (User Equipment) mobility, user data congestion, NF (Network Function) load, etc., as described in TS23.288. The types of analysis can be distinguished and selected by consumers using analysis IDs. Each NWDAF can support one or more analysis IDs and can have an inference role, called an NWDAF containing an AnLF (Analysis Logic Function) (or simply an AnLF), or a role of training an ML model, called an NWDAF containing an MTLF (Model Training Logic Function) (or simply an MTLF), or both roles. The AnLF subscription that supports the inference of a specific analysis ID subscribes to the corresponding MTLF responsible for ML model training.
[0003] Figure 1 is a block diagram that outlines various NWDAF styles, including potential input data sources and output consumers. The input data sources and output result consumers can include 5G core NFs, AFs (Application Functions), 5G core data repositories (e.g., ADRF (Analysis Data Repository Function)), and OAM (Operations, Administration, and Maintenance). Communication with untrusted AFs can be achieved via NEF (Network Exposure Function). OAM can include MnS (Management Service) consumers or MFs (Management Functions). MTLF and AnLF can exchange AI / ML models, for example, in terms of parameters or weights, or in a serialized or containerized manner. Optionally, DCCF (Data Collection Coordination Function) and MFAF (Message Framework Adapter Function) can be involved to distribute and collect duplicate data from various data sources.
[0004] Traditional analysis enablers are based on supervised / unsupervised learning, which may face some major challenges, including user data privacy and security, which may complicate the UE-level data collection of NWDAF. In addition, with the introduction of MTLF, various types of data from different fields are required to train the ML model of the NWDAF containing MTLF. However, it may be difficult to collect all the raw data from distributed data sources for the NWDAF containing MTLF.
[0005] To address these challenges, 3GPP has adopted Federated Learning (FL), also known as Federated Machine Learning. FL technology can be used in the NWDAF containing MTLF to train ML models without any transmission of the original data. Instead, only the ML model or the ML model parameters or weights (the ML model is considered to be fully characterized by its ML model parameters or ML model weights, where the latter are synonymous here) are transmitted between MTLFs that support FL capabilities. In the context of analysis, two FL capabilities are defined, namely the FL server and the FL client.
[0006] The FL server is responsible for handling the FL process, including selecting FL clients for training the ML model based on the training latency of its ML model, aggregating the ML model parameters received from the FL clients to generate an updated version of the ML model, determining when the updated ML model meets the specified target performance (e.g., based on ML model validation and testing) and / or when a certain confidence level is reached, and distributing the updated ML model among AnLFs that subscribe to receive the update or request retraining of the ML model.
[0007] The FL client is responsible for performing ML model training when the FL server requests ML model training using local data and sending the ML model parameters to the FL server when the local training is completed.
[0008] The FL capabilities related to the FL server and the FL client can be registered with each corresponding MTLF in the NRF (Network Repository Function) for a specific analysis ID and a specific ML model. The ML model training process using FL can be executed in multiple repeated iterations. In each iteration, the FL server selects FL clients and provides them with the ML model parameters. Each FL client trains the ML model using local data and returns an updated version of the ML model parameters to the FL server. The FL server aggregates the received ML model parameters and repeats the process by selecting FL clients and distributing the updated ML model parameters. When a certain target performance or prediction confidence level (which can be determined based on ML model validation and testing performed by the FL server) is reached, the ML model training can be completed.
[0009] In the traditional FL process, the FL server waits until responses from all FL clients are received, then aggregates the received ML model parameters and distributes an updated version of the ML model in the next iteration. However, each FL client can have its own schedule for collecting data and performing the training process using its own computing and communication resources. This means that the FL process will only execute at the speed of the slowest FL client. Changes in the schedules associated with each FL client can challenge the synchronization of the entire FL process and may introduce further latency for ML model training.
[0010] In this specification, the expressions "ML model distribution" and "distribution of ML model parameters" are synonyms. It should be understood that the term "distribution of ML model parameters" should not be construed as a limitation on the particular technology used to distribute the ML model. Those skilled in the art will understand that the ML model can equally be executed by sending the parameters or weights of the ML model, or by sending the ML model via serialization or containerization, or by any other means that can transfer ML model information between the FL server and the FL clients, all of which are considered to be included in the term "distribution of ML model parameters". Summary of the Invention
[0011] The object of the present invention is to overcome the above problems in traditional FL processing, and in particular to provide a server device, a client device and a method that can accelerate FL processing.
[0012] To achieve these objectives, a particular method of the present invention is to perform federated learning using an asynchronous process, where the FL server uses the received local ML model updates to generate a globally updated ML model and distributes the globally updated ML model before local updates are received from all FL clients.
[0013] According to a first aspect of the present invention, there is provided a server device for collaborative training of a machine learning ML model with a plurality of clients according to a federated learning scheme. The server device is configured to perform the following steps: select a subset of the plurality of clients available for federated learning; distribute the parameters of the ML model to each of the selected clients to allow the clients to train the ML model on local data; receive local updates for the ML model from the clients that have completed training the ML model on local data; generate a globally updated ML model by applying the received local updates to the ML model; and determine the model accuracy of the globally updated ML model. The server device is further configured to at least iterate the steps of distributing, receiving, generating and determining. When at least one of one or more predefined iteration conditions is satisfied, a new iteration is started by distributing the parameters of the globally updated ML model, and at least one of one or more predefined iteration conditions is independent of whether local updates have been received from all the selected clients.
[0014] Preferably, the server device is further configured to repeat the steps of receiving, generating and determining until the improvement in model accuracy caused by the last received local update does not exceed a predefined threshold, where once the improvement in model accuracy caused by the last received local update exceeds the predefined threshold, a new iteration is started.
[0015] In this way, the FL process is accelerated because a new iteration is started when the local update of the current iteration cannot provide a significant improvement. Therefore, there is no longer a need to wait for slow clients that provide ineffective updates.
[0016] Preferably, whenever a local update is received from one of the selected clients, a new iteration is started, and the above iteration is started by distributing the parameters of the global updated ML model to the client from which the local update was received.
[0017] In this way, the FL process is further accelerated because a new iteration is started individually for each client as soon as each client provides its local update.
[0018] Preferably, the server device is further configured to complete the training of the ML model when a target training time limit expires or when the model accuracy of the globally updated ML model reaches a target value.
[0019] In this way, the requirements regarding the duration of the training process and the desired model accuracy can be met.
[0020] Preferably, the server device is further configured to generate rating information in response to receiving a local update for the ML model from a client, the rating information indicating a measure of the impact of the local update received from the above client on the model accuracy and / or the time required for the above client to complete the training of the ML model on local data; and store the generated rating information in association with the identifier of the above client.
[0021] In this way, rating information that can be collected and saved for future reference can be obtained, and this rating information helps to select FL clients and / or allocate optimized training tasks for each of the selected clients.
[0022] Preferably, the server device is further configured to obtain, for each client available for federated learning, rating information indicating a measure of the impact of the local update previously received from the above client on the model accuracy and / or the time required for the above client to complete the previous training of the ML model on local data; and select a subset of the multiple clients available for federated learning at least partially based on the obtained rating information.
[0023] In this way, a set of clients that are particularly suitable for a given training task can be selected based on the previously collected rating information.
[0024] Preferably, the server device is further configured to, for each of a plurality of clients available for federated learning, obtain data statistics information indicating at least one of a range, a quantity, a mean, and a variability of local data provided by the client for training an ML model; and select a subset of the plurality of clients available for federated learning at least in part based on the obtained data statistics information. Preferably, the server device is further configured to, for each of a plurality of clients available for federated learning, obtain data context information indicating a context in which the local data is collected, the context including at least one of the following: network conditions, network load conditions, network use cases, radio access technologies, network slices, service IDs, application IDs, and the time when the local data is collected; and select a subset of the plurality of clients available for federated learning at least in part based on the obtained data context information.
[0025] In this way, a set of clients particularly suitable for a given training task can be selected based on data statistics and / or context information of local training data.
[0026] Preferably, the server device is further configured to generate, for each of the selected clients, target training data information indicating a subset of local data to be used for training the ML model; and provide the corresponding target training data information for each of the selected clients so that each client trains the ML model only on the corresponding subset of local data. Preferably, the local data includes a plurality of data samples, each data sample including a plurality of data features, wherein the target training data information indicates constraints on the data samples and / or data features to be used for training the ML model.
[0027] In this way, the model training process can be split into smaller tasks, each task focusing on a subset and / or a sub-range of local training data.
[0028] Preferably, the server device is further configured to, for each of the selected clients, obtain rating information indicating a measure of the impact of local updates previously received from the client on model accuracy and / or the time required for the client to complete previous training of the ML model on local data; and generate the target training data information at least in part based on the obtained rating information.
[0029] In this way, the training tasks assigned to individual clients can be designed specifically for the capabilities of the corresponding clients.
[0030] Preferably, the server device is further configured to: if the received local update includes an outdated update from a client that has performed training of the ML model based on stale parameters of a previous iteration of the ML model, apply a weight to the outdated update to reduce the impact of the outdated update on the ML model. Preferably, the server device is further configured to calculate the weight of the outdated update based on the correlation between the respective impacts of the outdated update and the non-outdated update on the ML model, where the non-outdated update is a local update received from a client that has performed training of the ML model based on the latest parameters of the ML model distributed during the current iteration.
[0031] In this way, compared to the impact of the outdated update, the impact of the non-outdated update can be enhanced, where the non-outdated update can be more relevant to updating the current ML model, while the outdated update can be less relevant.
[0032] Preferably, the server device is further configured to exclude the outdated update from generating the globally updated ML model if a predetermined condition is met, where the predetermined condition includes at least one of the following: the amount of time that has passed since the stale parameters were distributed or the number of iterations that have been completed, the number of non-outdated updates received within the current iteration, where the non-outdated update is a local update received from a client that has performed training of the ML model based on the latest parameters of the ML model distributed during the current iteration, the impact of the outdated update on the performance of the globally updated ML model, and the rating information of the client from which the outdated update has been received.
[0033] In this way, very old and thus presumably irrelevant outdated updates to the current version of the ML model can be discarded, thereby avoiding a reduction in model accuracy due to the application of inappropriate updates.
[0034] Preferably, the server device is further configured to, before starting a new iteration or completing the training of the ML model, send an early termination request to a selected client from which a local update has not been received, to request the client to stop training the ML model and send a temporary update based on the currently available training results; receive the temporary update sent by the client in response to the early termination request; and apply the received temporary update to the ML model.
[0035] According to a second aspect of the present invention, there is provided a client device for training an ML model on local data in a federated learning scenario. The client device is configured to: obtain the parameters of the ML model and target training data information indicating a subset of the local data to be used for training the ML model from a server; perform a training process on the ML model based on the subset of the local data indicated by the obtained target training data information; generate a local update for the ML model based on the result of the training process; and send the local update to the server.
[0036] Preferably, the local data includes a plurality of data samples, each data sample includes a plurality of data features, and the target training data information indicates constraints on the data samples and / or data features to be used in the training process.
[0037] In this way, the ML model training process can be split into smaller tasks, each task focusing on a subset and / or sub-range of the local training data.
[0038] Preferably, the client device is further configured to receive an early termination request from the server; and in response to receiving the early termination request, stop the training process of the ML model, generate a temporary update for the ML model based on the currently available training results, and send the temporary update to the server.
[0039] In this way, even if the local training process is not completed in time, the training results of slow clients can still be used to update the ML model.
[0040] Preferably, the client device is further configured to send the local update to the server together with information indicating the version of the ML model for which the training process has been performed on it.
[0041] In this way, the server can determine whether the received local update is an outdated update and / or how stale the update has become, and decide how / whether to apply the update to the ML model.
[0042] According to a third aspect of the present invention, there is provided a method for collaboratively training an ML model with a plurality of clients according to a federated learning scheme. The method includes the following steps: selecting a subset of a plurality of clients available for federated learning; distributing the parameters of the ML model to each of the selected clients to allow the clients to train the ML model on local data; receiving local updates for the ML model from the clients that have completed training the ML model on local data; generating a globally updated ML model by applying at least one received local update to the ML model; determining the model accuracy of the globally updated ML model; and at least iterating the steps of distributing, receiving, generating, and determining. When at least one of one or more predefined iteration conditions is satisfied, a new iteration is started by distributing the parameters of the globally updated ML model, and at least one of one or more predefined iteration conditions is independent of whether local updates have been received from all the selected clients.
[0043] The present invention content is provided to introduce a series of concepts in a simplified form, which will be further described in the following detailed description. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. Description of the Drawings
[0044] Embodiments and aspects of the present invention will be described in conjunction with the accompanying drawings in the following description, where
[0045] Figure 1 is a block diagram that outlines various NWDAF styles, including potential input data sources and output consumers.
[0046] Figure 2 is a schematic diagram illustrating an example of dividing training data for machine learning into vertical and horizontal data sets.
[0047] Figure 3 is a schematic diagram illustrating an asynchronous FL model training process according to a first embodiment of the present invention.
[0048] Figure 4 is a schematic diagram illustrating an asynchronous FL model training process according to a second embodiment of the present invention. Detailed Description
[0049] Reference will now be made in detail to the exemplary embodiments, which are illustrated in the accompanying drawings.
[0050] The present invention addresses the performance and latency issues in traditional ML model training by providing an asynchronous process for federated learning. The present invention also provides mechanisms for selecting appropriate FL clients in each iteration, for accelerating each iteration of the FL model training process, and for handling stale information, i.e., local updates derived by slow FL clients for an old version of the ML model, which are only received by the FL server after a new iteration starts and an updated version of the ML model is distributed to newly selected FL clients.
[0051] [Mechanism for Selecting Appropriate FL Clients]
[0052] Selecting appropriate FL clients requires the FL server to first discover potential FL clients. This process is described in TR 23.700-81, where the FL server registers its FL capabilities, i.e., FL server or FL client, in the NRF. From the NRF query, the FL server can discover NWDAFs with FL client capabilities, the geographical area (location) related to the FL client, and the FL client load that is prone to dynamic changes. From the list of FL clients received from the NRF, the FL server can ask individual FL clients about the data available for FL training, i.e., whether they have collected data that can be used for FL training and the time available for performing FL training.
[0053] The selection of FL clients can be performed at least in part based on statistical attributes of the local data that each FL client can provide for training the ML model, such as the range, quantity, mean, variability, etc. of the local data. When the FL server asks individual FL clients for more information during the preparation phase, the above statistical data on the locally available training data can be provided to the FL server.
[0054] Specifically, the FL server can associate data statistics related to each potential FL client to select FL clients with more diverse training data. This can also save network and FL client computing resources.
[0055] In addition, the selection of FL clients can also be performed at least in part based on the context of the corresponding local training data, particularly based on the context in which the local training data was collected. This includes but is not limited to information indicating network conditions, such as network load conditions, the use case in which the network was when the data was collected (e.g., energy saving), a specific RAT (radio access technology, such as 5G, LTE, WiFi, etc.), a network slice (i.e., S-NSSAI (single network slice selection assistance information)), a service and / or application ID, the time when the data was collected (e.g., morning, evening, etc.), or the freshness of the data.
[0056] [Mechanisms for accelerating each iteration]
[0057] The present invention accelerates the FL model training process by immediately updating the global ML model after the FL server receives an update from an FL client. In addition, by distributing the updated global ML model to a new set of selected FL clients before all updates have been received from the selected FL clients, i.e., when at least one of one or more predefined iteration conditions is satisfied, the FL server can start a new FL iteration, where at least one of one or more predefined iteration conditions is independent of whether local updates have been received from all selected clients. These iteration conditions can include a time limit or the number of received updates.
[0058] In addition, the FL model training process can be split into separate tasks so as to assign smaller but more diverse and targeted tasks to FL clients for ML model training. Due to reasons such as shortage of computing resources or high load, FL clients are expected to slow down, so smaller training tasks can be assigned to FL clients by indicating the desired level data range or vertical feature set that the training should focus on.
[0059] For example, data statistics can be used to divide the desired data space for FL model training into horizontal data sets by splitting the training data based on a specified sample range that covers all desired ML model features. Features are independent variables in the ML model. For example, a mobility analysis ML model can include the type of UE (e.g., car), UE speed and direction, and a specific location granularity (e.g., below cell level) as features for predicting mobility. Based on the values of these features, it can be predicted that the UE will, for example, switch to a different cell or be at a specific location. Vertical data sets can also be arranged based on the features included in a specific FL client across the entire data sample range. Figure 2 An example of dividing training data for machine learning into vertical and horizontal data sets is illustrated in
[0060] The FL training process can be divided by assigning training jobs for specified horizontal and / or vertical data to FL clients, which can speed up the FL process without sacrificing performance. Other data statistics (including data volume, mean, and variability) can also indicate the diversity and distribution of data in the target data space.
[0061] The present invention further speeds up the FL model training process by allowing the FL server to send a "last call message" to slow FL clients before distributing a new ML model version. The slow FL clients can then stop and / or abandon the training process and send an update of the current ML model information to the FL server.
[0062] The FL server can also send the new ML model version to the slow FL clients, which are required to respond with the current ML model information upon receiving the new ML model version and then either start the training process from scratch or continue training using the new ML model version.
[0063] [Mechanism for handling stale information]
[0064] The present invention also provides a mechanism for handling stale update information, i.e., local updates derived by slow FL clients for an old version of the ML model, which are only received by the FL server after a new iteration starts and an updated version of the ML model is distributed to newly selected FL clients.
[0065] For example, the FL server can apply a weight that controls (i.e., degrades or attenuates) the impact of stale ML model updates. Then, the stale updates with the applied weight can still be used to perform the global FL model update. The weight can be calculated or derived by correlating the impact of accurate information (i.e., FL client information using the latest data version of the ML model) and stale information (i.e., FL clients using a non-latest data version of the ML model) on the global ML model update.
[0066] The FL server can also introduce a limit on accepting out-of-date ML model updates. This limit can be defined according to the iteration time or period after the global FL model update is performed at the FL server. This limit can also be formulated according to the number of new updates that have been received from FL clients using the updated ML model version distributed from the FL server.
[0067] The FL server can also determine the impact of out-of-date information on the global ML model performance and use this impact as a basis for accepting or rejecting out-of-date updates. The impact of a specific update on the ML model can be determined by the FL server through an ML model verification and testing process. The ML model verification and testing process can be performed when an out-of-date update (i.e., an update generated using a non-up-to-date version of the ML model) is received. FL client model updates that degrade the global ML model performance can be rejected. After receiving a certain number of these useless updates, the FL server can reject all additional out-of-date updates.
[0068] The FL server can also generate rating information for FL clients. The rating information can indicate the performance, speed, or reliability of the clients performing the local training process. The rating information can also indicate the accuracy, usefulness, or impact of the updates generated by the corresponding FL clients. The rating information can be locally stored in the FL server during the duration of the FL process so that it can be used in each different iteration. The rating information can also be generated after each FL process and stored by the NRF to affect the priorities in the NF profile described in TS29.510, which may affect the selection of FL clients for future FL processes.
[0069] [First Embodiment]
[0070] Figure 3 FIG. is a schematic diagram illustrating an asynchronous FL model training process according to the first embodiment of the present invention. The following will explain this process on the premise that the FL capabilities (i.e., the FL server and FL clients) related to a specific NWDAF containing MTLF have been registered in the NRF.
[0071] The process starts from step 1, where a consumer (i.e., the NWDAF containing AnLF) issues an ML model subscription request to the FL server. The request can include an analysis ID and ML model filter information as described in TS23.288. The request can also include an indication of the desired time to execute or complete the ML model training update and / or the ML model performance target value.
[0072] In subsequent steps 2-4, the FL server selects a set of FL clients that will be used in the next iteration of the FL model training process.
[0073] Specifically, in step 2, the FL server sends a request to the NRF to discover potential FL clients in the area of interest (AoI). The NRF provides information on the FL clients. The above information may include the load of each FL client. The information provided by the NRF may also include rating information and / or priority information obtained from the previous use of the corresponding FL client during the FL model training, such as a measure of the impact of the local updates received from the above clients on the ML model accuracy and / or the time required for the above clients to complete the training of the ML model on the local data.
[0074] After the list of potential FL clients is received, in step 3, the FL server checks whether the potential FL clients have collected the data required for local training (data availability of the potential FL clients). The FL server further checks whether the potential FL clients have the computing resources required to complete the local training process in a timely manner, or whether no other very important tasks are scheduled (time availability of the potential FL clients). In addition, the FL server may check the data statistical information indicating the statistical attributes of the locally collected training data and the data context information indicating the conditions under which the local training data is collected.
[0075] In step 4, the FL server analyzes the responses and then selects the FL clients that will participate in the next FL iteration.
[0076] At this stage, the FL server may optionally split the ML model training into smaller tasks, for example, as described above, by requesting at least some FL clients to perform ML model training only on a limited local training dataset and / or a limited range of model features. In this way, the FL server can still select FL clients with only a limited amount of computing resources available, that is, FL clients that are expected to have an insufficient duration for performing local training but have valuable data for training the ML model.
[0077] In step 5, the FL server sends a request to each of the selected FL clients that will participate in the next FL iteration to perform ML model training using their local data. The request may optionally include target training data information that indicates the subset of the local data to be used for training the ML model. In this way, the ML model training can be split into smaller tasks, as described above.
[0078] In step 6, the FL server checks the remaining time for the ML model training target time imposed on the consumer (i.e., AnLF). If there is still time, the FL server waits for at least one FL client to provide an ML model update report (step 7). The report includes an indication of the result of the ML model training process performed by the corresponding FL client on its local data. The result can be provided in the form of a new set of ML model parameters, a set of ML model parameter gradients (weights), or any other suitable form.
[0079] In step 8, the FL server uses the results received from the FL clients to update the ML model. Thus, the step of updating the ML model can be performed at the FL server in response to receiving an ML model update report from a single FL client, without waiting until all FL clients have sent their ML model update reports.
[0080] Then in step 9, the FL server validates and / or tests the thus updated ML model to evaluate the ML model performance, especially to estimate the expected confidence.
[0081] In step 10, the FL server can rate the local ML model training reports from the FL clients from which they are received. The above rating can indicate a measure of the impact of the local update received from that particular client on the accuracy, performance, or confidence of the ML model, and / or the time required for the above client to complete the training of the ML model on the local data. The FL server can save the rating information locally in this case for further use in the current FL iteration.
[0082] In step 11, the FL server checks the ML model training progress. If the ML model performance does not reach the target performance provided by the consumer (i.e., AnLF), the FL server checks the number of reports received from the FL clients. If the FL server has received reports from all the FL clients selected in step 4, the FL server starts a new iteration by jumping to step 2. If the FL server has not received reports from all the selected FL clients, it checks the ML model improvement considering the last FL client update received. If the last update has a significant impact on the ML model, for example, the model performance improvement caused by the last update exceeds a certain threshold, the FL server waits to receive more FL client report updates and repeats steps 6 - 11. Otherwise (the impact of the last update on the ML model is negligible), the FL server abandons the current FL iteration and starts a new iteration by distributing the updated ML model to a newly selected set of FL clients, i.e., jumping to one of steps 2, 3, or 4.
[0083] On the other hand, if the FL server determines in step 6 that there is no remaining time for the ML model training target time to be applied to the consumer AnLF, or if the FL server has determined in step 11 that the ML model performance has reached the target performance, the FL server abandons the FL iteration loop. However, before providing the final result of the FL training process to the consumer AnLF, optional steps 12 - 16 can be performed to consider the interim results from FL clients that have not completed their local ML training processes, e.g., FL clients that have not used their entire local dataset for ML model training. To this end, the FL server can send an early termination request in step 12 to all FL clients that have not provided an ML model update report. The early termination request can prompt the FL client to abandon its ML training process and immediately report the current ML model update (step 13). In step 14, the FL server receives the interim update reports created by the FL clients in response to the early termination request. In step 15, the FL server aggregates all the received update reports and updates the global ML model accordingly.
[0084] In step 16, the FL server can rate the FL clients from which it has received local ML model training reports. Here, the process can be similar to the process described above in connection with step 8. However, in step 17, the FL server can configure NF profile variables related to priority and / or capability parameters for each FL client involved in the FL process using the rating information thus obtained, to reflect the rating in the NRF.
[0085] Finally, in step 18, the FL server provides an updated training notification to the consumer (i.e., AnLF), including the updated version of the ML model.
[0086] [Second Embodiment]
[0087] Figure 4 is a schematic diagram illustrating an asynchronous FL model training process according to a second embodiment of the present invention. The second embodiment is similar to the first embodiment, except that outdated update information is also considered. The second embodiment will be described on the premise that the FL capabilities (i.e., the FL server and FL clients) related to a specific NWDAF including MTLF have been registered in the NRF.
[0088] The process starts with selecting a subset of FL clients and allocating ML model parameters in steps 1 - 5, which are the same as steps 1 - 5 of the process according to the first embodiment described above.
[0089] In step 6, the FL server checks the remaining time for the ML model training target time imposed on the consumer (i.e., AnLF). If there is still time, the FL server waits for at least one FL client to provide an ML model update report (step 7). The report includes an indication of the result of the ML model training process performed by the corresponding FL client on its local data. The result can be provided in the form of a new set of ML model parameters, a set of ML model parameter gradients (weights), or any other suitable form.
[0090] Contrary to the first embodiment, the ML model update report received in step 7 also includes information indicating the version of the ML model on which the ML model training process was performed. This information can be provided, for example, in the form of an iteration ID or any other suitable identifier. The FL server can then use this information to distinguish between outdated and non-outdated updates. The FL server can also distinguish between different outdated updates based on their respective ages or iterations to which they belong.
[0091] In step 8, the FL server updates the ML model based on the ML model update report received from a single FL client, i.e., without waiting for responses from all FL clients.
[0092] When updating the ML model based on the received ML model update report, the FL server also takes into account the ML model version information or iteration ID included in the received ML model update report. Specifically, the FL server can apply weights to the update information to reduce the impact of outdated updates in the globally updated ML model. The FL server can also discard outdated updates that are considered too old, for example, based on the difference between the model version or iteration indicated in the model update report and the current model version or iteration.
[0093] For example, if the iteration ID in the received ML model update report indicates that the FL client used a non-up-to-date ML model, the FL server applies weights to moderate the impact of the outdated information. Otherwise, the FL server uses the ML model update report directly, i.e., without applying weights.
[0094] The FL server can determine the weights empirically or by correlating the reports received from FL clients with the updated ML model and / or by retaining statistical data from previous ML model reports and validation / test results.
[0095] Then in step 9, the ML model updated thereby is verified and / or tested by the FL server to evaluate the ML model performance, particularly to estimate the expected confidence level.
[0096] In step 10, the FL server can rate the local ML model training reports from the FL clients from which they are received. The aforementioned rating can indicate a measure of the impact of the local updates received from that particular client on the accuracy, performance, or confidence of the ML model, and / or the time required for the aforementioned client to complete the training of the ML model on local data. The FL server can save the rating information locally in this case for further use in the current FL iteration.
[0097] In step 11, the FL server checks whether the ML model training progress, particularly the ML model performance, has reached the target performance provided by the consumer (i.e., AnLF).
[0098] If it is determined that the ML model performance has not reached the target performance, the FL server proceeds to step 12, where the FL server provides the updated ML model parameters to the FL client that sent the ML model update report (i.e., the ML model update report received in step 7) used to update the ML model in step 8. Thereafter, a new iteration is started by jumping to step 6, i.e., the FL process is repeated starting from step 6.
[0099] On the other hand, if it is determined that the ML model performance has reached the target performance, the repeating loop is abandoned, and the FL server proceeds to step 13, where the FL server configures the NF profile variables related to the priority and / or capability parameters for each FL client involved in the FL process using the rating information obtained in step 10 to reflect the rating in the NRF.
[0100] Finally, in step 14, the FL server provides an updated training notification to the consumer (i.e., AnLF), including the updated version of the ML model.
[0101] The above embodiments are described only as examples. Combinations and variations are possible within the scope of the appended claims. For example, in terms of sending an early termination request to slow FL clients before providing the updated notification to the consumer to prompt them to send temporary updates ( Figure 3 steps 12 - 15 of the process shown) can also be applied to the second embodiment described in combination with Figure 4 Furthermore, in terms of considering stale updates (or discarding updates considered too stale) by providing information indicating the version or iteration from which the corresponding update was derived to each ML model update report and controlling the impact of each update on the global ML model based on whether the corresponding update is stale ( Figure 4 steps 7 - 8 of the process shown) can also be applied to the first embodiment described in combination with Figure 3 described.
[0102] The above features can be implemented as computer-implemented methods in a mobile telecommunications network. All of the above network functions can be computer functions running on independent computer servers that implement the corresponding functions, or different network functions can share one or more computer servers. The computer servers include corresponding computer hardware, such as at least one processor, at least one memory for storing computer-readable instructions that can be executed by the processor. The computer servers can also include components for communicating with other computers or network devices, such as network interface cards.
Claims
1. A server device for collaborating with multiple clients to train a machine learning (ML) model according to a federated learning scheme, the server device being configured to perform the following steps: Select a subset of multiple clients available for federated learning; Distribute the parameters of the ML model to each of the selected clients to allow the clients to train the ML model on local data; Receive local updates for the ML model from clients that have completed training the ML model on local data; Generate a globally updated ML model by applying the received local updates to the ML model; and Determine the model accuracy of the globally updated ML model, wherein the server device is further configured to at least iterate the steps of distributing, receiving, generating, and determining, and wherein when at least one of one or more predefined iteration conditions is satisfied, a new iteration is started by distributing the parameters of the globally updated ML model, and at least one of the one or more predefined iteration conditions is independent of whether local updates have been received from all of the selected clients.
2. The server device according to claim 1, further configured to: repeat the receiving step, the generating step, and the determining step until the improvement in the model accuracy caused by the last received local update does not exceed a predefined threshold, wherein once the improvement in the model accuracy caused by the last received local update exceeds the predefined threshold, a new iteration is started.
3. The server device according to claim 1, wherein whenever a local update is received from one of the selected clients, a new iteration is started, and the iteration is started by distributing the parameters of the globally updated ML model to the client from which the local update was received.
4. The server device according to claim 1, further configured to: complete the training of the ML model when a target training time limit expires or when the model accuracy of the globally updated ML model reaches a target value.
5. The server device according to claim 1, further configured to: Generate rating information in response to receiving a local update for the ML model from a client, the rating information indicating a measure of the impact of the local update received from the client on the model accuracy and / or the time required for the client to complete the training of the ML model on the local data; and Store the generated rating information in association with the identifier of the client.
6. The server device according to claim 1, further configured to: Obtain rating information for each of the clients available for federated learning, the rating information indicating a measure of the impact of previously received local updates from the client on the model accuracy and / or the time required for the client to complete previous training of the ML model on local data; and Select the subset of the multiple clients available for federated learning at least partially based on the obtained rating information.
7. The server device according to claim 1 is further configured to: For each of the multiple clients available for federated learning, obtain data statistical information, where the data statistical information indicates at least one of a range, quantity, mean, and variability of the local data provided by the client for training the ML model; and Select the subset of the multiple clients available for federated learning at least in part based on the obtained data statistical information.
8. The server device according to claim 1 is further configured to: For each of the multiple clients available for federated learning, obtain data context information, where the data context information indicates the context in which the local data was collected, and the context includes at least one of the following: network conditions, network load conditions, network use cases, radio access technologies, network slices, service IDs, application IDs, and the time when the local data was collected; and Select the subset of the multiple clients available for federated learning at least in part based on the obtained data context information.
9. The server device according to claim 1 is further configured to: Generate target training data information for each of the selected clients, where the target training data information indicates a subset of the local data to be used for training the ML model; and Provide the corresponding target training data information for each of the selected clients so that each client trains the ML model only on the corresponding subset of the local data.
10. The server device according to claim 9, where the local data includes a plurality of data samples, and each data sample includes a plurality of data features, where the target training data information indicates constraints on the data samples and / or the data features to be used for training the ML model.
11. The server device according to claim 8 is further configured to: For each of the selected clients, obtain rating information, where the rating information indicates a measure of the impact of a previously received local update from the client on the model accuracy and / or the time required for the client to complete a previous training of the ML model on local data; and Generate the target training data information at least in part based on the obtained rating information.
12. The server device according to claim 1 is further configured to: If the received local update includes an outdated update from a client that has performed training of the ML model based on stale parameters of the ML model in a previous iteration, apply a weight to the outdated update to reduce the impact of the outdated update on the ML model.
13. The server device according to claim 12 is further configured to: calculate the weight of the stale update based on a correlation between the respective impacts of the stale update and the non-stale update on the ML model, the non-stale update being a local update received from a client that has performed training of the ML model based on the latest parameters of the ML model distributed during the current iteration.
14. The server device according to claim 12 is further configured to: exclude the stale update from generating the global updated ML model if a predetermined condition is satisfied, the predetermined condition including at least one of the following: the amount of time that has passed since the stale parameters were distributed or the number of iterations that have been completed, the number of non-stale updates received within the current iteration, the non-stale update being a local update received from a client that has performed training of the ML model based on the latest parameters of the ML model distributed during the current iteration, the impact of the stale update on the performance of the global updated ML model, and the rating information of the client from which the stale update has been received.
15. The server device according to claim 1 is further configured to: before starting a new iteration or completing the training of the ML model, send an early termination request to a selected one of the clients that has not received a local update, to request the client to stop the training of the ML model and send a temporary update based on the currently available training results; receive the temporary update sent by the client in response to the early termination request; and apply the received temporary update to the ML model.
16. A client device for training a machine learning (ML) model on local data in a federated learning scenario, the client device being configured to: obtain parameters of the ML model and target training data information indicating a subset of the local data to be used for training the ML model from a server; perform a training process on the ML model based on the subset of the local data indicated by the obtained target training data information; generate a local update for the ML model based on the result of the training process; and send the local update to the server.
17. The client device according to claim 16, wherein the local data includes a plurality of data samples, each data sample including a plurality of data features, wherein the target training data information indicates constraints on the data samples and / or the data features to be used in the training process.
18. The client device according to claim 16 is further configured to: receive an early termination request from the server; and in response to receiving the early termination request, stop the training process of the ML model, generate a temporary update for the ML model based on the currently available training results, and send the temporary update to the server.
19. The client device according to claim 16, further configured to: send the local update to the server together with information indicating a version of the ML model for which the training process has been performed on it.
20. A method for collaboratively training a machine learning (ML) model with multiple clients according to a federated learning scheme, the method comprising the steps of: selecting a subset of multiple clients available for federated learning; distributing parameters of the ML model to each of the selected clients to allow the clients to train the ML model on local data; receiving local updates for the ML model from clients that have completed training of the ML model on local data; generating a globally updated ML model by applying at least one of the received local updates to the ML model; determining a model accuracy of the globally updated ML model; and iterating at least the steps of distributing, receiving, generating, and determining, wherein when at least one of one or more predefined iteration conditions is satisfied, a new iteration is started by distributing parameters of the globally updated ML model, and at least one of the one or more predefined iteration conditions is independent of whether local updates have been received from all of the selected clients.