Avoiding interruptions during model training when using federated learning

The method addresses interruptions in federated learning by selecting and configuring alternative devices for training based on data distribution and communication quality, enhancing training efficiency and resilience.

JP2025528082APending Publication Date: 2025-08-26NOKIA TECHNOLOGIES OY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025505993
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-08-04
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Federated learning systems experience interruptions during model training due to device failures, resource limitations, or communication quality issues, leading to inefficiencies and incomplete training processes.

Method used

A method for selecting and configuring alternative devices for model training based on data distribution similarity, location, proximity, and communication quality, with mechanisms for device state identification and dynamic reconfiguration to ensure continuous training.

Benefits of technology

Ensures uninterrupted and efficient model training by leveraging alternative devices, improving training resilience and reducing interruptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025528082000001_ABST
    Figure 2025528082000001_ABST
Patent Text Reader

Abstract

An apparatus configured to train a model in a communications network using federated learning, the apparatus comprising: means for selecting at least two further devices for training a local model; means for selecting at least one alternative device of the at least two selected further devices; means for configuring each of the at least two further devices for training the local model and for configuring at least one alternative device of the two selected further devices for training the local model; means for receiving local training results from at least one of the at least two further devices and from at least one alternative device of the two selected further devices; and means for combining the local training results to generate an aggregated training result for the model.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present application relates to methods, apparatus, systems, and computer programs for performing training of models using federated learning, and in particular, but not exclusively, to methods, apparatus, systems, and computer programs for avoiding interruptions during training of models using federated learning. [Background technology]

[0002] A communication system can be thought of as a facility that enables communication between two or more entities, such as terminals and / or other nodes, or provides connected services to the entities. A communication system can include a communication network and one or more compatible terminals (also known as communication devices). The communications may carry, for example, voice, video, electronic mail (email), text messages, multimedia data, and / or content data, etc. Non-limiting examples of connected services provided by a communication system include enhanced mobile broadband, ultra-reliable low-latency communications, mission-critical communications, massive Internet of Things (IoT), and multimedia services.

[0003] In a communication system, at least a portion of the communication between at least two entities occurs over wireless links. Examples of networks in a communication system include radio access networks, such as public land mobile networks (PLMNs), terrestrial radio access networks or non-terrestrial radio access networks (e.g., satellite networks), and various wireless local networks, such as wireless local area networks (WLANs). Radio access networks can include cells and are therefore often referred to as cellular networks.

[0004] A terminal may be referred to as user equipment (UE) or user device. A terminal is equipped with appropriate signal receiving and transmitting equipment to enable communication, e.g., to enable access to a communication network or direct communication with other terminals. A terminal may access a carrier provided by a base station, e.g., a base station of a radio access network, and transmit and / or receive communications on that carrier.

[0005] A communication system and associated compatible terminals typically operate according to a given standard or specification that specifies what is permitted for various network entities of the communication system and how this should be achieved. The communication protocols and / or parameters used for communication are also typically defined. One example of a communication system is a Universal Mobile Telecommunications System (UMTS) system (e.g., a communication system using 3G radio access technology). Other examples of communication systems include so-called 4G systems (e.g., communication systems operating using 4G radio access technology) and 5G or New Radio (NR) systems (e.g., communication systems operating using 5G or NR radio access technology). Radio access technologies used by communication systems are standardized by the 3rd Generation Partnership Project (3GPP). [Prior art documents] [Non-patent literature]

[0006] [Non-Patent Document 1] TR23.501 [Non-patent document 2] TR23.502 [Non-patent document 3] Konecny, Jakub, H.B. McMahan, and Daniel Ramage, "Federated Optimization: Distributed Optimization Beyond the Datacenter," ArXiv abs / 1511.03575 (2015). Summary of the Invention [Problem to be solved by the invention]

[0007] According to one aspect, an apparatus configured to train a model in a communication network using federated learning is provided, the apparatus comprising: means for selecting at least two further devices for training a local model; means for selecting an alternative device for at least one of the at least two selected further devices; means for configuring each of the at least two further devices for training the local model and for configuring an alternative device for at least one of the two selected further devices for training the local model; means for receiving local training results from at least one of the at least two further devices and local training results from an alternative device for at least one of the two selected further devices; and means for combining the local training results to generate an aggregated training result for the model.

[0008] The means for further selecting at least one alternative device of the at least two selected further devices may be for selecting the alternative device based on information indicative of at least one of: a similarity between a data distribution of data in a local dataset of the at least one further device and a data distribution of data in a local dataset of the alternative device; a location of the further device; a location of the alternative device; a proximity between the further device and the alternative device; a mobility pattern of the alternative device relative to the further device; a quality of communication on a sidelink between the further device and the alternative device; and at least one characteristic of a wireless link between the further device and a base station of the radio access network.

[0009] The means may be means for receiving information indicative of one or more candidate alternative devices from the at least one further device, and the means for selecting an alternative device for at least one of the at least two selected further devices may be for selecting the alternative device from the one or more candidate alternative devices identified by the further device.

[0010] The means may further be for generating and transmitting an FL report configuration to each of the at least two further devices, the FL report configuration comprising an indicator that enables the two further devices to generate an FL report comprising information identifying one or more potential alternate devices.

[0011] The means for configuring each of the at least two further devices for training a local model on the at least two further devices and for configuring each alternative device for training a local model on the alternative device may be for generating an alternative training UE configuration for at least one of the at least two further devices and the alternative device, the alternative training UE configuration comprising at least one of: a further device identifier configured to uniquely identify at least one of the at least two further devices; an alternative device identifier configured to uniquely identify the alternative further device; and a state identifier configured to identify a state in which at least one of the at least two further devices is unable to train the local model, causing the alternative device to run the local training model.

[0012] The conditions may comprise at least one of a minimum quality of the Uu link between the further device and a base station of the radio access network, a minimum availability of computational resources in the further device, a minimum availability of power resources in the further device, and a minimum security / integrity level associated with the local dataset of the further device.

[0013] The means for obtaining local training results from the at least two further devices may further be means for receiving an indicator from the alternative device when at least one of the at least two further devices is unable to train a local model, the indicator identifying that the trained local model of the local training results should be used as a replacement for the trained local model of the local training results.

[0014] The means for configuring each of the at least two further devices for training the local model and for configuring an alternative device for at least one of the two selected further devices for training the local model may be means for generating a global model and a training configuration for the at least two selected further devices and the alternative device, and the training of the local model is based on the global model and the training configuration.

[0015] The means may further be for receiving an indication from at least one of the at least two further devices that at least one of the at least two further devices is unable to train the local model, and for generating a request for an alternative device to perform the local model training.

[0016] The request may comprise at least one of an indicator indicating a cause for which at least one of the at least two further devices is unable to train the local model and a time indicator indicating a time when the alternative device will perform the local model training.

[0017] The means for configuring each of the at least two further devices for training the local model and for configuring at least one alternative device of the two selected further devices for training the local model may further include means for receiving an approved or rejected alternative training UE configuration from at least one of the at least two further devices, means for receiving an approved or rejected alternative training UE configuration from the alternative device, and means for reselecting and reconfiguring at least one further alternative device of the at least two further devices based on receiving the at least one rejected alternative training configuration from at least one of the at least two further devices or the alternative device.

[0018] The device may be any of: a base station of a radio access network, wherein the at least two further devices and the alternative devices are user equipments; a Network Data Analytics entity, wherein the at least two further devices and the alternative devices are distributed Network Data Analytics entities; and an Operations, Administration and Maintenance entity, wherein the at least two further devices and the alternative devices are base stations.

[0019] According to a second aspect, an apparatus configured to train a local model during federated learning is provided, the apparatus comprising: means for receiving an alternative training UE configuration from a further apparatus configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: an apparatus identifier configured to uniquely identify the apparatus for training the local model; an alternative apparatus identifier configured to uniquely identify an alternative apparatus for the apparatus; and a state identifier configured to identify a state in which the apparatus is unable to train the local model, causing the alternative apparatus to train the local model on the alternative apparatus; and means for training the local model and sending the local training results to the further apparatus, or determining that the apparatus is unable to train the local model based on the state in which the apparatus is unable to train the local model and sending a local model training request to one of the further apparatus or the alternative apparatus to cause the alternative apparatus to perform local training on the alternative apparatus using a local dataset.

[0020] The means may further be for generating information indicative of the candidate alternative devices based on information indicative of at least one of a data distribution of the data in the local dataset at the device and the data distribution of the data in the local dataset at the candidate alternative devices, a dissemination / distribution of local data of the device and the candidate alternative devices, a range of data in the local dataset at the device and the candidate alternative devices, an interquartile range of data in the local dataset at the device and the candidate alternative devices, a standard deviation of data in the local dataset at the device and the candidate alternative devices, a variance of data in the local dataset at the device and the candidate alternative devices, a proximity between the device and the candidate alternative devices, and a mobility pattern between the device and the candidate alternative devices.

[0021] The means may further be for receiving a request from the further device and generating information indicative of alternative candidate devices.

[0022] The at least one condition may comprise at least one of a minimum quality of the Uu link between the device and a base station of the radio access network, a minimum availability of computational resources in the device, a minimum availability of power resources in the device, and a minimum security / integrity level associated with the local dataset of the further device.

[0023] The local model training request may comprise an indicator identifying the condition that caused the device to be unable to train the local model and a time indicator indicating when the alternate device will train the local model.

[0024] The means may further be for generating an approval or rejection of the alternative training UE configuration to the further device, and the further device may be caused to reselect and reconfigure the further alternative device to the device.

[0025] The device may be a user equipment, the alternative device may be a user equipment and the further device may be a base station of a radio access network.

[0026] The apparatus may be a wireless communication device, the alternative apparatus may be a wireless communication device, and the further apparatus may be a base station of a radio access network.

[0027] The device may be a distributed Network Data Analytics entity, the further device may be a centralized Network Data Analytics entity, and the alternative device may be a distributed Network Data Analytics entity.

[0028] The device may be a base station of a radio access network, the further device may be an Operations, Administration and Maintenance entity, and the alternative device may be a base station of a radio access network.

[0029] The device may be an open radio access network function, the further device may be an open radio access network function, and the alternative device may be an open radio access network function.

[0030] According to a third aspect, there is provided an apparatus configured to train a local model for federated learning, the apparatus comprising: means for receiving an alternative training UE configuration from a further apparatus configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: an apparatus identifier configured to uniquely identify the other apparatus as an apparatus for training the local model using a local dataset; an alternative apparatus identifier configured to uniquely identify the apparatus as an alternative training apparatus for the other apparatus; and a state identifier configured to identify a state in which the other apparatus is unable to train the local model, the state identifier causing the apparatus to train the local model at the apparatus; means for receiving a local model training request from the other apparatus or the further apparatus when the other apparatus is unable to train the local model; and means for training the local model using the local dataset and transmitting the training results to the further apparatus after receiving the request.

[0031] The local model training request may comprise at least one of an indicator identifying a condition that prevents another device from training the local model and a time indicator indicating a time during which the device will train the local model.

[0032] The means may further be for generating an approval or denial message to the further device, and the further device may be caused to reselect and reconfigure the further alternative device relative to another device.

[0033] The means for training the local model and transmitting the trained local model to the further device after receiving the request may further be for transmitting an indicator to identify that the update to the parameters of the trained local model is to be used as an alternative update to the parameters of the trained local model.

[0034] The device may be a user equipment, another device may be a user equipment and the further device may be a base station of a radio access network.

[0035] The apparatus may be a wireless communication device, another apparatus may be a wireless communication device, and the further apparatus may be a base station of a radio access network.

[0036] A device may be a distributed Network Data Analytics entity, a further device may be a centralized Network Data Analytics entity, and another device may be a distributed Network Data Analytics entity.

[0037] The device may be a base station of a radio access network, a further device may be an Operations, Administration and Maintenance entity, and another device may be a base station of a radio access network.

[0038] The device may be an open radio access network entity, the further device may be an open radio access network entity, and the other device may be an open radio access network entity.

[0039] According to a fourth aspect, there is provided a method for an apparatus configured to train a model in a communication network using federated learning, the method comprising the steps of selecting at least two further apparatuses for training a local model, selecting an alternative apparatus for at least one of the at least two selected further apparatuses, configuring each of the at least two further apparatuses for training the local model and configuring an alternative apparatus for at least one of the two selected further apparatuses for training the local model, receiving local training results from at least one of the at least two further apparatuses and from the alternative apparatus for at least one of the two selected further apparatuses, and combining the local training results to generate an aggregated training result for the model.

[0040] The step of selecting at least one alternative device of the at least two selected further devices may comprise selecting the alternative device based on information indicative of at least one of the following: a similarity between a data distribution of data in the local dataset of the at least one further device and a data distribution of data in the local dataset of the alternative device; a location of the further device; a location of the alternative device; a proximity between the further device and the alternative device; a mobility pattern of the alternative device relative to the further device; a quality of communication on a sidelink between the further device and the alternative device; and at least one characteristic of a Uu link between the further device and a base station of the radio access network.

[0041] The method may comprise receiving information indicating one or more potential alternative devices from the at least one further device, and selecting an alternative device for at least one of the at least two selected further devices may comprise selecting an alternative device from the potential alternative devices identified by the further device.

[0042] The method may further comprise generating and transmitting an FL report configuration to each of the at least two further devices, the FL report configuration comprising an indicator that enables the two further devices to generate an FL report comprising information identifying one or more potential alternate devices.

[0043] The step of configuring each of the at least two further devices to train a local model on the at least two further devices and configuring each alternative device to train a local model on the alternative device may comprise generating an alternative training UE configuration for at least one of the at least two further devices and the alternative device, the alternative training UE configuration comprising at least one of: a further device identifier configured to uniquely identify at least one of the at least two further devices; an alternative device identifier configured to uniquely identify the alternative further device; and a state identifier configured to identify a state in which at least one of the at least two further devices is unable to train a local model, causing the alternative device to train the local model on the alternative device.

[0044] The conditions may comprise at least one of a minimum quality of the Uu link between the further device and a base station of the radio access network, a minimum availability of computational resources in the further device, a minimum availability of power resources in the further device, and a minimum security / integrity level associated with the local dataset of the further device.

[0045] The step of obtaining local training results from the at least two further devices may further comprise the step of receiving an indicator from the alternative device that identifies the trained local model of the local training results as being to be used as an alternative to the trained local model of the local training results when at least one of the at least two further devices is unable to train a local model.

[0046] The steps of configuring each of the at least two further devices for training the local model and configuring an alternative device for at least one of the two selected further devices for training the local model may further comprise generating a global model and training configuration for the at least two selected further devices and the alternative device, wherein training of the local model is based on the global model and training configuration.

[0047] The method may further comprise receiving an indication from at least one of the at least two further devices that at least one of the at least two further devices is unable to train the local model, and generating a request to an alternative device to have the local model trained on the alternative device.

[0048] The request may comprise at least one of an indicator indicating a cause for which at least one of the at least two further devices is unable to train the local model and a time indicator indicating a time for the alternative device to train the local model training on the alternative device.

[0049] The steps of configuring each of the at least two further devices for training the local model and configuring an alternative device for at least one of the two selected further devices for training the local model may further comprise the steps of receiving an approved or rejected alternative training UE configuration from at least one of the at least two further devices, receiving an approved or rejected alternative training UE configuration from the alternative device, and reselecting and reconfiguring a further alternative device for at least one of the at least two further devices based on receiving the at least one rejected alternative training UE configuration from at least one of the at least two further devices or the alternative device.

[0050] The device may be any of: a base station of a radio access network, wherein the at least two further devices and the alternative devices are user equipments; a Network Data Analytics entity, wherein the at least two further devices and the alternative devices are distributed Network Data Analytics entities; and an Operations, Administration and Maintenance entity, wherein the at least two further devices and the alternative devices are base stations.

[0051] The device may be an open wireless access network application function, and at least two further devices and an alternative device are open wireless access network applications.

[0052] According to a fifth aspect, there is provided a method for an apparatus configured to train a local model during federated learning, the method comprising: receiving an alternative training UE configuration from a further apparatus configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: an apparatus identifier configured to uniquely identify the apparatus for training the local model; an alternative apparatus identifier configured to uniquely identify an alternative apparatus for the apparatus; and a state identifier configured to identify a state in which the apparatus is unable to train the local model, the state identifier causing the alternative apparatus to train the local model on the alternative apparatus; and training the local model and sending the local training results to the further apparatus, or determining that the apparatus is unable to train the local model based on the state in which the apparatus is unable to train the local model and sending a local model training request to one of the further apparatus or the alternative apparatus to cause the alternative apparatus to perform local training on the alternative apparatus using a local dataset.

[0053] The method may further comprise generating information indicative of the candidate alternate devices based on information indicative of at least one of a data distribution of the data in the local dataset of the device and the data distribution of the data in the local dataset of the candidate alternate devices, a dissemination / distribution of the local data of the device and the candidate alternate devices, a range of the data in the local dataset at the device and the candidate alternate devices, an interquartile range of the data in the local dataset at the device and the candidate alternate devices, a standard deviation of the data in the local dataset at the device and the candidate alternate devices, a variance of the data in the local dataset at the device and the candidate alternate devices, a proximity between the device and the candidate alternate devices, and a mobility pattern between the device and the candidate alternate devices.

[0054] The method may further comprise receiving a request from the further device to generate information indicative of alternative candidate devices.

[0055] The at least one condition may comprise at least one of a minimum quality of the Uu link between the device and a base station of the radio access network, a minimum availability of computational resources in the device, a minimum availability of power resources in the device, and a minimum security / integrity level associated with the local dataset of the further device.

[0056] The local model training request may comprise an indicator identifying the condition that caused the device to be unable to train the local model and a time indicator indicating when the alternate device will train the local model.

[0057] The method may further comprise generating an acceptance or rejection of the alternative training UE configuration to the further device, and the further device may be caused to reselect and reconfigure the further alternative device to the device.

[0058] The device may be a user equipment, the alternative device may be a user equipment and the further device may be a base station of a radio access network.

[0059] The apparatus may be a wireless communication device, the alternative apparatus may be a wireless communication device, and the further apparatus may be a base station of a radio access network.

[0060] The device may be a distributed Network Data Analytics entity, the further device may be a centralized Network Data Analytics entity, and the alternative device may be a distributed Network Data Analytics entity.

[0061] The device may be a base station of a radio access network, the further device may be an Operations, Administration and Maintenance entity, and the alternative device may be a base station of a radio access network.

[0062] The device may be an open radio access network function, the further device may be an open radio access network function, and the alternative device may be an open radio access network function.

[0063] According to a sixth aspect, there is provided a method for an apparatus configured to train a local model for federated learning, the method comprising: receiving an alternative training UE configuration from a further apparatus configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: an apparatus identifier configured to uniquely identify the other apparatus as an apparatus for training the local model using a local dataset; an alternative apparatus identifier configured to uniquely identify the apparatus as an alternative training apparatus for the other apparatus; and a state identifier configured to identify a state in which the other apparatus is unable to train the local model, the state identifier causing the apparatus to train the local model at the apparatus; receiving a local model training request from the other apparatus or the further apparatus when the other apparatus is unable to train the local model; and after receiving the request, training the local model using the local dataset and transmitting the training results to the further apparatus.

[0064] The local model training request may comprise at least one of an indicator identifying a condition that prevents another device from training the local model and a time indicator indicating a time during which the device will train the local model.

[0065] The method may further comprise the step of generating an acceptance or rejection message to the further device, and the further device may be caused to reselect and reconfigure the further alternative device relative to the other device.

[0066] After receiving the request, the step of training the local model and transmitting the trained local model to the further device may further comprise the step of transmitting an indicator to identify that the update to the parameters of the trained local model is to be used as an alternative update to the parameters of the trained local model.

[0067] The device may be a user equipment, another device may be a user equipment and the further device may be a base station of a radio access network.

[0068] The apparatus may be a wireless communication device, another apparatus may be a wireless communication device, and the further apparatus may be a base station of a radio access network.

[0069] A device may be a distributed Network Data Analytics entity, a further device may be a centralized Network Data Analytics entity, and another device may be a distributed Network Data Analytics entity.

[0070] The device may be a base station of a radio access network, a further device may be an Operations, Administration and Maintenance entity, and another device may be a base station of a radio access network.

[0071] The device may be an open radio access network entity, the further device may be an open radio access network entity, and the other device may be an open radio access network entity.

[0072] According to a seventh aspect, there is provided an apparatus configured to train a model in a communications network using federated learning, the apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: select at least two further devices for training a local model; select an alternative device for at least one of the at least two selected further devices; configure each of the at least two further devices for training the local model and configure an alternative device for at least one of the two selected further devices for training the local model; receive local training results from at least one of the at least two further devices and from the alternative device for at least one of the two selected further devices; and combine the local training results to generate an aggregated training result for the model.

[0073] The device adapted to select an alternative device for at least one of the at least two selected further devices may further be adapted to select the alternative device based on information indicative of at least one of the following: a similarity between the data distribution of data in the local dataset of the at least one further device and the data distribution of data in the local dataset of the alternative device; a location of the further device; a location of the alternative device; a proximity between the further device and the alternative device; a mobility pattern of the alternative device relative to the further device; a quality of communication on a sidelink between the further device and the alternative device; and at least one characteristic of a Uu link between the further device and a base station of the radio access network.

[0074] The device may be further adapted to receive information indicating one or more potential alternative devices from the at least one further device, and the device adapted to select an alternative device for at least one of the at least two selected further devices may be adapted to select the alternative device from the potential alternative devices identified by the further device.

[0075] The device may be further adapted to generate and transmit an FL report configuration to each of the at least two further devices, the FL report configuration comprising an indicator that enables the two further devices to generate an FL report comprising information identifying one or more potential alternate devices.

[0076] An apparatus configured to configure each of the at least two further devices to train a local model on the at least two further devices and to configure each alternate device to train a local model on the alternate device may be configured to generate an alternate training configuration for at least one of the at least two further devices and the alternate device, the alternate training configuration comprising at least one of: a further device identifier configured to uniquely identify at least one of the at least two further devices; an alternate device identifier configured to uniquely identify the alternate further device; and a state identifier configured to identify a state in which at least one of the at least two further devices is unable to train a local model, causing the alternate device to train the local model on the alternate device.

[0077] The conditions may comprise at least one of a minimum quality of the Uu link between the further device and a base station of the radio access network, a minimum availability of computational resources in the further device, a minimum availability of power resources in the further device, and a minimum security / integrity level associated with the local dataset of the further device.

[0078] The device adapted to obtain local training results from the at least two further devices may be further adapted to receive an indicator from the alternative device that identifies the trained local model of the local training results as being to be used as a replacement for the trained local model of the local training results if at least one of the at least two further devices is unable to train a local model.

[0079] The device adapted to configure each of the at least two further devices for training a local model and to configure an alternative device of at least one of the two selected further devices for training a local model may be further adapted to generate a global model and a training configuration for the at least two selected further devices and the alternative device, wherein the training of the local model is based on the global model and the training configuration.

[0080] The device may be further configured to receive an indication from at least one of the at least two further devices that at least one of the at least two further devices is unable to train the local model, and to generate a request to an alternative device to train the local model on the alternative device.

[0081] The request may comprise at least one of an indicator indicating a cause for which at least one of the at least two further devices is unable to train the local model and a time indicator indicating a time for the alternative device to train the local model training on the alternative device.

[0082] The device adapted to configure each of the at least two further devices for training the local model and to configure at least one alternative device of the two selected further devices for training the local model may be further adapted to receive an approval or rejection of the alternative training configuration from at least one of the at least two further devices, receive an approval or rejection of the alternative training configuration from the alternative device, and reselect and reconfigure at least one further alternative device of the at least two further devices based on receiving the at least one rejection of the alternative training configuration from at least one of the at least two further devices or the alternative device.

[0083] The device may be any of: a base station of a radio access network, wherein the at least two further devices and the alternative devices are user equipments; a Network Data Analytics entity, wherein the at least two further devices and the alternative devices are distributed Network Data Analytics entities; and an Operations, Administration and Maintenance entity, wherein the at least two further devices and the alternative devices are base stations.

[0084] The device may be an open wireless access network application function, and at least two further devices and an alternative device are open wireless access network applications.

[0085] According to an eighth aspect, there is provided an apparatus configured to train a local model during federated learning, the apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: receive an alternative training configuration from a further device configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: a device identifier configured to uniquely identify the device for training the local model; an alternative device identifier configured to uniquely identify an alternative device for the apparatus; and a state identifier configured to identify a state in which the apparatus is unable to train the local model, causing the alternative device to train the local model on the alternative device; and train the local model and send the local training results to the further device; or determine that the apparatus is unable to train the local model based on the state in which the apparatus is unable to train the local model and send a local model training request to one of the further device or the alternative device to cause the alternative device to perform local training on the alternative device using the local dataset.

[0086] The device may further be configured to generate information indicative of the candidate alternate devices based on information indicative of at least one of: a data distribution of the data in the device's local dataset and the data distribution of the data in the local datasets of the candidate alternate devices; a dissemination / distribution of the local data of the device and the candidate alternate devices; a range of the data in the local datasets at the device and the candidate alternate devices; an interquartile range of the data in the local datasets at the device and the candidate alternate devices; a standard deviation of the data in the local datasets at the device and the candidate alternate devices; a variance of the data in the local datasets at the device and the candidate alternate devices; a proximity between the device and the candidate alternate devices; and a mobility pattern between the device and the candidate alternate devices.

[0087] The device may further be adapted to receive requests from further devices to generate information indicative of alternative candidate devices.

[0088] The at least one condition may comprise at least one of a minimum quality of the Uu link between the device and a base station of the radio access network, a minimum availability of computational resources in the device, a minimum availability of power resources in the device, and a minimum security / integrity level associated with the local dataset of the further device.

[0089] The local model training request may comprise an indicator identifying the condition that caused the device to be unable to train the local model and a time indicator indicating when the alternate device will train the local model.

[0090] The device may further be adapted to generate an approval or rejection of the alternative training configuration to the further device, and the further device may be adapted to reselect and reconfigure further alternative devices for the device.

[0091] The device may be a user equipment, the alternative device may be a user equipment and the further device may be a base station of a radio access network.

[0092] The apparatus may be a wireless communication device, the alternative apparatus may be a wireless communication device, and the further apparatus may be a base station of a radio access network.

[0093] The device may be a distributed Network Data Analytics entity, the further device may be a centralized Network Data Analytics entity, and the alternative device may be a distributed Network Data Analytics entity.

[0094] The device may be a base station of a radio access network, the further device may be an Operations, Administration and Maintenance entity, and the alternative device may be a base station of a radio access network.

[0095] The device may be an open radio access network function, the further device may be an open radio access network function, and the alternative device may be an open radio access network function.

[0096] According to a ninth aspect, there is provided an apparatus configured to train a local model for federated learning, the apparatus comprising: at least one processor; and at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus to at least: receive an alternative training configuration from a further device configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: a device identifier configured to uniquely identify the other device as a device for training the local model using the local dataset; an alternative device identifier configured to uniquely identify the device as an alternative training device for the other device; and a state identifier configured to identify a state in which the other device is unable to train the local model, causing the apparatus to train the local model at the device; receive a local model training request from the other device or the further device when the other device is unable to train the local model; and, after receiving the request, train the local model using the local dataset and send the training results to the further device.

[0097] The local model training request may comprise at least one of an indicator identifying a condition that prevents another device from training the local model and a time indicator indicating a time during which the device will train the local model.

[0098] The device may further be adapted to generate an accept or reject alternate training configuration message to the further device, and the further device may be adapted to reselect and reconfigure the further alternate device relative to another device.

[0099] After receiving the request, the device adapted to train a local model and transmit the trained local model to the further device may be further adapted to transmit an indicator to identify that the update to the parameters of the trained local model is to be used as an alternative update to the parameters of the trained local model.

[0100] The device may be a user equipment, another device may be a user equipment and the further device may be a base station of a radio access network.

[0101] The apparatus may be a wireless communication device, another apparatus may be a wireless communication device, and the further apparatus may be a base station of a radio access network.

[0102] A device may be a distributed network data analytics entity, a further device may be a centralized Network Data Analytics entity, and another device may be a distributed Network Data Analytics entity.

[0103] The device may be a base station of a radio access network, a further device may be an Operations, Administration and Maintenance entity, and another device may be a base station of a radio access network.

[0104] The device may be an open radio access network entity, the further device may be an open radio access network entity, and the other device may be an open radio access network entity.

[0105] According to a tenth aspect, there is provided an apparatus configured to train a model in a communication network using federated learning, the apparatus comprising: means for selecting at least two further devices for training a local model; means for selecting an alternative device for at least one of the at least two selected further devices; means for configuring each of the at least two further devices for training the local model and for configuring an alternative device for at least one of the two selected further devices for training the local model; means for receiving local training results from at least one of the at least two further devices and local training results from an alternative device for at least one of the two selected further devices; and means for combining the local training results to generate an aggregated training result for the model.

[0106] According to an eleventh aspect, there is provided an apparatus configured to train a local model during federated learning, comprising: means for receiving an alternative training configuration from a further device configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: an apparatus identifier configured to uniquely identify the apparatus for training the local model; an alternative device identifier configured to uniquely identify an alternative device for the apparatus; and a state identifier configured to identify a trigger condition under which the apparatus is unable to train the local model, the state identifier causing the alternative device to perform the local model training; and means for training the local model and sending the local training results to the further device, or determining that the apparatus is unable to train the local model based on the condition under which the apparatus is unable to train the local model and sending a local model training request to one of the further device or the alternative devices to cause the alternative device to perform the local model training.

[0107] According to a twelfth aspect, there is provided an apparatus configured to train a local model for federated learning, comprising: means for receiving an alternative training configuration from a further device configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: a device identifier configured to uniquely identify the other device as an apparatus for training the local model using a local dataset; an alternative device identifier configured to uniquely identify the apparatus as an alternative training device for the other device; and a state identifier configured to identify a state in which the other device is unable to train the local model, causing the apparatus to train the local model at the apparatus; means for receiving a local model training request from the other or further device to train the local model if the other device is unable to train the local model; and means for, after receiving the request, training the local model using the local dataset and transmitting the local training results to the further device.

[0108] According to a thirteenth aspect, there is provided a computer program (or a computer-readable medium comprising program instructions) comprising instructions to cause an apparatus configured to train a model in a communications network using federated learning to at least select at least two further apparatuses for training a local model; select an alternative apparatus for at least one of the at least two selected further apparatuses; configure each of the at least two further apparatuses for training the local model and configure an alternative apparatus of at least one of the two selected further apparatuses for training the local model; receive local training results from at least one of the at least two further apparatuses and from the alternative apparatus of at least one of the two selected further apparatuses; and combine the local training results to generate an aggregated training result of the model.

[0109] According to a fourteenth aspect, there is provided a computer program (or a computer-readable medium comprising program instructions) comprising instructions to cause an apparatus configured to train a local model during federated learning to receive an alternative training configuration from at least a further apparatus configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: an apparatus identifier configured to uniquely identify the apparatus for training the local model; an alternative apparatus identifier configured to uniquely identify an alternative apparatus for the apparatus; and a state identifier configured to identify a state in which the apparatus is unable to train the local model, the state identifier causing the alternative apparatus to train the local model on the alternative apparatus; and to train the local model and send the local training results to the further apparatus, or determine that the apparatus is unable to train the local model based on the state in which the apparatus is unable to train the local model and send a local model training request to one of the further apparatus or the alternative apparatus to cause the alternative apparatus to perform local training on the alternative apparatus using the local dataset.

[0110] According to a fifteenth aspect, there is provided a computer program (or a computer-readable medium comprising program instructions) comprising instructions to cause an apparatus configured to train a local model for federated learning to receive an alternative training UE configuration from at least a further apparatus configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: an apparatus identifier configured to uniquely identify the other apparatus as an apparatus for training the local model using a local dataset; an alternative apparatus identifier configured to uniquely identify the apparatus as an alternative training apparatus for the other apparatus; and a state identifier configured to identify a state in which the other apparatus is unable to train the local model, the state identifier causing the apparatus to train the local model at the apparatus; receive a local model training request from the other apparatus or the further apparatus to train the local model when the other apparatus is unable to train the local model; and, after receiving the request, train the local model using the local dataset and send the local training results to the further apparatus.

[0111] According to a sixteenth aspect, there is provided a non-transitory computer-readable medium comprising program instructions to cause an apparatus configured to train a model in a communications network using federated learning to at least select at least two further apparatuses for training a local model; select an alternative apparatus for at least one of the at least two selected further apparatuses; configure each of the at least two further apparatuses for training the local model and configure an alternative apparatus of at least one of the two selected further apparatuses for training the local model; receive local training results from at least one of the at least two further apparatuses and from the alternative apparatus of the at least one of the two selected further apparatuses; and combine the local training results to generate an aggregated training result of the model.

[0112] According to a seventeenth aspect, there is provided a non-transitory computer-readable medium comprising program instructions to cause an apparatus configured to train a local model during federated learning to receive an alternative training configuration from a further apparatus configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: an apparatus identifier configured to uniquely identify the apparatus for training the local model; an alternative apparatus identifier configured to uniquely identify an alternative apparatus for the apparatus; and a state identifier configured to identify a state in which the apparatus is unable to train the local model, the state identifier causing the alternative apparatus to train the local model on the alternative apparatus; and train the local model and send the local training results to the further apparatus; or determine that the apparatus is unable to train the local model based on the state in which the apparatus is unable to train the local model and send a local model training request to one of the further apparatus or the alternative apparatus, causing the alternative apparatus to perform local training on the alternative apparatus using the local dataset.

[0113] According to an eighteenth aspect, there is provided a non-transitory computer-readable medium comprising program instructions to cause an apparatus configured to train a local model for federated learning to receive at least an alternative training configuration from a further apparatus configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: an apparatus identifier configured to uniquely identify the other apparatus as an apparatus for training the local model using the local dataset; an alternative apparatus identifier configured to uniquely identify the apparatus as an alternative training apparatus for the other apparatus; and a state identifier configured to identify a state in which the other apparatus is unable to train the local model, the state identifier causing the apparatus to train the local model at the apparatus; receiving a local model training request from the other apparatus or the further apparatus to train the local model if the other apparatus is unable to train the local model; and, after receiving the request, training the local model using the local dataset and transmitting the local training results to the further apparatus.

[0114] According to a 19th aspect, there is provided an apparatus configured to train a model in a communication network using federated learning, the apparatus comprising: a selection circuit configured to select at least two further devices for training a local model; a selection circuit configured to select an alternative device for at least one of the at least two selected further devices; a selection circuit configured to configure each of the at least two further devices for training the local model and to configure an alternative device for at least one of the two selected further devices for training the local model; a selection circuit configured to receive local training results from at least one of the at least two further devices and local training results from an alternative device for at least one of the two selected further devices; and a selection circuit configured to combine the local training results to generate an aggregated training result for the model.

[0115] According to a twentieth aspect, there is provided an apparatus configured to train a local model during federated learning, the apparatus comprising: a receiving circuit configured to receive an alternative training configuration from a further device configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: an apparatus identifier configured to uniquely identify the apparatus for training the local model; an alternative device identifier configured to uniquely identify an alternative device for the apparatus; and a state identifier configured to identify a state in which the apparatus is unable to train the local model, the state identifier causing the alternative device to train the local model on the alternative device; a training circuit configured to train the local model; a control circuit configured to control sending of the local training results to the further device; a decision circuit configured to determine that the apparatus is unable to train the local model based on the state in which the apparatus is unable to train the local model; and a control circuit configured to control sending a local model training request to one of the further device or the alternative device, and to cause the alternative device to perform local model training on the alternative device using a local dataset.

[0116] According to a twenty-first aspect, there is provided an apparatus configured to train a local model for federated learning, the apparatus comprising: a receiving circuit configured to receive an alternative training configuration from a further device configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: a device identifier configured to uniquely identify the other device as an apparatus for training the local model using a local dataset; an alternative device identifier configured to uniquely identify the apparatus as an alternative training device for the other device; and a state identifier configured to identify a state in which the other device is unable to train the local model, causing the apparatus to train the local model at the apparatus; a receiving circuit configured to receive a local model training request from the other or further device when the other device is unable to train the local model; a training circuit configured to train the local model using the local dataset after receiving the local model training request; and a control circuit configured to control transmission of the local training results to the further device.

[0117] According to a twenty-second aspect, there is provided a computer-readable medium comprising instructions to cause an apparatus configured to train a model in a communications network using federated learning to at least select at least two further apparatuses for training a local model; select an alternative apparatus for at least one of the at least two selected further apparatuses; configure each of the at least two further apparatuses for training the local model and configure an alternative apparatus of at least one of the two selected further apparatuses for training the local model; receive local training results from at least one of the at least two further apparatuses and from the alternative apparatus of at least one of the two selected further apparatuses; and combine the local training results to generate an aggregated training result of the model.

[0118] According to a twenty-third aspect, there is provided a computer-readable medium comprising program instructions to cause an apparatus configured to train a local model during federated learning to receive an alternative training configuration from at least a further apparatus configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: an apparatus identifier configured to uniquely identify the apparatus for training the local model; an alternative apparatus identifier configured to uniquely identify an alternative apparatus for the apparatus; and a state identifier configured to identify a state in which the apparatus is unable to train the local model, causing the alternative apparatus to train the local model on the alternative apparatus; and to train the local model and send the local training results to the further apparatus, or determine that the apparatus is unable to train the local model based on the state in which the apparatus is unable to train the local model and send a local training model request to one of the further apparatus or the alternative apparatus, causing the alternative apparatus to perform local training on the alternative apparatus using the local dataset.

[0119] According to a twenty-fourth aspect, there is provided a computer-readable medium comprising program instructions to cause an apparatus configured to train a local model for federated learning to receive an alternative training configuration from a further apparatus configured to train the local model in a communication network comprising the apparatus, the alternative configuration comprising: an apparatus identifier configured to uniquely identify the other apparatus as an apparatus for training the local model using a local dataset; an alternative apparatus identifier configured to uniquely identify the apparatus as an alternative training apparatus for the other apparatus; and a state identifier configured to identify a state in which the other apparatus is unable to train the local model, the state identifier causing the apparatus to train the local model at the apparatus; receiving a local model training request from the other apparatus or the further apparatus when the other apparatus is unable to train the local model; and, after receiving the request, training the local model using the local dataset and transmitting the training results to the further apparatus.

[0120] An apparatus comprising means for performing the actions of the above-described method.

[0121] An apparatus configured to perform the actions of the above-described method.

[0122] A computer program comprising program instructions for causing a computer to carry out the method described above.

[0123] A computer program product stored on the medium can cause an apparatus to perform the methods described herein.

[0124] According to an aspect, a non-transitory computer-readable medium is provided comprising program instructions for causing an apparatus to perform at least a method according to any of the preceding aspects.

[0125] A number of different embodiments have been described above, and it should be understood that any two or more of the above-described embodiments may be combined to provide further embodiments.

[0126] Embodiments will now be described, by way of example only, with reference to the accompanying drawings, in which: [Brief explanation of the drawings]

[0127] [Figure 1] FIG. 1 illustrates a network system, according to some example embodiments. [Figure 2] FIG. 1 illustrates a control device, according to some exemplary embodiments. [Figure 3] FIG. 1 illustrates an apparatus, according to some exemplary embodiments. [Figure 4] FIG. 1 is a flow diagram of an example of training a model using federated learning (FL). [Figure 5] FIG. 1 is a flow diagram of an example of interruption avoidance during training of a model using federated learning (FL), according to some exemplary embodiments. [Figure 6] FIG. 1 is a flow diagram of an example of interruption avoidance during training of a model using federated learning (FL), according to some exemplary embodiments. [Figure 7] FIG. 1 is a flow diagram of an example of interruption avoidance during training of a model using federated learning (FL), according to some exemplary embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0128] In the following, certain embodiments are described with reference to devices being able to communicate with a communication system that provides service to such devices. Before describing the exemplary embodiments in detail, certain general principles of a communication system, e.g., a 5G communication system, including an access network (AN) and a core network, and devices (e.g., terminals served by the communication system), will be briefly described with reference to Figures 1, 2, and 3 to aid in understanding the technology underlying the described examples.

[0129] FIG. 1 shows a schematic diagram of a 5G wireless communications system (5GS). The 5GS may consist of a radio access network (RAN) (e.g., a 5G radio access network (5G-RAN) or a next-generation radio access network (NG-RAN)), a 5G core network (5GC), one or more application functions (AFs), and a further data network (DN). In some embodiments, the AF is a customer of the 5GC and is connected to the 5GC user plane function (UPF) via the DN and to the 5GC network function (NF) via the 5GC network exposure function (NEF). In some embodiments, the AF is a trusted application function, and therefore, the trusted AF is implemented in the 5GC and directly connected to other NFs of the 5GC. The AF may include functionality to perform training using federated learning and select 5G-connected terminals to participate in the FL, as described in more detail below. While only one UPF is shown in FIG. 1, it will be understood that the 5GS may consist of a chain of UPFs, including a UPF anchor that connects to a DN. The connection between the AF, NEF and UPF (or AF and NF in 5GC) is via interfaces defined in the 3GPP standard.

[0130] The 5GC may comprise, for example, the following Network Functions (NFs) (also referred to as network entities): Network Slice Selection Function (NSSF), Network Exposure Function (NEF), Network Repository Function (NRF), Network Data Analytics Function (NWDAF), Policy Control Function (PCF), Unified Data Management (UDM), Authentication Server Function (AUSF), Access and Mobility Management Function (AMF), and Session Management Function (SMF). The 5GC NFs may have a service-based architecture as described in 3GPP standard TR 23.501. The NF services that may be provided by the 5GC NFs and the service-based interfaces of the 5GC NFs are described in 3GPP standards, particularly 3GPP standards TR 23.501 and 23.502.

[0131] Access by a terminal to the 5GC may occur more generally via an access network, such as a 5G Radio Access Network (5G-RAN). The 5G-RAN may comprise one or more base stations (e.g., gNodeBs (gNBs)). A gNB of the 5G-RAN may include gNB distributed units connected to a gNB central unit and remote radio heads connected to the gNB distributed units. In some embodiments, one or more base stations of the 5G-RAN may be an Evolved NodeB eNodeB (eNB). In some embodiments, the 5G-RAN may be a 3GPP radio access network (e.g., a RAN operating using NR or LTE radio access technology defined in 3GPP standards). While Figure 1 illustrates a 5G-RAN, those skilled in the art will appreciate that access to 5GC may occur via any wireless or wired access network, such as a non-3GPP access network (e.g., an untrusted Wireless Local Area Network (WLAN) accessing 5GC via a Non-3GPP Interworking Function (N3IWF), a trusted WLAN accessing 5GC via a Trusted Non-3GPP Gateway Function (TNGF), or a wired network accessing 5GC via a Wireline Access Gateway function (W-AGF). A non-3GPP access network is an access network that is not configured to communicate directly with a core network with the architecture and operation defined in 3GPP standards.

[0132] FIG. 2 illustrates an example of an apparatus 200 that may implement one or more NFs of the 5GC illustrated in FIG. 1. The apparatus 200 may include at least one random access memory (RAM) 211a, at least one read-only memory (ROM) 211b, at least one processor 212, 213, and a network interface 214. The at least one processor 212, 213 may be coupled to the RAM 211a and the ROM 211b. The at least one processor 212, 213 may be configured to execute software code 215. The software code 215 may include, for example, instructions for performing actions or operations of one or more NFs of the 5GC. In some embodiments, the software code 215 may include instructions for performing one or more actions or operations of a federated learning (FL) aggregator according to aspects of the present disclosure. The software code 215 may be stored in the ROM 211b. Apparatus 200 may implement one or more NFs of 5GC and may be interconnected with other apparatus 200 implementing one or more other NFs of 5GC. In such embodiments, apparatus 200 may be part of a distributed computing system. In some embodiments, each NF of 5GC may be implemented on a single apparatus 200. In such embodiments, apparatus 200 may be a cloud computing system.

[0133] FIG. 3 illustrates an example of the device 300 shown in FIG. 1. The device 300 may be any wireless communication device capable of transmitting and receiving radio signals. Non-limiting examples of the device 300 include a terminal, a wireless communication device, a user equipment, a mobile device such as a mobile station (MS) or what is known as a mobile phone or “smartphone,” a computer with a wireless interface card or other wireless interface equipment (e.g., a USB dongle), a personal data assistant (PDA) or tablet with wireless communication capabilities, a machine-type communication (MTC) device, an Internet of Things (IoT) communication device, or any combination thereof. The device 300 may be configured to communicate with a base station (e.g., an NG-eNB or gNB) of an access network such as 5G-RAN and 5GC via a base station of the 5G-RAN using, for example, non-access stratum (NAS) signaling to communicate data. The communication may include and carry one or more of voice, electronic mail (email), text message, multimedia, data, machine data, etc.

[0134] The apparatus 300 may receive wireless signals (e.g., radio or cellular signals) over an air or radio interface 307 (commonly referred to as a Uu interface) via a suitable device 306 for receiving wireless signals, and may transmit wireless signals (e.g., radio or cellular signals) via a suitable device for transmitting wireless signals. In FIG. 3, the apparatus 306 includes one or more antennas (or an antenna array comprising multiple antennas) and a transceiver, and is represented schematically by block 306. The apparatus 306 may be provided, for example, by a radio unit and an associated antenna arrangement comprising one or more antennas. The antenna arrangement may be located inside or outside the mobile device.

[0135] The device 300 may include at least one processor 301, at least one memory ROM 302a, at least one RAM 302b, and other possible components 303 for use in software and hardware-assisted execution of tasks designed to perform, such as controlling access to and communication with an access network, such as 5G-RAN, and other devices 300. The at least one processor 301 is coupled to RAM 311a and ROM 311b. The at least one processor 301 may be configured to execute appropriate software code 308. The software code 308 may include, for example, instructions that, when executed by the at least one processor 301, perform one or more actions or operations of aspects of the present invention. The software code 308 may be stored in ROM 311b.

[0136] At least one processor 301, storage, and other related controls may be provided on a suitable circuit board and / or chipset, this functionality being indicated by reference numeral 304. The terminal 300 may optionally have a user interface, such as a keypad 305, a touch-sensitive display screen, or a touch-sensitive pad, or a combination thereof. Optionally, one or more of a display, a speaker, and a microphone may be provided, depending on the type of device.

[0137] Control or configuration of such communication systems has traditionally been performed by control mechanisms that operate based on defined rules. To improve network performance (i.e., the performance of a network in a communication system such as 5GS), control mechanisms that implement machine learning (ML) models have been proposed, where network data and / or management data from many network entities of the communication system can be processed by the ML models to generate control outputs appropriate for such communication systems. Training of the machine learning (ML) models is centralized: network data and / or management data are collected by network entities (commonly referred to as distributed nodes) and provided to a single network entity (commonly referred to as a central node), which uses the received data to train the ML model.

[0138] To minimize the amount of data exchanged between the distributed nodes and the central node (where model training (hereinafter generally referred to as model training) is typically implemented) and prevent the loss of privacy of each node's data, models may be trained using federated learning (FL). In FL, instead of training a model at the central node using centralized data, the central node provides a global model comprising parameters or data to the distributed nodes, and each of the distributed nodes performs local training of a local model (hereinafter referred to as local model training) using a dataset comprising the distributed node's data during an iteration of FL. In other words, each distributed node has a dataset (hereinafter referred to as local dataset) and trains a local model (i.e., performs local model training) using its own local dataset. In the following disclosure, the terms local training and local model training are interchangeable. Then, at each iteration of FL, each distributed node sends local training results (e.g., the "learned" parameters of the local model) to the central node.

[0139] During each iteration of FL, the central node receives local training results of the local models from the distributed nodes and combines or aggregates the local training results to obtain new global model parameters or data (global training results of the global model).

[0140] The local training results may comprise values ​​of parameters of the local model, and the global training results may comprise aggregate values ​​of parameters of the global model.

[0141] The new global model parameters or data (e.g., aggregated received local training parameter values) are then sent to the distributed nodes in further iterations of FL, and the distributed nodes use these “new” global model parameters as parameters of their local models and perform further local training of the local models to learn new local model parameters. This process may be repeated until the values ​​of the parameters of the global model are optimized. The process of training local models at different distributed nodes is known to those skilled in the art and will not be described in further detail. For example, federated learning, and local and global models, are described in detail in Konecny, Jakub, H.B. McMahan, and Daniel Ramage, “Federated Optimization: Distributed Optimization Beyond the Datacenter,” ArXiv abs / 1511.03575 (2015).

[0142] An example of communication between a central node and distributed nodes for training a model (e.g., a global model) using FL when the central node and distributed nodes are part of a communication system or connected by a communication system such as the 5GS in Figure 1 is shown in Figure 4 and summarized below.

[0143] The example shown in FIG. 4 illustrates communication between a central node (referred to as FL aggregator 400) and three distributed nodes (i.e., three devices 300 shown in FIG. 4 as UE1 300a, UE2 300b, and UE3 300c) for two sets of repetitions of the FL: repetition set N 451 and repetition set N+1 491. However, there may be more than two repetition sets of the FL. There may be an initial sequence of communication for each repetition set of the FL. The communication between FL aggregator 400 and the three distributed nodes includes signaling and / or data communication. Signaling may be sent by FL aggregator 400 to the distributed nodes via the NEF and AMF and 5G-RAN, or by the FL aggregator via the AMF and 5G-RAN. Data communication may be sent by FL aggregator 400 to the distributed nodes via the 5G-RAN, UPF, and DN, or may be sent by the wireless communication device. 4, for simplicity, the 5G-RAN, AMF, NEF, UPF, and DN, as well as communications between these network entities, are omitted. For example, FL aggregator 400 may be configured to generate and send an FL report configuration to each of UE1 300a, UE2 300b, and UE3 300c. The FL report configuration comprises information defining the content to be included in the FL report, the format of the FL report, and how the FL report is generated. For example, as shown in first iteration set 451, FL aggregator 400 may be configured to generate and send an FL report configuration at 401 to UE1 300a, UE2 300b, and UE3 300c.

[0144] Each of UE1 300a, UE2 300b, and UE3 300c may then be configured to generate and send an FL report to FL aggregator 400 based on the FL report configuration. For example, UE1 300a is shown generating and sending an FL report to the FL aggregator at 403, UE2 300b is shown generating and sending an FL report to the FL aggregator at 405, and UE3 300c is shown generating and sending an FL report to the FL aggregator at 407. The FL report may include information useful to assist FL aggregator 400 in determining which UEs to select for FL (i.e., which UEs to select for training a local model in the current iteration of FL, abbreviated as local training in the following description). For example, such information may comprise information indicative of resource availability at the UE, such as processing resource availability (e.g., number of CPUs available at the UE), memory resource availability (amount of memory available at the UE), energy available at the UE, measurements indicative of the quality of a radio link (e.g., a Uu link) between the UE and a 5G-RAN base station (e.g., signal-to-noise ratio (SNR), reference signal received power (RSRP)), which may be part of the link between the UE and the FL aggregator 400, and data availability at the UE for training a local model.

[0145] The FL aggregator 400 is then configured to select K UEs (two UEs in this example shown in FIG. 4) for training the local model (also referred to herein as local training), as shown at 409 in FIG. 4. The selection of UEs by the FL aggregator 400 may be random in some circumstances or may be based on any suitable UE selection scheme that takes into account the obtained FL reports (provided by the UEs).

[0146] In this example, the FL aggregator 400 (for iteration N 451) selects UE1 300a and UE3 300c for the FL.

[0147] 4 , the FL aggregator 400 is configured to transmit the global model and the training configuration. In other words, the FL aggregator 400 is configured to broadcast the global model (e.g., aggregated training results or aggregated model parameter values ​​from previous iterations) and the training configuration to the K selected UEs. Furthermore, the FL aggregator 400 can transmit the training configuration to the K selected UEs to enable the selected UEs to perform local training. The FL aggregator 400 is then configured to transmit a signal to each of the selected UEs (e.g., UE1 and UE3) to perform local training of their local models. For example, in iteration N 451, FL is performed with UE1 and UE3, as indicated by dashed box 411. In other words, the FL aggregator 400 transmits the configuration to the selected UEs.

[0148] Each selected UE is then configured to perform local training. For example, UE1 is configured to train a local model based on a local dataset stored in UE1, as shown at 415 in FIG. 4, and UE3 is configured to train a local model based on a local dataset stored in UE3, as shown at 417 in FIG. 4. Each of the K selected UEs receives the global model and the training configuration, configures its local model based on the global model and the training configuration, and updates parameters of its local model to the received aggregate parameters. The K selected UEs are then configured to perform local training.

[0149] After performing the local training, the K selected UEs are then configured to report the results of their local training (also referred to herein as local training results). For example, each UE is configured to report its local training results to FL aggregator 400. For example, with respect to UE1 300a, the local training results are reported (i.e., transmitted) by UE1 300a to FL aggregator 400 at 419 in FIG. 4, and UE3 300c reports (i.e., transmits) its local training results to FL aggregator 400 at 423 in FIG. 4. The local training results reported by a particular UE may comprise parameter updates of the local model at the particular UE. In some embodiments, the parameter updates of the local model are values ​​of the parameters of the local model when training of the local model is completed at the UE. In some embodiments, the parameter updates of the local model are gradients of the parameters of the local model when training of the local model is completed at the UE.

[0150] Further, the FL aggregator 400 is configured to perform aggregation of the local training results received from each of the K UEs (e.g., UE1 and UE3) (i.e., aggregate or combine parameter updates of the local models), as shown at 423 in FIG. 4.

[0151] After performing aggregation of the local training results, FL is then repeated a predetermined number of times using UE1 and UE3. In other words, in 411, FL is repeated multiple times using UE1 and UE3.

[0152] 4 further illustrates a further example set of iterations N+1 491 in which UEs for FL are reselected. In the set of iterations N+1 491, UE2 300b and UE3 300c are selected for FL by FL aggregator 400. For example, at 461, FL aggregator 400 is shown configured to generate and send FL report configurations to UE1 300a, UE2 300b, and UE3 300c.

[0153] Each of the UEs may then be configured to generate and send an FL report based on the FL report configuration to the FL aggregator 400. For example, UE1 300a is shown generating and sending an FL report to the FL aggregator at 463, UE2 300b is shown generating and sending an FL report to the FL aggregator at 465, and UE3 300c is shown generating and sending an FL report to the FL aggregator at 467.

[0154] The FL aggregator 400 is then configured to select K UEs (2 UEs in this example) for local training, as shown at 469 in FIG.

[0155] In this example, the FL aggregator 400 selects UE2 300b and UE3 300c for the FL and implements the FL for the set of iterations N+1 491 using UE2 and UE3, as indicated by the dashed box 471.

[0156] The FL aggregator 400 is then configured to perform broadcasting of the global FL model and training configuration, as shown at 473 in Figure 4. Broadcasting the global FL model and training configuration comprises transmitting the global model and training configuration to each of UE2 300b and UE3 300c.

[0157] Each of the selected UEs is then configured to perform local training. For example, UE2 300b is configured to train a local model based on UE2 data, as shown by step 475 in Figure 4, and UE3 is configured to train a local model based on UE3 data, as shown by 477 in Figure 4. Local training for each of the selected UEs is performed.

[0158] After performing local training, the K selected UEs then report their local training results to the FL aggregator 400. For example, for UE2 300b, the local training results are provided (e.g., transmitted) to the FL aggregator 400 at 479 in Figure 4, and for UE3 300c, the local training results are provided (e.g., transmitted) to the FL aggregator 400 at 481 in Figure 4. The local training results may be values ​​of parameters of the local model after training has been performed.

[0159] Further, the FL aggregator 400 is then configured to perform a combination or aggregation of the local training results received from the K UEs to obtain an aggregated training result of the global model, as shown at 483 in FIG. 4 .

[0160] After performing the FL aggregate operation to combine or aggregate the local training results and obtain the model parameters of the local models to generate aggregated training results for the global model, the “local” training process is repeated a (further) predetermined number of times using UE2 and UE3. In other words, the training of the local UE2 and UE3 models indicated by box 461 is repeated multiple times.

[0161] The FL iteration set may then be repeated or continued until the FL model converges (i.e., until the parameters of the global model are optimized).

[0162] When FL is implemented in a communication system such as a 5G system, during each FL iteration, the UE receives an indication to perform training of a local model using a local dataset from an FL aggregator 400, for example, located in a network entity (also known as a network function) of a core network of the 5G system. The UE then reports the local training results (e.g., values ​​of parameters of the local model after local training is completed, updates of parameters of the local model after local training is completed, or gradients of parameters of the local model after local training is completed) to the FL aggregator 400. The FL aggregator 400 then combines or aggregates the local training results of the local models received from the UE to obtain aggregated training results of the global model (e.g., obtain values ​​of parameters of the global model). The aggregated training results (e.g., values ​​of parameters of the global model) are then included in the global model that is transmitted to the UE in subsequent iterations of FL. The UE can then start training for the next FL iteration.

[0163] An exemplary use case for the Federated Learning embodiments described below includes training a machine learning model (i.e., a global model) configured to predict Quality of Service (QoS) of device-to-device (D2D) communications over a sidelink (SL) between two devices, such as two apparatuses 300. In some embodiments, the global model may require data such as the geographic locations (or zones) of the two devices, the SL Received Signal Strength Indicator (SL RSSI) of the SL between the two devices, the SL Reference Signal Received Power (SL RSRP) of the SL between the two devices, the SL Channel State Information (SL CSI) of the SL between the two devices, the SL Channel Busy Ratio (SL CBR) of the SL between the two devices, SL transmission parameters (e.g., the modulation and coding scheme (MCS) of the SL between the two devices, the transmit power of the SL communications between the two devices, the number of retransmissions, priority, etc. of the SL between the two devices), and / or SL Hybrid Automatic Repeat Request (SL HARQ) feedback to train the global model.

[0164] Conventionally, for a model that is trained centrally (i.e., trained at a central node), this data needs to be transmitted from two devices to the central node via a communication system, such as a 5G system. In the aforementioned use case, an ML model configured to predict the QoS of D2D communication over SL between two devices may be implemented in a network entity of a core network of the 5G system. This ML model requires data for SL measurements (e.g., to predict the QoS of D2D communication over SL between two devices). However, the data for SL measurements is typically only available at the devices, and the data for SL measurements is typically not transmitted to the core network of the 5G system. Transmitting the data for SL measurements to the core network of the 5G system may significantly increase the signaling transmitted over the air interface (e.g., Uu interface) between the devices and the base station of the 5G-AN, which may be undesirable. Therefore, training an ML model using federated learning as described in embodiments herein may be beneficial in scenarios where signaling between the device (e.g., UE) and the core network is reduced because the device (e.g., UE) does not need to send local training data to the core network, and only local training results (e.g., values ​​of model parameters of a locally trained model, or gradients of parameters of a locally trained model) are sent from the device (e.g., UE) to the core network, eliminating privacy concerns.

[0165] The embodiments described herein further aim to improve the performance of FL in 5GS. Current FL involves reselection of UEs for FL (hereafter referred to as training UEs), which increases the signaling transmitted over the air interface between FL aggregator 400 and the UEs selected for FL, as well as the delay associated with model training using FL. It should be noted that selecting and configuring training UEs for FL requires an increased amount of signaling exchanged between FL aggregator 400 and the UEs.

[0166] Furthermore, if an already selected training UE cannot effectively participate in the training of the global model using FL (also referred to as model training), for example, due to poor quality of the air interface (e.g., Uu link) between the UE and the base station of the RAN or due to (temporary) unavailability of local resources at the UE (e.g., computational resources such as the amount of memory or processing resources available at the UE, power, etc.), FL aggregator 400 may need to reselect training UEs to participate in model training using FL. This reselection of training UEs may result in additional signaling being transmitted between FL aggregator 400 and the selected UEs, which may introduce delays and further increase the delay in model training using FL.

[0167] These problems may be particularly severe when UEs are frequently unavailable (intermittently unavailable) and the FL aggregator 400 frequently reselects training UEs to participate in model training using FL. Furthermore, among UEs having local datasets with similar data distributions, the FL aggregator 400 may select only one (a few) UEs to participate in model training using FL so that the trained global model is not biased toward a particular data distribution. This naturally puts the UE selected as the training UE at a disadvantage because the selected UE must participate in model training of the global model using FL and incurs an associated cost for applying resources (power usage, and / or processor and / or memory) required to perform local training of the local model using the local dataset.

[0168] Therefore, the following embodiments aim to provide an FL aggregator that minimizes reselection of UEs to participate in model training using FL, minimizes signaling sent to UEs selected for model training using FL (i.e., signaling overhead between the FL aggregator and UEs), avoids interruptions during model training using FL, and avoids unfair utilization of UEs during model training using FL.

[0169] For example, in some embodiments, the FL aggregator configures a training UE (referred to herein as a first UE or primary UE) with an alternative training UE (also referred to herein as a second UE or secondary UE). In these embodiments, the second UE is activated as a training UE (i.e., performs local training) only if the first UE is unavailable for local training.

[0170] Communication (i.e., signaling) between FL aggregator 500 and a first UE and a second UE according to one embodiment of the present disclosure is shown in Figure 5. In the embodiment shown in Figure 5, each of the UEs, UEs 300a, 300b, 300c, communicates with FL aggregator 500 via a link comprising a wireless link (e.g., a Uu link) between the UE and a base station of a RAN of the communication system and a wired link between the base station and FL aggregator 500.

[0171] The initial operations performed by FL aggregator 500 are similar to the initial operations performed by FL aggregator 400 shown in FIG. 4 and described above.

[0172] The FL aggregator 500 may be configured to generate and send an FL report configuration to each of the UEs. The FL report configuration may comprise information defining the content to be included in the FL report, the format of the FL report, and how the FL report is generated. For example, the FL aggregator 500 may be configured to generate and send 401 an FL report configuration to UE1 300a, UE2 300b, and UE3 300c.

[0173] Each of the UEs, UE 300a, UE 300b, UE 300c, may then be configured to generate and send an FL report to the FL aggregator 500 based on the FL report configuration. For example, UE1 300a is shown generating and sending an FL report to the FL aggregator at 403, UE2 300b is shown generating and sending an FL report to the FL aggregator at 405, and UE3 300c is shown generating and sending an FL report to the FL aggregator at 407. The FL report may include information useful to assist the FL aggregator 500 in determining which UEs to select for FL. For example, such information may comprise the availability of training resources at the UEs, such as CPU, energy, measurements related to the quality of a radio link (e.g., RSRP) that may be part of the link between the UE and the FL aggregator 500, the availability of local data at the UEs for training a local model at the UE (also referred to as training data availability).

[0174] The FL aggregator 500 is then configured to select K UEs for local training, as shown at 409 in FIG. 5. The selection of the K UEs for local training may be random in some circumstances or based on any suitable UE selection scheme that considers information contained in the received FL reports (provided by the UEs). In the following example, one of the selected UEs is UE1 300a. In the following description, this selected UE is the first UE. In the following example, FL signaling is shown for the UE selected for local training, UE1 300a. It will be understood that a similar signaling flow is implemented for other selected UEs, such as, for example, UE3 300c, which may be selected as the first UE along with UE1 300a.

[0175] Furthermore, the FL aggregator 500 is configured to select a second UE as an alternative training UE for the first UE, as shown at 501 in Figure 5. The selection of the second UE, which is UE2 300b in the example shown herein, as an alternative training UE for the first UE is to select one or more candidate UEs based on at least one of the following selection criteria: training UE selection assistance information received from a first UE (e.g., UE1 300a), the first UE including information identifying candidate alternative training UEs in the training UE selection assistance information; a data distribution of the data in the local dataset at the first UE and a data distribution of the data in the local dataset at each candidate alternative training UE; mobility patterns of candidate alternative training UEs for the first UE; At least one characteristic of a wireless link (i.e., a Uu link) between each candidate alternate training UE and a base station of the RAN; The locations of the primary UE and each of the candidate alternative training UEs.

[0176] For example, with respect to the data distribution of the local dataset of a candidate alternative training UE, it should be noted that FL aggregator 500 typically does not select two training UEs (within the same iteration) with very similar data distributions, so as to avoid bias in the trained global model. Thus, a UE that was not selected as a training UE in 409 because it has a data distribution similar to that of the local dataset of an already-selected training UE can be selected as an alternative training UE for the already-selected training UE. In this manner, even if the training UE is unavailable for local training, the alternative training UE can perform local training using a local dataset similar to the local dataset of the already-selected training UE. Furthermore, the selection of the alternative or second UE based on the respective locations of the first UE and the candidate alternative training UE can be based on whether a candidate alternative training UE located near (e.g., within a predetermined distance) the first UE is likely to have similar data, particularly when the training data in the local dataset comprises a measurement of the quality of a Uu link between the candidate alternative training UE and a base station of the RAN or a measurement of the quality of a sidelink between the first UE and the candidate alternative training UE. Therefore, a nearby candidate alternative training UE may be selected as the alternative training UE or second UE.

[0177] Selection of an alternative training UE or selection of a second UE based on the mobility pattern of the candidate alternative training UE relative to the first UE may be based on whether the nearby candidate alternative training UE or candidate alternative training UE is moving with the first UE (e.g., whether the first UE and the candidate alternative UE are in formation). The mobility pattern may consider, for example, the trajectory and / or speed of the candidate alternative training UE relative to the first UE. Selecting a second UE based on proximity and relative mobility may produce a second UE that is likely to have similar data, especially if the training data set includes training data indicative of wireless link (i.e., Uu link) measurements or sidelink measurements. Therefore, a nearby UE may be selected as the alternative training UE.

[0178] In some embodiments, the selection of an alternate training UE or a second UE from the candidate alternate training UEs can be based on a sidelink reachability parameter value. The sidelink reachability parameter value may relate to the ability of the first UE to communicate directly with the candidate alternate training UE. In some embodiments, the sidelink reachability parameter value may be determined based on information available at a base station (e.g., gNB) of the RAN from SL-related measurement reports in SL communication and SL relay-related scenarios. For example, the selection of an alternate training UE or a second UE from the candidate alternate training UEs can be configured to select an alternate training UE with good quality of the communication link between the first UE and the candidate alternate training UE, such that when selected, handover between the first (training) UE and the second (alternate training) UE is less likely to fail.

[0179] The selection criterion for at least one characteristic of the wireless link (e.g., Uu link) between the candidate alternative training UE and the base station or RAN may be based on ensuring that the candidate alternative training UE can participate in the FL even if the training UE is unable to participate in the FL due to a failure of the wireless link (e.g., Uu link).

[0180] The FL aggregator 500 is then caused to configure the selected second UE, in this example, UE2 300b, as an alternative training UE for the first UE, UE1 300a, as shown at 503 in Figure 5. The operation of configuring the second UE as an alternative training UE for the first UE may comprise generating an alternative training UE configuration and transmitting the alternative training UE configuration to the first UE and the second UE, where the alternative training UE configuration transmitted to the first UE and the second UE comprises identifiers of the second UE and the first UE, respectively, as shown at 505 in Figure 5. In some embodiments, the alternative training UE configuration transmitted to the first UE and the second UE may comprise the following: A unique ID for the training UE, Alternative training UE unique ID, A condition identifier configured to identify one trigger condition that triggers a transition of local model training from the training UE to the alternate training UE.

[0181] In some exemplary embodiments, the trigger conditions that trigger the transition of local model training from the training UE to the alternate training UE may include: a minimum quality of the wireless link (e.g., Uu link) between the training UE and a base station (e.g., gNB) of the RAN that the training UE must have to remain a training UE; the minimum computational and power resources that a training UE must have to remain a training UE; The minimum security / integrity level that must be met, associated with the training UE's local dataset.

[0182] Upon transmitting the alternative training UE configuration to the first UE, UE1 300a, and the second UE, UE2 300b, such that the second UE is an alternative training UE for the first UE, UE1 300a, the FL aggregator 500 can perform FL using training of the local model at the first UE, which may be implemented as shown at 507 in FIG. 5 .

[0183] In some embodiments, training the model at the first UE, UE1 300a, during an FL with the first UE, UE1 300a (or an FL with UE1) may comprise the FL aggregator 500 broadcasting or transmitting the global model and training configuration to the selected first UE and the selected second UE, as shown at 509 in Figure 5. In other words, the FL aggregator 500 is configured to broadcast the global model and training configuration to the K selected first UEs (and further to selected second UEs associated with the selected first UEs).

[0184] Then, each of the selected first UEs is configured to perform local training. For example, UE1 performs local training as shown in 511 of Figure 5. Local training is performed in all of the K selected first UEs that have received the global model and the training configuration, and the first UEs are configured to update their respective local models based on the global model and the training configuration.

[0185] After performing the local training, each of the K selected first UEs is configured to report the results of their local training (i.e., local training results). For example, the selected first UEs are configured to transmit the local training results to the FL aggregator 500. For UE1 300a, the local training results are transmitted to the FL aggregator 500 at 513 in FIG. 5. The local training results transmitted by each of the selected first UEs may be the values ​​of the parameters of the local model when the local training is completed, updates to the parameters of the local model, or gradients of the parameters of the local model. The updates to the parameters of the local model may be the difference between the original values ​​of the parameters and the final values ​​of the parameters after the local training is completed.

[0186] Furthermore, the FL aggregator 500 is configured to perform aggregation of training results received from the distributed nodes (selected first UEs), as shown at 515 in Figure 5. In other words, the FL aggregator 500 aggregates or combines the local training results to obtain aggregated training results for the global model, and updates parameters of the global model based on the aggregated training results.

[0187] In some embodiments, the second UE is configured to perform local model training when it is determined that the first UE is unavailable for local model training. In some embodiments, the second UE (also known as an alternate UE or secondary UE) is configured to perform local model training based on receiving a local training activation request (comprising an indication that the first UE is unable to perform local model training) from the selected first UE or FL aggregator 500.

[0188] 5, a first UE, UE1 300a, is configured to determine that the first UE is experiencing a local training unavailable event, as shown at 517 in FIG. 5. For example, in some embodiments, the training UE (first UE) is configured to evaluate at least one trigger condition that triggers transitioning local model training to an alternate training UE (included in an alternate training UE configuration). An example of a monitored trigger condition that triggers transitioning local model training to an alternate training UE includes when the quality of the wireless link (e.g., Uu link) between the base station and the first UE falls below a threshold, causing the first UE, UE1 300a, to initiate transitioning local model training to a second UE, UE2 300b.

[0189] Accordingly, the first UE, UE1 300a, may be configured to generate and send a request to perform local model training (or a model training transition request) to the second UE, UE2 300b, as shown at 519 in Figure 5. In some embodiments, the request to perform local model training (also referred to herein as a local model training request) may further include an indication of the reason for transitioning the local model training to the second UE, UE2 300b, and / or the time period for which the local model training will be performed at the second UE, UE2 300b.

[0190] The second UE, UE2 300b, may then be configured to generate a local training authorization message (which may be an acknowledgement message or a response message in some embodiments) in response to authorizing the second UE, UE2 300b, to perform local model training and transmit the local training authorization message to the first UE, as shown at 521 in Figure 5. The local training authorization message generated by the second UE, UE2 300b, indicates to the FL aggregator 500 that the second UE, UE2 300b, will perform local model training using the local data set.

[0191] In some embodiments, the first UE, e.g., UE1, is configured to send an indication of its unavailability to the FL aggregator 500, which is configured to forward this indication to the second UE or explicitly activate the second UE, UE2 300b, for local model training during the FL iteration (by generating an appropriate trigger message or model training transition request). Note that the FL aggregator 500 does not perform training UE reselection at this stage but simply sends an activation message to the second UE (or an alternative UE). In such embodiments, the second UE, UE2 300b, is configured to send a local training acknowledgement message (or acknowledgement or response) directly to the FL aggregator 500 rather than to the first UE, as shown in FIG. 5.

[0192] The first UE, UE1 300a, is then configured to stop monitoring the global model and training configuration transmitted by the FL aggregator 500, as shown at 523 in FIG. 5, and the second UE, UE2 300b, is configured to start monitoring the global model and training configuration transmitted by the FL aggregator 500, as shown at 525 in FIG. 5.

[0193] This effectively allows the second UE (UE2 300b) to begin local model training, which is shown in Figure 5 at 527 as local training with UE2 300b.

[0194] Local training at UE2 300b during the FL iteration may, in some embodiments, comprise the FL aggregator 500 performing broadcasting of the global model and training configuration, as shown at 529 in FIG.

[0195] Each of the "active" second UEs is then configured to perform local model training. For example, UE2 300b performs local model training (i.e., local training of a local model using a local data set), as shown at 531 in FIG. 5.

[0196] After performing local training, the "active" second UEs are then configured to report their local training results to the FL aggregator 500. For example, for UE2 300b, the local training results are reported to the FL aggregator 500 at 533 in FIG. 5, e.g., by sending a message including the results of the local training.

[0197] In an embodiment in which the second UE (e.g., UE2 300b) is configured to perform local model training based on receiving a local training activation request (comprising an indication that the first UE (e.g., UE1 300a) is unable to perform local model training) from the selected first UE or FL aggregator 500, the second UE (e.g., UE2 300b) may transmit an indication to the FL aggregator 500 indicating that the second UE is participating in the FL (i.e., an indication indicating that the second UE is an alternate training UE for the first UE) in addition to the local training results, so that the FL aggregator 500 does not discard the local training results transmitted by the second UE. The local training results and the indication may be included in a message transmitted by the second UE (e.g., UE2 300b) to the FL aggregator 500. The indication to the FL aggregator 500 that the second UE is participating in the FL is also referred to herein as a training handover indication and is shown at 533 in FIG. 5 .

[0198] Furthermore, the FL aggregator 500 is then configured to perform aggregation of the received local training results, as shown at 535 in FIG.

[0199] Thus, in these embodiments illustrated above, FL aggregator 500 is configured to use local training results sent by the second UE (or alternate UEs) in an FL iteration if the first UE is unable to perform local model training (or is unable to report local training results to FL aggregator 500). After receiving the local training results, FL aggregator 500 is configured to perform aggregation or combination of the local training results to obtain aggregated training results for the global model. Furthermore, in some embodiments, FL aggregator 500 is configured to further indicate or register changes in training UEs performing local model training to maintain a log of the availability of training UEs for local model training or the transition of local model training to alternate training UEs.

[0200] In some embodiments, FL aggregator 500 is further configured to request and obtain candidate alternative training UEs (UEs that are preferred as alternative training UEs) determined by each selected training UE. In other words, when a UE is selected as a first UE (i.e., a training UE), the UE may be requested by FL aggregator 500 to provide information identifying a second UE or other UE that may be selected as an alternative training UE.

[0201] Communication (i.e., signaling) between the FL aggregator 600 and the first and second UEs according to further embodiments of the present disclosure is shown in further detail in Figure 6. In the embodiment shown in Figure 6, the communication (e.g., signaling) is modified to incorporate preferred candidate information. In the following example, a candidate node (e.g., a candidate first UE or a candidate second UE) is a UE that can potentially be selected as the first UE or the second UE, respectively.

[0202] In the embodiment shown in FIG. 6, each of the UEs, UEs 300a, 300b, 300c, communicates with the FL aggregator 600 via a link comprising a wireless link (e.g., a Uu link) between the UE and a base station of the RAN of the communication system and a wired link between the base station and the FL aggregator 600.

[0203] The initial operations performed by FL aggregator 600 are similar to the initial operations performed by FL aggregator 400 shown in FIGS. 4 and 5 and described above.

[0204] The FL aggregator 600 may be configured to generate and transmit an FL report configuration to each of the UEs. The FL report configuration, as described above, may comprise information defining the content to be included in the FL report, the format of the FL report, and how the FL report is generated. The FL report configuration (e.g., information defining the content to be included in the FL report) may, in some embodiments, further require that the FL report comprise information indicating candidate (or preferred) alternative training UEs. For example, the FL aggregator 600 may be configured to generate and transmit 601 an FL report configuration comprising a request indicating candidate alternative training UEs to UE1 300a, UE2 300b, and UE3 300c.

[0205] In some embodiments, each candidate first UE (if indicated) may be configured to determine one or more candidate alternative training UEs (or candidate second UEs). In some embodiments, the determination of candidate alternative training UEs may be based on at least the following: The data distribution of the data in the local dataset at the candidate first UE and the data distribution of the data in the local dataset at the candidate alternative training UE (for reasons similar to those described above with respect to FIG. 5). In such an embodiment, the candidate first UE, e.g., UE1 300a, may be configured to request neighboring UEs to report the data distribution of their data in their local datasets at the neighboring UEs. The data distribution may be, for example, a range, an interquartile range, a standard deviation, a variance, or any combination thereof.

[0206] The proximity of locations between the candidate first UE and the candidate alternative training UEs and / or the mobility pattern of the candidate alternative training UE relative to the candidate first UE (for the same reasons as described above with respect to FIG. 5). For example, the candidate first UE may be configured to request neighboring candidate alternative training UEs to report their locations and mobility patterns, or the candidate first UE may be configured to calculate the proximity and relative mobility of neighboring candidate alternative training UEs by monitoring parameters such as cooperation acknowledgement messages sent by neighboring candidate alternative training UEs on the sidelink or by monitoring SL-RSRP.

[0207] Thus, for example, in FIG. 6 , UE1 300a determines a candidate alternative training UE list (for UE1 300a) at 603, UE2 300b determines a candidate alternative training UE list (for UE2 300b) at 605, and UE3 300c determines a candidate alternative training UE list (for UE3 300c) at 607.

[0208] Each of the UEs may then be configured to generate and send an FL report to the FL aggregator 600 based on the FL report configuration, the FL report configuration further comprising the identified candidate second UEs or candidate alternate training UEs. For example, UE1 300a is shown generating and sending an FL report comprising the candidate alternate training UEs at 609 to the FL aggregator, UE2 300b is shown generating and sending an FL report to the FL aggregator at 611, and UE3 300c is shown generating and sending an FL report to the FL aggregator at 613.

[0209] The FL aggregator 600 is then configured to select K UEs for local training, as shown at 615 in Figure 6. In the following example, one of the selected first UEs is UE1 300a, and another of the selected first UEs is UE3 300c. In the following exemplary drawing, FL signaling is shown for the UE selected for local training, UE1 300a. It will be understood that a similar signaling flow is implemented for the other selected UE, UE3 300c.

[0210] Additionally, the FL aggregator 600 is configured to select a second UE as an alternative training UE for the first UE, as shown at 617 in FIG. 6. The selection of the second UE, which in the example shown herein is UE2 300b, as an alternative training UE for the first UE may be based on a candidate alternative list included in the FL report received from UE1 300a. In some embodiments, the selection of the second UE as an alternative training UE for the first UE is based on the candidate list and further on any of the other selection criteria for the alternative training UE described above.

[0211] The signaling flow between the FL aggregator 600 and the training UE (first UE) and the alternate training UE (second UE) may then be similar to that described above.

[0212] Then, for example, the FL aggregator 600 is caused to configure the selected second UE, in this example, UE2 300b, as an alternative training UE for the first UE, UE1 300a, as shown at 619 in Figure 6. The operation of configuring the second UE as an alternative training UE for the first UE may comprise generating an alternative training UE configuration and transmitting the alternative training UE configuration to the first UE and the second UE, where the alternative training UE configuration transmitted to the first UE and the second UE comprises identifiers of the second UE and the first UE, respectively, as shown at 621 in Figure 6.

[0213] Upon transmitting the alternative training UE configuration to the first UE, UE1 300a, and the second UE, UE2 300b, such that the second UE is an alternative training UE for the first UE, UE1 300a, the FL aggregator 600 can then perform FL using training of the local model at the first UE, which may be implemented as shown at 623 in FIG. 6 .

[0214] In some embodiments, training the model at the first UE, UE1 300a, during an FL with the first UE, UE1 300a (or an FL with UE1) may comprise the FL aggregator 600 broadcasting or transmitting the global model and training configuration to the selected first UE and the selected second UE, as shown at 625 in Figure 6. In other words, the FL aggregator 600 is configured to broadcast the global model and training configuration to the K selected first UEs (and further to selected second UEs associated with the selected first UEs).

[0215] Each of the selected first UEs is then configured to perform local model training. For example, UE1 300a performs local model training as shown in 626 of Figure 6. Local model training is performed in all of the K selected UEs that received the global model and training configuration, and the first UEs are configured to update their respective local models based on the global model and training configuration.

[0216] After performing the local model training, each of the K selected first UEs is configured to report (their local training results) to the FL aggregator 600. For example, the selected first UEs are configured to transmit the local training results to the FL aggregator 600, as shown at 627 in FIG. 6, and the local training results are transmitted by UE1 300a to the FL aggregator 600. (Although not shown, other selected first UEs, e.g., UE3 300c, are configured to report their local training results.) The local training results transmitted by each selected first UE may be the value of the parameter of the local model when the local training is completed, an update to the parameter of the local model, or a gradient of the parameter of the local model. The update to the parameter of the local model may be the difference between the original value of the parameter and the final value of the parameter after the local training is completed.

[0217] Furthermore, the FL aggregator 600 is configured to perform aggregation of the training results received from the distributed nodes to obtain an aggregated training result of the global model, as shown at 629 in Figure 6. In other words, the FL aggregator 600 aggregates or combines the local training results to obtain an aggregated training result of the global model, and updates the parameters of the global model based on the aggregated training result.

[0218] In some embodiments, the second UE is configured to perform local model training when it is determined that the first UE is unavailable for local training. In some embodiments, the second UE (also known as an alternate training UE or secondary UE) is configured to perform local training of the local model based on receiving a local training activation request (comprising an indication that the first UE is unavailable to perform local training) from the selected first UE, UE1 300a, or FL aggregator 600.

[0219] Thus, for example, as shown in FIG. 6, a first UE, UE1 300a, is configured to determine that the first UE is experiencing a local training unavailable event, as shown at 631 in FIG.

[0220] The first UE, UE1 300a, is then configured to generate and send a request to perform local model training (or a model training transition request) to the second UE, UE2 300b, as shown at 633 in Figure 6. In some embodiments, the request to perform local model training (also referred to herein as a local model training request) may further include an indication of the reason for transitioning the local model training to the second UE, UE2 300b (or an alternative training UE) and / or the time period for which the local model training will be performed at the second UE, UE2 300b.

[0221] The second UE, UE2 300b, may then be configured to generate a local model training approval message (which may be an acknowledgement message or a response message in some embodiments) in response to the approval to perform local model training and transmit the local model training approval message to the first UE, as shown at 635 in Figure 6. The local model training approval message generated by the second UE, UE2 300b, indicates to the first UE, UE1 300a that the second UE, UE2 300b will perform local model training (i.e., perform local training of the local model using the local dataset).

[0222] As mentioned above, in some embodiments, a first UE, e.g., UE1, is configured to send an indication of its unavailability to FL aggregator 600, which is configured to forward this indication to the second UE or explicitly activate the second UE, UE2 300b, for local model training (by generating an appropriate trigger message or model training transition request). In such embodiments, the second UE, UE2 300b, is then configured to send a local model training approval message (e.g., an appropriate acknowledgement message or response message) to FL aggregator 600 in response to approving to perform local model training.

[0223] The first UE, UE1 300a, is then configured to stop monitoring the global model and training configuration transmitted by the FL aggregator 600, as shown at 637 in FIG. 6, and the second UE, UE2 300b, is configured to start monitoring the global model and training configuration transmitted by the FL aggregator 600, as shown by step 639 in FIG. 6.

[0224] This effectively allows the second UE, UE2 300b, to begin local model training, as shown at 641 in FIG.

[0225] Local model training is performed in UE2 300b, and during the FL iteration, in some embodiments, the FL aggregator 600 may be configured to perform broadcasting of the global model and training configuration, as shown at 643 in FIG. 6 .

[0226] Each of the "active" second UEs is then configured to perform local model training (i.e., local training of the local model using the local data set). For example, UE2 300b performs local model training (i.e., local training of the local model using the local data set) as shown at 645 in FIG. 6.

[0227] After performing local model training, the "active" second UEs are then configured to report their local training results to the FL aggregator 600. For example, for UE2 300b, the local training results are reported to the FL aggregator 500 at 647 in FIG. 6, e.g., by sending a message comprising the results of the local training.

[0228] In an embodiment in which the second UE (e.g., UE2 300b) is configured to perform local model training based on receiving a local training activation request (comprising an indication that the first UE (e.g., UE1 300a) is unable to perform local model training) from the selected first UE or FL aggregator 600, the second UE (e.g., UE2 300b) may send an indication to FL aggregator 600 indicating that the second UE is participating in the FL (i.e., an indication indicating that the second UE is an alternative training UE for the first UE) in addition to the local training results, so that FL aggregator 600 does not discard the local training results sent by the second UE. The local training results and the indication may be included in a message sent by the second UE (e.g., UE2 300b) to FL aggregator 600. The indication to FL aggregator 600 indicating that the second UE is participating in the FL is also referred to herein as a training handover indication and is shown at 647 in FIG. 6 . Furthermore, the FL aggregator 600 is then configured to perform aggregation of the received local training results, as shown at 649 in FIG.

[0229] In some embodiments, the training UE and / or the alternate training UEs may be configured to reject an alternate training UE configuration provided by FL aggregator 600. For example, in a situation where a first UE cannot reach a second UE via a sidelink (SL), the first UE may reject the alternate training UE configuration. Similarly, if the first UE cannot be reached via a sidelink from the second UE, the second UE may reject the alternate training UE configuration provided by FL aggregator 600.

[0230] In other words, the embodiments illustrate that based on a determination of whether the second UE receives (or does not receive) a local model training request from the first UE or whether the first UE receives (or does not receive) a local model training approval from the second UE, the first UE or second UE can accept or reject the alternative training UE configuration received from FL aggregator 600. Rejection of the alternative training UE configuration can then cause or trigger FL aggregator 600 to provide a new alternative training UE configuration.

[0231] This signaling associated with an exemplary rejection of an alternative training UE configuration is shown in connection with FIG. 7, which builds on the example shown in FIG. 5 and modifies it to incorporate the ability of the first UE and second UE to accept or reject the alternative UE and training configuration received from the FL aggregator.

[0232] The initial operation is similar to the implementation shown in FIG.

[0233] The FL aggregator 700 may be configured to generate and send an FL report configuration to each of the UEs. The FL report configuration may comprise information that defines the content to be included in the FL report, the format of the FL report, and how the FL report (and, in some embodiments, the alternative candidate list) is generated. For example, at 701, the FL aggregator 700 may be configured to generate and send an FL report configuration to UE1 300a, UE2 300b, and UE3 300c.

[0234] Each of the UEs, UE300a, UE300b, UE300c, may then be configured to generate and transmit an FL report to the FL aggregator 700 based on the FL report configuration (and, in some embodiments, further determine and output a candidate list).

[0235] For example, UE1 300a is shown generating and sending an FL report to the FL aggregator at 703, UE2 300b is shown generating and sending an FL report to the FL aggregator at 705, and UE3 300c is shown generating and sending an FL report to the FL aggregator at 707.

[0236] The FL aggregator 700 is then configured to select K UEs for local model training, as shown in 709 of Figure 7. The selection of K UEs for local model training can be based on any of the selection schemes described above.

[0237] Furthermore, the FL aggregator 700 is configured to select a second UE as an alternative training UE for the first UE, as shown in 711 of Figure 7. The selection of the alternative training UE can be based on any of the selection schemes described above.

[0238] The FL aggregator 700 is then caused to configure the selected second UE, in this example UE2 300b, as an alternative training UE for the first UE, UE1 300a, as shown at 713 in FIG.

[0239] Configuring the second UE as an alternative training UE for the first UE, UE1 300a, may comprise generating an alternative training UE configuration and transmitting the alternative training UE configuration to the first UE and the second UE, where the alternative training UE configuration transmitted to the first UE and the second UE comprises identifiers of the second UE and the first UE, respectively, as shown at 715 in FIG. 7 .

[0240] In these embodiments, the first UE and the second UE are configured to evaluate whether the alternative training UE configuration can be implemented or is acceptable to the UE receiving the alternative training UE configuration.

[0241] Thus, UE1 300a is configured to evaluate whether an alternative training UE configuration can be implemented in or acceptable to UE1 300a at 717, and UE2 300b is configured to evaluate whether an alternative training UE configuration can be implemented in or acceptable to UE2 300b at 719. As noted above, the evaluation of whether an alternative training UE configuration can be implemented in or acceptable to a UE can be based on whether a local model training request or a local model training acknowledgment can be sent from the first UE to the second UE or from the second UE to the first UE, respectively, via a sidelink between the first and second UEs.

[0242] The UEs (e.g., a first UE, UE1 300a, and a second UE, UE2 300b) may be configured to generate and transmit an approval message if the alternative training UE configuration can be implemented in the UE or is acceptable to the UE, and to generate and transmit a rejection message if the alternative training UE configuration cannot be implemented in the UE or is not acceptable to the UE. For example, UE1 300a is configured to generate and transmit an approval or rejection message at 721 regarding whether the alternative training UE configuration can be implemented in UE1 300a or is acceptable to UE1 300a, and UE2 300b is configured to generate and transmit an approval or rejection message at 723 regarding whether the alternative training UE configuration can be implemented in UE2 300b or is acceptable to UE2 300b.

[0243] The FL aggregator 700 may then optionally (in response to receiving a rejection message from at least one of the first UE or the second UE) generate and transmit a new alternative training UE configuration to the selected first UE and second UE, as shown at 725 in FIG. 7 .

[0244] These steps of selecting and configuring the first and second UEs may be repeated until the configuration is approved. Communications or signaling associated with the local model training described above may then be implemented.

[0245] Thus, once the first UE, UE1 300a, and the second UE, UE2 300b, are configured such that the second UE is an alternate training UE for the first UE, UE1 300a, FL repetitions can be performed using the first UE, UE1 300a, as shown at 731 in FIG. 7 .

[0246] The repetition of FL training with the first UE, UE1 300a, at 731 may, in some embodiments, comprise the FL aggregator 700 performing a broadcast of the global model and training configuration, as shown at 733 in FIG. 7 .

[0247] Each of the selected first UEs is then configured to perform local model training, for example, UE1 300a performs local model training as shown in 735 of FIG.

[0248] Upon implementing local training, the K selected first UEs are then configured to report the results of the local model training. For example, UE1 300a transmits the local training results to the FL aggregator 700 at 737 in FIG. 7 (although not shown, other selected first UEs, e.g., UE3 300c, are configured to report their local training results). The local training results transmitted by each selected first UE may be the value of the parameter of the local model when the local training is completed, an update to the parameter of the local model, or a gradient of the parameter of the local model. The update to the parameter of the local model may be the difference between the original value of the parameter and the final value of the parameter after the local training is completed.

[0249] Furthermore, the FL aggregator 700 is configured to perform aggregation of the received local training results to obtain aggregated training results of the global model, as shown at 739 in FIG.

[0250] In some embodiments, the second UE is configured to perform local model training when it is determined that the first UE is unavailable for local model training. In some embodiments, the second UE (also known as an alternate UE or secondary UE) is configured to perform local model training based on receiving a request to perform local model training (or a model training transition request). The request to perform local model training (also referred to herein as a local model training request) includes an indication that the first UE, UE1 300a, is unable to perform local model training. The local model training request is received by the second UE from a selected first UE, UE1 300a, or the FL aggregator 700.

[0251] Thus, for example, as shown in Figure 7, the first UE, UE1 300a, is configured to determine that the first UE, UE1 300a, is experiencing a local training unavailable event, as shown at 741 in Figure 7. In other words, the first UE, UE1 300a, determines that it cannot perform local model training based on, for example, a lack of computational or power resources available at the first UE.

[0252] The first UE, UE1 300a, is then configured to generate and send a request to transition the local model training (also referred to as a model training transition request) to the second UE, UE2 300b, as shown at 743 in Figure 7. The local model training request (or model training transition request) in some embodiments may further include information indicating the reason for transitioning the local model training to the second UE or an alternative training UE and / or the time period for which the local model training will be performed at the second UE.

[0253] The second UE, UE2 300b, may then be configured to generate and send a local training approval message (e.g., an appropriate acknowledgement message or response message) to the first UE in response to the approval to perform the local model training, as shown at 745 in FIG. 7 .

[0254] In some embodiments, the first UE, e.g., UE1 300a, is configured to send an indication of its unavailability directly to the FL aggregator 700, which is configured to forward this indication to the second UE or explicitly activate the second UE, UE2 300b, for local training of the local model (by generating an appropriate trigger message, local model training request, or model training transition request and sending the trigger message, local model training request, or model training transition request to the second UE). In such embodiments, the second UE, UE2 300b, is configured to send a local model training approval message (e.g., an appropriate acknowledgement message or response message) to the FL aggregator 700 in response to approving to perform local model training.

[0255] The first UE, UE1 300a, is then configured to stop monitoring the global model and training configuration transmitted by the FL aggregator 700, as shown at 747 in FIG. 7, and the second UE, UE2 300b, is configured to start monitoring the global model and training configuration transmitted by the FL aggregator 700, as shown at 749 in FIG. 7.

[0256] This effectively allows a second UE, UE2 300b, to participate in the FL repetitions, as shown at 751 in FIG.

[0257] In an FL iteration (also referred to as an FL iteration), in some embodiments, the FL aggregator 700 may be configured to perform a broadcast of the global model and training configuration, as shown at 753 in Figure 7. In other words, the FL aggregator 700 may broadcast the global model and training configuration to UE1 300a and UE2 300b.

[0258] Each of the "active" second UEs is then configured to perform local model training (i.e., local training of a local model using a local data set). For example, UE2 300b performs local model training, as shown at 755 in FIG. 7.

[0259] After performing the local model training, the "active" second UEs are then configured to report their local training results back to the FL aggregator 700. For example, for UE2 300b, the local training results are transmitted to the FL aggregator 700 at 757 in FIG. 7, e.g., by transmitting a message comprising the results of the local training.

[0260] In an embodiment in which the second UE (e.g., UE2 300b) is configured to perform local model training based on receiving a local training activation request (comprising an indication that the first UE (e.g., UE1 300a) is unable to perform local model training) from the selected first UE or FL aggregator 700, the second UE (e.g., UE2 300b) may send an indication to the FL aggregator 700 indicating that the second UE is participating in the FL (i.e., an indication indicating that the second UE is an alternate training UE for the first UE) in addition to the local training results, so that the FL aggregator 700 does not discard the local training results sent by the second UE. The local training results and the indication may be included in a message sent by the second UE (e.g., UE2 300b) to the FL aggregator 700. The indication to the FL aggregator 700 indicating that the second UE is participating in the FL is also referred to herein as a training handover indication and is shown at 757 in FIG. 7 .

[0261] Furthermore, the FL aggregator 700 is configured to perform aggregation of local training results received from the active second UE, UE2 300b, and other first and second UEs, as shown at 759 in FIG.

[0262] Thus, the above-described implementations and embodiments aim to provide uninterrupted model training using federated learning, even when one or more distributed nodes (e.g., training UEs) are unable to perform local training (e.g., if a first UE is temporarily unavailable, a second substitute UE can serve as the training UE until the first UE is back online). They also minimize signaling and computation associated with reselecting training UEs (UEs that participate in model training using FL) when one or more training UEs are unable to perform local training. This is particularly beneficial in scenarios where training UEs are frequently (intermittently) unavailable (which would otherwise trigger frequent reselection of training UEs for model training using FL at the FL aggregator).

[0263] Additionally, the above embodiments can provide load balancing for local model training to avoid misuse of a UE or small group of UEs in local model training using federated learning. Thus, embodiments allow local model training to be temporarily shifted to another UE, so that an "over-utilized" or "over-utilized" UE is not always running local model training.

[0264] The FL aggregators 500, 600, 700 of the present disclosure described herein may be implemented in the AF of the 5GS shown in Figure 1. Alternatively, the FL aggregators 500, 600, 700 of the present disclosure described herein may be implemented in a Network Data Analytics Function of the 5GC, an Operations, Administration and Maintenance entity of the 5GS, or any Network Function (NF) of the 5GC, such as an AMF, SMF, or PCF.

[0265] It is to be understood that the apparatus may comprise or be coupled to other units or modules, such as radio components or radio heads used in or for transmitting and / or receiving. Although the apparatus is described as one entity, the different modules and memories may be implemented in one or more physical or logical entities.

[0266] It should be noted that although some embodiments are described in the context of 5G systems, similar principles may be applied in the context of other networks and communication systems. Thus, although particular embodiments are described above by way of example with reference to particular exemplary architectures of radio access and core networks, radio access technologies and standards, the embodiments may be applied to any other suitable type of communication system implementing radio access technologies other than those shown and described herein.

[0267] It should also be noted that although the above describes exemplary embodiments, several variations and modifications may be made to the disclosed solutions without departing from the scope of the present invention.

[0268] In general, in various embodiments, the FL aggregator may be implemented in hardware or special-purpose circuits, software, logic, or any combination thereof. Some aspects of the present disclosure may be implemented in hardware, and other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device, but the present disclosure is not limited thereto. While various aspects of the present disclosure may be illustrated and described using block diagrams, flowcharts, or other graphical representations, it will be appreciated that these blocks, apparatus, systems, techniques, or methods described herein may be implemented in, by way of non-limiting example, hardware, software, firmware, special-purpose circuits or logic, general-purpose hardware or controller, or other computing device, or some combination thereof.

[0269] As used herein, the term "circuitry" may refer to one or more or all of the following: (a) hardware-only circuit implementations (e.g., analog and / or digital-only implementations); and (b) (if applicable): (i) a combination of analog and / or digital hardware circuitry and software / firmware; and (ii) A hardware processor with software (including a digital signal processor), software, and portions of memory that work together to cause a device, such as a mobile phone or server, to perform various functions. A combination of hardware circuits and software, such as (c) Hardware circuitry and / or processors, such as microprocessors or portions of microprocessors, that require software (e.g., firmware) to operate, although the software may be absent if not necessary for operation.

[0270] This definition of circuit applies to all uses of the term in this specification, including any claims. As a further example, as used herein, the term circuit covers implementations of only a hardware circuit or processor (or processors), or of portions of a hardware circuit or processor together with its (or their) accompanying software and / or firmware. The term circuit also covers, for example, baseband or processor integrated circuits for mobile devices, or similar integrated circuits in servers, cellular network devices, or other computing or network devices, if applicable to certain claim elements.

[0271] Embodiments of the present disclosure may be implemented by computer software executable by a data processor of a mobile device, such as a processor entity, or by hardware, or a combination of software and hardware. Computer software or programs, also referred to as program products, including software routines, applets, and / or macros, may be stored on any device-readable data storage medium and comprise program instructions for performing specific tasks. A computer program product may comprise one or more computer-executable components configured to perform embodiments when the program is executed. The one or more computer-executable components may be at least one software code or portion thereof.

[0272] Further, in this regard, it should be noted that the logic flow blocks as shown in the figures may represent program steps, or interconnected logic circuits, blocks and functions, or combinations of program steps and logic circuits, blocks and functions. The software may be stored on physical media such as memory chips or memory blocks implemented within a processor, magnetic media such as hard disks or floppy disks, and optical media such as DVDs and their data variants, CDs, etc. The physical media are non-transitory media.

[0273] The memory may be of any type suitable for the local technology environment and may be implemented using any suitable data storage technology, such as semiconductor-based memory devices, magnetic memory devices and systems, optical memory devices and systems, fixed memory, and removable memory. The data processor may be of any type suitable for the local technology environment and may comprise, by way of non-limiting examples, one or more of a general purpose computer, a special purpose computer, a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), an FPGA, a gate-level circuit, and a processor based on a multi-core processor architecture.

[0274] Embodiments of the present disclosure may be implemented in a variety of components, such as integrated circuit modules. Integrated circuit design is generally a highly automated process. Complex and powerful software tools are available for converting logic-level designs into semiconductor circuit designs ready to be etched onto semiconductor substrates.

[0275] The scope of protection sought for various embodiments of the present disclosure is defined by the independent claims. To the extent that some embodiments and features described herein do not fall within the scope of the independent claims, they are to be interpreted as examples that serve to understand various embodiments of the present disclosure.

[0276] The foregoing description has provided a complete and informative description of the exemplary embodiments of the present disclosure, by way of non-limiting example. However, various modifications and adaptations will become apparent to those skilled in the art upon reading the foregoing description in conjunction with the accompanying drawings and the appended claims. However, all such and similar modifications of the teachings of this disclosure are included within the scope of the present invention as defined by the appended claims. Indeed, further embodiments exist that comprise a combination of one or more of the embodiments with any of the other embodiments described above.

Claims

1. 1. An apparatus configured to train a model in a communication network using federated learning, comprising: means for selecting at least two additional devices for training the local model; means for further selecting an alternative device for at least one of the at least two selected further devices; means for configuring each of the at least two further devices for training a local model and for configuring an alternative device for at least one of the two selected further devices for training a local model; means for receiving local training results from at least one of the at least two further devices and local training results from an alternative device for at least one of the two selected further devices; A means for combining local training results to produce aggregated training results for the model. An apparatus comprising:

2. means for further selecting an alternative device for at least one of the at least two selected further devices, a similarity between the data distribution of the data of the local dataset of the at least one further device and the data distribution of the data of the local dataset of the alternative device; Further device locations; the location of the alternative device; the proximity between the further device and the alternative device; Mobility patterns of alternative devices relative to further devices; - the quality of communication on the sidelink between the further device and the alternative device; at least one characteristic of a wireless link between the further device and a base station of the radio access network; 10. The apparatus of claim 1, for selecting an alternative device based on information indicative of at least one of:

3. 3. The apparatus of claim 1, wherein the means is further for receiving information indicating one or more candidate alternative devices from the at least one further device, and the means for selecting an alternative device for at least one of the at least two selected further devices is for selecting an alternative device from the one or more candidate alternative devices identified by the further device.

4. 4. The apparatus of claim 3, wherein the means is further for generating and transmitting an FL report configuration to each of the at least two further devices, the FL report configuration comprising an indicator that enables the two further devices to generate an FL report comprising information identifying one or more candidate alternate devices.

5. The means for configuring each of the at least two further devices for training the local model on the at least two further devices, and for configuring each alternative device for training the local model on the alternative device, are further for generating an alternative training UE configuration for at least one of the at least two further devices and the alternative device, the alternative training UE configuration comprising: a further device identifier configured to uniquely identify at least one of the at least two further devices; an alternative device identifier configured to uniquely identify an alternative further device; a condition identifier configured to identify a trigger condition in which at least one of the at least two further devices is unable to train a local model, causing an alternative device to perform the local model training; and 5. The device according to claim 1, further comprising at least one of:

6. The trigger condition is a minimum quality of the wireless link between the further device and a base station of the radio access network, the availability of minimum computing resources in the further device; the availability of minimum power resources in the further device; a minimum security level associated with the local data set of the further device; and a minimum integrity level associated with the local data set of the further device; The apparatus of claim 5 , comprising at least one of:

7. 7. The apparatus of claim 1, wherein the means is further for receiving an indication from the alternative device that the alternative device is an alternative to one of the at least two selected further devices, and wherein combining the local training results comprises combining the local training results received from the alternative device with local training results from at least one of the at least two further devices.

8. The means are further receiving an indication from at least one of the at least two further devices that at least one of the at least two further devices is unable to train the local model; 8. The device of claim 1, for sending a request to an alternative device to cause the alternative device to perform local model training.

9. The request is an indicator indicating a cause for the inability of at least one of the at least two further devices to train a local model; a time indicator indicating when the alternate device will perform local model training; The apparatus of claim 8 , comprising at least one of:

10. The device, a base station of a radio access network, wherein at least two further devices and an alternative device are user equipments; a Network Data Analytics entity of a communications system, wherein at least two further devices and an alternative device are distributed Network Data Analytics entities of the communications system; an Operations, Administration and Maintenance entity of a communication system, wherein at least two further devices and an alternative device are base stations; 10. The device according to claim 1, wherein:

11. 1. An apparatus configured to train a local model during associative learning, comprising:

1. A means for receiving an alternative training UE configuration from a further device configured to train a local model in a communication network comprising the device, the alternative training UE configuration comprising: a device identifier configured to uniquely identify a device for training the local model; and a second device identifier configured to uniquely identify an alternative device for the device; and a condition identifier configured to identify a trigger condition in which at least one of the at least two further devices is unable to train a local model, causing an alternative device to perform the local model training; and means for receiving, means for training the local model and transmitting the local training results to the further device, or means for determining that the device cannot train the local model based on a trigger condition and transmitting a local model training request to the further device or one of the alternative devices to cause the further device or one of the alternative devices to perform the local model training; An apparatus comprising:

12. The means are further a data distribution of the data in the local dataset at the device and a data distribution of the data in the local dataset at the candidate alternate devices; Data distribution of local data for the device and candidate replacement devices; the extent of data in the local datasets at the device and at the candidate alternate devices; the interquartile range of the data in the local dataset at the device and candidate alternative devices; the standard deviation of the data in the local data sets at the device and at the candidate alternative devices; differences in data in the local data sets at the device and candidate alternate devices; the proximity between the device and candidate alternative devices; Mobility patterns between the device and candidate alternative devices 12. The apparatus of claim 11, for generating information indicative of at least one candidate alternative device based on information indicative of at least one of:

13. 13. The apparatus of claim 12, wherein the means is further for receiving a request from a further device and generating information indicative of at least one candidate alternative device.

14. 13. The device of claim 12, wherein the trigger conditions comprise at least one of a minimum quality of a wireless link between the device and a base station of a radio access network, a minimum availability of computational resources at the device, a minimum availability of power resources at the device, a minimum security level associated with a local dataset of the further device, and a minimum integrity level associated with a local dataset of the further device.

15. A local model training request is an indicator of why the device is unable to train a local model; a time indicator indicating when the alternate device will perform local model training; 15. The apparatus of claim 11, comprising:

16. 16. The apparatus of claim 11, wherein the means is further configured to generate one of an accept message when the alternative training UE configuration is acceptable to the apparatus and a reject message when the alternative training UE configuration is not acceptable to the apparatus, and to send the one of the accept message and the reject message to the further apparatus to cause the further apparatus to reselect or reconfigure the alternative training UE configuration.

17. 17. An apparatus according to any one of claims 11 to 16, wherein the apparatus is a first user equipment, the alternative apparatus is a second user equipment and the further apparatus is a base station of a radio access network.

18. 17. Apparatus according to any one of claims 11 to 16, wherein the apparatus is a distributed Network Data Analytics entity, the further apparatus is a centralized Network Data Analytics entity, and the alternative apparatus is a distributed Network Data Analytics entity.

19. 17. The apparatus of any one of claims 11 to 16, wherein the apparatus is a first base station of a radio access network, the further apparatus is an Operations, Administration and Maintenance entity, and the alternative apparatus is a second base station of the radio access network.

20. 1. An apparatus configured to train a local model during associative learning, comprising: A means for receiving an alternative training UE configuration from a further device configured to train a local model in a communication network comprising the device, the alternative configuration comprising: a device identifier configured to uniquely identify another device as a device for training a local model using the local dataset; and an alternate device identifier configured to uniquely identify the device as an alternate training device relative to another device; a condition identifier configured to identify a trigger condition in which another device cannot train a local model, the condition identifier causing the device to train a local model at the device; and means for receiving, means for receiving a local model training request from another or further device to perform local model training when the other device is unable to train the local model; means for, in response to receiving a local model training request, training the local model using the local dataset, generating a local training result, and transmitting the local training result to the further device; An apparatus comprising:

21. A local model training request is an indicator that identifies conditions that would prevent another device from training the local model; a time indicator that indicates when the device will perform local model training; 21. The apparatus of claim 20, comprising at least one of:

22. 22. The apparatus of claim 20 or 21, wherein the means is further for generating one of an accept message when the alternative training UE configuration is acceptable to the apparatus and a reject message when the alternative training UE configuration is not acceptable to the apparatus, and for transmitting the one of the accept message and the reject message to the further apparatus to cause the further apparatus to reselect or reconfigure the alternative training UE configuration.

23. 23. Apparatus according to any one of claims 20 to 22, wherein the means is further for transmitting an indication that the apparatus is an alternative training apparatus for another apparatus.

24. 23. An apparatus according to any one of claims 20 to 22, wherein the apparatus is a first user equipment, the further apparatus is a second user equipment and the further apparatus is a base station of a radio access network.

25. 23. Apparatus according to any one of claims 20 to 22, wherein the apparatus is a distributed Network Data Analytics entity, a further apparatus is a centralized Network Data Analytics entity and another apparatus is a distributed Network Data Analytics entity.

26. 23. The apparatus of any one of claims 20 to 22, wherein the apparatus may be a first base station of a radio access network, a further apparatus may be an Operations, Administration and Maintenance entity, and another apparatus is a second base station of the radio access network.

27. 1. A method for an apparatus configured to train a model in a communication network using federated learning, comprising: selecting at least two further devices for training the local model; selecting an alternative device for at least one of the at least two selected further devices; configuring each of at least two further devices for training a local model, and configuring an alternative device for at least one of the two selected further devices for training a local model; receiving local training results from at least one of the at least two further devices and local training results from an alternative device for at least one of the two selected further devices; combining the local training results to generate an aggregated training result for the model; A method comprising:

28. 1. A method for an apparatus configured to train a local model during associative learning, comprising: receiving an alternative training configuration from a further device configured to train a local model in a communication network comprising the device, the alternative configuration comprising: a device identifier configured to uniquely identify a device for training the local model; and an alternate device identifier configured to uniquely identify an alternate device for the device; a state identifier configured to identify a state in which a device is unable to train a local model, the state identifier causing an alternate device to train the local model on the alternate device; and receiving the signal; training the local model and transmitting the local training results to the further device, or determining that the device is unable to train the local model based on the condition that the device is unable to train the local model and transmitting a local model training request to one of the further device or an alternative device to cause the alternative device to perform local model training on the alternative device using the local dataset; A method comprising:

29. 1. A method for an apparatus configured to train a local model for associative learning, comprising: receiving an alternative training UE configuration from a further device configured to train a local model in a communication network comprising the device, the alternative configuration comprising: a device identifier configured to uniquely identify another device as a device for training a local model using the local dataset; and an alternate device identifier configured to uniquely identify the device as an alternate training device relative to another device; a state identifier configured to identify a state in which another device cannot train a local model, causing the device to train a local model at the device; and receiving the signal; receiving a local model training request from the other or further device to train the local model when the other device is unable to train the local model; After receiving the request, training a local model using the local dataset and transmitting the training results to the further device; A method comprising:

Citation Information

Patent Citations

  • Methods, apparatus and machine-readable media relating to machine-learning in a communication network

    WO2021032497A1

  • Enablement of federated machine learning for terminals to improve their machine learning capabilities

    WO2022156910A1

  • TR23.501

  • TR23.502