Contribution aware federated unlearning
By tracking and proactively managing participant contributions within the Unified Data Management function, the federated learning system effectively addresses the challenges of incomplete data exclusion and storage costs, ensuring efficient and accurate machine unlearning.
Patent Information
- Application Number
- PCT/EP2024/080898
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-16
- Filing Date
- 2024-10-31
- Publication Date
- 2025-05-22
AI Technical Summary
Existing federated learning techniques do not adequately address the issue of participant contributions and their influence on the model, leading to incomplete data exclusion and increased storage costs.
The proposed solution extends the Unified Data Management function to track participant contributions, allowing for proactive identification and exclusion of dominant or weakest participants, thereby reducing storage costs and maintaining model accuracy.
This approach enables efficient machine unlearning by quantifying participant contributions, reducing network transfer and compute costs, and maintaining high accuracy of the global model.
Smart Images

Figure EP2024080898_22052025_PF_FP_ABST
Abstract
Description
CONTRIBUTION AWARE FEDERATED UNLEARNINGTECHNICAL FIELD
[0001] Disclosed are embodiments related to federated learning, machine unlearning, Network Exposure Function (“NEF”), Unified Data Management (UDM), Network Data Analytics Function (NWDAF), and determining contributions or influence of participants.BACKGROUND
[0002] Federated learning was introduced in 2017 as an approach to enabling machine learning ("ML”) models to learn without directly accessing private data. Instead, learning is performed in a federated learning environment by sharing and merging neural networks as obtained from several participants over several communication rounds until the training process converges. By design, the original focus of such federated learning techniques was to approximate the performance of centralized training. As such, the ability for a participant to opt-out from the learning process, as constituted by the right-to-be-forgotten, was not considered.
[0003] To overcome the opt-out problem, the authors in reference [1] give an overview for machine unlearning, which is a technique that allows for participants to opt-out from the learning process and even from learned machine learning models. Reference [1] necessitates the need for sharding the different participants in different bins (accumulations of data), and then training different models to capture their corresponding information. The different models are then ensembled to allow for learning a global model. This approach differs from federated learning since it makes use of multiple models instead of creating a singular model. When one or more participants decides to opt-out, the corresponding models are excluded from the ensemble.
[0004] In reference [2] the authors explicitly investigate the problem of unlearning in a horizontal / vertical federated learning setting via the approach of deletion, which allows for omitting one or more datasets from different contributors to be removed during the training process. Given that datasets are known and preserved, certain datasets can be omitted and therefore the training process can be reproduced without the data that needs to be removed.
[0005] In reference [3], three approaches to federated unlearning are described: (i) FedEraser, which necessitates storing every update from every worker on the parameter server ("PS”) and a calibration process which uses these updates along with the information of which worker to exclude to reconstruct the unlearned model, (ii) FedRetrain, which is introduced as a baseline and unlearns a model by re-running the federation from scratch, and (iii) FedAccum, which retrains the model but instead weighs out the participants that decided to opt out exclusively in the accumulation process..SUMMARY
[0006] A challenge in reference [1] is how to place the different participants in their corresponding shards in a way that preserves privacy. In practice that can be impossible since the different shards are required to be homogeneous ( / .e. contain similar data) and the inspection to prove (or identify) such homogeneity requires direct access to the data to further analyze the distribution of the different features in the data.
[0007] A limitation with reference [2] is that the technique requires preserving the data for each participant, which is not only problematic from a privacy preservation perspective but also because it necessitates excessive storage, which can be costly.
[0008] A limitation with reference [3] is that it does not take into consideration the contributions of the different participants or how they influence the model. Instead, the calibration step is applied while reconstructing the model by excluding the input parameters of the participant to be removed. However, because in federated learning such contributions are aggregated in every round, despite of the calibration process, there is still leakage of data that is not entirely excluded throughout the training.
[0009] Third Generation Partnership Project ("3GPP”) networks are privacy aware since rel-16 with the introduction of the General Data Protection Regulation (“GDPR”), and mechanisms have been placed to enable users of a 3GPP system to know what data is being recorded and how it is used. But currently there exists no appropriate mechanism in 3GPP to allow for participants to opt-out from data collections that can be used for machine learning purposes.
[0010] Aspects of the present disclosure extend the Unified Data Management (“UDM”) function to track the contributions of each participant in a federation in order to allow for unlearning a participant's input when they request to opt-out. In addition two strategies are introduced to proactively omit the dominant / weakest participant(s) based on their propensity to pursue all iterations in the federation and not drop out (e.g., a "forecasted contribution”). These strategies produce unlearned models that maintain accuracy in the absence of corresponding types of participants and reduce the cost of learning the federated model.
[0011] According to one aspect, a computer-implemented method for machine unlearning in a collaborative learning environment is provided. The method includes training a global model in a collaborative learning environment comprising a plurality of participants. The training includes: (i) identifying an initial set of participants based on forecasted contributions of the participants, (ii) determining, for each participant in the initial set of participants, a contribution of a respective participant to the global model based on the forecasted contributions of the participants, (iii) selecting at least a first participant from the initial set of one or more participants based on the determined contribution of the at least first participant, (iv) generating a reduced set of participants for training the global model by removing the selected at least first participant from the initial set of participants, and (v) training the global model based on aggregated local model parameters from the reduced set of participants. The method includes obtaining an indication that a second participant has requested to opt-out of the training process. The method includes determining whether the second participant was included in the reduced set of participants. The method includes, in response to a determination that the second participant was included in the reduced set of participants, repeating the training steps(i)-(v) with the second participant excluded from the initial set of participants. The method includes transmitting the trained global model towards the reduced set of participants.
[0012] In some embodiments, the selecting the at least first participant is based on a level of contribution of the first participant to the global model. In some embodiments, the level corresponds to a highest or lowest contribution to the global model.
[0013] In some embodiments, the contribution comprises a measure of efficacy of the global model based on including or excluding a respective participant from training of the global model.
[0014] In some embodiments, the method includes producing n combinations of initial participants, wherein n is a number of participants in the initial set of participants, r is a number of participants to exclude in each combination, and r is less than n-1 ; aggregating local parameters of each respective combination of initial participants; measuring an accuracy of the global model based on the aggregated local parameters of each respective combination of initial participants; and determining the contribution of a respective participant to the global model based on the measured efficacy of the global model based on the aggregated local parameters of respective combinations of initial participants that do not include the respective participant. In some embodiments, the measuring comprises determining the efficacy of the global model based on the aggregated local parameters of each respective combination of initial participants against a validation set.
[0015] In some embodiments, each of the forecasted contributions comprises a likelihood that a respective participant will request to opt-out of the training process. In some embodiments, the likelihood that the respective participant will request to opt-out of the training process is based on one or more of: a number of previous requests of the respective participant to opt-out of other training processes, a consistency of parameters that the respective participant provides during training, a latency of parameter updates from the respective participant, a divergence of weights of the respective participants from an average model, a similarity of weights of the respective participant from the average model, a number of data points the respective participant trained its local model on, or metadata describing the respective participant.
[0016] In some embodiments, the identifying the initial set of participants comprises: transmitting a request to a plurality of participants, the request comprising an invitation for the plurality of participants to participate in training the global model; and receiving a response to the request, the response indicating that the initial set of participants elected to participate in training the global model. In some embodiments, the method further includes fine-tuning the forecasted contributions of the participants based on the response.
[0017] In another aspect there is provided a node with processing circuitry adapted to perform the method described above. In another aspect there is provided a computer program comprising instructions which when executed by processing circuity of a node causes the node to perform the methods described above. In another aspect there is provided a carrier containing the computer program, where the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium.BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate various embodiments.
[0019] FIG. 1 illustrates a network architecture, according to some embodiments.
[0020] FIG. 2 is a flow diagram, according to some embodiments.
[0021] FIG. 3 is a flow diagram, according to some embodiments.
[0022] FIG. 4 is a flow diagram, according to some embodiments.
[0023] FIGs. 5-12 illustrate error bar plots comparing efficacy of different learning frameworks, according to some embodiments.
[0024] FIG. 13 is a method, according to some embodiments.
[0025] FIG. 14 a block diagram of a node, according to some embodiments.DETAILED DESCRIPTION
[0026] Aspects of the present disclosure provide a leave-out aggregation that allows for measuring a contribution of the weakest / dominant participant or group of participants. The contribution of each participant is quantified and used for the purpose of constructing the global or aggregated model efficiently. In some embodiments, elimination of the participants allows for other participants to compensate, thereby improving the accuracy of the aggregated model.
[0027] Aspects of the present disclosure offer several technological advantages over prior federated or coordinated ML systems as well as ML unlearning techniques. These advantages include, but are not limited to: (i) tight integration with 3GPP's mechanism for user consent; (ii) overcoming the need to identify how to include participants in multiple shards and corresponding models, which can hurt unlearning performance if not done properly, may require reorganizing shards which can be costly, and increases complexity as it requires maintaining multiple models; (iii) identifying the contributions of each participant, thereby allowing for building better performing models by enabling remaining participants to achieve higher efficacy or accuracy while the model is reconstructed; (iv) proactively generating "unlearned” models, thereby reducing time to wait until the model is reconstructed; (v) because a contribution is identified, low contributing participants can be removed ahead of time, thereby reducing network transfer / compute costs, and conversely, removal of dominant contributors allows for lower contributors to compensate for missing this enabling on-par high efficacy / accuracy; (vi) no storage costs of intermediate updates on the parameter server; and / or (vii) contribution of each participant is quantified (positive / negative).
[0028] FIG. 1 illustrates a network architecture 100, according to some embodiments. In some embodiments, FIG. 1 illustrates a typical 5G network architecture 100. FIG. 1 illustrates a plurality of participants, or user equipment(“UE”) 102A-C, Access & Mobility Function 104, gNB or base station 106, Unified Data Management 108, Network Function Repository 110, User Plane Function 112, System Management Facility 114, Network Exposure Function (“NEF”) application programming interface (API) 116, and data network 118.
[0029] The NEF 116 was introduced in Release 15 with the main goal of enabling third parties to access Network Functions that were previously not accessible outside of the network via popular APIs typically used by developers, such as HTTP. Aspects of the present disclosure extend the role of the NEF 116 to allow for treating the 3GPP network as mechanism for training / learning any machine learning model by receiving as input a task from a 3rdparty via an NEF function and then orchestrating the entire process of learning / training the model including the selection of the UEs 102. The learning process also may consider user consent as input, which may be stored in the Unified Data Management (or “UDM”) 108.
[0030] Even though it is possible in a 3GPP network to capture whether a user consents to the collection of their data, it is not possible to remove a user's data from a trained model, such as in the case where a user has opted out from the learning process.
[0031] FIG. 2 is a flow diagram, according to some embodiments. At step 201, a third party 220 issues a machine learning task (ml_model_task) to the NEF 216 node. In some embodiments, the machine learning task is sent using an HTTP-based API. In some embodiments, the task contains metadata related to the training task, such as the model architecture, geographical region that the participants (202A-202C) might occupy and the relevant information which is expected to be used to train the corresponding model.
[0032] At step 203, the NEF 216 proxies the task further to an Operations / Administration and Management node (GAM) 222. The 0AM 222 may be used internally to track the learning process. The 0AM node 222 selects the participants (202A-C) using the UDM function 208, which has information of the participants that have given consent for certain data. At 205, the 0AM 222 requests the list of participants from the UDM 208. One goal of step 205 is to find the participants or users that are needed for the training process. The list of users is provided from the UDM 208 to the 0AM 222 at 207. The list may be, for example, a list of consenting participants. Afterwards every user (in the scope of this FIG. 2, UEs are illustrated as w1 , w2, w3) 202B-C receives an invitation to participate in a federation at 209, 211, and 213. The invite may include a set of one or more hyper parameters and a federation identifier. Without loss of generality, w1 , w2, w3 202A-C may also be group of different users, such as users served by the same UE or users in the same geographical area, instead of single UEs.
[0033] The training process takes place in steps 215, 217, 219, 221 , 223, 225, 227, 229, and 231. At 215, 217, and 219, neural parameters from w1 202A, w2 202B, and w3 202C are aggregated at a parameter server 224. These steps may take place as part of a typical federated learning or coordinated learning environment.
[0034] At 221 , the contributions of each participant are identified or measured and a strategy for including or excluding each participant based on their contributions is selected. At 223, the updated contributions of each participant may be sent to the UDM 208. At 225, the global model, or aggregated model, is created based on theincluded participants. Different strategies as described below may employed at 221 for including of excluding each participant based on their contributions.
[0035] Strategy 1 : one strategy is to identify the "dominant” user or the user whose input yields the highest increase (or smaller loss) in efficacy / accuracy during the training process for a validation set that is hosted on the parameter server. To identify this user, n! / r!(n-r)i combinations of different aggregations may be produced where a user is excluded, remaining users are aggregated, and then validation takes place. This step is repeated for all combinations. Without loss of generality one such example is provided below:
[0036] Here there are four participants (0, 1 , 2, 3), and four combinations are tried where one participant is excluded while the others are aggregated: (1) when aggregating with (0, 1 , 2) while excluding participant 3 the accuracy is 50.75; (2) when aggregating with (0, 1 , 3) while excluding participant 2 the accuracy is 49.43; (3) when aggregating with (0, 2, 3) while excluding participant 1 the accuracy is 49.65; (4) when aggregating with (1 ,2,3) while excluding participant 0 the accuracy is 52.38.
[0037] Via this process (implemented in step 221 ), one may identify or determine that the exclusion of participant 2 yields the least accuracy, which indicates that participant 2 is the dominant participant.
[0038] Despite the high cost of this process empirically, in some embodiments the dominance of a participant is stable throughout the federation since their data distribution does not change during this process and as such this check does not need to be done every time.
[0039] Strategy 2: Another strategy here is to identify the weakest participant. This can be implemented in a similar fashion by excluding different participants but instead of finding the highest contribution the smallest contribution is identified (or the aggregation that yields the highest accuracy / efficacy given the exclusion of a participant).
[0040] Strategy 3: A third strategy is based on white list / black list approach where historical information is recorded from the different participants to estimate the likelihood that a participant opts out. The historical information may include input such as: (i) number of drops of a participant in previous federated learning training (e.g., a developed reputation of a client); (ii) consistency of model parameters that a participant sends during training; (iii) latency in the updates; (iv) divergence of the weights from the average model; (v) similarity of the weights from the average model; (vi) number of data points a participant trained on; and / or (vii) other meta-data describing client characteristics including location, distribution, etc.
[0041] Given the output of that model, instead of evicting dominant vs weak participant one may evict the participant that is likely to drop-out given past experiences. Since this approach is probabilistic in some embodiments, it may be important to re-adjust or retrain the model using input from past wrong predictions but also to update with ground truth data (i.e., whether the participant actually dropped at the end).
[0042] Using these strategies, two types of "unlearned” models may be obtained in advance.
[0043] The first model is built by excluding one or more dominant participants, which enables the other participants (and their corresponding optimizers) to learn as good as possible models leveraging the generalization obtained from the remaining participants.
[0044] The second model is built by excluding one or more of the weakest participants, and as such the model is expected to perform well using the input from the dominant contributions. In both approaches, a smaller number of participants is used for training.
[0045] One observation that can be made empirically when choosing between the different strategies is that the removal of the dominant contributor (thus federating exclusively relying on the weakest contributors) leads to longer time for the federation to converge since the corresponding optimizers for each participant require more iterations to compensate for the missing input. Yet the accuracy / efficacy obtained by the removal of the dominant participant vs the removal of the weakest participants is on-par.
[0046] Afterwards, at 223 the learned model / aggregated parameters are provided from the parameter server 224 to the 0AM 222. At 235, the learned model is made available as a service back to the NEF 216 or if requested by the 3rdparty 220 the aggregated parameters are provided back.
[0047] Steps 237, 239, 241, 243, 245, 247, 249, and 251 are the active unlearning process, according to some embodiments. At 237, where in a user (e.g., w2 202B) has decided to opt-out from the learning process and notifies the UDM 208 (e.g., a retract consent). If the user has already been excluded by the previous two strategies the task is already complete and there is no need to reconstruct the model. However, if that is not the case, the process is repeated now taking into consideration the user that has recently decided to opt out and by re-inviting the remaining participants. At 239, the UDM 208 instructs the 0AM 222 to initialize a federation. For each iteration, at 241 and 243 neural parameters are sent from the remaining participants (w1 202A and w3 2020) to the parameter server 224. The parameter server at 245 identifies the contribution of each participant and identifies or selects a strategy as described above for step 221 . At 247 the aggregated model is created based on the strategy, and at 249 and 251 the aggregated parameters are sent to the participants w1 202A and 2020. In some embodiments, the training loop may be repeated until model convergence.
[0048] FIG. 3 is a flow diagram, according to some embodiments. FIG. 3 illustrates an example implementation of the process in a Network Data Analytics Function (NWDAF) environment with a consumer 304, server 306, UDM 308, client 1 302A, client 2 302B, client 3 3020, and client 4 302D. At 301 , a subscription request for ML model provisioning / training is sent from the consumer 304 to the server 306. At 303, the server 306 sends a client selection request to the UDM 308 (e.g., client_nwdaf(s) selection, and the UDM 308 returns a list of one or more clients at 305.
[0049] The first box shows federated learning training. At 307, 309, and 311 , the server initiates federated learning parameters provisioning (e.g., nnwdaf_MLModelTraining_Subscribe) with clients 1 -3 (302A-C), respectively. The clients 302A-C engage in data collection, and at 313, 315, and 317 send their trained models (e.g., Nnwdaf_ML_ModelT r ai n I ng_Notify ) to the server 306. At 319, the server 306 identifies the client contributions. At step321, the client contributions are sent to the UDM 308. At 323, an update training status is sent to the consumer 304. At 325, a modified subscription for update or terminate is sent from the consumer 304 to the server 306. At 327, the server 306 updates or terminates the federated learning training process. At 329, 331, and 333, aggregated model information distribution is sent from the server 306 to clients 1-3 (302A-C), respectively. The clients 1-3 (302A-C) update their local models based on the aggregated model.
[0050] The second box shows federated unlearning. At 335, an unlearning request for a participant (e.g., client 2 302B) is sent from the consumer 304 to the server 306. At 337, the server 306 checks the participant's contribution at the UDM 308. If the participant was already excluded, no further action may be required.
[0051] In the alternative case where the participant was not excluded from training, the process as described above repeats without using data from the participant that opted out. At 339, 341, and 343, the server initiates federated learning parameters provisioning (e.g., nnwdaf_MLModelTraining_Subscribe) with participating Clients 1, 3, and 4 (302A and C-D) respectively. The participant clients engage in data collection, and at 345, 347, and 349 send their trained models (e.g., Nnwdaf_ML_ModelTraining_Notify) to the server 306. At 351, the server 306 identifies the client contributions. At step 53, the client contributions are sent to the UDM 308. At 323, an update training status is sent to the consumer 304. At 355, a modified subscription for update or terminate is sent from the consumer 304 to the server 306. At 357, the server 306 updates or terminates the federated learning training process. At 361 , 363, and 365, aggregated model information distribution is sent from the server 306 to clients 1 and 3-4 (302A, C-D), respectively. The clients update their local models based on the aggregated model.
[0052] FIG. 4 is a flow diagram, according to some embodiments. FIG. 4 illustrates implementation of the process in an open radio access network (O-RAN), according to some embodiments.
[0053] At 401, a consumer 404 transmits an onboard request to orchestration function 406 in the Service Management and Orchestration (SMC). The SMC may include the orchestration 406, data collector 408, resource controller, and non-real time (RT) RAN Intelligent Controller (RIC) 412. The onboard request may include a ML model description for training and a federation ID. At 403, participants are selected at the orchestration function 406. At 405, an initiate training message is sent to a Non-RT RIC 412. The initiate training message may include a description of the model and the selected participants. Data collection may commence in the O-RAN amongst the near-RT RIC 414, O-CU 416, and O-DU / O-RU 418.
[0054] The first loop includes training the ML model according to the techniques described herein. At 407, 409, and 411, the non-RT RIC 412 transmits a message to train the ML model to the ML training hosts 1-3 (402A-C), respectively. At 413, 415, and 417, the ML training hosts 1-3 (402A-C) transmit an updated ML model to the non-RT RIC 412, respectively. At 419, the non-RT RIC 412 identifies the contributions of the ML training hosts 1-3. At 421, the participant contributions are sent to the orchestration function 406. At 423, the global model parameters are aggregated based on the contributions. At 425, 427, and 429, the aggregated model is transmitted towards the ML training hosts 1-3 (402A-C), respectively.
[0055] At 431 , an unlearn notification (e.g., identifying participant 2 and the federated learning identifier) is transmitted from the consumer 404 towards the orchestration 420. If the participant was not included in the aggregated model, then the process concludes.
[0056] Otherwise, the alternative loop is performed repeating the loop without using the contributions of the identified participant (e.g., participant 2). At 433, participants are selected, excluding the identified participant. At 435, an initiate training message is sent to the Non-RT RIG 412. At 437 and 439, a message to train the ML model is transmitted to the participating ML training hosts 1 and 3 (403A and C). At 441 and 443, the ML training hosts 1 and 3 (403A and C) transmit an updated ML model to the non-RT RIG 412, respectively. At 445, the non-RT RIG 412 identifies the contributions of the ML training hosts 1 and 3. At 447, the participant contributions are sent to the orchestration function 406. At 449, the global model parameters are aggregated based on the contributions. At 451 , 453, and 455, the aggregated model is transmitted towards the ML training hosts 1 -3 (402A-C), respectively.
[0057] FIGs. 5-12 illustrate empirical evaluations of the techniques described herein, according to some embodiments. Several non-limiting empirical evaluations were performed to measure the feasibility and results obtained by the proposed approach. FIGs. 5-8 evaluate the removal of the weakest participant strategy (i.e., strategy 1) and FIGs. 9-12 evaluate the removal of the dominant participant strategy (i.e., strategy 2). Different tests were performed with 4, 6 and 8 participants and with two different datasets, the Modified National institute of Standards and Technology (MNIST) and Canadian Institute for Advanced Research (CIFAR) databases in independent and identically distributed (“iid") and non-iid distributions. As used herein, efficacy may mean a measure of accuracy of the model; suitability for a certain objective, purpose or result an optimization of a desired objective, purpose, or result; a minimization of an undesired objective, purpose, or result; or achieve a certain objective, purpose, or result within a certain number of iterations, among other suitable metrics for measuring a performance of the model.
[0058] FIG. 5 illustrates an empirical evaluation of removal of the weakest contributor based on MNIST in a non-iid setting. Starting with MNIST in a non-iid setting with 4, 6 and 8 participants, several observations can be made (1) Accuracy / efficacy is improved when removing the weakest participant while also the same learning trend is observed which appears to be increasing beyond the 20th iteration. (2) The removal of the participant is verified by the aggregated model's complete inability to predict the removed participant's validation set. (3) The dominant participant retains its dominance throughout the federation, thus reducing the need to re-compute its participation / contribution.
[0059] FIG. 6 illustrates an empirical evaluation of removal of the weakest contributor based on CIFAR (non- iid). Similar results are encountered when applying the proposed approach in the CIFAR dataset in a non-iid setting.
[0060] FIG. 7 illustrates an empirical evaluation of removal of the weakest contributor based on MNIST in an iid setting. In a non-iid setting, the removal of the weakest contributor did not hurt the accuracy / efficacy of the model since other participants were able to compensate for the missing data given that all participants share the same distribution. At the same time, it was hard to say if the participant had been removed by evaluating on their dataset because the other participants compensated for the missing information, and therefore the missing dataset can still be recognizedby the model. However, in some cases the accuracy / efficacy of the removal of the weakest participant is worse, which gives credence to the fact that at least some information from the weakest participant (excluded in this experiment) has been removed.
[0061] FIG. 8 illustrates an empirical evaluation of removal of the weakest contributor based on CIFAR in an iid setting. Similar results are encountered when applying the proposed approach in the CIFAR dataset in an iid setting.
[0062] FIG. 9 illustrates an empirical evaluation of removal of the dominant participant based on MNIST in a non iid setting. As shown in FIG. 9, the results are similar to the removal of the weakest participant since this case allows for remaining participants to compensate for the missing data. The accuracy / efficacy is similar and the dominance of the participant remains stable throughout the process. The removal of the participant also hurt the model's capability of recognizing its dataset.
[0063] FIG. 10 illustrates an empirical evaluation of removal of the dominant participant based on MNIST in an iid setting. In this case there was a similar accuracy / efficacy before and after the removal, and the dominant participant remains constant throughout the federation. Similar to the previous iid cases, there was a small drop in accuracy / efficacy once the dominant participant was removed, but it was hard to say that the participant was removed since input from other participants compensates and allows for the missing data to be recognized.
[0064] FIG. 11 illustrates an empirical evaluation of removal of the dominant participant based on CIFAR in a non-iid setting. The results obtained with CIFAR in a non-iid setting but now with the removal of the dominant participant are similar to what was seen in non iid cases when removing the weakest participant. Accuracy / efficacy is maintained if not improved after the removal, and since this is a non-iid setting, it is possible to verify the complete removal of the participant. The dominance of the participant as well is maintained throughout the process.
[0065] FIG. 12 illustrates an empirical evaluation of removal of the dominant participant based on CIFAR in an iid setting. In the case of removing the dominant participant in CIFAR in an iid setting, there was an impact in accuracy / efficacy in the test dataset of the dominant participant when it was removed from the training process. However, overall accuracy / efficacy is maintained before and after the removal.
[0066] FIG. 13 illustrates a method 1300, according to some embodiments. In some embodiments, method 1300 is a computer-implemented method for machine unlearning in a collaborative learning setting. In some embodiments, the collaborative learning setting is a federated learning setting. In some embodiments, a node performs method 1300.
[0067] At step s1301 , the method includes training a global model in a collaborative learning environment comprising a plurality of participants. The training includes: (i) identifying an initial set of participants based on forecasted contributions of the participants, (ii) determining, for each participant in the initial set of participants, a contribution of a respective participant to the global model based on the forecasted contributions of the participants, (iii) selecting at least a first participant from the initial set of one or more participants based on the determined contribution of the at least first participant, (iv) generating a reduced set of participants for training the global model by removing the selectedat least first participant from the initial set of participants, and (v) training the global model based on aggregated local model parameters from the reduced set of participants.
[0068] Step s1303 of the method includes obtaining an indication that a second participant has requested to opt-out of the training process.
[0069] Step s1305 of the method includes determining whether the second participant was included in the reduced set of participants.
[0070] Step s1307 of the method includes, in response to a determination that the second participant was included in the reduced set of participants, repeating the training steps (i)-(v) with the second participant excluded from the initial set of participants.
[0071] Step s1309 of the method includes transmitting the trained global model towards the reduced set of participants. Accordingly, in some embodiments, the trained global model is transmitted only towards the remaining participants, e.g., excluding the participants that have opted out and / or are no longer part of the federation. In other embodiments the trained global model is transmitted towards all participants.
[0072] FIG. 14 is a block diagram of a node 1400 according to some embodiments. As shown in FIG. 14, the node, or apparatus, may comprise: processing circuitry (PC) 1402, which may include one or more processors (P) 1455 (e.g., one or more general purpose microprocessors and / or one or more other processors, such as an application specific integrated circuit (ASIC), field-programmable gate arrays (FPGAs), and the like); communication circuitry 648, comprising a transmitter (Tx) 1445 and a receiver (Rx) 1447 for enabling the node to transmit data and receive data (e.g., wirelessly transmit / receive data) over network 1410; and a local storage unit (a.k.a., "data storage system”) 1408, which may include one or more non-volatile storage devices and / or one or more volatile storage devices. In embodiments where PC 1402 includes a programmable processor, a computer program product (CPP) 1441 may be provided. CPP 1441 includes a computer readable medium (CRM) 1442 storing a computer program (CP) 1443 comprising computer readable instructions (CRI) 1444. CRM 1442 may be a non-transitory computer readable medium, such as, magnetic media (e.g., a hard disk), optical media, memory devices (e.g., random access memory, flash memory), and the like. In some embodiments, the CRI 1444 of computer program 1443 is configured such that when executed by PC 1402, the CRI causes the node to perform steps described herein (e.g., steps described herein with reference to the flow charts). In other embodiments, the node may be configured to perform steps described herein without the need for code. That is, for example, PC 1402 may consist merely of one or more ASICs. Hence, the features of the embodiments described herein may be implemented in hardware and / or software.
[0073] While various embodiments are described herein, it should be understood that they have been presented by way of example only, and not limitation. Thus, the breadth and scope of this disclosure should not be limited by any of the above described embodiments. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the disclosure unless otherwise indicated herein or otherwise clearly contradicted by context.
[0074] Additionally, while the processes described above and illustrated in the drawings are shown as a sequence of steps, this was done solely for the sake of illustration. Accordingly, it is contemplated that some steps may be added, some steps may be omitted, the order of the steps may be re-arranged, and some steps may be performed in parallel.
[0075] REFERENCES
[0076] [1] https: / / arxiv.org / pdf / 1909.08525.pdf
[0077] [2] https: / / arxiv.org / pdf / 1912.03817.pdf
[0078] [3] https: / / arxiv.org / pdf / 2012.13891.pdf
[0079] ABBREVIATIONS
[0080] PS Parameter Server
[0081] UE User equipment
[0082] 0AM Operations / Administration and Maintenance
[0083] UDM Unified Data Management
[0084] NEF Network Exposure Function
Claims
CLAIMS1. A computer-implemented method (1300) for machine unlearning in a collaborative learning environment, the method comprising: training (s1301 ) a global model in a collaborative learning environment comprising a plurality of participants, wherein the training comprises:(i) identifying an initial set of participants based on forecasted contributions of the participants,(ii) determining, for each participant in the initial set of participants, a contribution of a respective participant to the global model based on the forecasted contributions of the participants,(iii) selecting at least a first participant from the initial set of one or more participants based on the determined contribution of the at least first participant,(iv) generating a reduced set of participants for training the global model by removing the selected at least first participant from the initial set of participants, and(v) training the global model based on aggregated local model parameters from the reduced set of participants; obtaining (s1303) an indication that a second participant has requested to opt-out of the training process; determining (s1305) whether the second participant was included in the reduced set of participants; in response to a determination that the second participant was included in the reduced set of participants, repeating (s1307) the training steps (i)-(v) with the second participant excluded from the initial set of participants; and transmitting (s1309) the trained global model towards the reduced set of participants.
2. The method of claim 1 , wherein the selecting the at least first participant is based on a level of contribution of the first participant to the global model.
3. The method of claim 2, wherein the level corresponds to a highest or lowest contribution to the global model.
4. The method of any one of claims 1-3, wherein the contribution comprises a measure of efficacy of the global model based on including or excluding a respective participant from training of the global model.
5. The method of any one of claims 1-4, further comprising: producing n combinations of initial participants, wherein n is a number of participants in the initial set of participants, r is a number of participants to exclude in each combination, and r is less than n-1; aggregating local parameters of each respective combination of initial participants;measuring an efficacy of the global model based on the aggregated local parameters of each respective combination of initial participants; and determining the contribution of a respective participant to the global model based on the measured efficacy of the global model based on the aggregated local parameters of respective combinations of initial participants that do not include the respective participant.
6. The method of claim 5, wherein the measuring comprises: determining the efficacy of the global model based on the aggregated local parameters of each respective combination of initial participants against a validation set.
7. The method of claim 1 , wherein each of the forecasted contributions comprises a likelihood that a respective participant will request to opt-out of the training process.
8. The method of claim 7, wherein the likelihood that the respective participant will request to opt-out of the training process is based on one or more of: a number of previous requests of the respective participant to opt-out of other training processes, a consistency of parameters that the respective participant provides during training, a latency of parameter updates from the respective participant, a divergence of weights of the respective participants from an average model, a similarity of weights of the respective participant from the average model, a number of data points the respective participant trained its local model on, or metadata describing the respective participant.
9. The method any one of claims 1-8, wherein the identifying the initial set of participants comprises: transmitting a request to a plurality of participants, the request comprising an invitation for the plurality of participants to participate in training the global model; and receiving a response to the request, the response indicating that the initial set of participants elected to participate in training the global model.
10. The method of claim 9, further comprising: fine-tuning the forecasted contributions of the participants based on the response.
11. A computing device (1400) comprising processing circuitry (1402) and a memory (1442) coupled to the processing circuitry, wherein the device is adapted to:train (s1301 ) a global model in a collaborative learning environment comprising a plurality of participants, wherein the training the global model comprises:(vi) identifying an initial set of participants based on forecasted contributions of the participants,(vii) determining, for each participant in the initial set of participants, a contribution of a respective participant to the global model based on the forecasted contributions of the participants(viii) selecting at least a first participant from the initial set of one or more participants based on the determined contribution of the at least first participant,(ix) generating a reduced set of participants for training the global model by removing the selected at least first participant from the initial set of participants, and(x) training the global model based on aggregated local model parameters from the reduced set of participants; obtain (s1303) an indication that a second participant has requested to opt-out of the training process; determine (s1305) whether the second participant was included in the reduced set of participants; in response to a determination that the second participant was included in the reduced set of participants, repeat (s1307) the training steps (i)-(v) with the second participant excluded from the initial set of participants; and transmit (s1309) the trained global model towards the reduced set of participants.
12. The computing device of claim 11, wherein the device is adapted to perform any one of the methods according to claims 2-10.
13. A computer program (1443) comprising instructions (1444) which when executed by processing circuity (1402) of a computing device (1400) causes the device to perform the method of any one of methods according to claims 1-10.
14. A carrier containing the computer program of claim 13, wherein the carrier is one of an electronic signal, an optical signal, a radio signal, and a computer readable storage medium (1442).
15. An apparatus (1400), the apparatus comprising: communication circuitry (1448); and processing circuitry (1402), wherein the apparatus is configured to: train (s1301 ) a global model in a collaborative learning environment comprising a plurality of participants, wherein the training the global model comprises:(xi) identifying an initial set of participants based on forecasted contributions of the participants,(xii) determining, for each participant in the initial set of participants, a contribution of a respective participant to the global model based on the forecasted contributions of the participants(xiii) selecting at least a first participant from the initial set of one or more participants based on the determined contribution of the at least first participant,(xiv) generating a reduced set of participants for training the global model by removing the selected at least first participant from the initial set of participants, and(xv) training the global model based on aggregated local model parameters from the reduced set of participants; obtain (s1303) an indication that a second participant has requested to opt-out of the training process; determine (s1305) whether the second participant was included in the reduced set of participants; in response to a determination that the second participant was included in the reduced set of participants, repeat (s1307) the training steps (i)-(v) with the second participant excluded from the initial set of participants; and transmit (s1309) the trained global model towards the reduced set of participants.
16. An apparatus (1400), the apparatus adapted to perform the method of any one of the methods of claims 1-10.
Citation Information
Patent Citations
Systems and methods for weighted federated learning in a hybrid operating room environment
US20230316141A1
Cited By
Intelligent manufacturing equipment fault diagnosis excitation method and system based on federal forgetting
CN121561571A