Techniques to guarantee model performance in secure federated learning

By masking only the output layer parameters during federated learning, the method addresses the issue of low model accuracy due to privacy preservation, achieving a balance between privacy and accuracy with a guaranteed model performance.

WO2025106074A1PCT designated stage expired Publication Date: 2025-05-22VISA INTERNATIONAL SERVICE ASSOCIATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/US2023/079910
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-15
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Federated learning methods often result in low model accuracy due to techniques used to preserve client dataset privacy, leading to models that are too inaccurate for real-world applications.

Method used

Client computers mask only the parameters corresponding to their respective model's output layer during federated learning, while keeping the parameters of non-output layers unmasked, to balance privacy and accuracy.

Benefits of technology

This approach provides a model performance guarantee, bounding the inference error and ensuring that the trained models are accurate enough for real-world use, while still preserving data privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2023079910_22052025_PF_FP_ABST
    Figure US2023079910_22052025_PF_FP_ABST
Patent Text Reader

Abstract

Methods and systems for performing secure federated learning on multilayer machine learning models are disclosed. Rather than masking model parameters corresponding to every layer of a local multilayer machine learning model, client computers can mask model parameters corresponding to only the output layer of their respective multilayer machine learning model. These output layer parameters can be unmasked by a server computer, which can combine sets of non-output layer parameters and unmasked output layer parameters to produce a combined set of parameters, which can be returned to the client computers. The client computers can update their local models using the combined set of parameters. While masking the output layer parameters introduces inference error, it introduces less inference error than fully-private federated learning methods, in which all parameters of a multilayer model are masked. Moreover, embodiments provide a performance guarantee, which provides a bound on the inference error introduced by parameter masking.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNIQUES TO GUARANTEE MODEE PERFORMANCE IN SECURE FEDERATED EEARNINGBACKGROUND

[0001] ‘‘Federated learning” can refer to a ty pe of collaborative machine learning that can be performed by some collection of entities (e.g., devices such as personal computers, server computers, smartphones, other smart devices, “Internet of Things” devices, wearable devices, etc.). Each entity may have its own local dataset, which the entities may use for model training during federated learning. Federated learning can be used in contexts in which the entities cannot directly share their local datasets with one another, due to privacy issues, security issues, or other rules and regulations.

[0002] In “centralized” federated learning, client computers, each possessing their own local datasets, can perform federated learning with a central server computer. As part of federated learning, the client computers may perform steps to prevent leakage of potentially sensitive data from their datasets. How ever, these steps may result in a loss of trained model accuracy. In some cases, models trained using federated learning may be too inaccurate for real-world applications. Embodiments address these and other problems, individually and collectively.SUMMARY

[0003] Embodiments are directed to novel methods for performing secure federated learning for multilayer machine learning models, as well as systems that can be used to perform such methods. As described above, some federated learning methods can result in low7model accuracy, as a result of techniques used during those federated learning methods to preserve the privacy of client datasets. In some cases, machine learning models trained using federated learning may be too inaccurate for real-world applications, limiting the usefulness of federated learning. As described in more detail in the detailed description below7, embodiments of the present disclosure address this problem. For example, in embodiments of the present disclosure, client computers can each mask the parameters corresponding to their respective model's output layer, and not mask the parameters corresponding to one or more non-output layers of their respective model.

[0004] In summary, one embodiment is directed to a method performed by a client computer for training a machine learning model using federated learning. The client computer can train the machine learning model using a first set of training samples, thereby generating (1) a set of client output layer parameters corresponding to an output layer of the machine learning model and (2) a set of non-output layer parameters corresponding to one or more non-output layer parameters of the machine learning model. The output layer can use one or more continuous activation functions, such as Lipschitz continuous activation functions. The client computer can mask the set of client output layer parameters, thereby producing a set of masked client output layer parameters. Masking the set of client output layer parameters can cause a loss of accuracy, e.g., due to inference error in the trained machine learning model upon completing secure federated learning. The client computer can transmit the set of non-output layer parameters and the set of masked client output layer parameters to a server computer. The server computer can be configured to determine a combined set of parameters based on a plurality of respective sets of non-output layer parameters and a plurality of respective sets of masked client output layer parameters from a plurality of client computers, including the client computer. The client computer can receive the combined set of parameters from the server computer, and the client computer can update the machine learning model using the combined set of parameters, thereby producing an updated machine learning model.

[0005] Another embodiment is directed to a method performed by a server computer for training a combined machine learning model using federated learning. The server computer can receive a plurality of sets of masked client output layer parameters from a plurality of client computers. Each set of masked client output layer parameters can correspond to an output layer of a respective client-trained machine learning model used to determine the combined machine learning model. Each respective client-trained machine learning model can be trained by a different client computer of the plurality of client computers. The server computer can additionally receive a plurality' of sets of non-output layer parameters from the plurality of client computers. Each set of non-output layer parameters can correspond to one or more non-output layers of the respective client-trained machine learning model. The server computer can unmask the plurality of sets of masked client output layer parameters. Inthis way, the server computer can produce a plurality of sets of server output layer parameters. The server computer can produce a combined set of parameters based on the plurality of sets of non-output layer parameters and the plurality of sets of server output layer parameters. The server computer can transmit the combined set of parameters to the plurality of client computers. The plurality of client computers can be configured to update the respective client-trained machine learning model using the combined set of parameters, thereby training the combined machine learning model.

[0006] Some other embodiments are directed to computer systems (e.g., client computers, server computers, etc.) or other devices that can be configured to perform either of the methods described above. For example, one embodiment is directed to a computer system comprising one or more processors and a non-transitory computer readable medium coupled to the one or more processors. The non-transitory computer readable medium can comprise instructions that, when executed by the one or more processors, cause the one or more processors to perform either of the methods described above (or other methods described in the detailed description below).TERMS

[0007] A “server computer’ may include a powerful computer or cluster of computers. For example, the server computer can include a large mainframe, a minicomputer cluster, or a group of servers functioning as a unit. In one example, a server computer can include a database server coupled to a web server. The server computer may comprise one or more computational apparatuses and may use any of a variety of computing structures, arrangements, and compilations for servicing the requests for one or more client computers.

[0008] A “memory” may include any suitable device or devices that may store electronic data. A suitable memory may comprise a non-transitory computer readable medium that stores instructions that can be executed by a processor to implement a desired method. Examples of memories include one or more memory chips, disk drives, etc. Such memories may operate using any suitable electrical, optical, and / or magnetic mode of operation. A "‘memory buffer” can include a region of memory used to temporarily' store data.

[0009] A “processor” may include any suitable data computation device or devices. A processor may comprise one or more microprocessors working together to accomplish a desired function. The processor may include a CPU that comprises at least one high-speeddata processor adequate to execute program components for executing user and / or system generated requests. The CPU may be a microprocessor such as AMD’s Athlon, Duron and / or Opteron; IBM and / or Motorola’s PowerPC; IBM’s and Sony’s Cell processor; Intel’s Celeron, Itanium, Pentium, Xenon, and / or XScale; and / or the like processor(s).

[0010] A “data set” may include any set of one or more “observations” or “data values.” A “data value” can include any data element. A data value can comprise a “data vector,” one or more values (represented in vector form) corresponding to a data element or observation.

[0011] A “machine learning model” can refer to a software module that can be trained, and which can be configured to be run on one or more processors to produce some output responsive to a property of one or more sample inputs, such as a classification or numerical value of a property of one or more samples. A supervised learning model is an example of a type of machine learning model. Example supervised learning models may include different approaches and algorithms including analytical learning, artificial neural network, backpropagation, boosting (meta-algorithm), Bayesian statistics, case-based reasoning, decision tree learning, inductive logic programming, Gaussian process regression, genetic programming, group method of data handling, kernel estimators, learning automata, learning classifier systems, minimum message length (decision trees, decision graphs, etc.), multilinear subspace learning, naive Bayes classifier, maximum entropy classifier, conditional random field, nearest neighbor algorithm, probably approximately correct learning (PAC) learning, ripple down rules, a knowledge acquisition methodology, symbolic machine learning algorithms, subsymbolic machine learning algorithms, minimum complexity machines (MCM), random forests, ensembles of classifiers, ordinal classification, data pre-processing, handling imbalanced datasets, statistical relational learning, or Proaftn, a multicriteria classification algorithm. The model may include linear regression, logistic regression, deep recurrent neural network (e.g., long short term memory, LSTM). hidden Markov model (HMM), linear discriminant analysis (LDA). k-means clustering, densitybased spatial clustering of applications with noise (DBSCAN), random forest algorithm, support vector machine (SVM), or any model described herein. Supervised learning models can be trained in various ways using various cost / loss functions that define the error from the known label (e.g., least squares and absolute difference from known classification) andvarious optimization techniques, e.g., using backpropagation, steepest descent, conjugate gradient, and Newton and quasi -Newton techniques.

[0012] The process of "‘training” a machine learning model may include any steps used to prepare a machine learning model to perform some task. Often training involves determining or optimizing a set of “parameters” (which characterize the machine learning model) which result in acceptable model performance. Training can be performed in a series of “training rounds” during which training data is used to update the parameters of the machine learning model, for example, based on a loss value.

[0013] A “loss value” or “error value” may include any value that indicates the deviation between a result of some process, method, or function and an expected, desired, or correct result. For example, if a machine learning model can detect anomalies in a data set comprising 100 data values, 17 of which are anomalous, if the machine learning model only detects 15 the 17 anomalous data values, the loss value could comprises, e.g., 2 (17 - 15). Loss values can be used to train and evaluate the training of machine learning models, e g., by optimizing machine learning model parameters by minimizing the loss value, using processes such as stochastic gradient descent or backpropagation.

[0014] A “hyperparameter” can include any value used to configure a machine learning model that is external to the machine learning model. Typically, a hyperparameter is set, and is not estimated or determined from the training data that is used to train the machine learning model.

[0015] A machine learning model may comprise multiple “sub-models” or “layers,” which may refer to parts of a larger machine learning system. For example, a machine learning model could comprise a long short-term memory layer (which itself can comprise multiple layers), in addition to an attention layer and a linear layer. Layers can sometimes be organized in series, such that the input to a machine learning system is processed by a first set of layers, which produces an output that is then processed by a subsequent set of layers, and so forth until the output of the machine learning model is produced by the final layer in the series.BRIEF DESCRIPTION OF THE DRAWINGS

[0016] FIG. 1 shows a diagram of an exemplary multilayer artificial neural network.

[0017] FIG. 2 shows a diagram of a system for performing federated learning and secure federated learning.

[0018] FIG. 3 shows a diagram of an exemplary multilayer artificial neural network, used to summarize some masking operations according to embodiments.

[0019] FIG. 4 shows a flowchart of an exemplary secure federated learning method according to embodiments.

[0020] FIG. 5 shows an exemplary' computer system according to some embodiments.DETAILED DESCRIPTION

[0021] As summarized above, embodiments of the present disclosure are directed to methods and systems for performing secure federated learning with a model performance guarantee. Federated learning, particularly centralized federated learning, is summarized in more detail further below with reference to FIG. 2. As a general synopsis however, in centralized federated learning, each client computer can train an agreed-upon machine learning model using their own local dataset. After training, each trained model can be defined by a corresponding set of parameters, and each client computer can transmit their respective set of parameters to the server computer, e.g.. over a network such as the Internet. The server computer can combine these parameters to produce a combined set of parameters, and can transmit the combined set of parameters back to the client computers. The client computers can then update their local models using the combined set of parameters. This process can be repeated any number of times, e.g., in order to achieve model convergence. For example, in a subsequent iteration, or "‘cycle,” each client computer can retrain their updated local model using their respective local data sets and transmit their respective sets of parameters (corresponding to the retrained local models) back to the server computer. This process can be repeated until some terminating condition has been met, at which point each entity will possess a trained local machine learning model.

[0022] Due to the parameter aggregation by the server computer, the local machine learning models trained using federated learning benefit from the data collectively possessed by the client computers. However, because no client computer transmits their local dataset to the server computer or any other client computer, data privacy is generally preserved. Unfortunately however, it may be possible for the server computer, its operator, or anymalicious entities that intercepted the set of model parameters (e.g., as they were transmitted over the Internet) to leam private information about the client datasets based on the sets of model parameters.

[0023] For this reason, in “secure federated learning,” the client computers can “mask” their sets of parameters prior to transmitting those sets of parameters to the server computer. In “fully -private” secure federated learning, when training a multilayer machine learning model, each set of parameters corresponding to each layer of the multilayer machine learning model may be masked. Masking (described in more detail in Section II below) the parameters in this way can prevent malicious entities from determining any private data contained in the clients’ local datasets. Unfortunately, masking the sets of parameters can result in greater “inference error” in the trained machine learning models. As a result, the trained machine learning models may be less accurate at performing their respective tasks.

[0024] In contrast to fully-private secure federated learning, in federated learning methods according to embodiments of the present disclosure, each client computer can mask the parameters corresponding to their respective model’s output layer, and not mask the parameters corresponding to one or more non-output layers of their respective machine learning model. While masking the output layer parameters still introduces some inference error, it introduces less inference error than masking the parameters corresponding to all model layers, as in fully-private federated learning. Moreover, methods according to embodiments can provide a model performance guarantee, which is not provided by fully- private federated learning methods. This performance guarantee can provide a bound on the inference error for a trained multilayer model (such as a deep neural network) for which activation functions in the output layer are Lipschitz continuous. As the majority of deep neural network models use Lipschitz continuous activation functions, methods according to embodiments can be used to train a large number of current deep neural networks using secure federated learning.

[0025] In broad terms, this model performance guarantee bounds the inference error of the trained model based on the “parameter error,” described in more detail in Sections I and II below. The parameter error itself is based on the extent to which model parameters (e.g., the output layer parameters) are masked, e.g., a greater degree of masking can result in greater parameter error. Generally, in federated learning, increasing parameter error by performingmore masking generally results in greater inference error in the trained machine learning models. However, unlike embodiments of the present disclosure, fully-private federated learning methods do not provide a model performance guarantee. In other words, in fully- private federated learning methods, it is generally impossible to determine how much inference error is introduced by masking model parameters.

[0026] Thus the model performance guarantee is an advantage of embodiments of the present disclosure, as it enables data scientists (or other practitioners of methods according to embodiments) to exercise control over the privacy-accuracy tradeoff inherent to secure federated learning. Although only the output layer parameter are masked, data privacy can be preserved by controlling the extent to which those output layer parameters are masked, as it is possible to mask the output layer parameters to a sufficient degree that an attacker is unable to determine private data from the client computers’ local datasets, even with access to the unmasked non-output layer parameters. As such, embodiments of the present disclosure enable practitioners of embodiments to design and train federated machine learning models that are both private and accurate. For example, if a practitioner has a particular accuracy requirement, the practitioner can design a federated learning model that maximizes data privacy while still achieving the required accuracy by using the model performance guarantee bound.

[0027] Because embodiments of the present disclosure relate to machine learning and federated learning, an overview of machine learning, inference error, parameter error, federated learning, and masking is provided in the following sections in order to facilitate a better understanding of embodiments of the present disclosure.I. MACHINE LEARNING AND ASSOCIATED ERRORS

[0028] Machine learning models are often defined by sets of parameters. A machine learning model’s set of parameters generally control or define how that machine learning model produces output data responsive to input data. As an example, a support vector machine (SVM) is a type of machine learning model that interprets input data as points in space (or, e.g., vectors that point to such points in space), and divides those points using a hyperplane. Data on one side of the hyperplane is classified as one class (e.g., normal data) and data on the other side of the hyperplane is classified as another class (e.g., anomalousdata). The parameters of the support vector machine can comprise the coefficients used to define the equation of the hyperplane. Changing these parameters changes the shape of the hyperplane, and thus changes how the support vector machine produces outputs (e.g., classifications) in response to input data. For other types of machine learning models, such as a deep neural network, the model’s parameters can comprise weights corresponding to the model’s activation functions, as described in more detail further below.

[0029] In broad terms, the process of training a machine learning model can involve determining the set of parameters that cause the model to achieve the "best” performance on the particular task for which the model is being trained. As an example, for a machine learning model that determines whether input images depict dogs or do not depict dogs, training can involve determining a set of model parameters that are believed to result in the highest image classification accuracy, e.g.. based on a high true positive and true negative rate, and a low false positive and false negative rate.

[0030] Training is often performed using a training dataset (which may comprise labeled training data points) and an error or loss function. The loss function can relate the expected or ideal performance of the machine learning model against its actual performance on the training data set. For example, the training dataset could comprise images that are labeled to indicate whether those images depict dogs or do not depict dogs. The machine learning model could be tasked to classify the images in the training dataset, and the output classifications could be compared against the know n labels (using the loss function) to produce a loss value or error value. A loss function typically decreases in value as a model’s performance improves. As such, training a machine learning model often involves determining a set of parameters that minimize the loss function used during training. Sometimes a random parameter estimate is generated to initialize the model, then an optimization process (such as gradient descent) is used to iteratively refine the parameter estimate, eventually resulting in a final set of parameters associated with the trained machine learning model.

[0031] Iterative training processes, such as gradient descent or other applicable training processes can be performed in a series of training “rounds”, “epochs”, or other appropriate divisions. In each round, a machine learning model's performance can be evaluated using the loss function, and then the model’s parameters can be updated based on this evaluation, e.g.,with the goal of reducing the resulting loss value over time. As an example, in gradient descent, the gradient of the loss function can be determined with respect to the parameters. Such a gradient corresponds to a change in model parameters that achieve the greatest immediate change (e.g., reduction) in the loss function. By changing the model parameters based on the gradient, the loss function can be reduced during each successive training round. This process can be repeated until a terminating condition has been met.

[0032] One example of a terminating condition is a defined number of training rounds. This terminating condition can be met if the number of training rounds performed (e.g., by a computer system training the machine learning model) equals or exceeds the defined number of training rounds, at which point the iterative training process has been completed. Another type of terminating condition is a convergence condition. This terminating condition can be met of the machine learning model's parameters “converge / ’ In broad terms, convergence is achieved when the value of the loss function and / or the value of the model parameters change in increasingly small amounts with each successive training round. For example, a convergence condition can be achieved if the value of the loss function decreases by less than 0.1% in two successive training rounds.

[0033] Some machine learning models, such as “hybrid” or “ensemble” models can comprise multiple sub-models, configured in some manner to collectively perform an assigned task. For example, an ensemble model can comprise two different image classifier sub-models and the output classification produced by each sub-model could be combined in some way (e.g., a weighted average) to produce a single output classification. Similarly, some machine learning models can be “multilayer” models comprising multiple layers. A layer can generally comprise an element of a machine learning model, usually with its own input and output. In a multilayer model, input data can be input into the first or “input” layer of the machine learning model. This first layer can produce an output, which can then be input into a subsequent layer of the machine learning model, and so on. until the last layer of the model (or “output layer”) produces the machine learning model’s output.

[0034] In some cases, each layer of a machine learning model may comprise a structure generally understood to correspond to a known type of machine learning model. For example, deep neural networks can comprise a series of layers comprising sets artificial neurons. In other cases, each layer of a machine learning model can comprise a distinct sub-model. For example, a hybrid machine learning model could comprise a transformer submodel layer and a convolutional neural network sub-model layer, and the output of the transformer sub-model layer could comprise the input of the convolutional neural network sub-model layer. In such a case, the convolutional neural network sub-model layer could itself comprise multiple layers, e.g., sets of artificial neurons of the convolutional neural network.

[0035] Each sub-model and / or layer in a machine learning model may have its own parameter set, and the parameters of the machine learning model as a whole can collectively comprise the parameter sets of each of the sub-models and / or layers. In some cases, each layer and / or sub-model can be trained simultaneously based on a single loss function (which may comprise a "combined" loss function), i.e., each set of parameters for each layer and / or sub-model can be updated in each training round based on the loss function.A. Inference Error and Parameter Error

[0036] Throughout the rest of this disclosure, the output of a machine learning model may generally be referred to as an “inference.” As defined above, the term “error” typically refers to a metric describing the difference between two things, e.g., two datasets or two elements of data. In many cases, one thing (e.g., one dataset) corresponds to an expected or idealized case, while another thing (e.g., a second dataset) corresponds to an actual case. As an example, for a labeled dataset, the “ground truth” classifications (or “labels”) corresponding to each element of data may be known. For example, for a labeled dataset of images, each label may indicate whether the corresponding image depicts a dog or does not depict a dog. A hypothetical idealized machine learning model may always produce correct inferences, i.e., inferences that exactly match the ground truth labels. By contrast, an actual machine learning model may produce both correct inferences and incorrect inferences. In such a case, "inference error” can refer to a metric that describes the difference between the inferences produced by the actual machine learning model and the idealized machine learning model, or more directly, the difference between the inferences produced by the actual machine learning model and the ground truth labels. An inference error of “0.25”, for example, could indicate that the actual machine learning model’s inferences differed from the ground truth labels for 25% of inferences.

[0037] However, it is possible to determine inference errors by comparing inferences other than actual inferences and ground truth labels. For example, an inference error could be determined by comparing inferences produced by a first machine learning model against inferences produced by a second machine learning model. Both machine learning models could be actual existing models that are imperfect, i.e., not 100% accurate. However, one of the models (e.g., the first machine learning model) may produce inferences that are good enough for a particular real-world task. As such, an inference error could be produced comparing the inferences produced by second machine learning model against the first machine learning model, as a way to “benchmark” the performance of the second machine learning model relative to the performance of the first machine learning model. An inference error such as “0.05” could indicate that 5% of the inferences produced by the second machine learning model were different than the inferences produced by the first machine learning model, or in other words, that 95% of inferences produced by the second machine learning model matched the corresponding inferences produced by the first machine learning model. A data scientist executing both models on a computer system could, for example, evaluate this inference error to determine whether or not to use the second machine learning model on real -world tasks.

[0038] There are a variety of metrics that can be used as the inference error. One example is the magnitude of the difference between two sets of inferences, determined, e.g., by interpreting each set of inferences as a vector, subtracting the two vectors from each other, then calculating the magnitude of the resulting vector. For a first machine learning model represented by a functionoperating on a set of inputs z, the expression couldrepresent a set of inferences produced by that first machine learning model. Likewise, for a second machine learning model represented by a function f2operating on the same set of inputs z, the expression could represent a set of inferences produced by the secondmachine learning model. In such a case, the inference error could be represented by an expression such as Other metrics, such as the cosine similarity between thefirst set of inferences / i(z) and the second set of inferences can also be used to derivean inference error metric.

[0039] Inference error can also be used to compare machine learning models that differ from one another based on some quality. For example, if two machine learning models aregiven identical inputs, are otherwise identical in structure and form, but have different respective sets of parameters, then presumably any differences in their respective inferences (outputs) is a result of the differences in their respective sets of parameters. As such, the inference error between these two models is also a result of the difference in the respective sets of parameters, and the inference error can be used as a metric to evaluate the effect of parametric differences. For a first parameter set w. a second parameter set w' , otherwise identical machine learning models / , and a set of inputs z, the inference error could be represented by an expression such as

[0040] In secure federated learning, a first parameter set w could comprise a set of parameters produced by a client computer during local model training. The client computer could mask that parameter set and transmit it to a server computer, which could then unmask the masked parameter set to produce a second parameter set w' Because of information loss due to masking, this second parameter set w' may be different than the first parameter set. In such a case, the inference errormay be a result of this information loss due to masking.

[0041] Similar to inference error, “parameter error” can comprise metrics describing the difference between two sets of parameters. One example is the magnitude of the difference between two sets of parameters, determined e.g., by interpreting each set of parameters as a vector, subtracting the two vectors from each other, then calculating the magnitude of the resulting vector. That is. for a first vector representing a first parameter set w and a second vector representing a second parameter set w'. the parameter error could be represented by the expression || w — w' ||. As with the inference error, other metrics (e.g., cosine similarity) can also be used to derive parameter error. As described above, in secure federated learning, parameter sets w masked by client computers may be distinct from parameter setsreceived and unmasked by server computers. In such a case, the parameter errormay be a measure of the difference introduced by such masking operations.

[0042] In general, for otherwise equivalent machine learning models with equivalent inputs and different sets of parameters w and w' , the parameter error corresponding tothose is expected to positively correlate with inference error Thegreater the parameter error between sets of parameters w and w'. the greater the expected inference error between sets of inferences produced by those otherwise equivalent machinelearning models. Similarly, as parameter error is reduced, it is expected that the inference error will also reduce. When the parameter error is zero, the sets of parameters are identical, and therefore all aspects of the two machine learning models are now identical, and therefore the inferences produced by the two models are expected to be identical (or identically distributed).

[0043] However, for pairs of otherwise equivalent multilayer models, even small parameter differences (and thus small parameter error) can lead to large differences in inferences and large inference error. If the parameters for the first layer of two multilayer models are different, the outputs of the first layer may be different. Because the outputs of each layer are provided as inputs to the models’ respective subsequent layers, each subsequent layer is effectively processing the error introduced by the previous layers, in addition to contributing more error due to its own erroneous parameter set. As a result, inference error may accumulate with each successive layer in a multilayer machine learning model, and even small parameter error can result in large inference error due to this accumulation.

[0044] This phenomena relates to the problems with fully -private federated learning described above. Due to masking operations, there is parameter error between setsof parameter w generated by client computers and sets of parameters w' unmasked by server computers. This parameter error can result in large inference error due to the structure of multilayer machine learning models, leading to high inference error, which may result in machine learning models that are too inaccurate to perform real-world tasks.B. Artificial Neural Network

[0045] As summarized above, embodiments are directed to novel methods (and associated systems) for performing secure federated learning for multilayer machine learning models. Particularly, for machine learning models with output layers comprising Lipschitz continuous activation functions, embodiments of the present disclosure can provide a model performance guarantee, which can bound the inference error of the multilayer machine learning model being trained, and can thus reduce the inference error relative to fully-private methods of secure federated learning for multilayer machine learning models.

[0046] Artificial neural networks, particularly deep artificial neural networks, are machine learning models that (at the time of writing) see considerable use within the field of machinelearning. Artificial neural networks often comprise multilayer models, and can suffer from accumulation of inference error due to parameter error. In most cases, the output layers of artificial neural networks use activation functions that are Lipschitz continuous with respect to the parameters corresponding to those output layers. As such, methods according to embodiments are well-suited to for training artificial neural networks using secure federated learning, as methods according to embodiments can provide the abovementioned model performance guarantee for such models. As such, artificial neural networks are summarized in some detail below.

[0047] FIG. 1 depicts an artificial neural network comprising model layers 104-110. Each model layer can comprise artificial neurons, such as artificial neuron 112. The input 102 can be applied to the first model layer (model layer 104), and the output of this model layer can be applied as the input to the subsequent model layer (i.e.. model layer 106), and so on. until the output layer (model layer 1 10) produces the output(s) 114 of the machine artificial neural network.

[0048] Each artificial neuron (such as artificial neuron 112) may have some number of inputs (which may be produced by artificial neurons in previous model layers), which it may combine in some manner to produce an output, which may be applied as an input to an artificial neuron in a subsequent layer. An artificial neuron may combine its inputs to produce an output according to a propagation function and an activation function. For example, a propagation function p(z) (e.g., propagation function 118 in FIG. 1) can comprise a weighted combination of m numerical inputs, where wtrefers to the weight corresponding to the Ithinput zL(or corresponding to a general weighing term). Such weights can comprise the parameters characterizing the propagation function, and sets of such parameters can thereby characterize the artificial neural network as a whole:

[0049] Additionally, and in broad terms, an activation function , activation function120 in FIG. 1) can be used in conjunction with the propagation function p(z) to determine an artificial neuron’s output , for example:

[0050] A large variety of activation functions are used in artificial neural networks, such as the sigmoid activation function:

[0051] Artificial neural networks may have a variety of different topologies, which may describe how artificial neurons within the network are “connected”, e g., which neuron’s outputs comprise inputs of other neurons. One example is a “fully-connected” artificial neural network. In such a topology, each artificial neuron in a given layer has, as inputs, outputs from each artificial neuron in the previous layer. For example, if in a hypothetical neural network “layer 1” comprises five artificial neurons, “layer 2” comprises seven artificial neurons, and “layer 3” comprises four artificial neurons, then each artificial neuron in layer 2 can have five inputs (corresponding to the five outputs of the five neurons in layer 1) and each artificial neuron in layer 3 can have seven inputs (corresponding to the seven outputs of the seven artificial neurons in layer 2).

[0052] As noted in FIG. 1, in fully -private federated learning, a participating client computer, training an artificial neural network would “mask” the parameters corresponding to each layer of the artificial neural network (model layers 104-110), prior to sending those masked parameters to a server computer. In broad terms, while such masking operations can preserve client data privacy, they can also introduce inference error into the artificial neural network, which may be exacerbated by the artificial neural network's multilayer structure. This inference error may be sufficiently large that such models cannot be used for many real- world tasks, as they cannot achieve sufficient accuracy.II. FEDERATED LEARNING

[0053] A brief description of federated learning and secure federated learning, particularly centralized federated learning is provided below to facilitate a better understanding of embodiments of the present disclosure.

[0054] FIG. 2 shows a diagram of a system for performing federated learning and secure federated learning. As depicted in FIG. 2, in centralized federated learning, a plurality of client computers (e g., client computers 202-206) can each train their own local model (e g., client machine learning models 208-212) using their own local dataset (e.g., client datasets214-218). In the context of supervised learning, client datasets 214-218 can comprise labeled training data, e.g., images along with associated labels indicating whether those images e.g., depict a dog or do not depict a dog. In most applications of federated learning, client machine learning models 208-212 may be the same type of machine learning model with the same structure. For example, client machine learning models 208-212 may all comprise artificial neural networks with the same number of layers, the same topology, the same activation functions, etc.

[0055] As a result of this local training, client computers 202-206 can determine sets of model parameters corresponding to client machine learning models 208-212 (e.g., model parameters 220-224). If, for example, client machine learning models 208-212 comprise artificial neural networks, model parameters 220-224 could comprise weights corresponding to propagation functions in those artificial neural networks. Because each client computer has access to a different client dataset, and because no client computer has access to the other client computers’ datasets, each client computer may produce a different set of model parameters corresponding to their respective machine learning model.

[0056] In insecure federated learning, the client computers 202-206 can transmit model parameters 220-224 to a server computer 226, without masking or otherwise obscuring model parameters 220-224. The client computers 202-204 can transmit these model parameters to the server computer 226 via a communication network 228. The client computers 202-206 and server computer 226 can also otherwise communicate with one another via communication network 228 if necessary, e.g., in order to establish a communication channel over which to transmit model parameters. Communications network 228 can take any suitable form, and may include any one and / or the combination of the following: a direct interconnection; the Internet; a Local Area Network (LAN); a Metropolitan Area Network (MAN); an Operating Missions as Nodes on the Internet (OMNI); a secured custom connection; a Wide Area Network (WAN); a wireless network (e.g., employing protocols such as, but not limited to a Wireless Application Protocol (WAP), I-mode, and / or the like); and / or the like. Messages between computers and devices may be transmitted using a secure communications protocol, such as, but not limited to, File Transfer Protocol (FTP);HyperText Transfer Protocol (HTTP); Secure HyperText Transfer Protocol (HTTPS); Secure Socket Layer (SSL), ISO (e.g., ISO 8583) and / or the like.

[0057] After receiving the model parameters 220-224, the server computer 226 can combine model parameters 220-224, e.g., by determining an average set of model parameters. The server computer 226 can transmit these combined model parameters 230 back to the client computers 202-206, e.g., via communications network 228. Client computers 202-206 can then update client machine learning models 208-212 using the combined model parameters 230 (e.g., by replacing model parameters 220-224 with combined model parameters 230, adding combined model parameters 230 to model parameters 220-224. performing a weighted average of combined model parameters 230 and model parameters 220-224, or any other suitable update method).

[0058] While it is possible to complete a federated learning process after a single “cycle” (e.g., the steps described above or similar steps), in practice, it is unlikely that the client machine learning models 208-212 will converge or will be sufficiently trained in order to be useful in for real-world applications. As such, the process described above can be repeated any number of times until client machine learning models 208-212 are sufficiently trained. For example, client computers 202-206 can perform a subsequent round of local training on client machine learning models 208-212 using client datasets 214-218, thereby further updating model parameters 220-224. Model parameters 220-224 can then be sent back to server computer 226 via the communication network 228. After combining model parameters 220-224, server computer 226 can again transmit the new combined model parameters 230 back to client computers 202-204 via communication network 228. Client computers 202-206 can then update model parameters 220-224 using combined model parameters 230, and so forth until a sufficient number of cycles of federated learning have been performed.

[0059] Decentralized federated learning is similar to centralized federated learning, except there is no centralized server computer 226. Instead, in decentralized federated learning, each client computer of client computers 202-206 can transmit their respective model parameters of model parameters 220-224 to each other client computer. Each client computer can then generate combined model parameters 230 using their respective set of model parameters and each other set of model parameters received from each other client computer, then update their respective client machine learning model using combined model parameters 230.

[0060] Federated learning can be used in situations where clients are unwilling or unable to share their respective datasets with one another or the operator of server computer 226. If the clients were willing or able to share their respective datasets with one another or the operator of the server computer 226, it may be more efficient to join or merge client datasets 214-218 and perform more non-federated machine learning training using the combined client dataset as the training dataset. As an example, client datasets 214-218 could correspond to patient medical data, which may be protected by laws preventing the disclosure of such data without consent of the patients in question. As another example, certain countries may prevent the export of data corresponding to their citizens, which may prevent e.g., client computer 202 from transmitting client dataset 214 to server computer 226. By transmitting model parameters, rather than the client datasets themselves, federated learning offers a workaround for these issues.

[0061] However, simply transmitting model parameters instead of the client datasets used to generate those model parameters may be insufficient to address the security and privacy needs of clients. For example, in some cases it is possible to use a set of model parameters to learn information about the dataset used to generate those model parameters. If for example, client dataset 214 comprises private medical data, the owner or operator of client computer 202 may be obligated to make sure that model parameters 220 do not leak any private medical data from client dataset 214. Additionally, clients may wish to protect the model parameters 220-224 themselves. For example, if these model parameters 220-224 have some value to the clients and the operator of the server computer 226. these entities may not wish for these parameters to be stolen or intercepted over a potentially insecure communication network 228.

[0062] As such, instead of performing insecure federated learning, as described above, the client computers 202-206 and server computer 226 can perform secure federated learning. Secure federated learning is similar to insecure federated learning, except the client computers 202-206 can “mask” or otherwise obscure model parameters 220-224 before transmitting them to server computer 226 over a potentially insecure communication network 228. Server computer 226 can then unmask the masked model parameters, combine the unmasked model parameters to produce the combined model parameters 230. then transmit the combined model parameters 230 back to client computers 202-206. In some cases, servercomputer 226 may mask the combined model parameters 230 before transmitting them to client computers 202-204. The client computers 202-206 can then update client machine learning models 208-212 using the combined model parameters 230 (unmasking the combined model parameters 230 beforehand, if necessary), as described above.A. Masking

[0063] "‘Masking” can generally comprise operations or sets of operations that, either intentionally or unintentionally (although usually intentionally), obscure or hide the nature of the data being masked. Masking can generally protect both the model parameters 220-224 (e.g., from attackers or other malicious entities that might intercept the model parameters 220-224 as they are transmitted over communication network 228) and protect the client datasets 214-218 used to generate model parameters 220-224 (e.g., from not only such attackers, but also from other participants in the secure federated learning process, e g., the server computer 226).

[0064] One form of masking is encryption. Client computers 202-206 can encrypt their respective model parameters 220-224 using any appropriate encryption technique. Client computers 202-206 can then transmit their encrypted model parameters to the server computer 226, which can then unmask the model parameters by decrypting the model parameters before combining them. For example, each client computer can encrypt their respective model parameters using a public key corresponding to the server computer 226. The server computer 226 can then decrypt the encrypted model parameters using its corresponding private key. Provided the server computer’s private key remains private, malicious entities cannot determine model parameters 220-224 from any intercepted encrypted model parameters. However, while encryption may prevent outside observers from learning the model parameters 220-224, the server computer 226 still learns model parameters 220-224 via this decryption. As such, encryption alone would not prevent a malicious operator of server computer 226 from, e.g., attempting to leam information about client datasets 214-218 by analyzing the decry pted model parameters.

[0065] Another form of masking is compression. While often performed for practical or performance reasons (e.g., in order to reduce the amount of data that needs to be encrypted and transmitted over communication network 228), lossy compression techniques can resultin a loss of information. As such, any model parameters possessed by the server computer after decompressing compressed model parameters 220-224 may be different from model parameters 220-224. While this may result in a loss in accuracy (as described in further detail below), it has the privacy benefit of making it more difficult for the operator of the server computer 226 from using the decompressed model parameters to leam information about client datasets 214-218, and thus compression can potentially preserve client privacy.

[0066] Another form of masking is noise addition. Noise addition generally involves adding randomly (or pseudorandomly) generated values to data in order to mask the underlying data. As a consequence, the exact value of any given data point is indeterminate, as it is impossible for an outside observer to conclude exactly how much noise was added to each data point. If each client computer 202-206 adds noise to their respective model parameters 220-224, then an operator of the server computer 226 cannot determine the exact value of each parameter in model parameters 220-224. Like lossy compression, noise addition may result in a loss in accuracy, but it has the benefit of making it more difficult for the operator of the server computer 226 to use the noisy model parameters to leam information about client datasets 214-218, and thus noise addition can preserve client privacy.

[0067] Client computers 204-206 can perform any number of masking operations, including those described above and others, in any order as part of a masking process. For example, client computer 202 could mask model parameters 220 by encry pting those model parameters 220. Alternatively, client computer 202 could first compress the model parameters 220, then encrypt the compressed model parameters, then add noise to the encry pted compressed model parameters. As yet another alterative, client computer 202 could add noise to the model parameters, compress the noisy model parameters, then encry pt the compressed noisy model parameters.

[0068] The server computer 226 can unmask the masked model parameters by (in very general terms) reversing the masking operation performed by client computers 202-206. For example, if client computer 202 first compressed model parameters 220 then encrypted the compressed model parameters, server computer 226 could first decrypt the masked model parameters, then decompress the decrypted masked model parameters. Some operations, such as noise addition, may not be “reversible’7in this manner. However, due to the law oflarge numbers, for a sufficient number of client computers, the average of the noise across all sets of masked model parameters may be close to the expected value of the random process used to generate that noise. As such, the server computer 226 can potentially "denoise” the combined model parameters 230 by subtracting the expected value of the random noise generation process from combined model parameters 230. However, in some cases, particularly when there are a small number of client computers, there may be greater deviation between the average value of the noise across model parameters 220-224 and the expected value of the noise. As a consequence, the ‘'denoised” combined model parameters may be different than the combined model parameters 230 that would be generated in the absence of noise.B. Parameter Error Due to Masking

[0069] Masking operations generally introduce parameter error, as model parameters 220- 224 (represented by w) may be different from the unmasked model parameters (represented by w') produced by the server computer 226, e.g., due to noise addition and lossy compression as described above, or due to other sources of error (e.g., transmission error). As stated above, parameter error generally correlates to inference error. Because the combined model parameters 230 are generated by combining the (potentially erroneous) unmasked model parameters w' , the combined model parameters 230 may themselves be erroneous. Because the client computers 202-206 update their respective client machine learning models 208-212 using these potentially erroneous combined model parameters, inference error may be introduced into client machine learning models 208-212 during secure federated learning. Briefly referring back to FIG. 1, for an artificial neural network comprising multiple model layers 104-110, in fully -private secure federated learning, all model parameters corresponding to all of these model layers 104-110 would be masked. As a consequence, during fully -private secure federated learning, inference error may be introduced to every model layer in an artificial neural network.

[0070] As described above, for multilayer models, even relatively small parameter error (due, e.g., to a small difference between the model parameters 220-224 w produced by client computers 202-206 and the unmasked model parameters w' produced by server computer 226) can lead to large inference error, due to the accumulation of errors resulting from the sequential layered structure of multilayer models. As such, fully-private secure federatedlearning methods may be unsuitable for training multilayer models (such as deep artificial neural networks), due to parameter error introduced by masking, which may result in trained models that are too inaccurate to apply to real-world tasks due to inference error resulting from parameter error.III. SECURE FEDERATED LEARNING METHOD

[0071] As summarized further above, some embodiments of the present disclosure are directed to methods for performing secure federated learning. These methods address some of the inference error related problems associated with fully -private secure federated learning, and further provide a model performance guarantee. Some differences between secure federated learning methods according to embodiments and fully -private secure federated learning methods are illustrated by the artificial neural networks depicted in FIGs. 1 and 3. As depicted in FIG. 1, in fully -private secure federated learning, all parameters corresponding to all model layers 104-110 are masked during fully -private secure federated learning. This can result in parameter error in each model layer 104-106, which can result in an accumulation of inference error.

[0072] By contrast, as depicted in FIG. 3, in methods according to embodiments, only the parameters corresponding to output model layer 310 are masked. Non-output model layer parameters (e.g., corresponding to non-output layers 304-308) are not masked. In general, masking only the output model layer parameters is sufficient to preserve client privacy. Further, because only the output model layer parameters are masked, parameter error is limited to the output model layer 310, and thus there is no accumulation of inference error as data passes through the non-output model layers 304-308.

[0073] Further, if the output model layer 310 uses activation functions (e.g., activation function 320) that are continuous (e.g., Lipschitz continuous) with respect to output layer parameters 316, then a bound relating parameter error and inference error can be demonstrated (see Section IV further below). This bound forms the basis of the model performance guarantee. Provided that a practitioner of methods according embodiments can control or estimate the amount of parameter error introduced by masking, that practitioner can estimate the amount of inference error induced by that parameter error. This enables a practitioner of embodiments to exercise greater control over the privacy-accuracy tradeoffpresent in secure federated learning. A practitioner could, for example, mask the model parameters such that privacy is maximized for a given accuracy criteria.

[0074] Fully-private methods of secure federated learning do not provide this model performance guarantee. As a consequence, practitioners of fully-private secure federated learning cannot determine the exact impact of parameter error (or masking processes that introduce such parameter error) on model inference error. While much research has been performed to discover methods of reducing parameter error, considerable inference error can still occur even with reduced parameter error, and as such fully-private secure federated learning techniques may be unsuitable for training machine learning models for real-world tasks. Further, the lack of a model performance guarantee makes it difficult for practitioners to estimate or otherwise determine the amount of inference error introduced by parameter error.A. Method Flowchart

[0075] Some methods of performing secure federated learning according to embodiments are described below with reference to the flowchart of FIG. 4. Such methods can be performed to train a machine learning model (or multiple machine learning models) using federated learning. Such methods can be performed by a plurality of client computers and a server computer performing the secure federated learning process. The client computers and the server computer can be configured to perform their respective steps, e.g., using computer programming or other suitable computer configuration processes.

[0076] At step 402, a plurality of client computers and a server computer can perform any steps associated with setting up the secure federated learning process. These steps can include, e.g., exchanging cryptographic keys, thereby enabling the client computers to mask sets of output layer parameters (via encry ption), and further enabling the client computers to securely transmit those model parameters via a network such as the Internet. Further, the client computers and server computer can determine or otherwise agree upon a particular type of multilayer machine learning model (e.g., an artificial neural network) to train. In some cases, the server computer can initialize the multilayer machine learning model (e.g., by generating a random parameter set corresponding to that model) and transmit the initialized machine learning model to each of the client computers. The client computers and servercomputer (or their operators) can agree upon training criteria (e.g., the number of cycles, number of training rounds within a cycle, any terminating conditions, optimization methods, batch sizes, etc.) and any hyperparameters associated with model training.

[0077] At step 404, each client computer of the plurality of client computers can train their local multilayer machine learning model using a set of training samples, thereby each generating a set of client model parameters. In some embodiments, in which the client computers perform multiple cycles of federated learning, the set of training samples may be sampled from a larger set of training samples (e.g., a client dataset) and may comprise a first set of training samples. In subsequent cycles of federated learning, each client computer may sample the same set of training samples or a different set of training samples (e.g., a second set of training samples, one or more additional sets of training samples, etc.). The client computers can train their local machine learning models using any appropriate model training procedure, provided it is consistent with the conditions and criteria established by the plurality of client computers and the server computer at step 402.

[0078] Each set of client model parameters can comprise a set of client output layer parameters corresponding to an output layer of the machine learning model, and a set of nonoutput layer parameters corresponding to one or more non-output layers of the machine learning model. The output layer of the machine learning model can use one or more continuous activation functions, which may be continuous with respect to the output layer parameters corresponding to the output layer. In some embodiments, the one or more continuous activation functions may be Lipschitz continuous with respect to the set of client output layer parameters. As described in more detail in Section IV below, if the output layer activation functions are Lipschitz continuous with respect to the set of client output layer parameters, then a bound on inference error can be demonstrated, resulting in a model performance guarantee.

[0079] Methods according to embodiments can be practiced with any appropriate multilayer machine learning model. In some embodiments, the local machine learning models trained by the client computers can comprise neural networks. In such a model, the output layer can comprise an output neural network layer, and the one or more non-output layers can comprise one or more non-output neural network layers. These one or more non- output neural network layers can comprise an input neural network layer and one or morehidden neural network layers. In such embodiments, each set of client output layer parameters can comprise a set of output layer weights characterizing the output neural network layer, e.g., a set of weights corresponding to one or more propagation functions used in the output neural network layer. Likewise, the one or more sets of non-output layer parameters corresponding to each client-trained machine learning model can comprise one or more sets of non-output layer weights characterizing the one or more non-output neural network layers, e.g., one or more sets of weights corresponding to one or more propagation functions used in the non-output neural network layers.

[0080] At step 406, each client computer can mask its respective set of client output layer parameters, thereby producing a set of masked client output layer parameters. These sets of masked client output layer parameters (which can be represented by w in the description below) may later be transmitted to the server computer (e.g., in step 408). The server computer can unmasked these masked client output layer parameters to produce sets of server output layer parameters (iv'), which can be combined to produce a combined set of parameters. As described above, masking a set of client output layer parameters may cause a loss of accuracy, e.g., due to inference error introduced by the masking process. As a result, the server output layer parameters w' may be different from the client output layer parameters w. Unlike in fully-private secure federated learning, each client computer may mask its respective set of client output layer parameters, and may not mask its set of non- output layer parameters.

[0081] As described above, there are a variety of masking operations that can be used, alone or in combination, to mask data. As such, the client computers can use any appropriate masking operations to mask the sets of client output layer parameters. In some embodiments, masking a set of client output layer parameters can comprise encry pting the set of client output layer parameters, thereby producing the set of masked client output layer parameters. In some embodiments, masking the set of client output layer parameters can comprise compressing the set of client output layer parameters, thereby producing the set of masked client output layer parameters. In some embodiments, masking the set of client output layer parameters can comprise adding noise to the set of client output layer parameters, thereby producing the set of masked client output layer parameters.

[0082] In other embodiments, masking a set of client output layer parameters can comprise compressing the set of client output layer parameters, thereby producing a set of compressed client output layer parameters. A client computer can then encrypt the set of compressed client output layer parameters, thereby producing a set of encrypted compressed client output layer parameters. The client computers can add random noise (e.g., random values sampled from a Gaussian distribution) to the set of encrypted compressed client output layer parameters, thereby producing the set of masked client output layer parameters. However, it should be understood that other masking procedures are also possible. A client computer could, for example, add noise to their set of output layer parameters prior to compressing and encrypting those output layer parameters.

[0083] In some embodiments, the masking process may be controlled, so that the sets of client output layer parameters w are masked such that an absolute difference between a set of client output layer parameters w and a corresponding set of server output layer parameters w' is less than a difference threshold T. In some embodiments, the absolute difference can comprise an absolute value of a Euclidean distance between a first vector representing the set of client output layer parameters w and a second vector representing the corresponding set of server output layer parametersIn some cases, the difference threshold T mav be of the order — In such cases, n mavcomprise a number of training data points in a client computer’s set of training samples (or e.g., the average number of training data points across all client computers’ sets of training samples) and p may comprise an integer greater than one. As such, in some embodiments, the difference threshold T may be proportional to a number of training samples (i.e., n) in the set of training samples used by a client computer to train their respective machine learning model. Secure federated learning theory may guarantee that it is possible to determine an integer p greater than one, e.g., by controlling the masking process (e.g., by adding more or less noise). This aspect of secure federated learning theory is described in more detail in Chen et al.’s “The Fundamental Price of Secure Aggregation in Differentially Private Federated Learning.”

[0084] At step 408, each client computer can transmit their respective set of non-output layer parameters and respective set of masked output layer parameters to the server computer. The client computers can transmit their respective sets of non-output layer parameters andmasked output layer parameters to the server computer via any appropriate means (e.g., via a communication network such as the Internet) and using any appropriate communication protocol. In this way the server computer can receive a plurality of respective sets of nonoutput layer parameters and a plurality of respective sets of masked client output layer parameters from the plurality of client computers. As described above, each set of masked client output layer parameters can correspond to a respective client trained machine learning model, and each respective client trained machine learning model can be trained by a different client computer of the plurality of client computers. As described in more detail in steps 410-412 below, the server computer can be configured to determine a combined set of parameters based on the sets of non-output layer parameters and the sets of masked client output layer parameters. This set of combined parameters can correspond to a combined machine learning model, i.e., a machine learning model generated using the sets of parameters received from the client computers.

[0085] At step 410, the server computer can unmask the plurality of respective sets of masked client output layer parameters, thereby producing a plurality of sets of server output layer parameters (w'). As described above, the server computer's unmasking process may depend on the masking process used by the client computers to mask the plurality’ of respective sets of client output layer parameters. If, for example, a client computer generated a respective set of masked client output layer parameters by encrypting a set of client output layer parameters, then the server computer can be configured to unmask the plurality of respective sets of masked client output layer parameters by decrypting the plurality of respective sets of masked client output layer parameters, thereby producing the plurality of sets of server output layer parameters. Alternatively, if the masking process comprises compression, the serv er computer can be configured to unmask the plurality of respective sets of masked client output layer parameters by decompressing the plurality of respective sets of masked client output layer parameters, thereby producing the plurality of sets of server output layer parameters.

[0086] In some embodiments, in which the masking process comprises both compression and encryption (and optionally noise addition), the server computer can unmask the plurality of sets of masked client output layer parameters (thereby producing the plurality of sets of server output layer parameters), by performing both decryption and decompressionoperations. The server computer can decry pting the plurality of respective sets of masked client output layer parameters, thereby producing a plurality of sets of decrypted masked client output layer parameters. For example, if each client computer encrypted their respective set of client output layer parameters using the server computer’s public key, then the server computer could decrypt the plurality of respective sets of masked client output layer parameters using the server computer’s private key. Afterwards, the server computer can decompress the plurality of sets of decrypted masked client output layer parameters, thereby producing a plurality of sets of server output layer parameters.

[0087] At step 412, the server computer can produce the combined set of parameters based on the plurality' of sets of non-output layer parameter and the plurality' of sets of server output layer parameters. The server computer can do this by producing a combined set of nonoutput layer parameters and a combined set of non-output layer parameters. The combined set of parameters can comprise both the combined set of non-output layer parameters and the combined set of server output layer parameters. This combined set of parameters can comprise a complete set of parameters that can be used to characterize the combined machine learning model. The server computer can produce the combined set of non-output layer parameters by combining the plurality of respective sets of non-output layer parameters, thereby producing the combined set of non-output layer parameters. Likewise, the server computer can produce the combined set of server output layer parameters by combining the plurality of sets of server output layer parameters, thereby producing the combined set of server output layer parameters.

[0088] The server computer can combine or aggregate the plurality of sets of server output layer parameters to produce the combined set of server output layer parameters using any appropriate combination or aggregation technique, and the same is true for the plurality' of sets of non-output layer parameters and the combined set of non-output layer parameters. As an example, in some embodiments, the server computer can produce a set of combined server output layer parameters by performing a weighted average of the plurality' of sets of server output layer parameters, e.g., by averaging each server output layer parameter with the corresponding output layer parameters (e.g., corresponding to the same weight of the same propagation function in an artificial neural network) in each other set of server output layer parameters. In some cases, the weighted average may be equally weighted. As a simplifiedexample, if there were three sets of server output layer parameters, each comprising five parameters, the server computer could compute five parameter averages, each parameter average corresponding to three corresponding parameters from the three sets of server output layer parameters (e g., one average corresponding to the “first parameter” from each of the three sets of server output layer parameters, another average corresponding to the “second parameter” from each of the three sets of server output layer parameters, and so on), and those five averages could comprise the set of combined server output layer parameters.

[0089] Likewise, the server computer can produce the set of combined non-output layer parameters by performing a weighted average of the plurality of sets of non-output layer parameters, or using any other appropriate combination technique. As a simplified example, if there were three client computers, each training an artificial neural netw ork w ith four non- output layers, and each set of non-output layer parameter comprised five parameters, then each client computer may produce four sets of non-output layer parameters corresponding to each layer of their local artificial neural network model, each set of non-output layer parameters comprising five parameters. In such a case, the server computer could determine four subsets of combined non-output layer parameters corresponding to each of the four layers of the artificial neural network. The set of combined non-output layer parameters could comprise these four subsets of combined non-output layer parameters collectively. The server computer could produce each subset of non-output layer parameters by averaging (or otherwise aggregating) the non-output layer parameters corresponding to each layer of the artificial neural network. For example, the server computer could average the sets of non- output layer parameters corresponding to the first layer of the artificial neural to produce a subset of combined non-output layer parameters corresponding to that first layer, and so on for the second, third, and fourth layers.

[0090] At step 414, the server computer can transmit the combined set of parameters to the plurality of client computers. The server computer can transmit the combined set of parameters to the plurality of client computers using any appropriate communication protocol and transmission medium, e.g., via a network such as the Internet. In this way, the plurality of client computers can receive the combined set of parameters from the server computer. As described in step 416 below, the plurality of client computers can be configured to updatetheir respective client-trained machine learning models using the combined set of parameters, thereby training the combined machine learning model.

[0091] At step 416, each client computers can update its respective machine learning model using the combined set of parameters, thereby producing an updated machine learning model. The client computers can update their respective machine learning model using any appropriate method consistent with federated learning. For example, a client computer can update a machine learning model by replacing the model parameters associated with that model with the combined set of parameters. Alternatively, a client computer can update a machine learning model by averaging the combined set of parameters and the existing set of parameters associated with that machine learning model. In some embodiments, the client computer can update the machine learning model using the set of training samples in addition to the combined set of parameters, e.g., by retraining the machine learning model using the set of training samples and the combined set of parameters as an initial parameter estimate.

[0092] As describe above, multiple cycles of federated learning may be needed to achieve model convergence or produce a useful machine learning model. As such, at step 418, the client computers and server computer can perform additional cycles of federated learning as necessary, i.e., additional training processes on the updated machine learning model produced at step 418. As such, in some embodiments, the set of training samples used by a client computer can comprise a first set of training samples, and in a subsequent round of training a second set of training samples and / or any number of additional sets of training samples (e.g., a third set of training samples, etc.) can be sampled by the client computer and used to train the machine learning model. Likewise, the set of client output layer parameters described above can comprise a first set of client output layer parameters, the set of non-output layer parameters can comprise a first set of non-output layer parameters, the combined set of parameters comprises a first combined set of parameters, and the method can further comprise performing an additional training process on the updated machine learning model.

[0093] The additional training process can involve each client computer training their respective updated machine learning model using the first set of training samples and / or a second set of training samples. In this way, the client computers can generate a plurality7of second sets of client output layer parameters corresponding to the output layer and a second set of non-output layer parameters corresponding to the one or more non-output layers. Theclient computers can mask the second sets of client output layer parameters, thereby producing second sets of masked client output layer parameters, using, e.g.. the masking techniques described above. The client computers can transmit the second plurality of sets of non-output layer parameters and the second plurality of sets of masked client output layer parameters to the server computer. The server computer can determine a second combined set of parameters based on a second plurality of respective sets of non-output layer parameters and a second plurality of respective sets of masked client output layer parameters from the plurality of client computers. This second combined set of parameters can be transmitted to the plurality of client computers, which can be configured to update their respective client-trained machine learning models using the second combined set of parameters (e.g., as described above), thereby training the combined machine learning model.

[0094] If necessary, one or more additional training processes can be performed by repeating the additional training process (described above) one or more times, using the first set of training samples and / or the second set of training samples and / or one or more additional sets of training samples (e.g., a third set of training samples). The additional training processes can be performed until a terminating condition has been met. e.g., the combined model convergence or a defined number of cycles have been performed by the client computers and server computer.IV. MODEL PERFORMANCE GUARANTEE

[0095] As mentioned above, methods according to embodiments provide a model performance guarantee by only masking the output layer parameters during federated learning, particularly when that output layer corresponds to a Lipschitz continuous activation function. This performance guarantee bounds the inference error resulting from parameter error introduced during the federated learning process (e.g., due to masking operations). The following paragraphs briefly describe the mathematical basis for this performance guarantee.

[0096] Let refer to a plurality of sets of client output layerparameters, which (as described above) a plurality' of client computers can determine or produce when training their respective local machine learning models. Each client computer can mask their respective set of client output layer parameters (e.g., by compressing and / or encrypting their respective sets of client output layer parameters and / or adding noise), andproduce a set of masked client output layer parameters, which can be transmitted to the server computer. The server computer can unmask the sets of masked client output layer parameters (e.g., by decrypting and decompressing the masked client output layer parameters) to produce a plurality of sets of server output layer parameters, referred to by w' = (IVQ, W[, w^).

[0097] The masking process used by the client computers may result in a loss of accuracy, and as such, the server output layer parameters w' may not be equal to the client output layer parameters w, for example, due to the addition of noise or a lossy compression system. As described above, this difference in the server output layer parameters w' and client output layer parameters w. can result in inference error. However, secure federated learning theory enables the difference between w and w' to be controlled with some upper bound, e.g., by controlling the masking methods used to mask the client output layer parameters. This aspect of secure federated learning theory is described in more detail in Chen et al ’s “The Fundamental Price of Secure Aggregation in Differentially Private Federated Learning.”

[0098] If it is assumed that the parameter error between w and w' can be controlled such that , and it is assumed that all activation functions fkcorresponding tothe output layers of the client-trained models are Lipschitz continuous (with respect to the client output layer parameters w or server output layer parameters w'), then it can be shown that the inference error corresponding to a given activation function is bounded:

[0099] In the formulas above, n is the number of training data points in a given client’s training data set (or e.g., the average number of training data points among all client’s training data sets), p is an integer greater than 1 from federated learning theory7(which may depend on the masking methods used), and M can comprise the maximum slope of the output layer activation function, which can comprise a number independent of the client output layer parameters w and the server output layer parameters w' . This formula demonstrates that the error in model outputs (the inference error) can be well-controlled, aswhen n is large.

[0100] It should be understood that the formula above demonstrates a bound on the inference error under the assumption that the activation function isLipschitz continuous with respect to the client output layer parameters w and server output layer parameters w'. It does not demonstrate a bound on the inference error if the activation function fkis not Lipschitz continuous. However, despite being yet unproven, a similar bound may exist for broader classes of continuous functions, such as Holder continuous functions or uniformly continuous functions.

[0101] Regardless, with few exceptions (e g., exponential activation functions), most activation functions in deep neural networks are Lipschitz continuous wi th respect to the model parameters (weights) associated with those activation functions. As such the formula above demonstrates that for most deep neural networks, methods according to embodiments can be used to control inference errorresulting from parameter error This is shown below for the sigmoid activation function which isLipschitz continuous with respect to the model parameters:

[0102] As with the more general formula presented further above, one interpretation of this formula is that the inference error is bounded by the parameter errorsuch that when the parameter error || w' — w|| is small, the inference error is also small. As such, if the parameter error can becontrolled by controlling the masking process, then the inference errorcan likewise be controlled.

[0103] Methods according to embodiments can be practiced using any machine learning model with multiple layers and a Lipschitz continuous activation layer, including hybrid machine learning models and ensemble machine learning models. For example, a hybrid machine learning model could comprise a transformer model, a convolutional neural network, and an linear output layer with a softmax activation function. The input to the model (e.g., an image that either depicts a dog or does not depict a dog) could be input into the transformer model, which can produce a first embedding. This first embedding could then be input into the convolutional neural network to produce a second embedding. The second embedding can then be input into the output layer, which can then produce an inference, indicating whether the input image depicts a dog or does not depict a dog. Each sub-model (transformer, convolutional neural network, and linear output layer) could each have its own set of sub-model parameters. However, using methods according to embodiments, only the parameters corresponding to the linear output layer could be masked during federated learning, the parameters corresponding to the transformer sub-model and convolutional neural network could be transmitted in unmasked form. Because the output linear layer has a softmax activation function (which is Lipschitz continuous), the performance guarantee described above will still hold, and the inference error will be bounded. By contrast, in a fully-private federated learning method, in which all the parameters corresponding to all the sub-models are masked, no performance guarantee can be provided, and the depth of the model may result in considerable inference error.V. COMPUTER SYSTEM

[0104] Any of the computer systems mentioned herein may utilize any suitable number of subsystems. Examples of such subsystems are shown in FIG. 5 in computer system 500. In some embodiments, a computer system includes a single computer apparatus, where the subsystems can be the components of the computer apparatus. In other embodiments, a computer system can include multiple computer apparatuses, each being a subsystem, with internal components. A computer system can include desktop and laptop computers, tablets, mobile phones and other mobile devices.

[0105] The subsystems shown in FIG. 5 are interconnected via a system bus 512. Additional subsystems such as a printer 508, keyboard 518, storage device(s) 520, monitor 524 (e.g., a display screen, such as an LED), which is coupled to display adapter 514, andothers are shown. Peripherals and input / output (I / O) devices, which couple to I / O controller 502, can be connected to the computer system by any number of means known in the art such as input / output (I / O) port 516 (e.g., USB, FireWire®). For example, I / O port 516 or external interface 522 (e.g. Ethernet, Wi-Fi, etc.) can be used to connect computer system 500 to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via system bus 512 allows the central processor 506 to communicate with each subsystem and to control the execution of a plurality of instructions from system memory 504 or the storage device(s) 520 (e.g., a fixed disk, such as a hard drive, or optical disk), as well as the exchange of information between subsystems. The system bus 512 may “couple” the system memory 504 to the central processor 506. The system memory' 504 and / or the storage device(s) 520 may embody a computer readable medium, which may be non-transitory. Another subsystem is a data collection device 510, such as a camera, microphone, accelerometer, and the like. Any of the data mentioned herein can be output from one component to another component and can be output to the user.

[0106] A computer system can include a plurality of the same components or subsystems, e.g., connected together by external interface 522, by an internal interface, or via removable storage devices that can be connected and removed from one component to another component. In some embodiments, computer systems, subsystem, or apparatuses can communicate over a network. In such instances, one computer can be considered a client and another computer a server, where each can be part of a same computer system. A client and a server can each include multiple systems, subsystems, or components.

[0107] Any of the computer systems mentioned herein may utilize any suitable number of subsystems. In some embodiments, a computer system includes a single computer apparatus, where the subsystems can be components of the computer apparatus. In other embodiments, a computer system can include multiple computer apparatuses, each being a subsystem, with internal components.

[0108] A computer system can include a plurality of the components or subsystems, e.g., connected together by external interface or by an internal interface. In some embodiments, computer systems, subsystems, or apparatuses can communicate over a network. In such instances, one computer can be considered a client and another computer a server, where eachcan be part of a same computer system. A client and a server can each include multiple systems, subsystems, or components.

[0109] It should be understood that any of the embodiments of the present invention can be implemented in the form of control logic using hardware (e.g., an application specific integrated circuit or field programmable gate array) and / or using computer software with a generally programmable processor in a modular or integrated manner. As used herein a processor includes a single-core processor, multi-core processor on a same integrated chip, or multiple processing units on a single circuit board or networked. Based on the disclosure and teachings provided herein, a person of ordinary skill in the art will know and appreciate other ways and / or methods to implement embodiments of the present invention using hardware and a combination of hardware and software.

[0110] Any of the software components or functions described in this application may be implemented as software code to be executed by a processor using any suitable computer language such as, for example, Java, C, C++, C#, Objective-C, Swift, or scripting language such as Perl or Python using, for example, conventional or object-oriented techniques. The software code may be stored as a series of instructions or commands on a computer readable medium for storage and / or transmission, suitable media include random access memory (RAM), a read only memory (ROM), a magnetic medium such as a hard-drive or a floppy disk, or an optical medium such as a compact disk (CD) or DVD (digital versatile disk), flash memory, and the like. The computer readable medium may be any combination of such storage or transmission devices.

[0111] Such programs may also be encoded and transmitted using carrier signals adapted for transmission via wired, optical, and / or wireless networks conforming to a variety of protocols, including the Internet. As such, a computer readable medium according to an embodiment of the present invention may be created using a data signal encoded with such programs. Computer readable media encoded with the program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer readable medium may reside on or within a single computer product (e.g. a hard drive, a CD, or an entire computer system), and may be present on or within different computer products within a system or network. A computer system may include a monitor, printer or other suitable display for providing any of the results mentioned herein to a user.

[0112] Any of the methods described herein may be totally or partially performed with a computer system including one or more processors, which can be configured to perform the steps. Thus, embodiments can be involve computer systems configured to perform the steps of any of the methods described herein, potentially with different components performing a respective steps or a respective group of steps. Although presented as numbered steps, steps of methods herein can be performed at a same time or in a different order. Additionally, portions of these steps may be used with portions of other steps from other methods. Also, all or portions of a step may be optional. Additionally, and of the steps of any of the methods can be performed with modules, circuits, or other means for performing these steps.

[0113] The specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of embodiments of the invention. However, other embodiments of the invention may be involve specific embodiments relating to each individual aspect, or specific combinations of these individual aspects. The above description of exemplary embodiments of the invention has been presented for the purpose of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise form described, and many modifications and variations are possible in light of the teaching above. The embodiments were chosen and described in order to best explain the principles of the invention and its practical applications to thereby enable others skilled in the art to best utilize the invention in various embodiments and with various modifications as are suited to the particular use contemplated.

[0114] The above description is illustrative and is not restrictive. Many variations of the invention will become apparent to those skilled in the art upon review of the disclosure. The scope of the invention should, therefore, be determined not with reference to the above description, but instead should be determined with reference to the pending claims along with their full scope or equivalents.

[0115] One or more features from any embodiment may be combined with one or more features of any other embodiment without departing from the scope of the invention.

[0116] A recitation of “a”, '‘an’’ or “the” is intended to mean “one or more” unless specifically indicated to the contrary. The use of “or” is intended to mean an “inclusive or,” and not an “exclusive or” unless specifically indicated to the contrary .

[0117] All patents, patent applications, publications and description mentioned herein are incorporated by reference in their entirety for all purposes. None is admitted to be prior art.

Claims

WHAT IS CLAIMED IS:

1. A method performed by a client computer for training a machine learning model using federated learning, the method comprising: training the machine learning model using a set of training samples, thereby- generating (1) a set of client output layer parameters corresponding to an output layer of the machine learning model and (2) a set of non-output layer parameters corresponding to one or more non-output layers of the machine learning model, wherein the output layer uses one or more continuous activation functions; masking the set of client output layer parameters, thereby producing a set of masked client output layer parameters, wherein masking the set of client output layer parameters causes a loss of accuracy; transmitting the set of non-output layer parameters and the set of masked client output layer parameters to a server computer, wherein the server computer is configured to determine a combined set of parameters based on a plurality of respective sets of non-output layer parameters and a plurality of respective sets of masked client output layer parameters from a plurality of client computers, including the client computer; receiving, from the server computer, the combined set of parameters; and updating the machine learning model using the combined set of parameters, thereby producing an updated machine learning model.

2. The method of claim 1, wherein the plurality of respective sets of non- output layer parameters and the plurality of respective sets of masked client output layer parameters correspond to a plurality of machine learning models trained by the plurality of client computers.

3. The method of claim 1, wherein masking the set of client output layer parameters comprises adding random noise to the set of client output layer parameters, thereby producing the set of masked client output layer parameters.

4. The method of claim 1, wherein the client computer updates the machine learning model using the set of training samples in addition to the combined set of parameters.

5. The method of claim 1, wherein the server computer is configured to produce the combined set of parameters by being configured to: unmask the plurality of respective sets of masked client output layer parameters, thereby producing a plurality of sets of server output layer parameters; combine the plurality of respective sets of non-output layer parameters, thereby producing a combined set of non-output layer parameters: and combine the plurality of sets of server output layer parameters, thereby producing a combined set of server output layer parameters, wherein the combined set of parameters comprises the combined set of non-output layer parameters and the combined set of server output layer parameters.

6. Thet method of claim 5, wherein the set of client output layer parameters are masked such that an absolute difference between the set of client output layer parameters and a corresponding set of server output layer parameters is less than a difference threshold.

7. The method of claim 6, wherein the absolute difference comprises an absolute value of a Euclidean distance between a first vector representing the set of client output layer parameters and a second vector representing the corresponding set of server output layer parameters.

8. The method of claim 6, wherein the difference threshold is proportional to a number of training samples in the set of training samples.

9. The method of claim 5, wherein: masking the set of client output layer parameters comprises encrypting the set of client output layer parameters, thereby producing the set of masked client output layer parameters; and the server computer is configured to unmask the plurality of respective sets of masked client output layer parameters by decry pting the plurality of respective sets of masked client output layer parameters, thereby producing the plurality of sets of server output layer parameters.

10. The method of claim 5, wherein:masking the set of client output layer parameters comprises compressing the set of client output layer parameters, thereby producing the set of masked client output layer parameters; and the server computer is configured to unmask the plurality of respective sets of masked client output layer parameters by decompressing the plurality of respective sets of masked client output layer parameters, thereby producing the plurality of sets of server output layer parameters.

11. The method of claim 5, wherein: masking the set of client output layer parameters comprises: compressing the set of client output layer parameters, thereby producing a set of compressed client output layer parameters, encrypting the set of compressed client output layer parameters, thereby producing a set of encry pted compressed client output layer parameters, and adding random noise to the set of encrypted compressed client output layer parameters, thereby producing the set of masked client output layer parameters; and the server computer is configured to unmask the plurality of respective sets of masked client output layer parameters by being configured to: decrypt the plurality of respective sets of masked client output layer parameters, thereby producing a plurality’ of sets of decrypted masked client output layer parameters, and decompress the plurality of sets of decry pted masked client output layer parameters, thereby producing the plurality of sets of server output layer parameters.

12. The method of claim 1, wherein: the machine learning model comprises a neural network; the output layer comprises an output neural network layer; the one or more non-output layers comprise one or more non-output neural network layers comprising an input neural network layer and one or hidden neural network layers;the set of client output layer parameters comprise a set of output layer weights characterizing the output neural network layer; and the one or more sets of non-output layer parameters comprise one or more sets of non-output layer weights characterizing the one or more non-output neural network layers.

13. The method of claim 1, wherein the one or more continuous activation function are Lipschitz continuous with respect to the set of client output layer parameters.

14. The method of claim 1, wherein the set of training samples comprises a first set of training samples, wherein the set of client output layer parameters comprises a first set of client output layer parameters, wherein the set of non-output layer parameters comprises a first set of non-output layer parameters, wherein the combined set of parameters comprises a first combined set of parameters, and wherein the method further comprises performing an additional training process on the updated machine learning model, the additional training process comprising: training the updated machine learning model using the first set of training samples and / or a second set of training samples, thereby generating (3) a second set of client output layer parameters corresponding to the output layer and (4) a second set of non-output layer parameters corresponding to the one or more non-output layers; masking the second set of client output layer parameters, thereby producing a second set of masked client output layer parameters; transmitting the second set of non-output layer parameters and the second set of masked client output layer parameters to the server computer, wherein the server computer is configured to determine a second combined set of parameters based on a second plurality of respective sets of non-output layer parameters and a second plurality of respective sets of masked client output layer parameters from the plurality of client computers; receiving, from the server computer, the second combined set of parameters; and updating the updated machine learning model using the second combined set of parameters.

15. The method of claim 14, further comprising performing one or more additional training processes by repeating the additional training process one or more timesusing the first set of training samples and / or the second set of training samples and / or one or more additional sets of training samples.

16. A method performed by a server computer for training a combined machine learning model using federated learning, the method comprising: receiving a plurality of sets of masked client output layer parameters from a plurality of client computers, each set of masked client output layer parameters corresponding to an output layer of a respective client-trained machine learning model used to determine the combined machine learning model, wherein each respective client-trained machine learning model is trained by a different client computer of the plurality of client computers; receiving a plurality of sets of non-output layer parameters from the plurality of client computers, wherein each set of non-output layer parameters corresponds to one or more non-output layers of the respective client-trained machine learning model; unmasking the plurality of sets of masked client output layer parameters, thereby producing a plurality of sets of server output layer parameters; producing a combined set of parameters based on the plurality of sets of non- output layer parameters and the plurality of sets of server output layer parameters; and transmitting the combined set of parameters to the plurality of client computers, wherein the plurality of client computers are configured to update the respective client-trained machine learning model using the combined set of parameters, thereby training the combined machine learning model.

17. The method of claim 16, wherein the plurality of sets of masked client output layer parameters comprise a first plurality of sets of masked client output layer parameters, wherein the plurality of sets of non-output layer parameters comprise a first plurality of sets of non-output layer parameters, wherein the combined set of parameters comprises a first combined set of parameters, and wherein the method further comprises performing an additional training process on the combined machine learning model, the additional training process comprising: receiving a second plurality’ of sets of masked client output layer parameters from the plurality of client computers, each set of masked client output layer parameters corresponding to the output layer of a respective client-trained machine learning model;receiving a second plurality of sets of non-output layer parameters from the plurality’ of client computers, wherein each set of non-output layer parameters corresponds to one or more non-output layers of a respective client trained machine learning model; producing a second combined set of parameters based on the plurality of sets of non-output layer parameters and the plurality of sets of masked client output layer parameters; and transmitting the second combined set of parameters to the plurality’ of client computers, wherein the plurality of client computers are configured to update the respective client-trained machine learning model using the second combined set of parameters, thereby training the combined machine learning model.

18. The method of claim 16, wherein: unmasking the plurality’ of sets of masked client output layer parameters, thereby producing the plurality of sets of server output layer parameters comprises: decrypting the plurality of sets of masked client output layer parameters, thereby producing a plurality of sets of decrypted masked client output layer parameters, decompressing the plurality of sets of decrypted masked client computer output layer parameters, thereby producing a plurality of sets of server output layer parameters; and producing a combined set of parameters based on the plurality of sets of non- output layer parameters and the plurality' of sets of server output layer parameters comprises: producing a set of combined server output layer parameters by performing a weighted average of the plurality of sets of server output layer parameters, and producing a set of combined non-output layer parameters by performing a weighted average of the plurality of sets of non-output layer parameters, wherein the combined set of parameters comprises the set of combined server output layer parameters and the set of combined non-output layer parameters.

19. The method of claim 16, wherein each client computer is configured to produce a corresponding set of masked client output layer parameters by:compressing a set of client output layer parameters corresponding to the output layer of a machine learning model, thereby producing a set of compressed client output layer parameters; encrypting the set of compressed client output layer parameters, thereby producing a set of encrypted compressed client output layer parameters; and adding random noise to the set of encrypted compressed client output layer parameters, thereby producing the set of masked client output layer parameters.

20. A computer system comprising: one or more processors: and a non-transitory computer readable medium coupled to the one or more processors, the non-transitory computer readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform the method of any of claims 1-19.

Citation Information

Patent Citations

  • Edge computing privacy protection system and method based on joint learning

    CN110719158A

  • Cloud edge-end collaborative efficient federated learning privacy protection method based on model segmentation and homomorphic encryption

    CN116980107A

  • Apparatuses, computer program products, and computer-implemented methods for privacy-preserving federated learning

    US20210256309A1

  • Resource-limited federated learning using dynamic masking

    US20230334346A1