Bidirectional Compression and Privacy for Efficient Communication in Federated Learning

JP2024522050A5Active Publication Date: 2025-05-22QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023564579
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-26
Filing Date
2022-05-31
Publication Date
2025-05-22
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

Existing federated learning methods face challenges in balancing communication efficiency and privacy preservation, often compromising model performance through lossy compression techniques that do not adequately address data privacy concerns.

Method used

Implementing relative entropy coding for bidirectional compression in federated learning, using shared random seeds and probability distributions to encode model updates, ensuring differential privacy and reducing communication overhead without quantization or pruning.

Benefits of technology

This approach significantly reduces communication costs and maintains model performance while ensuring differential privacy, enabling efficient and secure decentralized machine learning across resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Some aspects of the present disclosure provide techniques for performing federated learning, including receiving a global model from a federated learning server, determining an updated model based on the global model and local data, and transmitting the updated model to the federated learning server using relative entropy coding.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to PCT Application No. PCT / US2022 / 072599, filed May 26, 2022, and claims the benefit of and priority to Greek Patent Application No. 20210100355, filed May 28, 2021, the contents of each of which are incorporated herein by reference in their entirety.

[0002] Aspects of the present disclosure relate to machine learning. [Background technology]

[0003] Machine learning is generally the process of creating a trained model (e.g., an artificial neural network, tree, or other structure) that exhibits a generalized fit to a set of training data that is known a priori. Applying the trained model to new data produces inferences that may be used to gain insight into the new data.

[0004] The widespread use of machine learning in various technology domains, sometimes referred to as artificial intelligence tasks, has created a need for more efficient processing of machine learning model data. For example, "edge processing" devices such as mobile devices, always-on devices, and Internet of Things (IoT) devices must balance the implementation of advanced machine learning capabilities with a variety of interrelated design constraints, such as packaging size, native computational capabilities, power storage and usage, data communication capabilities and cost, memory size, and heat dissipation.

[0005] Federated learning is a distributed machine learning framework that allows several clients, such as edge processing devices, to collaboratively train a shared global model without transferring local data to a remote server. Typically, a central server coordinates the federated learning process, and each participating client communicates only model parameter information with the central server while keeping its local data private. This distributed approach helps with the problem of limited client device capabilities (because training is federated) and often alleviates data privacy concerns.

[0006] Although federated learning generally limits the amount of model data in any single transmission between the server and client (or vice versa), the iterative nature of federated learning still generates a significant amount of data transmission traffic during training, which can be quite costly depending on the device and connection type. Therefore, it is generally desirable to try and reduce the size of the data exchange between the server and client during federated learning. However, conventional methods for reducing data exchange, such as lossy compression of model data being used to limit the amount of data exchanged between the server and client, result in inadequate models. Furthermore, conventional federated learning has been shown to not preserve privacy. Summary of the Invention [Problem to be solved by the invention]

[0007] Therefore, there is a need for improved ways to perform federated learning where privacy is improved without compromising model performance in favor of communication efficiency. [Means for solving the problem]

[0008] Some aspects provide a method for performing federated learning, the method including receiving a global model from a federated learning server, determining an updated model based on the global model and local data, and transmitting the updated model to the federated learning server using relative entropy coding.

[0009] A further aspect provides a method for performing federated learning, the method including sending a global model to a client device, determining a random seed, receiving an updated model from the client device using relative entropy coding, and determining an updated global model based on the updated model from the client device.

[0010] Other aspects provide a processing system configured to perform the methods described above and further described herein, a non-transitory computer readable medium comprising instructions that, when executed by one or more processors of the processing system, cause the processing system to perform the methods described above and further described herein, a computer program product embodied on a computer readable storage medium comprising code for performing the methods described above and further described herein, and a processing system comprising means for performing the methods described above and further described herein.

[0011] The following description and the related drawings set forth in detail certain illustrative features of the one or more aspects.

[0012] The accompanying drawings illustrate some aspects of the one or more aspects and therefore should not be considered limiting of the scope of the disclosure. [Brief description of the drawings]

[0013] [Figure 1] FIG. 1 illustrates an exemplary federated learning architecture. [Figure 2A]FIG. 1 illustrates an exemplary Algorithm 1 for a sender-side implementation of lossy relative entropy coding. [Figure 2B] FIG. 2 illustrates an exemplary algorithm for receiver-side implementation of lossy relative entropy coding. [Diagram 3] FIG. 1 is a schematic diagram of performing relative entropy coding for joint learning updates. [Figure 4] FIG. 1 illustrates an exemplary server-side algorithm for applying relative entropy coding to federated learning. [Diagram 5] FIG. 1 illustrates an exemplary client-side algorithm for applying relative entropy coding to federated learning. [Figure 6A] FIG. 1 illustrates an exemplary client-side algorithm for applying differentially private relative entropy coding to federated learning. [Figure 6B] FIG. 1 illustrates an exemplary server-side algorithm for applying differentially private relative entropy coding to federated learning. [Figure 7] FIG. 1 is a schematic diagram of performing differentially private relative entropy coding for joint learning updates. [Figure 8] FIG. 1 illustrates an example method for performing federated learning according to aspects described herein. [Figure 9] FIG. 1 illustrates another example method for performing federated learning according to aspects described herein. [Figure 10A] FIG. 1 illustrates an example processing system that may be configured to perform the methods described herein. [Figure 10B] FIG. 1 illustrates an example processing system that may be configured to perform the methods described herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0014] For ease of understanding, wherever possible, identical reference numbers have been used to designate identical elements common to the figures. It is contemplated that elements and features of one embodiment may be beneficially incorporated into other embodiments without further recitation.

[0015] Aspects of the present disclosure provide apparatus, methods, processing systems, and non-transitory computer-readable media for machine learning, and in particular, bidirectional compression for efficient private communication in federated learning.

[0016] The performance of modern neural network-based machine learning models scales very well with the amount of data they are trained on. At the same time, industry, legislators, and consumers are becoming more conscious of the need to protect the privacy of data that may be used in training such models. Federated learning describes a machine learning principle that aims to enable learning on distributed data by computing updates on the device. Instead of sending their data to a central location, "clients" in a federation of devices send model updates computed on their data to a central server. Such an approach to learning from distributed data promises to unlock the computing capabilities of billions of "edge" devices, enabling personalized models and enabling new applications, for example in healthcare, due to the inherently more private nature of the approach.

[0017] On the other hand, the federated learning paradigm introduces challenges along many dimensions, including learning from non-independent and identically distributed data, resource-constrained devices, heterogeneous computational and communication capabilities, fairness and representation issues, and communication overhead. In particular, neural networks require many passes over the data, repeatedly communicating the latest server-side model to clients and their updates back to the server, which significantly increases communication overhead. Compression of updates in federated learning therefore reduces such overhead and is a key step in, for example, "untethering" edge devices from Wi-Fi.

[0018] Conventional approaches to mitigate this problem include compression of uplink and / or downlink messages, either through quantization or pruning of the messages to be transmitted. However, these techniques can lead to data loss and performance degradation of the trained global model.

[0019] To overcome these technical problems with conventional approaches, aspects described herein implement a compression scheme, relative entropy coding, for transmitting models and model updates between clients and servers and vice versa, that does not rely on quantization or pruning and is adapted to work for federated learning settings.

[0020] In the aspects described herein, communication from a client to a server may be accomplished through a series of steps. In one example, the server and client first agree on a particular random seed R and prior distribution p (e.g., the last model the server sent to the client). The client then forms a probability distribution q centered on the model update it wishes to send to the server. The client then draws K random samples from p according to the random seed R. In particular, K can be determined by the data, by measuring the discrepancy between p and q. The client then assigns a ratio

[0021]

number

[0022] Probability π proportional to k The client then assigns π1,…,π K The client selects a random sample according to p and records its index k. The client then communicates the index k in log2K bits to the server. The server can decrypt the message by extracting random samples from p using the random seed R until it recovers the kth random sample.

[0023] In particular, this procedure can be implemented flexibly. For example, the procedure can be performed per parameter (e.g., communicating log2K bits per parameter), per layer (e.g., communicating log2K bits per layer in the network), or per network (e.g., communicating log2K bits overall). Any intermediate vector sizes are possible.

[0024] In the embodiment described herein, the communication from the server to the client can be compressed as well. First, the server tracks the last time each client was selected to participate in training, along with all model updates it received from the client. Then, whenever a client is selected to participate in a round (or instance) of federated learning, the server communicates all model updates required to update the old local global model copy in the client to the current one, instead of sending the current state of the global model. Since each of the model updates can be generated by a specific random seed R and log2K bits, the overall message length can be dramatically reduced compared to sending the entire floating-point model, especially when aggressive compression is used for model updates from the client to the server, as described above. In particular, it is possible to always compare the two message sizes (communicating the model in floating-point and communicating compressed past model updates) and select the format with the lower cost.

[0025] As mentioned above, the aspects described herein advantageously work without imposing quantization / pruning on messages sent between the client and the server. Moreover, the compression ratio when using the aspects described herein can be much higher than traditional scalar quantization, especially when performing a layer-by-layer scheme. Moreover, the bit-width of the messages can advantageously be determined / adapted on the fly.

[0026] Thus, aspects described herein provide a technical solution to the technical problem described above with respect to communication overhead. Aspects described herein beneficially improve the performance of any device participating in federated learning, such as by reducing the total communication cost, such as how many units (e.g., GB) of data are communicated between the client and the server during federated learning. The communication cost can be dramatically smaller compared to conventional scalar compression methods, especially when layer-wise compression is used according to the methods described herein.

[0027] Further aspects described herein relate to improving privacy while performing the aforementioned communication-efficient federated learning. Although federated learning provides an intuitive and practical notion of privacy by keeping data on the device, it has nevertheless been shown that client updates reveal sensitive information while allowing reconstruction of the client's training data. Conventional approaches have generally traded off between limiting compression and enhancing privacy, or vice versa, but conventional approaches have not achieved both simultaneously.

[0028] On the other hand, aspects described herein can implement modified relative entropy coding in a federated learning context to make it differentially private. In doing so, aspects described herein provide a differentially private federated learning algorithm that achieves extreme compression of client-to-server updates (e.g., down to 7 bits per tensor) at a privacy level (ε<1, ε quantifies how private the learning algorithm is and refers to how easily a hypothetical adversary can identify where individuals and their data participated in training the model) with minimal impact on model performance.

[0029] Thus, aspects described herein provide a simultaneous solution to privacy and communication efficiency using differentially private and coding-efficient compression of messages communicated during federated learning.

[0030] Example of a Federated Learning Architecture FIG. 1 illustrates an exemplary federated learning architecture 100 .

[0031] In this example, the mobile devices 102A-102C are examples of edge processing devices, each having a respective local data store 104A-104C and a respective local machine learning model instance 106A-106C. For example, the mobile device 102A is loaded with an initial machine learning model instance 106A (or receives the initial machine learning model instance 106A, for example, from a global machine learning model coordinator 108, which may be a software provider in some examples). Each of the mobile devices 102A-102C may use the respective machine learning model instance (106A-106C) for some useful task, such as processing the local data 104A-104C, and may further perform local training and optimization of its respective machine learning model instance.

[0032] For example, the mobile device 102A may use its machine learning model 106A to perform face recognition on photos stored as data 102B on the mobile device 102A. Because such photos may be considered personal, the mobile device 102A may not want to or may be prevented from sharing its photo data with the global model coordinator 108. However, the mobile device 102A may want to or be allowed to share local model updates, such as updates to model weights and parameters, with the global model coordinator 108. Similarly, the mobile devices 102B and 102C may similarly use their local machine learning model instances 106B and 106C, respectively, and share local model updates with the global model coordinator 108 without sharing the underlying data used to generate the local model updates.

[0033] A global model coordinator 108 (alternatively called a federated learning server) may use all of the local model updates to determine a global (or consensus) model update, which may then be distributed to the mobile devices 102A-102C. In this way, machine learning can leverage the mobile devices 102A-102C without centralizing training data and processing.

[0034] Thus, the federated learning architecture 100 enables decentralized deployment and training of machine learning models, which advantageously reduces latency, network utilization, and power consumption, while maintaining data confidentiality and security. Additionally, the federated learning architecture 100 allows models to evolve differently on different devices, but ultimately allows that distributed learned knowledge to be combined back into a global model.

[0035] In particular, the local data stored on the mobile devices 102A-102C, respectively, and used by the machine learning models 106A-106C may be referred to as individual data shards (e.g., data 104A-104C) and / or federated data. Because these data shards are generated by different users on different devices and are not mixed, they cannot be assumed to be independent and identically distributed (IID) with respect to one another. This applies more generally to any kind of device-specific data that is not combined to train a machine learning model. Only by combining the individual data sets 104A-104C of the mobile devices 102A-102C, respectively, can a global data set be generated where the IID assumption is valid.

[0036] Federated learning, general Federated learning is described in the form of the FedAvg algorithm, which is explained as follows: In each communication round t, the server (e.g., 108 in FIG. 1) obtains the current model parameters w(t) to a subset S′ of all S clients participating in the training (e.g., mobile devices 102A, 102B, and / or 102C in FIG. 1). Each selected client s receives the server-provided model w (t) We can, for example, compute a training set of size N via stochastic gradient descent using a given loss function such as s That local dataset D s (eg, data 104A, 104B, and / or 104C of FIG. 1, respectively) to better fit.

[0037]

number

[0038] After E epochs of optimization on the local dataset, the client-side optimization procedure produces an updated model

[0039]

number

[0040] Based on this, the client

number

[0041]

number

[0042] A generalization of this server-side averaging method is

[0043]

number

[0044] as the "gradient" for the server-side model and introduce more advanced update schemes such as adaptive momentum (e.g., the Adam algorithm).

[0045] Federated training involves repeated communication of model updates from the client to the server and vice versa. The total communication cost of this procedure can be significant, and thus federated learning is typically constrained to the use of unmetered channels such as Wi-Fi networks. Compression of the communicated messages therefore plays a key role in transferring federated learning to truly mobile use cases. To this end, aspects described herein extend a lossy version of relative entropy coding (REC) to the federated environment, e.g.

[0046]

number

[0047] Compress model updates from the client to the server, like this:

[0048] Relative Entropy Coding By using the information “shared” between the sender and the receiver, we can obtain a distribution q parameterized by φ. φ A random sample w from (w), i.e., w~q φ Lossy relative entropy coding, preceded by minimum random code learning, was first proposed as a method to compress (w). This information is stored as a shared prior distribution p θ (w) and a shared random seed R.

[0049] The sender selects a prior distribution p according to a random seed R. θ (w) K independent random samples w1,…,w K Then, for K samples, the probability of each sample is calculated by the likelihood ratio π k ∝qφ (w=w k ) / p θ (w=w k ) categorical distribution proportional to

[0050]

number

[0051] Finally, a random sample w k* of

[0052]

number

[0053] , which corresponds to the k*th sample extracted from the shared prior distribution. The sender can then communicate the index k* to the receiver in log2K bits. Figure 2A shows an exemplary Algorithm 1 for a sender-side implementation of lossy relative entropy coding.

[0054] On the receiver side, w k* initializes the random number generator in R and runs p up to the k*th sample. θ (w) can be reconstructed by sampling (w). Figure 2B shows an exemplary Algorithm 2 for a receiver-side implementation of lossy relative entropy coding.

[0055] In some cases, K is q φ (w) to the prior distribution p θ may be set equal to the exponential of the Kullback-Leibler (KL) divergence to (w) plus a constant t, i.e.,

[0056]

number

[0057] In this case, the message length is at least O(KL(q φ (w)||pθ (w)). Thus, under some assumptions, when the sender and receiver share a source of randomness, this KL divergence is a lower bound on the expected length of a communicated message.

[0058] This means that the compression ratio is p θ q for (w) information φ This gives us an intuitive notion of compression, which relates the amount of extra information encoded in (w). Thus, the less the amount of extra information, the shorter the message length, and q φ (w)=p θ In the extreme case where (w), the message length is O(1). Of course, if the bias of this procedure is high, it makes no sense to achieve this efficiency. Fortunately, for a suitable value of t, and under mild assumptions, for any function f, the bias, i.e.

[0059]

number

[0060] Therefore, in some aspects, K may be parameterized as a function of the binary bit width b, e.g., K=2 b where b can be treated as a hyperparameter.

[0061] Relative Entropy Coding for Efficient Communication in Associative Learning Aspects described herein include client to server messages, e.g., model updates.

[0062]

number

[0063] We use the distribution over

[0064]

number

[0065] We adapt lossy relative entropy coding to the associative environment by appropriately selecting , and , which may be defined as follows:

[0066]

number

[0067] In other words, a Gaussian distribution centered at zero is used for the prior distribution with an appropriately chosen σ, and a Gaussian distribution with the same standard deviation centered on the model update is used for the message distribution. The form of q is chosen to be implementable on resource-constrained devices and to satisfy the differential privacy constraints discussed below. Note that, in contrast to the FedAvg client update definition in Eq. (2), here we use

number

[0068] Therefore, the length of the federated learning message from the client to the server is the local dataset D s How much "extra" information about x is measured via the KL divergence,

[0069]

number

[0070] is a function of how much information is encoded in each update. This has good interactions with differential privacy (DP) since the differential privacy constraint bounds the amount of information encoded in each update, resulting in highly compressible messages. Also, in particular, this procedure can be done per parameter (e.g., communicate log2K bits per parameter), per layer (e.g., communicate log2K bits per layer in the global model), or per network (e.g., communicate log2K bits total). Any intermediate vector size is also possible. This means that

[0071]

number

[0072] This is achieved by splitting,x,into,M,independent groups (which is easy due to the,assumption of factorial distribution over the vector dimensions) and,applying the compression mechanism to each group,independently.

[0073] 3 illustrates a schematic diagram of an example 300 of communication from a client 302 to a server 304. In the illustrated example, the client 302 receives a distribution q φ and the shared prior distribution p θ The client 302 then transmits the index k to the server 304, which then calculates the shared prior distribution p θ The model update 308 can be recovered based on decrypting the index using shared information such as and the random seed R.

[0074] The compression procedure described with respect to federated learning messaging from client to server is a particular example of (probabilistic) vector quantization, where the shared codebook is determined by a shared random seed R. Beneficially, the principle of communicating an index into such a shared codebook also enables compression of server to client communications.

[0075] For example, instead of sending the complete server-side model to a particular client, the server can choose to collect all updates to the global model during two subsequent rounds that the client participates in. Based on this history of the codebook indexes, the client can deterministically reconstruct the current state of the server model before starting the local optimization.

[0076] Obviously, the expected length of the history is proportional to the total number of clients and the amount of subsampling of the clients performed during training. Thus, at the beginning of any round, the server compares the bit sizes of the clients' histories and instead chooses to use the full precision model w (t) Taking a model with 1k parameters as an example, when using an 8-bit codebook compression of the entire model, a single uncompressed model update is roughly equivalent to 4k communication indices. Importantly, compressing server-to-client messages in this manner does not affect the differentially private nature of the aspects described below, since any information released by the client is private, according to these aspects.

[0077] For clients participating in the first round of training, the first seed without index can be understood as a seed for random initialization of the server-side model. Algorithms 3 and 4 shown in Figures 4 and 5, respectively, show examples of the server-side and client-side procedures. Note that the client-side update rule must be equal to the server-side update rule (*). In other words, in the generalized FedAvg, the current global model w (t) When sending , it may be necessary to additionally send the optimizer state.

[0078] Differentially Private Relative Entropy Coding for Private and Efficient Communication in Federated Learning The relative entropy coding training compression schemes described above beneficially enable significant reductions in communication costs, often by several orders of magnitude, compared to conventional methods. However, model updates can still reveal sensitive information about the client's local dataset, and at least from a theoretical point of view, compressed model updates leak as much information as full-precision updates.

[0079] To mitigate privacy risks, differential privacy can be employed during training. Traditional differential privacy mechanisms for federated learning involve each client clipping the norm of the full-precision model updates before sending them to the server. The server then averages the clipped model updates, possibly using a secure aggregation protocol, and adds Gaussian noise with a certain variance. However, traditional applications of differential privacy do not work with compression.

[0080] Therefore, various aspects may modify the relative entropy coding training compression scheme described above to ensure privacy. Specifically, to ensure differential privacy of the relative entropy coding described above, it is necessary to bound its sensitivity to quantify the inherent noise. The bounding of the sensitivity is performed by the client update

[0081]

number

[0082] In the context of relative entropy coding, this consists of clipping the norm of the client message distribution

[0083]

number

[0084] At any round t, the prior distribution of the servers is

[0085]

number

[0086] This means that k cannot differ significantly from k. Note that the procedure itself is stochastic, so unlike traditional methods, no explicit injection of additional noise into the update is required. At each round t, two sources of randomness play a role: (1) We define a set of K samples as a prior distribution

[0087]

number

[0088] (2) sample from a critical sampling distribution

[0089]

number

[0090] Extract updates from.

[0091] Thus, Differentially Private Relative Entropy Coding (DP-REC) may generally be accomplished in two steps. First, each client may clip the norm of its model update before forming a probability distribution q centered on this clipped update. In one example, the clipping threshold is σ p The goal of this step is to ensure bounding of the Renyi divergence between the posterior distribution q and the server’s prior distribution p. The bounding is necessary to be sufficient to compute privacy guarantees.

[0092] Note that the Renyi divergence or alpha divergence of order α of a distribution P from a distribution Q is defined as follows:

[0093]

number

[0094] Second, the server records events that leak information about a client's data, e.g., a sampling of a particular client from the entire population, along with its probability in each round, or an importance distribution π q These events define a probability distribution over the possible model updates for all clients. The privacy accounting component uses this information in combination with a clipping limit to determine the maximum Renyi divergence between the update distributions of any two clients over the course of training, and then computes the ε, δ parameters of differential privacy by using the Chernoff bound. In probability theory, the Chernoff bound gives an exponentially decreasing bound on the tail distribution of a sum of independent random variables. Furthermore, ε declares the degree of "privacy" of a particular algorithm, and δ (usually considered small enough) is the probability that differential privacy fails (and thus does not give a private output).

[0095] 6A and 6B show Algorithms 5 and 6 for performing differentially private relative entropy coding (DP-REC) on the client side and server side, respectively.

[0096] 7 illustrates a schematic diagram of an example 700 of communication from a client 702 to a server 704. In the illustrated example, the client 702:

[0097]

number

[0098] Generate samples 1 to k based on the ratio 706 with (described above). However, in this example, the norm is clipped before generating the ratio, generating a clipped model update.

[0099] For example, m q is the model update,

[0100]

number

[0101] is the clipped model update computed according to

[0102]

number

[0103] Δ is the amount of clipping performed.

[0104] The index k is then sent from the client 702 to the server 704, which then calculates the shared prior distribution p θ The model update 708 can be recovered based on decrypting the index using shared information such as and the random seed R.

[0105] In particular, compared to conventional differential privacy techniques, the aspects described herein do not require injecting additional noise into updates at either the client or the server. Rather, randomness in the relative entropy coding procedure for federated learning updates is used. In that case, communication-efficient federated learning using relative entropy coding can be beneficially combined with the privacy-preserving aspects of differential privacy for a unified approach.

[0106] Exemplary Methods 8 illustrates an example method 800 for performing federated learning according to aspects described herein. The method 800 may be generally performed by a client in a federated learning scheme, such as the mobile device 102 of FIG.

[0107] The method 800 begins in step 802 by receiving a global model from a federated learning server, such as the global model coordinator 108 of FIG.

[0108] The method 800 then proceeds to step 804, where an updated model is determined based on the global model and the local data. For example, a local machine learning model, such as 106A in FIG. 1, may be trained on the local data 104A to generate an updated model. Determining the updated model may include generating updated model parameters, such as weights and biases, which may be determined as direct values ​​or as relative values ​​(e.g., delta). In some aspects, determining the updated model based on the global model and the local data includes performing gradient descent on the global model using the local data.

[0109] The method 800 then proceeds to step 806, where the updated model is transmitted to the federated learning server using relative entropy coding. In some aspects, the transmitting of the updated model to the federated learning server using relative entropy coding is performed according to the algorithm depicted and described with respect to FIG. 5 or FIG. 6A.

[0110] In some aspects, using relative entropy coding, sending the updated model to the federated learning server includes determining a random seed. In some aspects, determining the random seed includes receiving the random seed from the federated learning server. In other aspects, the client can determine the random seed and send it to the federated learning server, which may prevent any manipulation of the random seed by the federated learning server and improve privacy.

[0111] In some aspects, sending the updated model to the federated learning server using relative entropy coding further includes determining a first probability distribution based on the global model and a second probability distribution centered on the updated model.

[0112] In some aspects, sending the updated model to the federated learning server using relative entropy coding further includes determining a plurality of random samples from a first probability distribution according to a random seed, and assigning a probability to each respective random sample of the plurality of random samples based on a ratio of a likelihood of the respective random sample given the first probability distribution to a likelihood of the respective random sample given a second probability distribution.

[0113] In some aspects, determining a plurality of random samples from a first probability distribution according to a random seed is performed based on a difference between the first probability distribution and the second probability distribution. In some cases, the number of random samples (K) is calculated as K=exp(KL(q||p)+t), where KL is the Kulback-Leibler divergence between q and p, and t is an adjustment factor. In other cases, K=2 b where b is the number of bits allowed for a client-to-server message, such as the updated local model as shown and described with respect to FIG.

[0114] In particular, the ratio of the likelihood of each random sample given the second probability distribution to the likelihood of each random sample given the first probability distribution is The ratio can be determined for each parameter, such as q(w1) / p(w1), q(w2) / p(w2), etc. The ratio can be determined for each parameter, for example, (q(w1)×q(w2)×…×q(w k )) / (p(w1)×p(w2)×…×p(w kIt may also be determined for a given number of elements, which may represent a layer of the model to be updated, such as (k, n, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 112, 123, 134, 140, 141, 152, 163, 170, 184, 196, 197, 198, 199, 102, 103, 104, 10

[0115] In some aspects, sending the updated model to the federated learning server using relative entropy coding further includes selecting a random sample from the plurality of random samples according to a probability of each of the plurality of random samples.

[0116] In some aspects, sending the updated model to the federated learning server using relative entropy coding further includes determining an index associated with the selected random sample and sending the index to the federated learning server.

[0117] For example, suppose there are eight samples, then there is a probability distribution over these eight samples, and then a random sample can be drawn from this distribution that represents an index to one of the eight samples.

[0118] In some cases, the index is transmitted using log2K bits, where K is the number of multiple random samples from the first probability distribution.

[0119] In some aspects, the method 800 further includes clipping the updated model prior to determining the second probability distribution centered on the updated model, where the clipping is based on a standard deviation (σ) of the global model, and the second probability distribution is based on the clipped updated model. In one aspect, the clipping value is calculated as C×σ, where σ is the prior standard deviation of the global model. The complete paper includes these details, and can currently be found in line 4 of Algorithm 5.

[0120] In some aspects, clipping the updated model includes clipping a norm of the updated model.

[0121] 9 illustrates an example method 900 for performing federated learning according to aspects described herein. The method 900 may generally be performed by a server in a federated learning scheme, such as the global model coordinator 108 of FIG.

[0122] The method 900 begins in step 902 with transmitting a global model to a client device.

[0123] The method 900 then proceeds to step 904, where a random seed is determined.

[0124] The method 900 then proceeds to step 906, where it receives an updated model from the client device using relative entropy coding.

[0125] In some aspects, receiving the updated model from the client device using relative entropy coding is performed according to the algorithm depicted and described with respect to FIG. 4 or FIG. 6B.

[0126] The method 900 then proceeds to step 908, where it determines an updated global model based on the updated model from the client device.

[0127] In some aspects, receiving an updated model from the client device using relative entropy coding includes receiving an index from the client device, determining a sample from the probability distribution based on the global model, the random seed, and the index, and using the determined sample to determine the updated global model.

[0128] In some aspects, the index is received using log2K bits, where K is the number of random samples determined from a probability distribution based on a global model.

[0129] In some aspects, the determined samples are used to update parameters of an updated global model.

[0130] In some aspects, the determined samples are used to update a layer of an updated global model.

[0131] In some aspects, determining the random seed includes receiving the random seed from the client device, hi other aspects, determining the random seed is performed by a federated learning server, which transmits the random seed to the client device.

[0132] Exemplary Processing System for Implementing Sparsity-Aware Compute-in-Memory 10A illustrates an example processing system 1000 for performing federated learning, e.g., as described herein with respect to Figures 1-8. Processing system 1000 may be an example of a client device, such as client devices 102A-C of Figure 1.

[0133] The processing system 1000 includes a central processing unit (CPU) 1002, which in some examples may be a multi-core CPU. Instructions executed on the CPU 1002 may be loaded from a program memory associated with the CPU 1002 or may be loaded from a memory partition 1024, for example.

[0134] The processing system 1000 also includes additional processing components tailored to specific functions, such as a graphics processing unit (GPU) 1004, a digital signal processor (DSP) 1006, a neural processing unit (NPU) 1008, a multimedia processing unit 1010, and wireless connectivity components 1012.

[0135] An NPU, such as 1008, is generally a specialized circuit configured to implement all the necessary control and computational logic to execute machine learning algorithms, such as algorithms for processing artificial neural networks (ANN), deep neural networks (DNN), random forests (RF), etc. An NPU may alternatively be referred to as a neural signal processor (NSP), a tensor processing unit (TPU), a neural network processor (NNP), an intelligence processing unit (IPU), a vision processing unit (VPU), or a graph processing unit.

[0136] An NPU, such as 1008, is configured to accelerate the execution of common machine learning tasks, such as image classification, machine translation, object detection, and various other predictive models. In some examples, multiple NPUs may be instantiated on a single chip, such as a system-on-chip (SoC), while in other examples, they may be part of a dedicated neural network accelerator.

[0137] An NPU may be optimized for training or inference, or in some cases may be configured to balance performance between both. In an NPU capable of performing both training and inference, the two tasks may generally be performed independently.

[0138] NPUs designed to accelerate training are generally configured to accelerate the optimization of new models, which is a highly computationally intensive operation that involves inputting an existing dataset (often labeled or tagged), iterating over the dataset, and then adjusting model parameters such as weights and biases to improve model performance. In general, optimization based on erroneous predictions involves propagating backwards through layers of the model and determining gradients to reduce prediction errors. In some cases, NPUs may be configured to perform the federated learning methods described herein.

[0139] NPUs designed to accelerate inference are generally configured to operate on complete models. Thus, such NPUs may be configured to input new data and rapidly process the data through already trained models to generate model outputs (e.g., inferences).

[0140] In one implementation, the NPU 1008 is part of one or more of the CPU 1002, GPU 1004, and / or DSP 1006.

[0141] In some examples, the wireless connectivity component 1012 may include sub-components for, for example, third generation (3G) connectivity, fourth generation (4G) connectivity (e.g., 4G LTE), fifth generation connectivity (e.g., 5G or NR), Wi-Fi connectivity, Bluetooth connectivity, and other wireless data transmission standards. The wireless connectivity processing component 1012 is further connected to one or more antennas 1014. In some examples, the wireless connectivity component 1012 enables performing federated learning in accordance with the methods described herein over various wireless data connections, including cellular connections.

[0142] The processing system 1000 may also include one or more sensor processing units 1016 associated with any type of sensor, one or more image signal processors (ISPs) 1018 associated with any type of image sensor, and / or a navigation processor 1020, which may include satellite-based positioning system components (e.g., GPS or GLONASS), as well as inertial positioning system components.

[0143] The processing system 1000 may also include one or more input and / or output devices 1022, such as a screen, a touch-sensitive surface (including a touch-sensitive display), physical buttons, a speaker, a microphone, etc.

[0144] In some examples, one or more of the processors of processing system 1000 may be based on the ARM or RISC-V instruction set.

[0145] Processing system 1000 also includes memory 1024, which represents one or more static and / or dynamic memories, such as dynamic random access memory, flash-based static memory, etc. In this example, memory 1024 includes computer-executable components that may be executed by one or more of the above-mentioned processors of processing system 1000.

[0146] In particular, in this example, memory 1024 includes a receiving component 1024A, a model updating component 1024B, a transmitting component 1024C, and a model parameters component 1024D. The components shown and other components not shown may be configured to perform various aspects of the methods described herein.

[0147] In general, the processing system 1000 and / or its components may be configured to perform the methods described herein.

[0148] Notably, in other cases, aspects of the processing system 1000 may be omitted or added. For example, the multimedia components 1010, the wireless connectivity 1012, the sensors 1016, the ISP 1018, and / or the navigation components 1020 may be omitted in other aspects. Additionally, aspects of the processing system 1000 may be distributed among multiple devices.

[0149] Figure 10B illustrates another example processing system 1050 for performing federated learning, e.g., as described herein with respect to Figures 1-7 and 9. The processing system 1050 may be an example of a federated learning server, such as the global model coordinator 108 of Figure 1.

[0150] In general, the CPU 1052, GPU 1054, NPU 1058, and I / O 1072 are as described above with respect to the similar elements in FIG. 10A.

[0151] The processing system 1050 also includes memory 1074, which represents one or more static and / or dynamic memories, such as dynamic random access memory, flash-based static memory, etc. In this example, the memory 1074 includes computer-executable components that may be executed by one or more of the above-mentioned processors of the processing system 1050.

[0152] In particular, in this example, the memory 1074 includes a receiving component 1074A, a model updating component 1074B, a transmitting component 1074C, and a model parameters component 1074D. The components shown and other components not shown may be configured to perform various aspects of the methods described herein.

[0153] In general, the processing system 1050 and / or its components may be configured to perform the methods described herein.

[0154] In particular, in other cases, aspects of the processing system 1050 may be omitted or added. Further, aspects of the processing system 1050 may be distributed across multiple devices, such as a cloud-based service. The components shown are limited for clarity and brevity.

[0155] Example clauses Example implementations are described in the following numbered clauses.

[0156] Clause 1: A method comprising: receiving a global model from a federated learning server; determining an updated model based on the global model and local data; and transmitting the updated model to the federated learning server using relative entropy coding.

[0157] Clause 2: The method of clause 1, wherein the step of sending the updated model to the federated learning server using relative entropy coding includes the steps of determining a random seed, determining a first probability distribution based on the global model, determining a second probability distribution centered on the updated model, determining a plurality of random samples from the first probability distribution according to the random seed, assigning a probability to each respective random sample of the plurality of random samples based on a ratio of a likelihood of the respective random sample given the second probability distribution to a likelihood of the respective random sample given the first probability distribution, selecting one random sample from the plurality of random samples according to the probability of each of the plurality of random samples, determining an index associated with the selected random sample, and sending the index to the federated learning server.

[0158] Clause 3: The method of clause 2, wherein the step of determining a plurality of random samples from a first probability distribution according to a random seed is performed based on a difference between the first probability distribution and a second probability distribution.

[0159] Clause 4: The method of any one of clauses 2 to 3, wherein the index is transmitted using log2K bits, where K is a number of random samples from the first probability distribution.

[0160] Clause 5: The method of any one of clauses 2 to 4, wherein a plurality of random samples are associated with a plurality of parameters of a global model.

[0161] Clause 6: A method according to any one of clauses 2 to 4, wherein a plurality of random samples are associated with a layer of a global model.

[0162] Clause 7: The method of any one of clauses 2 to 4, wherein a plurality of random samples are associated with a subset of parameters of the global model.

[0163] Clause 8: The method of any one of clauses 2 to 7, further comprising the step of clipping the updated model before determining the second probability distribution centered on the updated model, wherein the clipping is based on a standard deviation of the global model and the second probability distribution is based on the clipped updated model.

[0164] Clause 9: The method of clause 8, wherein clipping the updated model comprises clipping a norm of the updated model.

[0165] Clause 10: The method of any one of clauses 1 to 9, wherein the step of determining an updated model based on the global model and the local data includes a step of performing gradient descent on the global model using the local data.

[0166] Clause 11: The method of any one of clauses 2 to 10, wherein the step of determining the random seed includes the step of receiving the random seed from a federated learning server.

[0167] Clause 12: A method comprising: transmitting a global model to a client device; determining a random seed; receiving an updated model from the client device using relative entropy coding; and determining an updated global model based on the updated model from the client device.

[0168] Clause 13: The method of clause 12, wherein the step of receiving an updated model from a client device using relative entropy coding includes the steps of receiving an index from the client device; determining a sample from a probability distribution based on the global model, a random seed, and the index; and using the determined sample to determine an updated global model.

[0169] Clause 14: The method of clause 13, wherein the index is received using log2K bits, where K is a number of random samples determined from a probability distribution based on a global model.

[0170] Clause 15: The method according to any one of clauses 13 to 14, wherein the determined samples are used to update parameters of an updated global model.

[0171] Clause 16: The method according to any one of clauses 13 to 15, wherein the determined samples are used to update a layer of an updated global model.

[0172] Clause 17: The method of any one of clauses 12 to 16, wherein determining the random seed comprises receiving the random seed from a client device.

[0173] Clause 18: A processing system comprising a memory containing computer-executable instructions and one or more processors configured to execute the computer-executable instructions to cause the processing system to perform a method according to any one of clauses 1 to 17.

[0174] Clause 19: A processing system, comprising means for carrying out the method according to any one of clauses 1 to 17.

[0175] Clause 20: A non-transitory computer-readable medium comprising computer-executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform the method of any one of clauses 1 to 17.

[0176] Clause 21: A computer program embodied on a computer readable storage medium comprising code for performing the method according to any one of clauses 1 to 17.

[0177] Additional Considerations The preceding description is provided to enable any person skilled in the art to practice the various aspects described herein. The examples discussed herein are not intended to limit the scope, applicability, or aspects described in the claims. Various modifications of these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. For example, changes may be made in the function and arrangement of the elements discussed without departing from the scope of the disclosure. Various examples may omit, substitute, or add various procedures or components, as appropriate. For example, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. Also, features described with respect to some examples may be combined in some other examples. For example, an apparatus may be implemented or a method may be practiced using any number of the aspects described herein. In addition, the scope of the disclosure is intended to encompass such apparatus or methods practiced using other structures, functions, or structures and functions in addition to or other than the various aspects of the disclosure described herein. It should be understood that any aspect of the disclosure disclosed herein may be embodied by one or more elements of a claim.

[0178] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any aspect described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects.

[0179] As used herein, a phrase referring to "at least one of" a list of items refers to any combination of those items, including single members. As an example, "at least one of a, b, or c" is intended to cover a, b, c, ab, ac, bc, and abc, as well as any combination having multiples of the same element (e.g., aa, aaa, aab, aac, abb, acc, bb, bbb, bbc, cc, and ccc, or any other permutation of a, b, and c).

[0180] The term "determining" as used herein encompasses a wide variety of actions. For example, "determining" may include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, database, or another data structure), ascertaining, and the like. "Determining" may also include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory), and the like. "Determining" may also include resolving, selecting, choosing, establishing, and the like.

[0181] The methods disclosed herein comprise one or more steps or actions for achieving the method. Method steps and / or actions may be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions may be modified without departing from the scope of the claims. Furthermore, various operations of the methods described above may be performed by any suitable means capable of performing the corresponding functions. The means may include various hardware and / or software components and / or modules, including but not limited to circuits, application specific integrated circuits (ASICs), or processors. In general, where there are operations illustrated in figures, those operations may have corresponding counterpart means-plus-function components with similar numbering.

[0182] The following claims are not limited to the embodiments set forth herein, but are to be accorded the full scope consistent with the language of the claims. Within the claims, reference to an element in the singular does not mean "one and only one" unless expressly stated as such, but means "one or more." Unless otherwise expressly stated, the term "several" refers to one or more. No element of a claim is to be construed under the provisions of 35 U.S.C. 112(f) unless the element is expressly recited using the phrase "means for" or, in the case of a method claim, unless the element is recited using the phrase "step for." All structural and functional equivalents of the elements of the various embodiments described throughout this disclosure that are known or later become known to those skilled in the art are expressly incorporated herein by reference and are intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be made public, regardless of whether such disclosure is expressly recited in the claims. [Explanation of symbols]

[0183] 100 Federated Learning Architecture 102A~102C Mobile Devices 102B Data 104A~104C Local Data Store 106A~106C Local Machine Learning Model Instances 108 Global Machine Learning Model Coordinator 302 Client 304 Server 306 Ratio 308 Model Updates 702 Client 704 Server 706 Ratio 708 Model Update 1000 Processing Systems 1002 Central Processing Unit (CPU) 1004 Graphics Processing Unit (GPU) 1006 Digital Signal Processor (DSP) 1008 Neural Processing Unit (NPU) 1010 Multimedia Processing Unit 1012 Wireless connectivity components 1016 Sensor Processing Unit 1018 Image Signal Processor (ISP) 1020 Navigation Processor 1022 input and / or output devices 1024 memory 1024A Receiving Component 1024B Model Update Components 1024C Transmission Components 1024D Model Parameters 1050 Processing System 1052 CPU 1054 GPU 1058 NPU 1072 Input / Output 1074 Memory 1074A Receiving Component 1074B Model Update Components 1074C Transmission Components 1074D Model Parameters

Claims

1. 1. A computer-implemented method comprising: receiving a global model from a federated learning server; determining an updated model based on the global model and local data; transmitting the updated model to the federated learning server using relative entropy coding; Including, transmitting the updated model to the federated learning server using relative entropy coding; determining a random seed, the random seed including receiving the random seed from the federated learning server; determining a first probability distribution based on the global model; determining a second probability distribution centered on the updated model; determining a plurality of random samples from the first probability distribution according to the random seed based on a difference between the first probability distribution and the second probability distribution, the plurality of random samples being associated with a plurality of parameters of the global model; assigning a probability to each respective random sample of the plurality of random samples based on a ratio of the likelihood of the respective random sample given the second probability distribution to the likelihood of the respective random sample given the first probability distribution; selecting a random sample from the plurality of random samples according to the probability for each of the plurality of random samples; determining an index associated with the selected random sample; transmitting the index to the federated learning server; A method comprising:

2. The index is log 2 It is transmitted using K bits, K is the number of random samples from the first probability distribution. The method of claim 1.

3. The method of claim 1 , wherein the plurality of random samples is associated with a layer of the global model.

4. The method of claim 1 , wherein the plurality of random samples is associated with a subset of parameters of the global model.

5. clipping the updated model before determining the second probability distribution centered on the updated model; the clipping is based on a standard deviation of the global model; the second probability distribution is based on a clipped update model; The method of claim 1.

6. The method of claim 5 , wherein clipping the updated model comprises clipping a norm of the updated model.

7. 2. The method of claim 1, wherein determining the updated model based on the global model and local data comprises performing gradient descent on the global model using the local data.

8. 1. A computer-implemented method comprising: transmitting the global model to a client device; determining a random seed, the determining step including receiving the random seed from the client device; receiving an updated model from the client device using relative entropy coding; determining an updated global model based on the updated model from the client device; Including, receiving the updated model from the client device using relative entropy coding, receiving an index from the client device; determining a sample from a probability distribution based on the global model, the random seed, and the index; using the determined samples to determine an updated global model, the determined samples being used to update parameters of the updated global model; A method comprising:

9. The index is log 2 It is received using K bits, K is the number of random samples determined from the global model-based probability distribution; The method of claim 8.

10. The method of claim 8 , wherein the determined samples are used to update a layer of the updated global model.

11. A processing system comprising a memory containing computer-executable instructions and one or more processors configured to execute the computer-executable instructions to cause the processing system to perform the method of any one of claims 1 to 10.

12. A processing system comprising means for carrying out the method according to any one of claims 1 to 10.

13. A non-transitory computer readable medium comprising computer executable instructions that, when executed by one or more processors of a processing system, cause the processing system to perform the method of any one of claims 1 to 10.

14. A computer program embodied on a computer readable storage medium comprising code for performing the method of any one of claims 1 to 10.