Sampling sub-models for federated learning in processor devices
By decomposing neural network weight matrices into orthogonal terms and sampling sub-models based on client device capacities, the method addresses the challenge of training heterogeneous devices in federated learning, achieving efficient and accurate model training across diverse client devices.
Patent Information
- Application Number
- PCT/US2024/051188
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-01
- Filing Date
- 2024-10-14
- Publication Date
- 2025-08-07
AI Technical Summary
Federated learning faces challenges in training machine learning models across heterogeneous client devices with varying processing, power, and communications capabilities, making it infeasible to train the same model on each device.
A processor-based server device decomposes a neural network weight matrix into orthogonal terms and generates inclusion probabilities based on client device capacities, randomly samples a subset of these terms as sub-models, and transmits them to client devices for local training, aggregating the updated sub-models into the weight matrix.
This approach optimizes the sampling of sub-models for federated learning, ensuring unbiased estimates and minimizing mean squared error, enabling effective training across diverse client devices while preserving privacy.
Smart Images

Figure US2024051188_07082025_PF_FP_ABST
Abstract
Description
SAMPLING SUB-MODELS FOR FEDERATED LEARNING IN PROCESSOR DEVICESPRIORITY APPLICATION
[0001] The present application claims priority to Greek Patent Application Serial No. 20240100067, filed February 1, 2024 and entitled “SAMPLING SUB-MODELS FOR FEDERATED LEARNING IN PROCESSOR DEVICES,” which is incorporated herein by reference in its entirety.BACKGROUNDI. Field of the Disclosure
[0002] The technology of the disclosure relates generally to machine learning, and, in particular, to mechanisms for training machine learning models.IL Background
[0003] Machine learning is a subfield of artificial intelligence (Al) that uses algorithms trained on data sets to create models that enable computers to perform tasks such as image categorization or data analysis. Because the process of training a machine learning algorithm can be resource-intensive in terms of processor, power, memory, and communications capacity, a number of approaches have been developed for training using multiple processor-based devices. One approach, known as distributed learning, employs parallel processing of a single model on multiple server devices using local datasets that are independent and identically distributed and that have similar sizes. In contrast, an alternative approach known as federated learning focuses on training a machine learning algorithm using multiple local datasets on client devices that do not explicitly exchange data samples with each other. With federated learning, each client performs local training, and then exchanges parameters such as weights and biases with a server device that aggregates the updates into a global model. Because client devices do not share data, federated learning approaches offer increased protection for user privacy.
[0004] However, one underlying assumption of federated learning is that the client devices may include various types of heterogeneous processor-based devices such as desktop computers, mobile phones, tablets, and the like. Consequently, the client devicesmay span a wide range of processing, power, memory, and communications capabilities, making training the same model for each client device infeasible. This issue may be addressed by having each client device perform local training using a smaller sub-model sampled from a larger global model, wherein the sub-model is determined based on the capabilities of the corresponding client device. It is thus desirable to optimize the sampling of sub-models for processing by client devices.SUMMARY OF THE DISCLOSURE
[0005] Aspects disclosed in the detailed description include sampling sub-models for federated learning in processor devices. Related apparatus and methods are also disclosed. In this regard, in some exemplary aspects disclosed herein, a processor-based server device is configured to receive a capacity indicator representing a processing capacity (e.g., in terms of processor, power, memory, and / or communications capability) from each client device of the plurality of client devices. At the start of a communication round, the processor-based server device decomposes a weight matrix of a neural network into a plurality of mutually orthogonal terms. The processor-based server device then generates a plurality of inclusion probabilities, each corresponding to a term of the plurality of mutually orthogonal terms, based on one or more estimates of the weight matrix for one or more client devices of the plurality of client devices. In some aspects, generating the plurality of inclusion probabilities may comprise generating the plurality of inclusion probabilities to keep an estimate of the weight matrix for the client device unbiased, while some aspects may provide that generating the plurality of inclusion probabilities comprises generating the plurality of inclusion probabilities to minimize a mean squared error between the weight matrix and an aggregation of estimates of the weight matrix for the plurality of client devices.
[0006] The processor-based server device next randomly samples a subset of the plurality of mutually orthogonal terms based on the capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities as a sub-model. According to some aspects, randomly sampling a subset of the plurality of mutually orthogonal terms comprises randomly sampling using a sampling design that preserves the plurality of inclusion probabilities. The processor-based server device then transmits the sub-model to the client device. In some aspects, the processor-based server devicesubsequently receives, from each client device of the plurality of client devices, an updated sub-model based on local training. The processor-based server device then aggregates the updated sub-model into the weight matrix, thus concluding the communication round.
[0007] In another aspect, a processor-based server device is disclosed. The processorbased server device comprises a processor device that is configured to receive, from a client device of a plurality of client devices, a capacity indicator representing a processing capacity of the client device. The processor device is further configured to decompose a weight matrix of a neural network into a plurality of mutually orthogonal terms. The processor device is also configured to generate a plurality of inclusion probabilities each corresponding to a term of the plurality of mutually orthogonal terms, based on one or more estimates of the weight matrix for one or more client devices of the plurality of client devices. The processor device is additionally configured to randomly sample a subset of the plurality of mutually orthogonal terms based on the capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities as a submodel. The processor device is further configured to transmit the sub-model to the client device.
[0008] In another aspect, a processor-based server device is disclosed. The processorbased server device comprises means for receiving, from a client device of a plurality of client devices, a capacity indicator representing a processing capacity of the client device. The processor-based server device further comprises means for decomposing a weight matrix of a neural network into a plurality of mutually orthogonal terms. The processorbased server device also comprises means for generating a plurality of inclusion probabilities each corresponding to a term of the plurality of mutually orthogonal terms, based on one or more estimates of the weight matrix for one or more client devices of the plurality of client devices. The processor-based server device additionally comprises means for randomly sampling a subset of the plurality of mutually orthogonal terms based on the capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities as a sub-model. The processor-based server device further comprises means for transmitting the sub-model to the client device.
[0009] In another aspect, a method for sampling sub-models for federated learning in processor devices is disclosed. The method comprises receiving, by a processor-basedserver device from a client device of a plurality of client devices, a capacity indicator representing a processing capacity of the client device. The method further comprises decomposing, by the processor-based server device, a weight matrix of a neural network into a plurality of mutually orthogonal terms. The method also comprises generating, by the processor-based server device, a plurality of inclusion probabilities each corresponding to a term of the plurality of mutually orthogonal terms, based on one or more estimates of the weight matrix for one or more client devices of the plurality of client devices. The method additionally comprises randomly sampling, by the processorbased server device, a subset of the plurality of mutually orthogonal terms based on the capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities as a sub-model. The method further comprises transmitting, by the processor-based server device, the sub-model to the client device.
[0010] In another aspect, a non-transitory computer-readable medium is disclosed. The non-transitory computer-readable medium stores computer-executable instructions that, when executed, cause a processor device of a processor-based server device to receive, from a client device of a plurality of client devices, a capacity indicator representing a processing capacity of the client device. The computer-executable instructions further cause the processor device to decompose a weight matrix of a neural network into a plurality of mutually orthogonal terms. The computer-executable instructions also cause the processor device to generate a plurality of inclusion probabilities each corresponding to a term of the plurality of mutually orthogonal terms, based on one or more estimates of the weight matrix for one or more client devices of the plurality of client devices. The computer-executable instructions additionally cause the processor device to randomly sample a subset of the plurality of mutually orthogonal terms based on the capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities as a sub-model. The computer-executable instructions further cause the processor device to transmit the sub-model to the client device.BRIEF DESCRIPTION OF THE FIGURES
[0011] Figure 1 is a block diagram illustrating an exemplary processor-based server device configured to sample sub-models for federated learning, according to some aspects;
[0012] Figures 2A and 2B provide a flowchart illustrating exemplary operations of the processor-based server device of Figure 1 for sampling sub-models for federated learning, according to some aspects; and
[0013] Figure 3 is a block diagram of an exemplary processor-based device that can include the processor-based server device of Figure 1.DETAILED DESCRIPTION
[0014] With reference now to the drawing figures, several exemplary aspects of the present disclosure are described. The word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other aspects. The terms “first,” “second,” and the like used herein are intended to distinguish between similarly named elements, and do not indicate an ordinal relationship between such elements unless otherwise indicated.
[0015] Aspects disclosed in the detailed description include sampling sub-models for federated learning in processor devices. Related apparatus and methods are also disclosed. In this regard, in some exemplary aspects disclosed herein, a processor-based server device is configured to receive a capacity indicator representing a processing capacity (e.g., in terms of processor, power, memory, and / or communications capability) from each client device of the plurality of client devices. At the start of a communication round, the processor-based server device decomposes a weight matrix of a neural network into a plurality of mutually orthogonal terms. The processor-based server device then generates a plurality of inclusion probabilities, each corresponding to a term of the plurality of mutually orthogonal terms, based on one or more estimates of the weight matrix for one or more client devices of the plurality of client devices. In some aspects, generating the plurality of inclusion probabilities may comprise generating the plurality of inclusion probabilities to keep an estimate of the weight matrix for the client device unbiased, while some aspects may provide that generating the plurality of inclusion probabilities comprises generating the plurality of inclusion probabilities to minimize a mean squared error between the weight matrix and an aggregation of estimates of the weight matrix for the plurality of client devices.
[0016] The processor-based server device next randomly samples a subset of the plurality of mutually orthogonal terms based on the capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities as a sub-model. According to some aspects, randomly sampling a subset of the plurality of mutually orthogonal terms comprises randomly sampling using a sampling design that preserves the plurality of inclusion probabilities. The processor-based server device then transmits the sub-model to the client device. In some aspects, the processor-based server device subsequently receives, from each client device of the plurality of client devices, an updated sub-model based on local training. The processor-based server device then aggregates the updated sub-model into the weight matrix, thus concluding the communication round.
[0017] In this regard, Figure 1 illustrates an exemplary processor-based server device 100 that includes a processor device 102 and a memory device 104. The processor device 102, which also may be referred to as a “processor core” or a “central processing unit (CPU) core,” may be an in-order or an out-of-order processor (OoP), and / or may be one of a plurality of processor devices 102 provided by the processor-based server device 100. The processor-based server device 100 of Figure 1 and the constituent elements thereof may encompass any one of known digital logic elements, semiconductor circuits, processing cores, and / or memory structures, among other elements, or combinations thereof. Embodiments described herein are not restricted to any particular arrangement of elements, and the disclosed techniques may be easily extended to various structures and layouts on semiconductor sockets or packages. It is to be understood that some embodiments of the processor-based server device 100 may include elements in addition to those illustrated in Figure 1. For example, the processor device 102 may further include one or more instruction caches, unified caches, controller circuits, interconnect buses, and / or additional memory devices, caches, and / or controller circuits.
[0018] The processor-based server device 100 of Figure 1 is configured to coordinate training of a neural network 106 by a plurality of client devices 108(0)-108(C). The client devices 108(0)- 108(C) each may comprise, as non-limiting examples, a desktop processor-based device, a mobile processor-based device, a tablet processor-based device, and the like. The neural network 106 includes a weight matrix 110 that stores weights (not shown) comprising numerical values that are associated with connectionsbetween neurons (not shown) across layers (not shown) of the neural network 106. It is to be understood that, while Figure 1 shows only one weight matrix 110 for the sake of clarity, the neural network 106 may comprise multiple weight matrices 110.
[0019] As noted above, the client devices 108(0)- 108(C) may vary widely in terms of processing, power, memory, and communications capabilities, making training the same model for each of the client devices 108(0)- 108(C) infeasible. Accordingly, in this regard, the processor-based server device 100 is configured to sample sub-models for federated learning. In exemplary operation, the processor-based server device 100 receives a capacity indicator 112 from, e.g., the client device 108(0). The capacity indicator 112 represents a processing capacity (e.g., in terms of processor, power, memory, and / or communications capability) of the client device 108(0). Although not shown in Figure 1, it is to be understood that each of the client devices 108(0)-108(C) provides a corresponding capacity indicator to the processor-based server device 100.
[0020] To perform federated learning, the processor-based server device 100 begins a communication round with the client devices 108(0)- 108(C). The processor-based server device 100 first decomposes the weight matrix 110 of the neural network 106 into a plurality of mutually orthogonal terms (captioned as “TERM” in Figure 1) 114(0)- 114(T). The processor-based server device 100 then generates a plurality of inclusion probabilities (captioned as “INC PROBABILITY” in Figure 1) 116(0)-l 16(T), each of which corresponds to a term of the plurality of mutually orthogonal terms 11 (0)- 114(T). The processor-based server device 100 generates the inclusion probabilities 116(0)- 116(T) based on one or more estimates of the weight matrix 110 for one or more client devices of the plurality of client devices 108(0)- 108(C). For example, in some aspects, the processor-based server device 100 may generate the plurality of inclusion probabilities 116(0)-116(T) to keep an estimate of the weight matrix 110 for the client device 108(0) unbiased (e.g., using a Horvitz-Thompson estimator). Some aspects may provide that the processor-based server device 100 generates the plurality of inclusion probabilities 116(0)-l 16(T) to minimize a mean squared error between the weight matrix 110 and an aggregation of estimates of the weight matrix 110 for the plurality of client devices 108(0)-108(C) (e.g., using a collective estimator).
[0021] Next, for each of the client devices 108(0)- 108(C), the processor-based server device 100 randomly samples a subset of the plurality of mutually orthogonal terms11 (0)- 11 (T) based on the corresponding capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities 116(0)-l 16(T) as a sub-model. As shown in Figure 1, the processor-based server device 100 uses the capacity indicator 112 and the inclusion probabilities 116(0)-l 16(T) to randomly sample terms 118(0)-l 18(S) (i.e., a subset of the terms 114(0)-114(T)) as a sub-model 120. According to some aspects, the processor-based server device 100 may randomly sample the subset of the plurality of mutually orthogonal terms 114(0)-l 14(T) using a sampling design that preserves the plurality of inclusion probabilities 116(0)-l 16(T) (such as, e.g. Conditional Poisson sampling (CPS), Brewer’s sampling, and MinSupport sampling, as non-limiting examples). The processor-based server device 100 transmits the sub-model 120 to the client device 108(0). While not shown in Figure 1 for the sake of clarity, it is to be understood that the processor-based server device 100 generates a sub-model for, and transmits the sub-model to, each of the client device 108(0)-108(C).
[0022] In some aspects, the processor-based server device 100 subsequently receives an updated sub-model from each of the client devices 108(0)- 108(C) (such as the submodel 120' received from the client device 108(0)), based on local training. The processor-based server device 100 then aggregates the updated sub-model 120' into the weight matrix 110. The communication round then concludes. The processor-based server device 100 may conduct multiple communication rounds as described above in the process of training the neural network 106.
[0023] To illustrate exemplary operations performed by the processor-based server device 100 of Figure 1 for sampling sub-models for federated learning according to some aspects, Figures 2A and 2B provide a flowchart showing exemplary operations 200. For the sake of clarity, elements of Figure 1 are referenced in describing Figures 2A and 2B. It is to be understood that some aspects may provide that some operations illustrated in Figures 2A and 2B may be performed in an order other than that illustrated herein, and / or may be omitted.
[0024] The exemplary operations 200 begin in Figure 2A with a processor-based server device (e.g., the processor-based server device 100 of Figure 1) receiving, from a client device of a plurality of client devices (such as the client device 108(0) of the plurality of client devices 108(0)-108(C) of Figure 1), a capacity indicator (e.g., the capacity indicator 112 of Figure 1) representing a processing capacity of the client device108(0) (block 202). The processor-based server device 100 decomposes a weight matrix of a neural network (such as the weight matrix 110 of the neural network 106 of Figure 1) into a plurality of mutually orthogonal terms (e.g., the terms 114(0)-l 14(T) of Figure 1) (block 204). The processor-based server device 100 then generates a plurality of inclusion probabilities (such as the inclusion probabilities 116(0)-l 16(T) of Figure 1) each corresponding to a term of the plurality of mutually orthogonal terms 114(0)- 114(T), based on one or more estimates of the weight matrix 110 for one or more client devices of the plurality of client devices 108(0)-108(C) (block 206). In some aspects, the operations of block 206 for generating the plurality of inclusion probabilities 116(0)- 116(T) may comprise generating the plurality of inclusion probabilities 116(0)-l 16(T) to keep an estimate of the weight matrix 110 for the client device 108(0) unbiased (block 208). Some aspects may provide that the operations of block 206 for generating the plurality of inclusion probabilities 116(0)- 116(T) comprise generating the plurality of inclusion probabilities 116(0)-l 16(T) to minimize a mean squared error between the weight matrix 110 and an aggregation of estimates of the weight matrix 110 for the plurality of client devices 108(0)- 108(C) (block 210).
[0025] The processor-based server device 100 next randomly samples a subset of the plurality of mutually orthogonal terms 114(0)- 114(T) based on the capacity indicator 112 and corresponding inclusion probabilities of the plurality of inclusion probabilities 116(0)-l 16(T) as a sub-model (e.g., the sub-model 120 of Figure 1) (block 212). According to some aspects, the operations of block 212 for randomly sampling a subset of the plurality of mutually orthogonal terms 114(0)-l 14(T) comprises randomly sampling using a sampling design that preserves the plurality of inclusion probabilities 116(0)-l 16(T) (block 214). The exemplary operations then continue at block 216 of Figure 2B.
[0026] Referring now to Figure 2B, the processor-based server device 100 transmits the sub-model 120 to the client device 108(0) (block 216). In some aspects, the processorbased server device 100 subsequently receives, from each client device 108(0) of the plurality of client devices 108(0)- 108(C), an updated sub-model (e.g., the sub-model 120' of Figure 1) based on local training (block 218). The processor-based server device 100 then aggregates the updated sub-model 120' into the weight matrix 110 (block 220).
[0027] The processor-based server device according to aspects disclosed herein and discussed with reference to Figures 1, 2A, and 2B may be provided in or integrated into any processor-based device. Examples, without limitation, include a set top box, an entertainment unit, a navigation device, a communications device, a fixed location data unit, a mobile location data unit, a global positioning system (GPS) device, a mobile phone, a cellular phone, a smart phone, a session initiation protocol (SIP) phone, a tablet, a phablet, a server, a computer, a portable computer, a mobile computing device, laptop computer, a wearable computing device (e.g., a smart watch, a health or fitness tracker, eyewear, etc.), a desktop computer, a personal digital assistant (PDA), a monitor, a computer monitor, a television, a tuner, a radio, a satellite radio, a music player, a digital music player, a portable music player, a digital video player, a video player, a digital video disc (DVD) player, a portable digital video player, an automobile, a vehicle component, an avionics system, a drone, and a multicopter.
[0028] In this regard, Figure 3 illustrates an example of a processor-based device 300, which corresponds in functionality to the processor-based server device 100 and the client devices 108(0)- 108(C) of Figure 1. The processor-based device 300 includes a processor device 302 which comprises one or more CPUs 304 coupled to a cache memory 306. The CPU(s) 304 is also coupled to a system bus 308 and can intercouple devices included in the processor-based device 300. As is well known, the CPU(s) 304 communicates with these other devices by exchanging address, control, and data information over the system bus 308. For example, the CPU(s) 304 can communicate bus transaction requests to a memory controller 310. Although not illustrated in Figure 3, multiple system buses 308 could be provided, wherein each system bus 308 constitutes a different fabric.
[0029] Other devices may be connected to the system bus 308. As illustrated in Figure 3, these devices can include a memory system 312, one or more input devices 314, one or more output devices 316, one or more network interface devices 318, and one or more display controllers 320, as examples. The input device(s) 314 can include any type of input device, including, but not limited to, input keys, switches, voice processors, etc. The output device(s) 316 can include any type of output device, including, but not limited to, audio, video, other visual indicators, etc. The network interface device(s) 318 can be any devices configured to allow exchange of data to and from a network 322. The network 322 can be any type of network, including, but not limited to, a wired or wirelessnetwork, a private or public network, a local area network (LAN), a wireless local area network (WLAN), a wide area network (WAN), a BLUETOOTH™ network, and the Internet. The network interface device(s) 318 can be configured to support any type of communications protocol desired. The memory system 312 can include the memory controller 310 coupled to one or more memory arrays 324.
[0030] The CPU(s) 304 may also be configured to access the display controller(s) 320 over the system bus 308 to control information sent to one or more displays 326. The display controller(s) 320 sends information to the display(s) 326 to be displayed via one or more video processors 328, which process the information to be displayed into a format suitable for the display(s) 326. The display(s) 326 can include any type of display, including, but not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, a light emitting diode (LED) display, etc.
[0031] Those of skill in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithms described in connection with the aspects disclosed herein may be implemented as electronic hardware, instructions stored in memory or in another computer readable medium and executed by a processor device. The master devices and slave devices described herein may be employed in any circuit, hardware component, integrated circuit (IC), or IC chip, as examples. Memory disclosed herein may be any type and size of memory and may be configured to store any type of information desired. To clearly illustrate this interchangeability, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. How such functionality is implemented depends upon the particular application, design choices, and / or design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0032] The various illustrative logical blocks, modules, and circuits described in connection with the aspects disclosed herein may be implemented or performed with a processor device, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A processordevice may be a microprocessor, but in the alternative, the processor device may be any conventional processor device, controller, microcontroller, or state machine. A processor device may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration).
[0033] The aspects disclosed herein may be embodied in hardware and in instructions that are stored in hardware, and may reside, for example, in Random Access Memory (RAM), flash memory, Read Only Memory (ROM), Electrically Programmable ROM (EPROM), Electrically Erasable Programmable ROM (EEPROM), registers, a hard disk, a removable disk, a CD-ROM, or any other form of computer readable medium known in the art. An exemplary storage medium is coupled to the processor device such that the processor device can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor device. The processor device and the storage medium may reside in an ASIC. The ASIC may reside in a remote station. In the alternative, the processor device and the storage medium may reside as discrete components in a remote station, base station, or server.
[0034] It is also noted that the operational steps described in any of the exemplary aspects herein are described to provide examples and discussion. The operations described may be performed in numerous different sequences other than the illustrated sequences. Furthermore, operations described in a single operational step may actually be performed in a number of different steps. Additionally, one or more operational steps discussed in the exemplary aspects may be combined. It is to be understood that the operational steps illustrated in the flowchart diagrams may be subject to numerous different modifications as will be readily apparent to one of skill in the art. Those of skill in the art will also understand that information and signals may be represented using any of a variety of different technologies and techniques. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referenced throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or particles, optical fields or particles, or any combination thereof.
[0035] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosurewill be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other variations. Thus, the disclosure is not intended to be limited to the examples and designs described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0036] Implementation examples are described in the following numbered clauses:1. A processor-based server device, comprising a processor device configured to: receive, from a client device of a plurality of client devices, a capacity indicator representing a processing capacity of the client device; decompose a weight matrix of a neural network into a plurality of mutually orthogonal terms; generate a plurality of inclusion probabilities each corresponding to a term of the plurality of mutually orthogonal terms, based on one or more estimates of the weight matrix for one or more client devices of the plurality of client devices; randomly sample a subset of the plurality of mutually orthogonal terms based on the capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities as a sub-model; and transmit the sub-model to the client device.2. The processor-based server device of clause 1, wherein the processor device is configured to generate the plurality of inclusion probabilities to keep an estimate of the weight matrix for the client device unbiased.3. The processor-based server device of clause 1, wherein the processor device is configured to generate the plurality of inclusion probabilities to minimize a mean squared error between the weight matrix and an aggregation of estimates of the weight matrix for the plurality of client devices.4. The processor-based server device of any one of clauses 1-3, wherein the processor device is configured to randomly sample the subset of the plurality of mutually orthogonal terms using a sampling design that preserves the plurality of inclusion probabilities.5. The processor-based server device of clause 4, wherein the sampling design comprises one of Conditional Poisson sampling (CPS), Brewer’s sampling, and MinSupport sampling.6. The processor-based server device of any one of clauses 1-5, wherein the processor device is further configured to: receive, from each client device of the plurality of client devices, an updated submodel based on local training; and aggregate the updated sub-model into the weight matrix.7. The processor-based server device of any one of clauses 1-6, integrated into a device selected from the group consisting of: a set top box; an entertainment unit; a navigation device; a communications device; a fixed location data unit; a mobile location data unit; a global positioning system (GPS) device; a mobile phone; a cellular phone; a smart phone; a session initiation protocol (SIP) phone; a tablet; a phablet; a server; a computer; a portable computer; a mobile computing device; a wearable computing device; a desktop computer; a personal digital assistant (PDA); a monitor; a computer monitor; a television; a tuner; a radio; a satellite radio; a music player; a digital music player; a portable music player; a digital video player; a video player; a digital video disc (DVD) player; a portable digital video player; an automobile; a vehicle component; avionics systems; a drone; and a multicopter.8. A processor-based server device, comprising: means for receiving, from a client device of a plurality of client devices, a capacity indicator representing a processing capacity of the client device; means for decomposing a weight matrix of a neural network into a plurality of mutually orthogonal terms; means for generating a plurality of inclusion probabilities each corresponding to a term of the plurality of mutually orthogonal terms, based on one or more estimates of the weight matrix for one or more client devices of the plurality of client devices;means for randomly sampling a subset of the plurality of mutually orthogonal terms based on the capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities as a sub-model; and means for transmitting the sub-model to the client device.9. A method for sampling sub-models for federated learning in processor devices, comprising: receiving, by a processor-based server device from a client device of a plurality of client devices, a capacity indicator representing a processing capacity of the client device; decomposing, by the processor-based server device, a weight matrix of a neural network into a plurality of mutually orthogonal terms; generating, by the processor-based server device, a plurality of inclusion probabilities each corresponding to a term of the plurality of mutually orthogonal terms, based on one or more estimates of the weight matrix for one or more client devices of the plurality of client devices; randomly sampling, by the processor-based server device, a subset of the plurality of mutually orthogonal terms based on the capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities as a sub-model; and transmitting, by the processor-based server device, the sub-model to the client device.10. The method of clause 9, comprising generating the plurality of inclusion probabilities to keep an estimate of the weight matrix for the client device unbiased.11. The method of clause 9, comprising generating the plurality of inclusion probabilities to minimize a mean squared error between the weight matrix and an aggregation of estimates of the weight matrix for the plurality of client devices.12. The method of any one of clauses 9-11, comprising randomly sampling the subset of the plurality of mutually orthogonal terms using a sampling design that preserves the plurality of inclusion probabilities.13. The method of clause 12, wherein the sampling design comprises one of Conditional Poisson sampling (CPS), Brewer’s sampling, and MinSupport sampling.14. The method of any one of clauses 9-13, further comprising: receiving, by the processor-based server device from each client device of the plurality of client devices, an updated sub-model based on local training; and aggregating, by the processor-based server device, the updated sub-model into the weight matrix.15. A non-transitory computer-readable medium, having stored thereon computerexecutable instructions that, when executed, cause a processor device of a processorbased server device to: receive, from a client device of a plurality of client devices, a capacity indicator representing a processing capacity of the client device; decompose a weight matrix of a neural network into a plurality of mutually orthogonal terms; generate a plurality of inclusion probabilities each corresponding to a term of the plurality of mutually orthogonal terms, based on one or more estimates of the weight matrix for one or more client devices of the plurality of client devices; randomly sample a subset of the plurality of mutually orthogonal terms based on the capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities as a sub-model; and transmit the sub-model to the client device.16. The non-transitory computer-readable medium of clause 15, wherein the computer-executable instructions cause the processor device to generate the plurality ofinclusion probabilities to keep an estimate of the weight matrix for the client device unbiased.17. The non-transitory computer-readable medium of clause 15, wherein the computer-executable instructions cause the processor device to generate the plurality of inclusion probabilities to minimize a mean squared error between the weight matrix and an aggregation of estimates of the weight matrix for the plurality of client devices.18. The non-transitory computer-readable medium of any one of clauses 15-17, wherein the computer-executable instructions cause the processor device to randomly sample the subset of the plurality of mutually orthogonal terms using a sampling design that preserves the plurality of inclusion probabilities.19. The non-transitory computer-readable medium of clause 18, wherein the sampling design comprises one of Conditional Poisson sampling (CPS), Brewer’s sampling, and MinSupport sampling.20. The non-transitory computer-readable medium of any one of clauses 15-19, wherein the computer-executable instructions further cause the processor device to: receive, from each client device of the plurality of client devices, an updated submodel based on local training; and aggregate the updated sub-model into the weight matrix.
Claims
What is claimed is:
1. A processor-based server device, comprising a processor device configured to: receive, from a client device of a plurality of client devices, a capacity indicator representing a processing capacity of the client device; decompose a weight matrix of a neural network into a plurality of mutually orthogonal terms; generate a plurality of inclusion probabilities each corresponding to a term of the plurality of mutually orthogonal terms, based on one or more estimates of the weight matrix for one or more client devices of the plurality of client devices; randomly sample a subset of the plurality of mutually orthogonal terms based on the capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities as a sub-model; and transmit the sub-model to the client device.
2. The processor-based server device of claim 1, wherein the processor device is configured to generate the plurality of inclusion probabilities to keep an estimate of the weight matrix for the client device unbiased.
3. The processor-based server device of claim 1, wherein the processor device is configured to generate the plurality of inclusion probabilities to minimize a mean squared error between the weight matrix and an aggregation of estimates of the weight matrix for the plurality of client devices.
4. The processor-based server device of claim 1, wherein the processor device is configured to randomly sample the subset of the plurality of mutually orthogonal terms using a sampling design that preserves the plurality of inclusion probabilities.
5. The processor-based server device of claim 4, wherein the sampling design comprises one of Conditional Poisson sampling (CPS), Brewer’s sampling, and MinSupport sampling.
6. The processor-based server device of claim 1, wherein the processor device is further configured to: receive, from each client device of the plurality of client devices, an updated submodel based on local training; and aggregate the updated sub-model into the weight matrix.
7. The processor-based server device of claim 1, integrated into a device selected from the group consisting of: a set top box; an entertainment unit; a navigation device; a communications device; a fixed location data unit; a mobile location data unit; a global positioning system (GPS) device; a mobile phone; a cellular phone; a smart phone; a session initiation protocol (SIP) phone; a tablet; a phablet; a server; a computer; a portable computer; a mobile computing device; a wearable computing device; a desktop computer; a personal digital assistant (PDA); a monitor; a computer monitor; a television; a tuner; a radio; a satellite radio; a music player; a digital music player; a portable music player; a digital video player; a video player; a digital video disc (DVD) player; a portable digital video player; an automobile; a vehicle component; avionics systems; a drone; and a multicopter.
8. A processor-based server device, comprising: means for receiving, from a client device of a plurality of client devices, a capacity indicator representing a processing capacity of the client device; means for decomposing a weight matrix of a neural network into a plurality of mutually orthogonal terms; means for generating a plurality of inclusion probabilities each corresponding to a term of the plurality of mutually orthogonal terms, based on one or more estimates of the weight matrix for one or more client devices of the plurality of client devices; means for randomly sampling a subset of the plurality of mutually orthogonal terms based on the capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities as a sub-model; and means for transmitting the sub-model to the client device.
9. A method for sampling sub-models for federated learning in processor devices, comprising: receiving, by a processor-based server device from a client device of a plurality of client devices, a capacity indicator representing a processing capacity of the client device; decomposing, by the processor-based server device, a weight matrix of a neural network into a plurality of mutually orthogonal terms; generating, by the processor-based server device, a plurality of inclusion probabilities each corresponding to a term of the plurality of mutually orthogonal terms, based on one or more estimates of the weight matrix for one or more client devices of the plurality of client devices; randomly sampling, by the processor-based server device, a subset of the plurality of mutually orthogonal terms based on the capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities as a sub-model; and transmitting, by the processor-based server device, the sub-model to the client device.
10. The method of claim 9, comprising generating the plurality of inclusion probabilities to keep an estimate of the weight matrix for the client device unbiased.
11. The method of claim 9, comprising generating the plurality of inclusion probabilities to minimize a mean squared error between the weight matrix and an aggregation of estimates of the weight matrix for the plurality of client devices.
12. The method of claim 9, comprising randomly sampling the subset of the plurality of mutually orthogonal terms using a sampling design that preserves the plurality of inclusion probabilities.
13. The method of claim 12, wherein the sampling design comprises one of Conditional Poisson sampling (CPS), Brewer’s sampling, and MinSupport sampling.
14. The method of claim 9, further comprising: receiving, by the processor-based server device from each client device of the plurality of client devices, an updated sub-model based on local training; and aggregating, by the processor-based server device, the updated sub-model into the weight matrix.
15. A non-transitory computer-readable medium, having stored thereon computerexecutable instructions that, when executed, cause a processor device of a processorbased server device to: receive, from a client device of a plurality of client devices, a capacity indicator representing a processing capacity of the client device; decompose a weight matrix of a neural network into a plurality of mutually orthogonal terms; generate a plurality of inclusion probabilities each corresponding to a term of the plurality of mutually orthogonal terms, based on one or more estimates of the weight matrix for one or more client devices of the plurality of client devices; randomly sample a subset of the plurality of mutually orthogonal terms based on the capacity indicator and corresponding inclusion probabilities of the plurality of inclusion probabilities as a sub-model; and transmit the sub-model to the client device.
16. The non-transitory computer-readable medium of claim 15, wherein the computer-executable instructions cause the processor device to generate the plurality of inclusion probabilities to keep an estimate of the weight matrix for the client device unbiased.
17. The non-transitory computer-readable medium of claim 15, wherein the computer-executable instructions cause the processor device to generate the plurality of inclusion probabilities to minimize a mean squared error between the weight matrix and an aggregation of estimates of the weight matrix for the plurality of client devices.
18. The non-transitory computer-readable medium of claim 15, wherein the computer-executable instructions cause the processor device to randomly sample the subset of the plurality of mutually orthogonal terms using a sampling design that preserves the plurality of inclusion probabilities.
19. The non-transitory computer-readable medium of claim 18, wherein the sampling design comprises one of Conditional Poisson sampling (CPS), Brewer’s sampling, and MinSupport sampling.
20. The non-transitory computer-readable medium of claim 15, wherein the computer-executable instructions further cause the processor device to: receive, from each client device of the plurality of client devices, an updated submodel based on local training; and aggregate the updated sub-model into the weight matrix.
Citation Information
Patent Citations
Method, system and apparatus for federated learning
EP4036806A1
Adaptive offloading of federated learning
US20230016827A1
Systems and methods for distributed learning for wireless edge dynamics
US20230068386A1
GR20240100067A