A controller and method to select clients in federated learning based system
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ROBERT BOSCH GMBH
- Filing Date
- 2025-11-24
- Publication Date
- 2026-06-04
Smart Images

Figure EP2025083997_04062026_PF_FP_ABST
Abstract
Description
FORM 2THE PATENTS ACT, 1970(39 of 1970) & The Patents Rules 2003COMPLETE SPECIFICATION(SECTION 10 and Rule 13)1. Title of the invention:A CONTROLLER AND METHOD TO SELECT CLIENTS IN FEDERATED LEARNING BASED SYSTEM2. Applicants:a. Name: Bosch Global Software Technologies Private Limited Nationality: INDIAAddress: 123, Industrial Layout, Hosur Road, Koramangala, Bangalore - 560095, Karnataka, Indiab. Name: Robert Bosch GmbHNationality: GERMANYAddress: Postfach 30 02 20, 0-70442, Stuttgart, GermanyComplete Specification:The following specification describes and ascertains the nature of this invention and the manner in which it is to be performed.Field of the invention:
[0001] The present invention relates to a controller and method to select clients in federated learning based system.Background of the invention:
[0002] With the increased pace of urbanization, we are witnessing tremendous growth in vehicles on road. More recently, due to the increased emission and sustainable environment growth, there is a push towards adopting Battery Electric Vehicles (BEV). This further includes commercial fleets, passenger vehicles, 2 / 3 wheelers to public transports. The number of on road vehicles are increasing day by day. On the one hand the number of on road vehicles are growing; on the other hand, the connected solution and services are on high demand. Vehicles are equipped with advanced sensors and communication capabilities which further generates high volumes of data. Considering the data generated from the vehicles, data driven analytics and solution are helping the fleet owners to end customer to take informed decisions. For example, understanding the usage pattern of the vehicle, predictive maintenance planning, location-based services, after sales support, etc. Different machine learning (ML) based solutions are being offered and under development to improve the overall quality of service (QoS) to the end users. However, recent data protection and privacy regulations mandate strict adoption of privacy preserved learning. Considering such market demand, privacy preserved on device learning is very critical.
[0003] From the real-world scenarios, on road vehicles distribution combines gasoline / petrol / diesel vehicles, electric vehicles (four wheelers), e-bikes ( two wheeler and powersports vehicles), battery operated public transports (Bus) to last-mile commercial fleets for deliveries (for example three-wheeled), etc. The energy consumption / usage for the EVs / battery operated vehicles are one of the most influential parameters for the end users. End user / customer is more than happy to optimize the energy consumption so that there is no overhead for frequent battery charging / swapping of the vehicles to maximize the battery health. While this holdstrue, the ML based solution which engages such mix of heterogeneous vehicles should address the challenges of residual energy available within vehicles, compute capabilities, data quality, network bandwidth, etc. Most of the current solution considers that the vehicles are heterogeneous in terms of compute power while ignoring the cost of energy consumption, data quality for participation in the privacy preserved learning, overall wasted energy while improving the participation ratio of different vehicles (also known as client in such learning set up).
[0004] In existing systems, the solutions are majorly on the data source selection in federated learning set up to check the data of the data source against the annotated datapoint within validation set. Some solution uses blockchain technology which eventually adds more energy. The application domain is to identify the malicious Internet-of-Things (IoT) nodes, and focused on IoT communication technology. In some, the method for data access rights in federated learning setup is disclosed. In other solutions, the method comprises the steps that terminal data sent by all terminal devices in a current communication round are obtained, the terminal data are clustered according to data similarity, and a plurality of clustering clusters are obtained; selecting a feasible data set from the clustering cluster according to feasible constraint conditions; and performing iterative updating on the feasible data set according to an overhead minimization criterion to obtain global model training data.
[0005] According to a prior art US2023128548, federated learning data source selection is disclosed. One embodiment provides a method, including: receiving, at a central server, data from each of a plurality of data sources, the plurality of data sources being within a plurality of data storage locations, wherein the central server includes a validation dataset having a plurality of annotated datapoints; computing, at the central server, an influential score for each of the plurality of data sources based upon the data provided to the central server from each of the plurality of data sources, wherein an influential score of a data source identifies an influence of the data source in accurately predicting annotations of the validation dataset; selecting,at the central server and based upon the influential score of the plurality of data sources, a subset of the plurality of data sources; and generating, at the central server, the training dataset utilizing the data of the data sources included within the subset.Brief description of the accompanying drawings:
[0006] An embodiment of the disclosure is described with reference to the following accompanying drawings,
[0007] Fig. 1 illustrates a block diagram of a federated learning-based system with a cloud controller to select clients, according to an embodiment of the present invention;
[0008] Fig. 2 illustrates a flow diagram of a method for selecting clients in the federated learning based system, according to the present invention, and
[0009] Fig. 3 illustrates a flow diagram of an example scenario, according to the present invention.Detailed description of the embodiments:
[0010] Fig. 1 illustrates a block diagram of a federated learning based system with a cloud controller to select clients, according to an embodiment of the present invention. The Federated Learning (FL) based setup / system 100 typically consists of a set ofN vehicular clients, V (134) = {V1(126), V2(128), …, VN(130)} training a shared global model 112 collaboratively without giving access to their local data. The collaborative knowledge sharing is orchestrated by a centralized entity called as aggregator server or cloud 120 which initiates the process by selecting appropriate clients 134 and disseminating the initial global model 112 to each. In Fig. 1, five clients 134 are shown, with two clients 134, i.e., a first vehicle 122 and a second vehicle 124 are dropped from selection, whereas a third vehicle 126, a fourth vehicle 128 and a fifth vehicle 130 are selected for federated learning. The clients 134 is possible to be a scooter, a motorcycle, a car, a truck, a bus, a train, and other types of vehicles either powered by electricity or fuel. Each client 134 is provided with a device 132 called as control unit or processing unit. The number ofclients 134 shown is just for explanation, and in practice the number is possible to be less or more than the five.
[0011] According to an embodiment of the present invention, the controller 110 (or cloud controller) to select clients 134 in federated learning based system 100 is disclosed. The system 100 comprises the aggregator server / cloud 120, comprising the controller 110, connectable to plurality of clients 134 for federated learning. The controller 110, in / of the aggregator server / cloud 120, configured to establish connection with at least two clients 134 available for participation, characterized in that, receive, using a resource extractor 102, the resource information from each of the at least two clients 134, calculate a Client Selection Score (CSS) for each of the at least two clients 134 based on respective resource information, and select clients according to the CSS score for federated learning. The CSS is calculated using a score estimator module 104 of the controller 110. The resource extractor 102 is shown outside the cloud 120 for simplicity, but is shown to be part of the controller 110 of the cloud 120 using the dotted callout box.
[0012] According to an embodiment of the present invention, the resource information is selected from group comprising data rate (Rd), Central Processing Unit (CPU) quality (Qc), system memory (Ms), survival rate (Ts), client utility score (Cus), and participation count (Ps). The data rate (Rd) comprises number of samples, sample bit size, model parameters size; the CPU quality (Qc) comprises number of CPUs, frequency of CPUs, and CPU capacitance coefficients; the survival rate (Ts) comprises residual energy of the battery, time consumed during training, energy consumption during each training round, communication time cost, communication energy cost; and the participation count (Ps) comprises number of times the client 134 has participated in federated learning.
[0013] According to an embodiment of the present invention, before the clients 134 are selected based on CSS, the controller 110 of the cloud 120 configured to optimize, using an optimizer module 106, the CSS based on Particle SwarmOptimization (PSO) and calculate fitness score for each of the at least two clients 134, and filter out, using a selector module 108 unsuitable clients 134 from the result of PSO using filter parameters. The filter parameter is at least one selected from a group comprising the CPU quality and the residual energy of each of the respective client 134. The selection of clients 134 is possible with or without implementation of PSO.
[0014] According to an embodiment of the present invention, the inner structure of the controller 110 of the FL system 100 is disclosed. A computation model is provided which consists of the resource usage information during the training phase. The inherent heterogeneity in local compute resources of each device 134 affects the local training process and the reliability of FL framework. The computation model consists of the measurement of the training time and the energy consumed during training of local model. Each participating client 134 has certain compute capacity which determines the time required to finish the local training and the energy consumed by each client 134 during local training. The time consumed during training for ithclient 134 is defined asTcomp — bi |Di| pi / fiwhere,Di is the ithclient dataset,bi is the bit size of input data sample Di,pi is the average CPU cycles required to process one data sample, fi is the CPU frequency of ithclient.
[0015] Similarly, based on the training time delay, the energy consumption of client i in each training round is calculated asEcomp= cibi|Di| ρifi2where,ci is the capacitance coefficient of ithclient device’s on-board equipment chip.
[0016] According to an embodiment of the present invention, a communication model is disclosed. In vehicular networks, FL system 100 is deployed in on-board devices 134 of vehicular clients 134 with limited communication resources. The devices 134 with limited bandwidth and poor connectivity become the bottlenecks since the model parameters need to be exchanged between the clients 134 and the server 120. The exchange happens for several rounds until convergence, which generates huge communication overhead. After the local training in FL system 100, the client devices 134 transmit the model updates to the cloud 120 and receive the updated global model 112 after aggregation. The time taken to upload the model parameters from the client device 134 to the cloud 120 ignoring the communication cost during the download from cloud 120 to the client 134 is measured / calculated. The communication cost calculation includes the time and energy consumed during transmission. The communication time can be calculated asB log2[(1 + hipi) / (Ni+ εi)]where,B is the subchannel bandwidth,hi is the channel gain,pi is the transmit power of vehicle i,Ni is the noise power, andei is the received interference power.The communication energy cost can be derived fromMipi / Tcommwhere,Mi is the size of model parameters.
[0017] The communication model and the computation model are used in but not limited to survival rate.
[0018] According to an embodiment of the present invention, a Client Quality Assessment model is disclosed. The quality of the participating clients 134 plays avital role in overall convergence of the FL process. The selection of high quality clients 134 significantly improve the convergence time and improve energy efficiency too. However, inadequate assessment of client 134 resources may result in dropout of respective clients 134 which leads to wastage of consumed resources. To address this and filter out high quality clients 134, a client selection problem under a highly dynamic real-world scenario is evaluated / investigated. It is also unrealistic to assume full participation from each client 134. In fact, clients 134 are free to join and leave the loose “federation” at any time they want. With these considerations, selection of a portion of available clients 134 based on assessment score calculated from the quality of resources they have is done. The following resources to calculate resource aware assessment score for each client device is considered and defined. The resources are CPU Status, CPU Quality, data rate, system memory, residual energy score, and client utility score.
[0019] According to the present invention, the resources are described for better understanding. The CPU status gives the information about the compute capacity of the client device. Here, the CPU load required to perform one round of training and CPU capacity to handle this load is measured. Specifically, values are assigned, such as cpu status = 0, if the number of CPU cores present are less than the number of CPU cores required to bear the training load, else cpu status = 7 is assigned. The metric indicates the feasibility of training round’s success or failure. The CPU quality is the metric which measures the quality of compute capacity of the client 134. A higher value indicates the higher performance by the client 134 in terms of computation cost. It is defined ascpu_quality (Qc) = nCores * CPU freqwhere,nCores is the number of CPU cores in the client device, andCPUfreq denotes the CPU frequency.cpu quality (Qc) indicates higher CPU quality will reduce the computation COSt (Tcomp and Ecomp)
[0020] The data rate (Rd) defines the rate of transmission of model updates between the client 134 and the cloud 120. A higher value indicates stable and lower communication cost. It is expected to include clients 134 with a higher data rate to minimize the communication time (Tcomm) and communication energy (Ecomm). With respect to system memory (Ms), the number of local samples in the client 134 has a direct impact on its personalized performance. A higher system memory facilitates seamless FL process. However, it is interesting to monitor the memory usage during the training round. The limited availability of system memory slows down the running processes, hence higher availability of system memory is desired / preferred.
[0021] The power related information, such as the residual charge of the battery helps to assess the success rate of the FL training. The average energy consumption during the computation and communication is added to get approximate estimate of the battery power usage in each training iteration. It becomes essential to predict the estimated energy consumption per round when dealing with the battery powered devices, especially when the battery status is Not Charging. If the estimated energy consumption per round comes out to be higher than the residual energy of the battery, the client 134 has high chance of dropout or training failure. In this case, the energy consumed during partial training is considered as wastage. Therefore, these highly susceptible clients 134 from the training process must be filtered out. To perform this, two factors are considered, i.e., number of training rounds possible and residual energy before and after the training round to update the former value. The number of training rounds is defined aswhere,REcurdenotes the current residual energy of the client. However, the residual energy is updated after the training asREcur— (Ecomp+ Ecomm)(Note: REnew becomes REcurin the next iteration)From this, it is concluded that the client 134 with higher residual energy sustains a greater number of training iterations without fail.
[0022] To assess the quality of client 134 and its contribution in FL process, client utility score (Cus) is used. To calculate Cus, two factors are considered, i.e., the computation efficiency of the client 134 and the contribution of the client 134 in overall performance. The computation efficiency of the client 134 is measured by the average time taken during local training and uploading the model parameters per sample, i.e.,|D | / (Tcoinp + Tcoinin)Further, the contribution of the client 134 is measured by the divergence in the local loss values calculated during two consecutive iterations, i.e.,dlloss^ - losst||2 / |D |)Now, combining these two equations, Cus is defined asCus = (| |losst_ [ — 10SSt| |2 / |D |) * |D | / (Tcomp + Tcomm)This illustrates that the client 134 with lower loss divergence guides the global model 112 to reach convergence faster. It follows from the conclusion that the local loss of the client 134 represents the learning degree of the dataset on the global model 112.
[0023] According to the present invention, the calculation of resource aware Client Selection Score (CSS) is disclosed. The relationship between the resource availability and the energy efficiency during training is evaluated. In addition, the impact of training on residual energy of the battery and to filter the vulnerable clients 134, preventing the resource wastage caused by the dropout clients 134 is also assessed. In summary, the objective is to minimize the impact due to following challenges heterogeneity in resources and local data of the clients 134, unstable andnon-uniform connectivity due to mobility of the clients 134, and insufficient resources for local training of clients 134.
[0024] To minimize the impact of the above challenges, the clients 134 are selected based on their resource availability which are further utilized to estimate their selection score. To calculate the Client Selection Score (CSS), the following resource information is considered which is already disclosed above, CPU related (CPU cores, CPU frequency, CPU capacitance coefficient), memory related (RAM), power related (battery residual energy, battery status), data and model related (number of samples, sample bit size, model parameters size), wireless channel related (data rate, channel gain, sub-channel bandwidth, transmit power etc.), and participation count (Ps). The CPU related information is used to calculate the CPU status and CPU quality to further estimate the computation time and computation energy. Similarly, the wireless channel related information is required to predict the connection stability and to estimate the communication energy and time. Most important parameters are power related which help to optimize the resource wastage. Based on the power related resource information, the sustainability of any client 134 is predicted. If the client 134 has sufficient resources, then only it could be considered to perform training, else it may shutdown in between and will add to energy wastage. The participation count of any client 134 illustrates the number of iterations the client 134 has participated. This information is required because it provides the fairness value in participation. To improve the participation count of each client 134, the clients with lower value are given preference. All these resource information and their related parameters contribute into the estimation of client selection score, given byCSS = [Rd, Qc, Ms, Ts, Cus, Ps]
[0025] The controller 110 employs a score estimator module 104 to estimate the client selection score for each client 134. Estimating the selection score before the start of each FL iteration is important as it predicts the most efficient clients 134 to minimize resource wastage. We make use of this information and maximize the selection score for each client 134 by applying optimization algorithm. The idea ofutilizing the optimization algorithm is to further extrapolate it into greedy selection of clients 134 based on the required parameter.
[0026] According to the present invention, the CSS value optimization using Particle Swarm Optimization (PSO) is disclosed. The PSO is widely used nature inspired optimization algorithm. After estimating the CSS, the controller 110 configured to maximize CSS of each client 134 to get optimized selection of clients 134. First, the client selection problem is formulated as optimization of a fitness function which is described as,Fitness function (F) = max (CSS(Rd,Before applying the PSO algorithm to maximize the fitness function, the above equation is rewritten asF = a Rd + PQc + yMs + 8Ts + oCus + x Ps where,{a, P, y, 5, c, x} are fitness parameters which maximize the weights of each resource value and lies in between 0 and 1.
[0027] To estimate the values of fitness parameters and to maximize the CSS values for each client 134, the PSO algorithm is applied. In PSO, each particle is characterized by its position (Xi) and velocity (Vi),Xi = (al, pi, yi, 61, cri, xi )Vi = (Vai, Vpi, Vyi, V81, Vol, Vxi )where, i represents the particle index.
[0028] The particle's position is updated iteratively using the following equations:Vai(t+1) = C0Vai(t) + Ciri(Pai - ai) + C2r2(Pga - ai) Vpi(t+1) = coVpi(t) + ciri(Ppi - pi) + c2r2(Pgp - pi) VYi(t+l) = coVYi(t) + Cin(PYi- yi) + c2r2(PgY- yi) Vgi(t+1) = coVsi(t) + ciri(Psi - 6i) + c2r2(Pg5 - 6i) Vai(t+1) = coVai(t) + ciri(Pai - oi) + c2r2(Pgo- ci) Vxi(t+1) = coVTi(t) + ciri(PTi- xi) + c2r2(PgT- xi)ai(t+l) = ai(t) + Vai(t+1)Pi(t+1) = Pi(t) + Vpi(t+1)Yi(t+1) = yi(t) + VYi(t+l)5i(t+ 1 ) = 5i(t) + Vgi(t+1)Oi(t+1) = Oi(t) + Vai(t+1)ri(t+l) = ri(t) + Vri(t+1)where,Vai, Vpi, VTi, V(ii, Vo,, VT, are the velocities, andPai, Ppi, PYi, Psi, Pai, Pzi are the personal best positions of particle i in the ai, Pi, Yi, bi, Si, and Ti dimensions, respectively at iteration t.Pga, Pgp, Pgy, Pgs, Pga, and PgTare the global best positions in the m, Pi, Yi, bi, a,, and Ti dimensions, respectively.co is the inertial weight, ci and C2 are the acceleration coefficients, and n and r2 are random numbers between 0 and 1.
[0029] The personal best position Pai, Ppi, PTi, Pgi, Pai, and PTiof particle i is updated asai, f(ai, pt, yi, Si, ct i, Ti) > f(Pai, Pp i, Pyi, PSi, Peri, PT!)Pai, otherwisef(ai, pi, yi, Si, ai, Ti) > f(Pai, Pp i, Pyi, PSi, Pai, PT!)otherwisef(ai, pi, yi, Si, ai, Ti) > f(Pai, Pp i, Pyi, PSi, Peri, PT!) pYi= f (pPYyti> otherwisef(ai, pi, yi, Si, ai, Ti) > f(Pai, Pp i, Pyi, PSi, Peri, PT!)D PbKl' = f ]n5si'.(P61, otherwiseTi, f(ai, pi, yi, Si, ai, Ti) > f(Pai, Pp i, Pyi, PSi, Pai, PT!)PTi, otherwise
[0030] The global best positions Pga, PgP, Pgy, Pgb, Pgs, and PgT is updated as ( al, f(«i, fi, yi, Si, ai, Ti) > f(Pga, Pgf, Pgy, PgS, Pga, PgT) Pga = ] (.DPga, otherwise( pi, f(ai, pi, yi, Si, ai, Ti) > f(Pga, Pgf, Pgy, PgS, Pga, PgT)PgP =(PgP, otherwisef yi, f(ai, pi, yi, Si, ai, Ti) > f(Pga, Pgf, Pgy, PgS, Pga, PgT)PgY“ lPgY> otherwise_ f 81, ffai, pi, yi. Si, oi, Ti) > f(Pga, Pgf, Pgy, PgS, Pga, PgT)PgS =I PgS. otherwisec Ti, f(ai, pi, yi, Si, ai, Ti) > f(Pga, Pgf, Pgy, PgS, Pga, PgT)PgT =(PgT, otherwise
[0031] The algorithm continues iterating until a stopping criterion is met, such as a maximum number of iterations or the achievement of a satisfactory solution. This algorithm effectively explores the search space and converges towards the optimal solution for maximizing the given function for CSS by estimating the following,F = a Rd + PQc + yMs + STs + CTCUS + T PS
[0032] According to the present invention, energy efficient client selection algorithm is disclosed. To derive an efficient resource aware FL process for Internet of Vehicles (loV), a novel client selection algorithm is disclosed to mitigate the impact imposed by the resource constraint clients 134. After getting the optimized estimates of the fitness score for all clients 134, the controller 110 in the cloud 120 is now able to select the optimal clients 134 among all. If the equation of CSS is closely monitored, it contains data rate, CPU quality, system memory, residual energy of the battery, client utility score, and the participation count which formulate the objective function. By understanding and examining these resources, the resource awareness is integrated into the proposed FL framework. Initially, the FL server / cloud 120 publishes a task with minimum system requirements. All the interested clients 134 acknowledge the cloud 120 by sending their respective resource-availability information. The controller 110 in the cloud 120 then formulates an optimization problem to fully utilize the resource information. To enable this, PSO optimization is leveraged. The controller 110 of the cloud 120 needs to filter out the inefficient clients 134 to prevent the resource wastage. However, the optimized selection score only suggests the probability of selection for each client 134. With this, it may happen that among two clients 134 with close values of fitness score, one client has high chances of getting dropout. But because of overall high availability of resources, it is preferred over the other client 134. Now, after few epochs of local training the client 134 gets dropped due to insufficient residual energy or may be due to poor compute capacity. Ultimately contributing into the resource wastage. Therefore, apart from optimizing the CSSvalue, further constraints are needed for a robust and efficient client selection algorithm.
[0033] At the beginning, with the shared resources the controller 110 of the cloud 120 aims to estimate and optimize the CSS. The coefficients in Fitness function (F), are used to set weight for each of the resource values to be considered in priority. A higher value of this index indicates a higher priority. The selection procedure begins at the cloud 120 or server end.
[0034] According to the present invention, an experiment result is disclosed for better clarity and understanding. To compare the impact of resource heterogeneity in FL system 100 and to quantify the wastage of resources while optimizing the energy consumption, the solution proposed in the present invention is tested on a dataset. Using the dataset, a Driver Behavior Monitoring (DBM) task is performed and classifies it among Drowsy, Aggressive, and Normal classes. Two partitions of the dataset was created, i.e., Independent and Identically Distributed (IID) dataset, and Non-IID dataset. The partitioning of the clients 134 is done using Dirichlet distribution with a = 1, to realize real-time scenarios and created 50 non IID clients 134. A CNN-LSTM model is built for DBM task which consists of 2-D convolutional layer with 16 filters, 2-D Max pooling layer, and two fully connected layers with 80 and 64 channels. Adam as optimizer is adopted, learning rate is fixed to 0.001, batch size to 32, local epochs are 5, and ran the experiment for 100 rounds. The below table details a comparative analysis result and demonstrate significant benefits of the present invention for energy efficient selection algorithm for driver behavior monitoring application under heterogeneous vehicles distribution.Parameter Random Selection Proposed Novel Method Selection Method Total number of rounds 100 100Total number of clients 50 50Participating clients in each 5 5roundTotal dropout clients 49 0Average dropout / round 0.49 0Total energy consumed (in joule) 606.83 12.14Total energy wastage (in joule) 32.98 0Average energy wastage / round 0.3298 0Convergence round for minimum 46 4275% Average accuracy for driverbehavior monitoringAverage accuracy for driver 77.61 % 77.15 % behavior monitoring application
[0035] According to an embodiment of the present invention, a brief working of the controller 110 of the cloud 120 is disclosed. The controller 110 of the cloud 120 receives willingness requests to participate in the training together with the respective resource information. With the resource information collected, the controller 110 uses a selection score estimator to calculate client selection score for each client 134 and is forwarded to an optimizer module 106 within the controller 110. The optimizer module 106 implements PSO and forwards to a selector module 108 of the controller 110. The selector module 108 selects a fraction of optimal clients and shares / sends the global model 112 along with the configuration details to the selected clients 134. The configuration details comprises learning rate (Zr), epoch (e), number of rounds (rds), batch size (bs), list of selected clients (k), etc. On receiving the global model 112, the clients 134 perform local training on their private data and send the model updates 118 to the cloud 120 once the training is completed. An aggregator module 114 of the controller 110 in the cloud 120 performs aggregation of model once the model updates 118 from each selected clients 134 is received, substituting the previous global model 112. Finally, the new iteration is started with the first step.
[0036] Fig. 2 illustrates a flow diagram of a method for selecting clients in the federated learning based system, according to the present invention. The system 100 comprises the cloud 120, comprising the controller 110, connectable to plurality of clients 134 for federated learning. The method comprises plurality of steps of which a step 202 comprises establishing, with the controller 110, connection with at least two clients 134 available for participation. The method characterized by, a step 204 which comprises receiving, by the controller 110, resource information from the at least two clients 134. The controller 110 uses the resource extractor 102 for the same. A step 206 comprises calculating, by the controller 110, a Client Selection Score (CSS) for each of the at least two clients 134 based on respective resource information. A step 208 comprises selecting, by the controller 110, clients 134 according to the CSS score for federated learning. The method is performed / executed by the controller 110 and internal modules and the same must not be understood in limiting manner.
[0037] According to the method, before step 208, i.e., before selecting the clients 134 based on CSS, the method comprises a step 210. The step involves application of Particle Swarm Optimization (PSO). The step 210 further comprises a step 212 and a step 214. The step 212 comprises optimizing the CSS based on Particle Swarm Optimization (PSO) and calculating fitness score for each client. The step 214 comprises filtering out unsuitable clients 134 from the result of PSO using filter parameters. According to the method, the filer parameter is at least one selected from the group comprising CPU quality and residual energy of each of the client 134. Please note that the method is implementable with or without the step of PSO.
[0038] According to the method, the resource information is selected from group comprising data rate (Rd), Central Processing Unit (CPU) quality (Qc), system memory (Ms), survival rate (Ts), Client utility score (Cus), and participation count (Ps). The data rate comprises number of samples, sample bit size, model parameters size. The CPU quality (Qc) comprises number of CPUs, frequency of CPUs, and CPU capacitance coefficients. The survival rate (Ts) comprises residual energy ofthe battery, time consumed during training, energy consumption during each training round, communication time cost, communication energy cost, and the participation count (Ps) comprises number of times the client 134 has participated in federated learning.
[0039] Fig. 3 illustrates a flow diagram of an example scenario, according to the present invention. The working of the Federated learning using the cloud 120 and the clients 134 is explained. A step 302 comprises receiving, by the controller 110 of the cloud 120, a request from clients 134 for participation in Federated Learning (FL) based system 100. The step 302 also comprises requesting the interested clients 134 to share respective resource information. A step 304 comprises calculating CSS using Rd, Qc, Ms, Ts, Cus, Ps once the cloud 120 receives the resource information from all clients 134. A step 306 comprises optimizing the selection score using PSO algorithm considering similar priority index for each factor for maximizing the chances of selection. The algorithm calculates a fitness score for each client to make a resource aware client selection (refer the revised fitness function). A decision step 308 comprises performing a sanity check on clients 134 by applying additional constraints, such as compute capacity and on survival rate of the clients 134. This is done to filter out the unsuitable clients 134 or clients 134 with high chances of getting dropout. For example, if CPU quality (Qc.i) < Qc,iTh (Threshold Value for client i) or if survival rate (Ts) = 0, then the client 134 is removed from the selection list. A step 310 indicates he removal of the client 134 from the list. If all the clients 134 fulfill the requirements, then a step 312 is executed directly after the step 308 or after removal in the step 310. The step 312 comprises obtaining final list of useful participating clients 134.
[0040] A step 314 comprises selecting top (k*p)% participating clients 134 from the list, where k is selected clients and p is the configuration details. Alternatively, to diversify the participation of clients 134, the fitness score of clients 134 selected in the previous round is reduced by 50%. This gives priority to non-selected clients 134 over selected clients 134. Otherwise as per step 326, update the respectiveresource information and repeat from step 302. A step 316 comprises updating and broadcasting the global model 112 for the clients 134. The selected clients 134 receive the global model 112 and configuration details to perform local training. A step 318 comprises performing the federated learning among the clients 134. A step 320 comprises updating participation count=participation count+1, CSS, and residual energy = residual energy - energy consumed, and loss (t). (In experiments, it was assumed the vehicle clients 134 are mobile, hence the residual energy for each client 134 is updated after every predetermined number of rounds). A step 322 comprises updating or transmitting model parameters or updates to the controller 110 of the cloud 120. The controller 110 aggregates the received data from multiple clients 134 and updates the global model 112 and repeats the step 314 through 322 until a stopping criteria is met or until convergence.
[0041] According to the present invention, the controller 110 of the cloud 120 is in communication with the clients 134 through Cellular Vehicle-to-Everything (C-V2X).
[0042] According to the present invention, the controller 110 is refers to computing devices / units comprising components such as memory element such as Random Access Memory (RAM) and / or Read Only Memory (ROM), Analog-to-Digital Converter (ADC), Digital-to-Analog Convertor (DAC), clocks, timers, and a processor (such as Central Processing Unit (CPU)) (capable of implementing machine learning) connected with each other and to other components through communication bus channels in a Printed Circuit Board (PCB). The components mentioned are just for understanding and may have more or less components as per requirement. The memory element is prestored with map, table, model, modules, logics, instructions, programs, applications, threshold deviation, or values, which is accessed by the at least one processor as per the defined routines. The internal components of the controller 110 are not explained for being state of the art, and the same must not be understood in a limiting manner. The controller 110 is capable to communicate through wired and wireless means such as but not limited to GlobalSystem for Mobile Communications (GSM), 3G, 4G, 5G, Wi-Fi, Bluetooth, Ethernet, serial networks, Universal Serial Bus (USB) cable, micro-USB, Wi-Fi and the like.
[0043] Further, the processor may be implemented as any or a combination of one or more microchips or integrated circuits interconnected using a parent board, hardwired logic, software stored in the memory element and executed by a microprocessor, firmware, an application specific integrated circuit (ASIC), and / or a field programmable gate array (FPGA). The processor is configured to exchange and manage the processing of various Artificial Intelligence (Al) modules.
[0044] According to the present invention, the energy efficient privacy preserved learning under heterogeneous vehicles distribution is disclosed. In the present invention, a novel approach to address the challenges, as discussed in background, is disclosed, while optimizing the energy consumption, overall user experience and enabling the privacy preserved learning under heterogeneous vehicles distribution. The present invention offers a solution which is outperforming and optimizing the overall energy consumption, energy wastage of the entire federated learning system 100. The solution also improves the overall quality of experience by optimizing the dropout in contrast to baseline approach. In addition, the present invention has no accuracy trade-off as compared to the baseline approach while offering faster convergence time i.e., 42 rounds only as compared to the baseline approach.
[0045] It should be understood that the embodiments explained in the description above are only illustrative and do not limit the scope of this invention. Many such embodiments and other modifications and changes in the embodiment explained in the description are envisaged. The scope of the invention is only limited by the scope of the claims.
Claims
We claim:
1. A controller (110) to select clients (134) in federated learning based system (100), said system (100) comprises a cloud (120), comprising said controller (110), connectable to plurality of clients (134) for federated learning, said controller (110) configured to,establish connection with at least two clients (134) available for participation, characterized in that,receive resource information from each of said at least two clients (134);calculate a Client Selection Score (CSS) for each of said at least two clients (134) based on respective resource information, and select clients (134) according to said CSS score for federated learning.
2. The controller (110) as claimed in claim 1, wherein before said clients (110) are selected based on CSS, said controller (110) configured to, optimize, using an optimizer module (106), said CSS based on Particle Swarm Optimization (PSO) and calculate fitness score for each of said at least two clients (134), andfilter out, using a selector module (108), unsuitable clients (134) from the result of PSO using filter parameters.
3. The controller (110) as claimed in claim 2, wherein said filter parameter is at least one selected from a group comprising CPU quality and residual energy of each of said client (134).
4. The controller (110) as claimed in claim 1, wherein said resource information is selected from group comprising data rate (Rd), Central Processing Unit (CPU) quality (Qc), system memory (Ms), survival rate (Ts), client utility score (Cus), and participation count (Ps).
5. The controller (110) as claimed in claim 4, wherein said data rate comprises number of samples, sample bit size, model parameters size, said CPU quality comprises number of CPUs, frequency of CPUs, and CPU capacitance coefficients, said survival rate comprises residual energy of a battery, time consumed during training, energy consumption during each training round, communication time cost, communication energy cost, and said participation count comprises number of times said client (134) has participated in federated learning.
6. A method for selecting clients (134) in federated learning based system (100), said system (100) comprises a cloud (120), comprising a controller (110), connectable to plurality of said clients (134) for federated learning, said method comprising the steps of:establishing connection with at least two clients (134) available for participation, characterized by,receiving resource information from said at least two clients (134); calculating a Client Selection Score (CSS) for each of said at least two clients (134) based on respective resource information, and selecting clients (134) according to said CSS score for federated learning.
7. The method as claimed in claim 6, wherein before selecting said clients (134) based on CSS, said method comprises:optimizing, using an optimizer module (106), said CSS based on Particle Swarm Optimization (PSO) and calculating fitness score for each client (134), andfiltering out, using a selector module (108), unsuitable clients (134) from the result of PSO using filter parameters.
8. The method as claimed in claim 7, wherein said filter parameter is at least one selected from a group comprising CPU quality and residual energy of each of said client (134).
9. The method as claimed in claim 6, wherein said resource information is selected from group comprising data rate (Rd), Central Processing Unit (CPU) quality (Qc), system memory (Ms), survival rate (Ts), client utility score (Cus), and participation count (Ps).
10. The method as claimed in claim 9, wherein said data rate comprises number of samples, sample bit size, model parameters size, said CPU quality comprises number of CPUs, frequency of CPUs, and CPU capacitance coefficients, said survival rate comprises residual energy of a battery, time consumed during training, energy consumption during each training round, communication time cost, communication energy cost, and said participation count comprises number of times said client (134) has participated in federated learning.Dated 27 November 2024 (Digitally signed)Siddharth Karkhanis (IN / PA- 1195) On-behalf of the Applicants