Method for federated learning

WO2026206061A1PCT designated stage Publication Date: 2026-10-01SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/004933
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-09-11
Filing Date
2026-03-27
Publication Date
2026-10-01

Smart Images

  • Figure KR2026004933_01102026_PF_FP_ABST
    Figure KR2026004933_01102026_PF_FP_ABST
Patent Text Reader

Abstract

In an embodiment, the method may include receiving, from a server, a plurality of sets of trainable parameters of the plurality of image generation ML models which have been pre-trained, identifying, using local training data stored on the user device, a set of trainable parameters, among the received plurality of sets of trainable parameters, to be updated using the local data of the user device, updating, using the local training data, the selected set of trainable parameters by training the stored image generation ML model configured with the selected set of trainable parameters, adding noise to the updated set of trainable parameters to generate a set of noisy trainable parameters, transmitting the set of noisy trainable parameters to the server for aggregation with other sets of noisy trainable parameters, received from other user devices, corresponding to the selected set of trainable parameters.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD FOR FEDERATED LEARNING

[0001] The present application generally relates to a method for performing federated learning of a machine learning, ML, model. In particular, the present techniques provide a method for training a plurality of image generation ML models using federated learning, in a way that preserves user data privacy and prevents semantic and concept drift.

[0002] In recent years, more and more mobile devices and home appliances rely on deep models for computer vision tasks. Users in different locations around the globe interact with different objects and different environments. For example, cars in Italy, US and India can be very different with some classic cars only common in Italy, some large cars mainly seen in the US and the style of cars in India not generally seen in Europe and / or the US. Similarly, images for busy streets in Rome or Mumbai are different as are types of food and home interiors. Some regions around the globe are also severely underrepresented in data collected for model training, e.g. data on African languages is 108 times smaller than data on European languages and similarly certain ethnicities are under-represented in most vision datasets. Training on user data without protection can lead to privacy risks, such as leaking of data and / or use of personal data in creating deep fakes which can be used for identity theft or other crimes.

[0003] Computer vision tasks are varied and include classification, object localization, segmentation and semantic segmentation. Semantic segmentation comprises assigning to each pixel in an image, a label which represents the class or object (e.g. background, road sign, people, sidewalk, road, cars etc.) to which the pixel belongs. This is a dense task which has been revolutionized by deep learning. Semantic segmentation has commercial importance in various commercial applications such as image editing (e.g., filtering, object removal), smart wearable devices (e.g. Metaverse), autonomous vehicles, robotics, video surveillance, or medical images.

[0004] Figures 1A and 1B schematically illustrate training and inference of text-to-image generation model. Such image generation involves creating an image which is conditioned to a given source information which is provided as text. The known state-of the art involves using a latent diffusion model. Figure 1C shows how low-rank adaptation (LoRA) may be used to fine-tune the model. The main model is frozen and trainable rank decomposition matrices are injected into each layer of the model. The trainable rank decomposition matrices are trained on the user data. LoRA fine-tuning greatly reduces the number of parameters which are trained. For image generation, LoRA fine-tuning is typically combined with the Dreambooth protocol.

[0005] Figure 2 illustrates a model for generating music. As shown, text to music (TTM), image to music (ITM) or any other modality to music can be used. In each example, a music track with a pre-defined length which is conditioned to a given source of information (e.g. text, image, melody) is generated. The main blocks of the music generation model include an encoder, a generator and an audio decoder. The encoder converts the input (e.g., text, image, video, humming) information into an embedding that is used to condition the creation of music. The generator creates a sequence of tokens based on the embedding. The audio decoder transforms the sequence of tokens into a sequence of audio samples to create the music track.

[0006] In an embodiment, a computer-implemented method, performed by a user device, for training a plurality of image generation machine learning(ML) models may be provided. In an embodiment, the method may include receiving, from a server, a plurality of sets of trainable parameters of the plurality of image generation ML models which have been pre-trained. In an embodiment, the method may include identifying, using local training data stored on the user device, a set of trainable parameters, among the received plurality of sets of trainable parameters, to be updated using the local data of the user device, by: calculating, using the local training data, a loss for each of the received plurality of sets of trainable parameters, by applying the set of trainable parameters to an image generation ML model stored on the user device; and selecting the set of trainable parameters which provide the stored image generation ML model with the lowest calculated loss. In an embodiment, the method may include updating, using the local training data, the selected set of trainable parameters by training the stored image generation ML model configured with the selected set of trainable parameters. In an embodiment, the method may include adding noise to the updated set of trainable parameters to generate a set of noisy trainable parameters. In an embodiment, the method may include transmitting the set of noisy trainable parameters to the server for aggregation with other sets of noisy trainable parameters, received from other user devices, corresponding to the selected set of trainable parameters.

[0007] In an embodiment, a user device for training a plurality of image generation machine learning(ML) models may be provided. In an embodiment, the user device may comprise memory storing instructions, local training data, and an image generation ML model and at least one processor operatively coupled to the memory and comprising processing circuitry. In an embodiment, the at least one processor may individually or collectively execute the instructions to cause the user device to receive, from a server, a plurality of sets of trainable parameters of the plurality of image generation ML models that have been pre-trained. In an embodiment, the at least one processor may individually or collectively execute the instructions to cause the user device to identify, using the local training data, a set of trainable parameters, among the received plurality of sets of trainable parameters, to be updated using the local data of the user device, by: calculating, using the local training data, a loss for each of the received plurality of sets of trainable parameters by applying the set of trainable parameters to the stored image generation ML model; and selecting the set of trainable parameters which provide the stored image generation ML model with the lowest calculated loss. In an embodiment, the at least one processor may individually or collectively execute the instructions to cause the user device to update, using the local training data, the selected set of trainable parameters by training the stored image generation ML model configured with the selected set of trainable parameters. In an embodiment, the at least one processor may individually or collectively execute the instructions to cause the user device to add noise to the updated set of trainable parameters to generate a set of noisy trainable parameters. In an embodiment, the at least one processor may individually or collectively execute the instructions to cause the user device to transmit the set of noisy trainable parameters to the server for aggregation with other sets of noisy trainable parameters, received from other user devices, corresponding the selected set of trainable parameters.

[0008] In an embodiment, a computer-implemented method, performed by a server, for training a plurality of image generation machine learning(ML) models may be provided. In an embodiment, the method may include transmitting, to a plurality of user devices, a plurality of sets of trainable parameters of the plurality of image generation ML models that have been pre-trained. In an embodiment, the method may include receiving, from at least a portion of the plurality of user devices, sets of noisy trainable parameters. In an embodiment, the method may include for each set of trainable parameters: aggregating the received sets of noisy trainable parameters that correspond to the set of trainable parameters, to generate an updated set of trainable parameters for an image generation ML model corresponding to the set of trainable parameters; and transmitting the updated set of trainable parameters to the at least a portion of the plurality of user devices.

[0009] An embodiment of the present disclosure will now be described, by way of example only, with reference to the accompanying drawings, in which:

[0010] Figures 1A to 1C are diagrams of exemplary image generation models;

[0011] Figure 2 is a block diagram of an exemplary music generation model;

[0012] Figure 3 is a graph comparing accuracy between an embodiment of the present disclosure and FedAvg;

[0013] Figure 4 is a schematic diagram illustrating stage 1 of an embodiment ("DSCFL");

[0014] Figure 5 shows pseudocode for Algorithm 1, indicating DSCFL;

[0015] Figure 6A is a schematic diagram illustrating stage 2 of DSCFL with UpdateOne for sharing model updates;

[0016] Figure 6B shows Algorithm 2, indicating how users share update for a single cluster model;

[0017] Figure 7A is a schematic diagram illustrating stage 2 of DSCFL with UpdateAll for sharing model updates;

[0018] Figure 7B shows Algorithm 3, indicating how users share update for all cluster models;

[0019] Figures 8A to 8F show graphs of results of experiments to compare model accuracy;

[0020] Figure 9 is a table showing results of experiments;

[0021] Figures 10A and 10B are graphs comparing clustering accuracy;

[0022] Figure 11 is a flowchart of a computer-implemented method, performed by a user device, for training a plurality of image generation machine learning, ML, models;

[0023] Figure 12 is a schematic illustration of training a model according to an embodiment of the present disclosure;

[0024] Figure 13A is a schematic illustration of the operations prior to training a model according to an embodiment of the present disclosure;

[0025] Figure 13B is a schematic illustration of the training "receive all, send ID, receive one and train and send one" scheme;

[0026] Figure 14 is a schematic illustration of warmup training during an initialization according to an embodiment of the present disclosure;

[0027] Figure 15 is a schematic illustration of an alternative warmup training during an initialization according to an embodiment of the present disclosure;

[0028] Figure 16 is a flowchart of example operations of a computer-implemented method, performed by a server, for training a plurality of image generation machine learning, ML, models;

[0029] Figure 17 is a schematic illustration of noisy weights projection during an initialization according to an embodiment of the present disclosure;

[0030] Figure 18 is a schematic illustration of a noisy data projection during an initialization according to an embodiment of the present disclosure;

[0031] Figure 19 is a schematic diagram illustrating the manifold spanning scheme;

[0032] Figure 20 is an example use case of an embodiment of the present disclosure;

[0033] Figure 21 is an example use case of an embodiment of the present disclosure; and

[0034] Figure 22 is a schematic block diagram of a system.

[0035] In a first approach, there may be provided a computer-implemented method, performed by a user device, for training a plurality of image generation machine learning, ML, models, the method comprising: receiving, from a server, a plurality of sets of trainable parameters of the plurality of (global) image generation ML models which have been pre-trained; identifying, using local training data stored on the user device, a set of trainable parameters, among the received plurality of sets of trainable parameters, which is most suitable for training using the local data of the user device, by: calculating, using the local training data, a loss for each of the received plurality of sets of trainable parameters, by applying the set of trainable parameters to an image generation ML model stored on the user device (e.g. during an initialisation stage or otherwise); and selecting the set of trainable parameters which provide the stored image generation ML model with the lowest calculated loss; updating, using the local training data, the selected set of trainable parameters by training the stored image generation ML model configured with the selected set of trainable parameters; adding noise to the updated set of trainable parameters to generate a set of noisy trainable parameters; and transmitting the set of noisy trainable parameters to the server for aggregation with other sets of noisy trainable parameters, received from other user devices, corresponding to the selected set of trainable parameters.

[0036] Advantageously, an embodiment of the present disclosure may provide a method for training a plurality of global image generation ML models using federated learning, in a way that preserves user data privacy and prevents semantic and concept drift. User devices may receive a plurality of sets of trainable parameters for global image generation models, which are all generated using different distributions of data, and each user device may be able to pick the set of trainable parameters which has been trained using data that is similar to the user device's own local data. Thus, instead of using all user devices to train a single global model via federated learning, an embodiment of the present disclosure may advantageously enable user devices to locally train a global model by applying to the model a set of trainable parameters that is best matched to the user devices' own data. This means each resulting trained model has superior performance compared to a single global model, because each global model is being locally trained using data that is semantically and conceptually close to the data used to initially train the global model. Global models perform poorly when user data drifts drastically from the initial data used to train the global model. An embodiment of the present disclosure avoid this issue.

[0037] The user devices may transmit data to the server for aggregation, as per the standard federated learning process. However, an embodiment of the present disclosure are advantageous over standard federated learning because the data that is sent from the user devices to the server may be made noisy, to ensure that user privacy is maintained. This is explained in more detail below.

[0038] The plurality of image generation ML models may all contain frozen, non-trainable parameters and a set of trainable parameters. This ensures that the training performed by the user devices only impacts certain model parameters, thereby preventing issues such as catastrophic forgetting and ensuring that the updated model parameters can be easily aggregated by the server. Instead of providing the user devices with multiple models, an embodiment of the present disclosure may simply provide each user device with one initial image generation ML model and multiple sets of trainable parameters, which reduces the amount of data that needs to be transmitted from server to user devices. Each set of trainable parameters can be applied to the initial image generation ML model (e.g. by replacing the model's existing trainable parameters with that of the selected set), so that the initial image generation ML model can be trained locally on the user device using local training data stored on the user device.

[0039] As noted above, each user device may only train a single ML model (e.g., a single set of the received sets of trainable parameters) out of all the received models. The un-selected ML models (e.g., unselected sets of trainable parameters) may be deleted by the user device once it is determined that they are not required, should on-device storage be limited. The term model may be used interchangeably herein with the term set of trainable parameters, but it will be understood that it is the sets of trainable parameters that are being selected and then used for training on the user device, and which are aggregated by the server.

[0040] The term "local training data" is used herein to mean data that has been acquired by the user device and is stored on the user device in storage. For example, the local training data may be videos or images captured by the user device (e.g. smartphone). Since users may capture images or videos of scenes and objects they encounter in their everyday lives, the data captured by different users / user devices may vary based on geography, region, as well as the users' own interests and hobbies. The local training data may be labelled by an on-device model used for object and scene detection and classification. The labelling may preferably occur while the user device is not being used, e.g. when the user device is being charged and / or overnight, so that the labelling does not impact usage of hardware resources (e.g., processor, memory) when the user is actively using the user device. The labelling may be performed on-device to prevent any user data from being sent off-device, which could compromise user privacy. The labelled user data may be stored on the user device, and can then be used for federated learning.

[0041] The step of receiving a plurality of sets of trainable parameters may comprise receiving a plurality of sets of trainable parameters for image generation ML models that have each been pre-trained using data obtained from any one or more of: users in different geographical regions; users of different age groups; users of different genders; users with different interests or hobbies. For example, one image generation ML model may be pre-trained using user devices in India, while another image generation ML model may be pre-trained using user devices in South Korea, while another image generation ML model may be pre-trained using user devices in the United Kingdom. However, it will be understood that this is a non-limiting example. The data used to pre-train the image generation ML models may be acquired from users around the world and may not be in geographical groupings. For example, users that capture images of dogs may be located around the world and dog breeds may not be geographically-constrained, and so their data could be used to pre-train a single model. In contrast, users that capture images of classic cars or city scenes may also be around the world but the classic cars or city scenes they capture may be quite different. For example, a busy street in Rome looks very different to a busy street in Mumbai or Los Angeles.

[0042] As noted above, the user devices may receive all the sets of trainable parameters for different image generation ML models. There may be C image generation ML models on the server, and therefore C sets of trainable parameters. The user devices may select one model from the C models and only send the trainable parameters associated with the selected model back to the server. Alternatively, the user devices may select one model from the C models but may send the trainable parameters associated with all C models back to the server.

[0043] User device side: "Receive all, train one, send one"

[0044] In an embodiment, the operation of adding noise to the updated set of trainable parameters may comprise: calculating a difference between values of the updated set of trainable parameters after training and values of the selected set of trainable parameters before training; generating a difference vector based on the calculated difference; and adding noise to the generated difference vector to generate a noisy difference vector.

[0045] Then, the operation of transmitting the set of noisy trainable parameters may comprise: transmitting the noisy difference vector.

[0046] Furthermore, prior to updating the selected set of trainable parameters, the method may comprise: transmitting an identifier indicating the selected set of trainable parameters to the server, to inform the server which set of trainable parameters of the plurality of sets of trainable parameters is being trained by the user device; and receiving, from the server, information indicating a number of user devices training the same selected set of trainable parameters for use in adding distributed differential privacy noise to the updated set of trainable parameters.

[0047] Preferably, the identifier may need to be noised (e.g., made noisy) before it is transmitted to the server because the selected set of trainable parameters are derived from user data and therefore reveals information about user data (in the same way that the trained parameters reveal information about user data). For example, with four global models (e.g., four sets of trainable parameters) and some trained parameters from a user, a malicious third party can learn more about user data if the third party knows exactly which model these trained parameters relate to. In contrast, if the third party is not certain about which model these trained parameters relate to, then it is harder to extrapolate information about user data.

[0048] Therefore, preferably, transmitting an identifier may comprise: adding noise to the identifier using local differential privacy; and transmitting the noised identifier. Local differential privacy, LDP, is a technology that enables differential privacy (DP) guarantees without requiring any trust in a central authority. Here, the user devices may add noise locally to the identifier of the selected set of trainable parameters. Noise introduced in this way can be very large, which can impact the usability of the noisy data. However, on the whole, the impact of noise on the identifiers may be minimal. That is, some IDs may be incorrectly understood and therefore, the number of user devices training a particular set of trainable parameters may not be completely accurate, but will be approximately good enough.

[0049] User device side: "Receive all, train one, send all"

[0050] In this alternative case, adding noise to the updated set of trainable parameters may comprise: calculating a difference between values of the updated set of trainable parameters after training and values of the selected set of trainable parameters before training; generating a difference vector based on the calculated difference; generating a zero vector comprising a plurality of segments, wherein each of the plurality of segments corresponds to each of the plurality of sets of trainable parameters; replacing a segment in the zero vector that corresponds to the selected set of trainable parameters with the generated difference vector, to generate a modified vector; and adding noise to the modified vector to generate a noisy modified vector.

[0051] Then, the operation of transmitting the set of noisy trainable parameters may comprise: transmitting the generated noisy modified vector. Thus, since the user device only trains the parameters for one of the ML models, the server may effectively receive trained values for the trainable parameters of the selected ML model, and zero values for the trainable parameters of the remaining, unselected ML models.

[0052] Whatever data the user devices send back to the server, the user devices may add noise to the data prior to transmission. In an embodiment, the operation of adding noise to the difference vector or the modified vector may comprise adding noise using distributed differential privacy, DDP.

[0053] Differential privacy, DP, is a mathematical framework that sets a limit on an individual user device's influence on the outcome of a computation, such as their influence on the trainable parameters of the ML model. This is achieved by bounding the contribution of any individual user device, and adding noise during the training process to produce a probability distribution over the locally-trained models. DP may come with a parameter that quantifies how much the distribution could change when adding or removing the trainable parameters received from an individual user device (the smaller the change, the better).

[0054] DDP is a technology that enables DP guarantees with respect to an honest-but-curious server that is aggregating the trainable parameters received from the user devices. DDP may work by first making each participating user device clip the trainable parameters it has trained and to add noise to the trained trainable parameters, and then aggregating, at the server, the noisy clipped trainable parameters. In this way, the server may only receive clipped updates which it then uses to generate a noisy sum of clipped updates, and each user device's own trainable parameters may not be discernible, thereby ensuring user data privacy. For additional security, the method may comprise: generating a secure mask; appending the secure mask to the noisy difference vector or the noisy modified vector; and transmitting the masked noisy difference vector or masked noisy modified vector to the server. The secure masks may be designed such that they cancel out once aggregated, such that the server is unable to view individual vectors when the vectors are aggregated.

[0055] Local differential privacy, LDP, is a technology that enables DP guarantees without requiring any trust in a central authority. Here, the user devices may add noise locally to the trainable parameters each user device has trained. Noise introduced in this way can be very large, which can impact the usability of the noisy data.

[0056] Generally speaking, the operation of training a set of trainable parameters and transmitting the set of noisy trainable parameters may be repeated for a predefined number of federated learning rounds. That is, a pre-defined number, R, of total federated learning rounds may be performed. In the first round, the server may send the initial plurality of sets of trainable parameters for the image generation ML models. In all subsequent rounds, the server may send the plurality of sets of trainable parameters that have been updated using the trainable parameters received from user devices at the end of the previous round. The user device may need to re-perform the identifying operation to select the relevant set of trainable parameters to train, as it is possible that the most relevant parameters for the user device to train changes in each federated learning round.

[0057] As noted above, the plurality of sets of trainable parameters of the plurality of image generation ML models that are received from the server may be pre-trained using different datasets. In an embodiment, the plurality of image generation ML models may be generated in response to determining how distributed user data may be, which impacts how many separate ML models may be needed to ensure a wide variety of user devices are able to participate in federated learning and to prevent data drift. This process may occur during an 'initialisation' stage of the overall method, which is now described.

[0058] Initialisation

[0059] In an example of the initialisation stage, user devices may train an initial (original) general image generation ML model using their local training data. Again, the user device may only train trainable parameters of the initial general image generation ML model. These trainable parameters of the initial general image generation ML model may be returned to the server, to enable the server to determine how many different ML models need to be generated to suitably capture the distributions of data across user devices.

[0060] Thus, in an example, prior to receiving the plurality of sets of trainable parameters, the method may comprise: receiving an initial general image generation ML model from the server; updating a first set of trainable parameters of the initial general image generation ML model using the local training data stored on the user device, to obtain a trained initial image generation ML model; adding noise to the updated first set of trainable parameters to generate a first set of noisy trainable parameters; and transmitting the first set of noisy trainable parameters to the server for use in generating the plurality of sets of trainable parameters.

[0061] In an example, adding noise to the updated first set of noisy trainable parameters may comprise: identifying a first pre-defined number of top (e.g., largest) positive parameters and a second pre-defined number of top (e.g., smallest) negative parameters from the updated first set of trainable parameters; retaining values of the identified positive parameters and the identified negative parameters and setting all the parameters of the updated first set of trainable parameters other than the identified positive parameters and the identified negative parameters to zero; and adding noise to the updated first set of trainable parameters. This process may be useful to limit the effect of local differential privacy noise, which, as noted above, can be very large. Thus, if the updated trainable parameters of the image generation ML model are, for example, [5, 0.2, -5, 0.2], the top-L positive and negative parameters may be retained and the rest set to zero, i.g. [5, 0, -5, 0].

[0062] In an example of the initialisation stage, user devices may train parameters of an initial (original) general image generation ML model using their local training data. The server may also provide a plurality of input images X. The user device may process the input images X with the trained image generation ML model to generate outputs Y. The outputs may be, for example, image classifications. Here, the outputs Y may be returned to the server, to enable the server to determine how many different ML models need to be generated to suitably capture the distributions of data across user devices.

[0063] Thus, in an example, prior to receiving the plurality of sets of trainable parameters, the method may comprise: receiving an initial general image generation ML model from the server together with a plurality of input images; updating a first set of trainable parameters of the initial general image generation ML model using the local training data stored on the user device, to obtain a trained initial image generation ML model; processing the plurality of input images with the trained initial image generation ML model to generate a plurality of outputs; adding noise to the plurality of outputs to generate noisy outputs; and transmitting the noisy outputs to the server.

[0064] In an example, adding noise to the first set of trainable parameters or the plurality of outputs may comprise adding noise using local differential privacy. In an embodiment, noise may be added locally, using any suitable LDP technique. In this way, user data privacy is maintained during the initialisation stage. Here, LDP may be sufficient because during initialisation, noise is being applied to a very small number of parameters. This ensures that the noisy parameters used for initialisation are not too noisy. Note, DDP and secure aggregation may not be used for initialisation since for initialisation it is necessary to collect individual parameter information from each user device, in order to generate the multiple image generation ML models. DDP and secure aggregation (described above) may only allow the server to utilise an aggregated vector.

[0065] The initialisation stage may be repeated for a predefined number of rounds. For example, if a pre-defined number, R, of total federated learning rounds need to be performed, then the initialisation stage may comprise a pre-defined number R0 of rounds, where R0 << R. In the first round, the server may send the initial image generation ML model to the plurality of user devices. In all subsequent rounds, the server may send the same trainable parameters back to the user devices for training using further local training data. This may enable more datapoints (e.g. more trainable parameters) to be obtained for use in the clustering process (described below). Thus, the operations of receiving an initial image generation ML model, training a first set of trainable parameters, and transmitting the first set of noisy trainable parameters or noisy outputs may be repeated for a predefined number of federated learning rounds.

[0066] In a second approach, there may be provided a user device for training a plurality of image generation machine learning, ML, models, the user device comprising: memory storing instructions, local training data and an image generation ML model; and at least one processor operatively coupled to memory and comprising processing circuitry, wherein the at least one processor individually or collectively executes the instructions to cause the user device to : receive, from a server, a plurality of sets of trainable parameters of the plurality of image generation ML models that have been pre-trained; identify, using the local training data, a set of trainable parameters, among the received plurality of sets of trainable parameters, which is most suitable for training using the local data of the user device, by: calculating, using the local training data, a loss for each of the received plurality of sets of trainable parameters by applying the set of trainable parameters to the stored image generation ML model; and selecting the set of trainable parameters which provide the stored image generation ML model with the lowest calculated loss; update, using the local training data, the selected set of trainable parameters by training the stored image generation ML model configured with the selected set of trainable parameters; add noise to the updated set of trainable parameters to generate a set of noisy trainable parameters; and transmit the set of noisy trainable parameters to the server for aggregation with other sets of noisy trainable parameters, received from other user devices, corresponding the selected set of trainable parameters, to thereby generate a plurality of image generation ML models.

[0067] The features described above with respect to the first approach apply equally to the second approach and therefore, for the sake of conciseness, are not repeated.

[0068] The user device may comprise at least one processor and memory storing instructions that, when executed by the at least one processor individually or collectively, cause the user device to perform the methods described herein.

[0069] The user device may be a smart device. The user device may be a smartphone. A smartphone is an example of a smart device. The user device may be a smart appliance. A smart appliance is another example of a smart device. An example of a smart appliance may include a smart television (TV), a smart fridge, a smart oven, a smart vacuum cleaner, a smart robotic device, a smart lawn mower, and so on. More generally, the user device may be a constrained-resource device, but which has the minimum hardware capabilities to train a merged ML model as described above. The user device may be any one of: a smartphone, tablet, laptop, computer or computing device, virtual assistant device, a vehicle, an autonomous vehicle, a robot or robotic device, a robotic assistant, image capture system or device, an augmented reality system or device, a virtual reality system or device, a gaming system, an Internet of Things device, or a smart consumer device (such as a smart fridge, smart vacuum cleaner, smart lawn mower, smart oven). It will be understood that this is a non-exhaustive and non-limiting list of example devices.

[0070] So far, a method performed by the user device during the federated learning process have been described. A method performed by the server are now explained.

[0071] In a third approach, there may be provided a computer-implemented method, performed by a server, for training a plurality of image generation machine learning, ML, models, the method comprising: transmitting, to a plurality of user devices, a plurality of sets of trainable parameters of the plurality of image generation ML models that have been pre-trained; receiving, from at least a portion of the user devices, sets of noisy trainable parameters; for each set of trainable parameters: aggregating the received sets of noisy trainable parameters that correspond to the set of trainable parameters, to generate an updated set of trainable parameters for an image generation ML model corresponding to the set of trainable parameters; and transmitting the updated set of trainable parameters to at least a portion of the plurality of user devices.

[0072] It will be understood that there is no limit on the number of user devices which may communicate with the server. It is desirable to have a large number of user devices participate in federated learning, to reduce bias and to increase the diversity of the data used to train the models.

[0073] Many of the features described above with reference to the first approach also apply to the third approach and are not repeated.

[0074] The operation of transmitting a plurality of sets of trainable parameters of image generation ML models may comprise transmitting a plurality of sets of trainable parameters that have each been pre-trained using the local data, as explained above.

[0075] Server side: "Receive all, train one, send one"

[0076] The operation of receiving sets of noisy trainable parameters may comprise: receiving, from a first user device among the plurality of user devices, a noisy difference vector that indicates a difference between values of a first set of trainable parameters among the plurality of sets of trainable parameters prior to training by the first user device and values of the first set of trainable parameters after training by the first user device.

[0077] In an embodiment, the method may further comprise: receiving, from each of the user devices, a noised identifier identifying a set of trainable parameters selected by the user device for training; and determining, based on the noised identifiers, a total number of user devices that are training each set of trainable parameters.

[0078] In an embodiment, the user devices may transmit identifiers, IDs, to the server so that the server is able to determine the number of user devices that are training each of the C image generation ML models. This may be useful because the server can then provide specific user devices with the data needed for distributed differential privacy. For example, the server can provide the user devices with the maximum number of samples for the specific image generation ML model they are training.

[0079] Server side: "Receive all, train one, send all"

[0080] In an embodiment, the operation of receiving sets of noisy trainable parameters may comprise: receiving, from a first user device among the plurality of user devices, a noisy modified vector comprising a plurality of segments corresponding to the plurality of sets of trainable parameters, wherein a first segment in the noisy modified vector corresponds to a difference between values of a first set of trainable parameters among the plurality of sets of trainable parameters prior to training by the first user device and values of a first set of trainable parameters after training by the first user device, and segments in the noisy vector other than the first segment are zero values. In an embodiment, each user device may select one model from the C image generation ML models but may send the trainable parameters for all C image generation ML models back to the server. However, since each user device only trains one of the ML models, the server may effectively receive, from each user device, trained values for the trainable parameters of the selected ML model, and zero values for the trainable parameters of the remaining, unselected ML models.

[0081] Initialisation

[0082] As noted above, the plurality of image generation ML models that are received from the server may be pre-trained using different datasets. In an embodiment, the plurality of image generation ML models may be generated in response to determining how distributed user data may be, which impacts how many separate image generation ML models may be needed to ensure a wide variety of user devices are able to participate in federated learning and to prevent data drift. This process occurs during an 'initialisation' stage of the overall method, which is now described.

[0083] In an example of the initialisation stage, user devices may train an initial (original) image generation ML model using their local training data. Parameters of the initial image generation ML model may be returned to the server, to enable the server to determine how many different image generation ML models need to be generated to suitably capture the distributions of data across user devices.

[0084] Thus, in an example, prior to transmitting the plurality of sets of trainable parameters, the method may comprise: transmitting an initial image generation ML model to the plurality of user devices; receiving first sets of noisy trainable parameters from at least a portion of the user devices; and generating the plurality of image generation ML models by training the initial image generation ML model using the received first sets of noisy trainable parameters.

[0085] Here, generating the plurality of image generation ML models may comprise: clustering the received first sets of noisy trainable parameters into a plurality of clusters; and generating for each cluster, an image generation ML model by training the initial image generation ML model using the sets of noisy trainable parameters in the cluster, thereby generating the plurality of image generation ML models.

[0086] In an example of the initialisation stage, user devices may train an initial (original) image generation ML model using their local training data. The server may also provide a plurality of input images X to the user device. The user device may process the input images X with the trained image generation ML model to generate outputs Y. The outputs may be, for example, image classifications. Here, the outputs Y may be returned to the server, to enable the server to determine how many different ML models need to be generated to suitably capture the distributions of data across user devices.

[0087] Thus, in an example, prior to transmitting the plurality of sets of trainable parameters, the method may comprise: transmitting an initial image generation ML model to the plurality of user devices, together with a plurality of input images; receiving noisy outputs corresponding to the plurality of input images from at least a portion of the user devices; and generating the plurality of image generation ML models by training the initial image generation ML model using the plurality of input images and the received noisy outputs.

[0088] Here, generating the plurality of image generation ML models may comprise: clustering the received noisy outputs into a plurality of clusters; and for each cluster: calculating a centroid corresponding to the cluster; obtaining training data based on the centroid of the cluster, the training data comprising pairs of data items, each pair of data items comprising an image and a classification for the image; and training the initial image generation ML model using the obtained training data corresponding to the cluster. The training data for each cluster may be obtained from user devices or other data sources, where the data may share similarity traits (e.g. geographic region, age, gender, interests / hobbies).

[0089] In an example, a clustering algorithm may be used to understand the distribution of the data received from the user devices and thereby determine how many image generation ML models need to be generated for the federated learning process to maximise the number of user devices that can participate in the federated learning (while also avoiding data drift problems). In an example, the number of clusters may be equal to the number of image generation ML models C.

[0090] The clustering may comprise using an unsupervised clustering algorithm, such as k-means clustering.

[0091] An embodiment may be to construct a manifold from the data received from the user devices. A manifold is a topological space. Thus, generating the plurality of image generation ML models may comprise: constructing a manifold using the received first set of noisy trainable parameters or the received noisy outputs; calculating a plurality of maximally separated points that span the manifold; obtaining training data for each calculated point in the manifold, the training data comprising pairs of data items, each pair of data items comprising an image and a classification for the image; and training the initial image generation ML model using the obtained training data for each calculated point, thereby generating the plurality of image generation ML models.

[0092] In a fourth approach, there may be provided a server, for training a plurality of image generation machine learning, ML, models, the server comprising: memory storing instructions and a plurality of image generation ML models that have each been pre-trained using different datasets; and at least one processor operatively coupled to the memory and comprising processing circuitry, wherein the at least one processor individually or collectively executes the instructions to cause the server to: transmitting, to a plurality of user devices, a plurality of sets of trainable parameters of the plurality of the plurality of image generation ML models; receiving, from at least a portion of the user devices, sets of noisy trainable parameters; for each set of trainable parameters: aggregating the received sets of noisy trainable parameters that correspond to the set of trainable parameters, to generate an updated set of trainable parameters for an image generation ML model corresponding to the set of trainable parameters; and transmitting the updated set of trainable parameters to at least a portion of the plurality of user devices.

[0093] The features described above with respect to the third approach apply equally to the fourth approach and therefore, for the sake of conciseness, are not repeated.

[0094] The server may comprise at least one processor and memory storing instructions that, when executed by the at least one processor individually or collectively, cause the server to perform the methods described herein.

[0095] In a related approach, there may be provided a computer-readable storage medium comprising instructions which, when executed by at least one processor, causes the at least one processor to carry out any of the methods described herein.

[0096] As will be appreciated by one skilled in the art, an embodiment of the present disclosure may be embodied as a system, method or computer program product. Accordingly, an embodiment of the present disclosure may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects.

[0097] Furthermore, an embodiment of the present disclosure may take the form of a computer program product embodied in a computer readable medium having computer readable program code embodied thereon. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable medium may be, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.

[0098] Computer program code for carrying out operations of an embodiment of the present disclosure may be written in any combination of one or more programming languages, including object oriented programming languages and conventional procedural programming languages. Code components may be embodied as procedures, methods or the like, and may comprise sub-components which may take the form of instructions or sequences of instructions at any of the levels of abstraction, from the direct machine instructions of a native instruction set to high-level compiled or interpreted language constructs.

[0099] An embodiment of the present disclosure may also provide a non-transitory data carrier carrying code which, when implemented on at least one processor, causes the at least one processor to carry out any of the methods described herein.

[0100] An embodiment may further provide processor control code to implement the above-described methods, for example on a general purpose computer system or on a digital signal processor (DSP). An embodiment may also provide a carrier carrying processor control code to, when running, implement any of the above methods, in particular on a non-transitory data carrier. The code may be provided on a carrier such as a disk, a microprocessor, CD- or DVD-ROM, programmed memory such as non-volatile memory (e.g. Flash) or read-only memory (firmware), or on a data carrier such as an optical or electrical signal carrier. Code (and / or data) to implement embodiments of the techniques described herein may comprise source, object or executable code in a conventional programming language (interpreted or compiled) such as Python, C, or assembly code, code for setting up or controlling an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), or code for a hardware description language such as Verilog (RTM) or VHDL (Very high speed integrated circuit Hardware Description Language). As the skilled person will appreciate, such code and / or data may be distributed between a plurality of coupled components in communication with one another. An embodiment may comprise a controller which includes a microprocessor, working memory and program memory coupled to one or more of the components of the system.

[0101] It will also be clear to one of skill in the art that all or part of a logical method according to an embodiment of the present disclosure may suitably be embodied in a logic apparatus comprising logic elements to perform the operations of the above-described methods, and that such logic elements may comprise components such as logic gates in, for example a programmable logic array or application-specific integrated circuit. Such a logic arrangement may further be embodied in enabling elements for temporarily or permanently establishing logic structures in such an array or circuit using, for example, a virtual hardware descriptor language, which may be stored and transmitted using fixed or transmittable carrier media.

[0102] An embodiment may be realised in the form of a data carrier having functional data thereon, said functional data comprising functional computer data structures to, when loaded into a computer system or network and operated upon thereby, enable said computer system to perform all the steps of the above-described method.

[0103] The method described above may be wholly or partly performed on an apparatus, e.g. an electronic device, using a machine learning or artificial intelligence model. The model may be processed by an artificial intelligence-dedicated processor designed in a hardware structure specified for artificial intelligence model processing. The artificial intelligence model may be obtained by training. Here, "obtained by training" means that a predefined operation rule or artificial intelligence model configured to perform a desired feature (or purpose) is obtained by training a basic artificial intelligence model with multiple pieces of training data by a training algorithm. The artificial intelligence model may include a plurality of neural network layers. Each of the plurality of neural network layers may include a plurality of weight values and may perform neural network computation by computation between a result of computation by a previous layer and the plurality of weight values.

[0104] As mentioned above, an embodiment may be implemented using an AI model. A function associated with AI may be performed through the non-volatile memory, the volatile memory, and the processor. The processor may include one or a plurality of processors. At this time, one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an AI-dedicated processor such as a neural processing unit (NPU). The one or a plurality of processors may control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model may be provided through training or learning. Here, being provided through learning means that, by applying a learning algorithm to a plurality of learning data, a predefined operating rule or AI model of a desired characteristic is made. The learning may be performed in a device itself in which AI according to an embodiment is performed, and / o may be implemented through a separate server / system.

[0105] The AI model may consist of a plurality of neural network layers. Each layer may have a plurality of weight values, and may perform a layer operation through calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks may include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.

[0106] The learning algorithm may be a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning algorithms may include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0107] Broadly speaking, an embodiment of the present disclosure may provide a method for training a plurality of image generation ML models using federated learning, in a way that preserves user data privacy and prevents semantic and concept drift. Thus, instead of using all user devices to train a single model via federated learning, an embodiment of the present disclosure advantageously enable user devices to train a model by applying to the model a set of trainable parameters that is best matched to the user devices' own data. This means each resulting trained model has superior performance compared to a single model, because each model is being trained using data that is semantically and conceptually close to the data used to initially train the model. Furthermore, each user device can be provided with a trained image generation ML model which is specific to their own data.

[0108] Federated learning (FL) is a distributed learning paradigm where users collaboratively train a model without revealing their data to a server. Despite this, clients' contributions can still leak sensitive information. Differential privacy (DP) addresses this by providing formal privacy guarantees achieved by adding randomness to clients' contributions. Local DP (LDP) requires clients' contributions to be privatized at the edge and therefore protects against a malicious server. Central DP (CDP) adds noise at the server and requires the server to be fully trusted. The higher privacy benefits of LDP may lead to more noisy contributions and may drop in model performance. Distributed DP (DDP) offers a decentralized approach to privacy preservation, matching the privacy guarantees of CDP at the same noise level without relying on a trusted server. This is achieved by leveraging cryptographic secure multiparty computation (SMPC) protocols to ensure that the server only has access to the aggregated contributions and is prevented from viewing individual updates.

[0109] With FL training, the user data may be naturally heterogeneous or non-independent and identically distributed (non-IID). Heterogeneity in training data has been shown to negatively impact model convergence and performance, and remains a key challenge with FL. Furthermore, adding DP to FL (DP-FL) may amplify the negative effect of data heterogeneity. Figure 3 is a graph comparing accuracy between an embodiment of the present disclosure ("ours") and FedAvg on the Fashion-MNIST dataset, with and without DP. In an embodiment, the method may narrow the gap between ideal (IID) and realistic (non-IID) data distributions, making it practical for widespread use. FedAvg is the de facto baseline for FL with / without DP and under IID / non-IID data scenarios. As described previously, FedAvg under non-IID and DP condition is the worst performing configuration.

[0110] Clustered FL was recently proposed to address data heterogeneity challenges in FL. Unlike certain related approaches such as FedProx and SCAFFOLD, which train a single global model across all clients, clustered FL works by grouping clients with similar data distributions into clusters, thereby making data within each cluster more IID. These algorithms also concurrently learn to assign clients to their corresponding cluster during the training process. Other related works have integrated DP constraints into FL algorithms to tackle the data heterogeneity challenge. However, studies that adapt clustered FL to DP either do not provide strong privacy guarantees or require high LDP noise in model training resulting in drastically reduced model utility.

[0111] An embodiment of the present disclosure provide a novel approach to ensuring privacy in clustered FL by combining SMPC-based secure vector sum with DP. Amethod according to an embodiment achieves superior utility compared to existing DP-FL methods for heterogeneous data, while providing full privacy protection against an untrusted (honest-but-curious) server and effectively addressing the cold-start problem inherent in such scenarios. These advancements position clustered FL at the forefront of practical adoption. Some contributions of the disclosure include:

[0112] A novel algorithm, Differentially private and Secure Clustered Federated Learning (DSCFL), to tackle the data heterogeneity challenge with practical federated learning while formally guaranteeing users' privacy against an untrusted (honest-but-curious) server.

[0113] A solution to initialize cluster models without proxy user data requirements at the server or random restarts costing additional privacy budgets.

[0114] A novel adaptation of a general secure sum protocol to enable communication-efficient training of cluster models.

[0115] Theoretical analysis of the privacy leakage and extensive experimental evaluation across several datasets, with results consistently outperforming SOTA DP-FL algorithms on heterogeneous data by up to 3.3%.

[0116] Preliminaries

[0117] Federated Learning (FL): At the start of each communication round t, a global model may be provided by the server and a randomly sampled user (e.g., user device) set may be constructed. Each user may train the model locally to obtain and may share the model difference back to the server. The server may aggregate the updates , and then proceed to the next round.

[0118] Differential Privacy: Differential privacy (DP) may provide a formal definition to quantify the amount of private information an algorithm leaks regarding its input data. DP may be defined as follows:

[0119] Definition (Differential Privacy) A randomized mechanism may satisfy -DP if for any pair of adjacent datasets D and D', and any subset of outputs , the following exists

[0120] (1)

[0121] where is known as the privacy budget.

[0122] LoRA: Low-Rank Adaptation (LoRA) may be a parameter-efficient fine-tuning (PEFT) method for transformer-based pre-trained models. Instead of training the entire weight matrix, it may freeze the pre-trained weights and introduce new trainable low-rank decomposition matrices B and A as

[0123] (2)

[0124] where is initialized to zeros, follows random Gaussian initialization and . This effectively reduces the number of trainable parameters by an order of .

[0125] Clustered Federated Learning: Clustered FL algorithms may group clients with similar distributions together. Thereby, clients within the same cluster suffer less from data heterogeneity and train via FL more effectively.

[0126] Clustered FL methods may use techniques like cosine similarity, empirical loss and singular value decomposition to assign clients to clusters and train models. Clients may respond with their cluster ID and trained model.

[0127] Some related works add DP to clustered FL. However, some of the related works use LDP to privatize user updates throughout the entire training process which significantly degrades the model's utility, making it infeasible for training large models. Others provide sample-level privacy instead of user-level privacy which is a weaker form of privacy protection than the latter. It has also been pointed out that the loss-based clustering breaks the sample-level privacy guarantees by leaking more information than allowed. In contrast, some related works add secure aggregation to clustered FL without incorporating DP constraints, relying on a fully trusted server.

[0128] Differentially Private and Secure Clustered Federated Learning (DSCFL)

[0129] Overview: An embodiment of the present disclosure, referred to herein as DSCFL, may comprise two stages: (1) Federated Cluster Model Initialization and (2) Federated Clustered Model Training. In (1), an embodiment of the present disclosure may privately initialize cluster models from user updates; and in (2) an embodiment of the present disclosure may perform cluster identification and model training in a federated setting and privately update global models.

[0130] In an embodiment, contemporary pre-trained transformer-based models may be used. These models may be particularly suitable for cross-device FL since the models can be pre-trained on large amount of server data and later fine-tuned on user data for downstream tasks in a distributed manner. The models may also provide other benefits such as improved robustness on heterogeneous data.

[0131] Workflow Stage 1: Figure 4 is a schematic diagram illustrating stage 1 of an embodiemnt (e.g., DSCFL), and Figure 5 shows pseudocode for Algorithm 1, indicating DSCFL. Global models of all clusters (cluster models) may be initialized without requiring the server to hold data with similar distribution as user data. As illustrated in Figure 4, the server may send pre-trained weights to sampled clients(e.g., a plurality of user devices), which train on local data, transform model updates as explained below, and apply LDP before sending them back. The server may then apply each received update to the pre-trained weights and run a clustering algorithm such as k-means on the updated weights, yielding C number of cluster centers . Global models of all clusters, which are called cluster models, may then be initialized to . For example, the cluster models may include image generation ML models.

[0132] Stage 2: In an embodiment, the server may send cluster models to each sampled client at the start of each communication round. Clients may perform cluster identification based on training loss and train the selected cluster model on local data. Let be the parameters of a pre-trained model and be the samples held by client . The empirical loss F associated with client k may be defined as in Equation (3) where is the loss function associated with sample z.

[0133] (3) (4)

[0134] Each sampled client may then either share the updates to the trained cluster model with the server using a combination of LDP, DDP and secure sum as in Algorithm 2 or update to all cluster models using DDP and secure sum as in Algorithm 3.

[0135] After receiving the aggregated noisy updates through secure sum, for a number of communication rounds, the server may normalize the magnitude of each aggregate to the one with the smallest norm as in Equation (4) with and where . The normalization may be done to stabilize early training of cluster models.

[0136] Federated Cluster Model Initialization: Clustered FL algorithms may either require data with similar distribution as user data to be available at server side for a good initialization, or random restarts which costs additional privacy under DP-FL. In contrast, the cluster model initialization approach illustrated as part of stage 1 of DSCFL may relax both requirements.

[0137] After training the pre-trained weights on local data, each client may share the two largest positive values and two largest negative values from the model updates both in terms of absolute value. It may be assumed that there are at least two positive values and two negative ones in the model updates. All the other values in the model updates may be converted to zeros. LDP may be then applied to the update vector using the Gaussian mechanism with clipping and noise addition. LDP may be used here because the server needs access to individual vectors for running the clustering algorithm (e.g. k-means). However, with LDP, the magnitude of the added noise grows with the number of parameters updated. Therefore, only four parameters may be updated in this step so that the clustering can still produce meaningful results for initializing cluster models. This filtering process is denoted as FILTER in Algorithm 3.

[0138] However, updates to four parameters may provide limited information. It is therefore necessary to only train a small number of parameters for this operation to maximize the ratio of the number of updates received by the server to the total number of trainable parameters. Hence, LoRA may be applied with r=1 to the value projection matrix of the last attention layer denoted by and both the down-projection matrix and the output layer are frozen. This may leave the up-projection matrix as the only trainable part in stage 1 with the number of trainable parameters being equal to the hidden dimension d of the model. The last attention layer may be chosen as it tends to contain less general knowledge and is more likely to catch information regarding the underlying clustering structure. It has been shown that applying LoRA to value projection matrix with r=2 leads to similar results as applying LoRA to query / value projection matrices with r=1. This is why LoRA is chosen to apply to value projection matrix instead of query. Additionally, since the up-projection matrix is initialized with zeros, may be frozen and only may be trained. After initializing cluster models in stages 1, LoRA may be applied with r=1 to query / value projection matrices of all attention layers with random Gaussian initialization for A and zeros for B except for the previously obtained .

[0139] User Updates Sharing: After initializing cluster models, users may start training the selected cluster model on local data and sharing updates to the trained model with the server. The server may then update global cluster models based on the received model updates. Here, two methods may be proposed for sharing model updates, namely, UpdateOne and UpdateAll.

[0140] Method 1 ("UpdateOne"): This is also referred to herein as "receive all, train one, send one". Figure 6A is a schematic diagram illustrating stage 2 of DSCFL with UpdateOne for sharing model updates, and Figure 6B shows Algorithm 2, indicating how users share update for a single cluster model (e.g., single set of trainable parameters of the cluster model).

[0141] As illustrated in Figure 6A and detailed in Figure 6B / Algorithm 2, clients (e.g., user devices) may convert target cluster IDs to one-hot vectors, apply LDP, and participate in secure sum (explained below). The server may then receive noisy IDs, return back the number of participation to each cluster for DDP noise calculation, and receive the aggregated model updates without accessing individual contributions.

[0142] Method 2 ("UpdateAll"): This is also referred to herein as "receive all, train one, send all". Figure 7A is a schematic diagram illustrating stage 2 of DSCFL with UpdateAll for sharing model updates, and Figure 7B shows Algorithm 3, indicating how users share update for all cluster models (e.g., all sets of trainable parameters available for training).

[0143] As illustrated in Figure 7A and detailed in Figure 7B / Algorithm 3, each sampled client may construct a vector indicating updates to all cluster models with previously obtained updates for the trained cluster model and zero updates for the rest. Clients may then apply DDP to model updates(e.g., the vector) and share them with the server via secure sum. In this way, the server may receive the aggregated updates to all cluster models at once. Compared to UpdateOne, UpdateAll removes the necessity to apply LDP to cluster IDs, leading to less noisy clustering. However, the client communication cost (disregarding secure sum) may increase from to with n being the number of trainable parameters. The server learning rate may be set to C instead of 1 as in UpdateOne to compensate for the zero updates in the aggregated updates.

[0144] Secure Sum with Modified Server-Client Communication for UpdateOne: After local training, clients may participate in secure sum so that the server receives the aggregated updates without individual access. However, naively applying secure vector sum such as those based on SMPC to the UpdateOne sharing method would introduce additional rounds of server-client communication which would harm system efficiency and robustness. Below, the modifications to a general SMPC protocol are explained, which may be necessary for an embodiment of the present disclodure without any additional rounds of server-client communication.

[0145] Advertise Public Keys: SMPC protocols may require secure communication channels that can be realised via public-key cryptography. This may be modified such that each client applies LDP to the target cluster ID and sends the noisy ID together with the generated public key to the server which are used to populate client-cluster membership. Separate secure sum instances (one per cluster ) with a threshold value and participants may then be initialized at the server. The server and clients may abort the instance for the current training round if at any stage the number of surviving clients for cluster c falls below .

[0146] Encrypt, Share & Decrypt: The subsequent operations may remain unchanged within each instance as in existing SMPC protocols. Since clients may drop at any point during the training round, client will not be able to infer the precise number of surviving clients in at the server. Therefore, client may use the threshold to compute the DDP noise, perturb model updates, encrypt the model updates using the shared key and upload the encrypted noist updates to the server. Upon receiving the encrypted noisy updates, the server may aggregate the encrypted noist updates for each cluster c and then query clients to decrypt the resulting aggregated updates. Since the server is only able to recover the final vector sum if at least clients remain, privacy is preserved when . On the other hand, when , the added noise may be no less than that required to satisfy DP.

[0147] Privacy Analysis

[0148] In the following privacy analysis, it is shown that the present proposed algorithm satisfies user-level -DP against an honest-but-curious server. The analysis can be readily extended to methods that address the finite precision and modular arithmetic challenges associated with combining DDP with secure vector sum protocols.

[0149] Lemma 1. For the sharing method UpdateOne, the cluster identity i shared by a client may satisfy -DP.

[0150] Proof. DSCFL may split the total privacy budget into for cluster identities and for model updates with under sequential composition.

[0151] Definition (Sequential Composition) For any and , if are mechanisms each satisfying -DP, then their composition ( ) may satisfy ( )-DP.

[0152] A moments accountant may be used with Renyi Differential Privacy (RDP) for a tight composition bound. For any and , a randomized mechanism M satisfies -RDP if for all neighboring datasets D and D', the following holds

[0153] (5)

[0154] The Gaussian mechanism may then be used with noise for achieving RDP as in Equation (6) where is used to compute noise for privatizing the cluster ID, is the noise multiplier for IDs calculated by moments accountant for privacy parameters, such as privacy budget and sampling rate, and is the clipping threshold for IDs. Since clients convert cluster IDs to one-hot vectors before clipping, these vectors may have a fixed norm of 1. Therefore, may be set as in Equation (7).

[0155] (6) (7)

[0156] Lemma 2. For the sharing method UpdateOne, the model updates shared by a client may satisfy -DP.

[0157] Proof. For model updates, moments accountant may be used again with RDP for a tight composition bound as in Equation (5).

[0158] The Gaussian mechanism may be used with noise for achieving RDP together with privacy amplification via sampling as in Equation (8) where is used to compute noise for cluster c, is the noise multiplier for privatizing model updates, S is the clipping threshold for model updates and is the number of clients contributing to cluster c at communication round t with cohort size .

[0159] Since DDP is used with secure sum for privatizing updates, the equivalent for client-side noising may be defined in Equation (9).

[0160] The noise level of a larger cohort size may be simulated with a smaller one as in Equation (10).

[0161] (8) (9) (10)

[0162] Lemma 3. For the sharing method UpdateAll, the model updates shared by a client may satisfy -DP.

[0163] Proof. Following the analysis above for Lemma 2, since each sampled client shares updates to all cluster models in UpdateAll, client-side noise may be as in Equation (11) with being used to compute noise for privatizing the model updates and z being the noise multiplier calculated by moments accountants.

[0164] Lemma 4. The model updates shared by a client for cluster model initialization may satisfy -DP.

[0165] Proof. The moments accountant may be used again with RDP and Gaussian mechanism to privatize the shared client updates in stage 1 for cluster model initialization as in Equation (12) where is used to compute noise, for UpdateOne, for UpdateAll and is the clipping threshold.

[0166] (11) (12)

[0167] Experiments

[0168] Experimental settings: A privacy budget of is used, which is commonly used in existing works, and . A cohort size of 10k is simulated with a smaller cohort size as in Equation (10) to achieve a more realistic signal-to-noise ratio which represents industry scale more closely. For the UpdateOne sharing method, the privacy budget is split in half and each half is allocated to IDs and model updates respectively using sequential composition. The mean and standard deviation is reported over three runs.

[0169] Rotated CIFAR-10 (C=2), rotated FMNIST (C=4) and FEMNIST (C=2) are used for the experiments. The first two are generated by applying the same rotation (0, 180 degrees for CIFAR-10 and 0, 90, 180, 270 degrees for FMNIST) to all images of a client. The total number of clients is set to 5,000 for CIFAR-10 / FMNIST and 2,840 for FEMNIST. For IFCA, C is set to the same value as the method according to an embodiment. It has been shown that IFCA could be combined with personalized FL to improve performance even further, but this is not considered here. For FedProx, is set adaptively which was shown to work well in the original paper. DP-SCAFFOLD is excluded from the present experiments since it is designed for the cross-silo setting, and the focus is only on the more challenging cross-device setting in this work. Client dropouts are not considered in the experiments. ViT-small of 22 million parameters is used as global models, which was considered a large model in previous works.

[0170] Comparing with related DP-FL methods: Figures 8A to 8F and 9 show results for different methods with and without DP. Specifically, Figures 8A to 8F show graphs of results of experiments to compare model accuracy on CIFAR-10, FMNIST and FEMNIST for , while Figure 9 is a table showing results of experiments. It can be seen that DSCFL consistently outperforms both DP-FedAvg and DP-FedProx on all three datasets with either sharing method. For , the method according to an embodiment shows up to 2.4%, 2.9% and 3.2% improvements over DP-FedAvg and DP-FedProx on rotated CIFAR-10, rotated FMNIST and FEMNIST, respectively. Similarly, the method according to an embodiment outperforms SOTA DP-FL algorithms by up to 2.2%, 2.9% and 3.3% for on these datasets. Overall, DSCFL shows 2.8% improvements over DP-FedAvg and DP-FedProx for both and . This indicates that DSCFL is more effective than related DP-FL methods in terms of learning on heterogeneous data. Additionally, Figures 10A and 10B, which are graphs comparing clustering accuracy, show that the method according to an embodiment achieves high clustering accuracy. In contrast, incorporating LDP into IFCA throughout the entire training process has a detrimental effect on the clustering structure.

[0171] Thus, an embodiment of the present disclosure provides a novel algorithm DSCFL to tackle the challenge of data heterogeneity in DP-FL. DSCFL enables training of cluster models while providing DP guarantees against an honest-but-curious server using a combination of local DP, distributed DP and secure vector sum. Empirical results show that DSCFL outperforms SOTA DP-FL algorithms by up to 3.3% for the same privacy guarantees.

[0172] Figure 11 is a flowchart of a computer-implemented method, performed by a user device, for training a plurality of image generation machine learning, ML, models. In other words, Figure 11 shows the client / user side of the overall DSCFL method illustrated in Figures 4 and 5. The method may comprise: receiving, from a server, a plurality of sets of trainable parameters of the plurality of image generation ML models which have been pre-trained (operation S100); identifying, using local training data stored on the user device, a set of trainable parameters, among the received plurality of sets of trainable parameters, which is most suitable for training using the local data of the user device (operation S102), by: calculating, using the local training data, a loss for each of the received plurality of sets of trainable parameters, by applying the set of trainable parameters to an image generation ML model stored on the user device; and selecting the set of trainable parameters which provide the stored image generation ML model with the lowest calculated loss; updating, using the local training data, the selected set of trainable parameters by training the stored image generation ML model configured with the selected set of trainable parameters (operation S104); adding noise to the updated set of trainable parameters to generate a set of noisy trainable parameters (operation S106); and transmitting the set of noisy trainable parameters to the server for aggregation with other sets of noisy trainable parameters, received from other user devices, corresponding to the selected set of trainable parameters (operation S108).

[0173] The operation (S100) of receiving a plurality of sets of trainable parameters may comprise receiving a plurality of sets of trainable parameters of the plurality of image generation ML models that have each been pre-trained using data obtained from any one or more of: users in different geographical regions; users of different age groups; users of different genders; users with different interests or hobbies. For example, one model may be pre-trained using user devices in India, while another model may be pre-trained using user devices in South Korea, while another model may be pre-trained using user devices in the United Kingdom. However, it will be understood that this is a non-limiting example. The data used to pre-train the ML models may be acquired from users around the world and may not be in geographical groupings. For example, users that capture images of dogs may be located around the world and dog breeds may not be geographically-constrained, and so their data could be used to pre-train a single model. In contrast, users that capture images of classic cars or city scenes may also be around the world but the classic cars or city scenes they capture may be quite different. For example, a busy street in Rome looks very different to a busy street in Mumbai or Los Angeles.

[0174] As noted above, the user devices may receive all the sets of trainable parameters for different image generation ML models. There may be C image generation ML models on the server, and therefore C sets of trainable parameters. The user devices may select one image generation ML model from the C image generation ML models and only send the trainable parameters for the selected model back to the server. Alternatively, the user devices may select one image generation ML model from the C image generation ML models but may send the trainable parameters for all C models back to the server.

[0175] User device side: "Receive all, train one, send all" / "UpdateAll"

[0176] Figure 12 is a schematic illustration of training a image generation ML model according to an embodiment of the present disclosure. In an example, the training may be termed "receive all, select one and train and send all". In a first operation (1), C image generation ML models may be transferred from the server to a random subset of clients. In a second operation (2), each client may locally calculate the loss on the entire dataset for all C image generation ML models and choose the image generation ML model with the lowest loss for its training. In a third operation (3), each client may then train its chosen image generation ML model for E local epochs to produce pseudo-gradients Gi. In a fourth operation (4), each client may create a long zeros vector of size L*C where L is the number of trainable parameters and C is the number of clusters (or the number of the image generation ML models). In a fifth operation (5), DDP noise may be added to this zero vector to ensure CDP privacy after aggregation. In a sixth operation (6), a segment (e.g., slice) of L parameters (corresponding to the chosen client index) may get appended the final (pseudo-)gradients after training. In a seventh operation (7), K long vectors may be sent back to server and in operation (8), K long vectors may securely be aggregated at the server. The server may only be able to access an aggregated form which is CDP protected to ensure privacy. The process of operations (1) to (8) may be repeated for r rounds or until convergence. Note that each L-segment (e.g., L-slice) after aggregation has ~ [0,K] updates, since DDP noise is added based on K, it is still possible to guarantee privacy.

[0177] Therefore, in an embodiment, adding noise (operation S106 in Figure 11) to the updated set of trainable parameters may comprise: calculating a difference between values of the updated set of trainable parameters after the training and values of the selected set of trainable parameters before the training; generating a difference vector based on the calculated difference; generating a zero vector comprising a plurality of segments, wherein each of the plurality of segments corresponds to each of the plurality of sets of trainable parameters; replacing a segment in the zero vector that corresponds to the selected set of trainable parameters with the generated difference vector, to generate a modified vector; and adding noise to the modified vector to generate a noisy modified vector.

[0178] Then, the operation (S108 in Figure 11) of transmitting the set of noisy trainable parameters may comprise: transmitting the generated noisy modified vector. Thus, since the user device only trains the parameters for one of the image generation ML models, the server may effectively receive trained values for the trainable parameters of the selected image generation ML model, and zero values for the trainable parameters of the remaining, unselected image generation ML models.

[0179] User device side: "Receive all, train one, send one" / "UpdateOne"

[0180] Figure 13A is a schematic illustration of the first part prior to training a model according to an embodiment of the present disclosure. In an example, the training may be termed "receive all, send ID, receive one and train and send one". Before the operations of training take place, there may be a first stage of getting client IDs (same as advertise keys w / SecAgg). In a first operation (1) of the training, C image generation ML models (M1.. ) may be transferred from the server to a random subset of clients (K0). In a second operation (2), each client may locally calculate the loss on the entire dataset for all C models and choose the model with the lowest loss for its cluster ID (IDmin). In a third operation (3), each client may apply LDP (local DP) so that there is a probabilistic error on the cluster ID values thus user privacy is protected (ID) even with a malicious server. In operation (4), a subset of these (non-dropout) clients ( ) may return their noisy cluster IDs to the server. In operation (5), the noisy cluster IDs collected on the server may be mapped to determine the number of users per cluster (S1.. ) and initialize DDP noising parameters.

[0181] Figure 13B is a schematic illustration of the training "receive all, send ID, receive one and train and send one" and may occur in a second stage after the operations of Figure 13A. The operations of Figure 13B may be considered to represent a second stage of getting the trained models (same as MaskedInputCollection v / SecAgg). At operation (6), each of the non-dropout clients may be sent the maximum number of samples (or, the number of users) in their cluster (for DDP noise calculation) (S1.. ). In operation (7), since each client already has all the cluster models, the non-dropout clients may choose the minimum loss model Mminand now train these using their own data. In operation (8), the DDP noise may be calculated on the client-side based the maximum number of samples. In operation (9), the DDP may be added to the trained model (plus security masking for SecAgg). In operation (10), a subset of these (non-dropout) clients ( ) may return their noisy models to the server. In operation (11), the server may aggregate the results per cluster (because noisy cluster ID is known) to get updated centroids (and masks can be recovered via SecAgg protocol). As indicated by operation (12), this process may be repeated for R rounds or until convergence.

[0182] In an embodiment, the operation (S106 in Figure 11) of adding noise to the updated set of trainable parameters may comprise: calculating a difference between values of the updated set of trainable parameters after training and values of the selected set of trainable parameters before training; generating a difference vector based on the calculated difference; and adding noise to the generated difference vector to generate a noisy difference vector.

[0183] Then, the operation (S108 in Figure 11) of transmitting the set of noisy trainable parameters may comprise: transmitting the noisy difference vector.

[0184] Furthermore, and as shown in Figure 13A, prior to training the selected set of trainable parameters, the method may comprise: transmitting an identifier indicating the selected set of trainable parameters to the server, to inform the server which set of trainable parameters of the plurality of sets of trainable parameters is being trained by the user device; and receiving, from the server, information indicating a number of user devices training the same selected set of trainable parameters for use in adding distributed differential privacy noise to the updated set of trainable parameters.

[0185] In an embodiment, transmitting the identifier may comprise: adding noise to the identifier using local differential privacy; and transmitting the noised identifier. Local differential privacy, LDP, is a technology that enables DP guarantees without requiring any trust in a central authority. The user devices may add noise locally to the identifier indicating the selected set of trainable parameters. Noise introduced in this way can be very large, which can impact the usability of the noisy data. However, on the whole, the impact of noise on the identifiers may be minimal. That is, some IDs may be incorrectly understood and therefore, the number of user devices training a particular set of trainable parameters may not be completely accurate, but will be approximately good enough (e.g., will provide sufficient statistical utility for subsequent processing).

[0186] As explained above, whether transmitting one set of trainable parameters or transmitting all sets of trainable parameters back to the server, the user devices may add noise to the data prior to transmission. Thus, the operation of adding noise to the difference vector or the modified vector may comprise adding noise using distributed differential privacy, DDP.

[0187] Generally speaking, the operations (S104 to S108) of updating the selected set of trainable parameters and transmitting the set of noisy trainable parameters may be repeated for a predefined number of federated learning rounds. For example, a pre-defined number, R, of total federated learning rounds may be performed. In the first round, the server may transmit the initial plurality of sets of trainable parameters of the plurality of image generation ML models to the user devices. In all subsequent rounds, the server may transmit, to the user devices, the plurality of sets of trainable parameters that have been updated using the trainable parameters received from the user devices at the end of the previous round. The user device may need to re-perform the identifying operation (e.g., S102 in Figure 11) to select the relevant set of trainable parameters to train, as it is possible that the most relevant parameters for the user device to train changes in each federated learning round.

[0188] As noted above, the plurality of sets of trainable parameters of the plurality of image generation ML models that are received from the server are pre-trained using different datasets. In an embodiment, the plurality of ML models may be generated in response to determining how distributed user data may be, which impacts how many separate ML models may be needed to ensure a wide variety of user devices are able to participate in federated learning and to prevent data drift. This process may occur during an 'initialisation' stage of the overall method, which is now described.

[0189] Initialisation

[0190] Figure 14 is a schematic illustration of warmup training during an initialization according to an embodiment of the present disclosure. The warmup training may comprise in operation (1) sending the initial model (e.g., an initial general image generation ML model, M0) to a large sample of clients number of clients per round K. The initial model may comprise both trainable and frozen variables to minimize loss of utility due to noise addition. Next in operation (2) the clients K0may train for E0local epochs ( standard local epochs E) on their own data and apply local differential privacy (LDP) to trainable parameters. The noisy trainable parameters may be then sent back to the server in operation (3). These operations of training and sending back parameters to the server can be repeated for R0warmup rounds ( where R is the total rounds of FL training). At the end of these loops, the server may have a number (R0x K0) of noisy client model updates with client-level privacy guarantees.

[0191] To limit the effect of LDP noise, the top-L positive and negative parameters per client may be retained and the parameters, other than the top-L positive and negative parameters, may be set to non-zero before adding LDP noise (e.g. [5, 0.2, -5, 0.2] -> [5, 0, -5, 0]). Such an operation should alleviate the adverse effects of LDP noise for the clustering process

[0192] In a first example of the initialisation stage, user devices may train an initial (original) ML model using their local training data. Again, the user device may only train trainable parameters of the initial ML model. These trainable parameters of the initial ML model may be returned to the server, to enable the server to determine how many different ML models need to be generated to suitably capture the distributions of data across user devices.

[0193] Thus, in an example shown in Figure 14, prior to receiving the plurality of sets of trainable parameters, the method may comprise: receiving an initial general image generation ML model from the server; updating a first set of trainable parameters of the initial general image generation ML model using the local training data stored on the user device, to obtain a trained initial image generation ML model; adding noise to the updated first set of trainable parameters to generate a first set of noisy trainable parameters; and transmitting the first set of noisy trainable parameters to the server for use in generating the plurality of sets of trainable parameters.

[0194] In an example, adding noise to the updated first set of noisy trainable parameters may comprise: identifying a first pre-defined number of largest positive parameters and a second pre-defined number of smallest negative parameters from the updated first set of trainable parameters; retaining values of the identified positive parameters and the identified negative parameters and setting all the parameters of the updated first set of trainable parameters other than the identified positive parameters and the identified negative parameters to zero; and adding noise to the updated first set of trainable parameters. This may be useful to limit the effect of local differential privacy noise, which, as noted above, can be very large. Thus, if the extracted parameters of the model are, for example, [5, 0.2, -5, 0.2], the top-L positive and negative parameters may be retained and the rest set to zero, e.g. [5, 0, -5, 0].

[0195] Figure 15 is a schematic illustration of an warmup training during an initialization according to an embodiment of the present disclosure. As in Figure 14, the warmup training may comprise in operation (1) sending the initial model (e.g., an initial general image generation ML model, M0) to a large sample of clients number of clients per round K. Additionally, in an example, public input data samples (X0) may be also sent from the server to the client. Next in operation (2) the clients K0may train for E0local epochs ( standard local epochs E) on their own data. Clients may then use their trained model and inputs X0to produce outputs Y1..K0. At operation (3), the outputs Y1..K0(one per client) may have LDP noise applied and may be then sent back to the server. These operations of training and sending back parameters to the server can be repeated for R0warmup rounds ( where R is the total rounds of FL training). At the end of these loops, the server may have a number (R0x K0) of noisy outputs Y1..K0(e.g. the multiple versions of the output based on models trained on client data given the input X0.) As in Figure 14, there may be client-level local DP guarantees. It is noted that X is a collection of input samples and Y is a collection of corresponding output predictions, e.g. X can be images and Y can be class indices.

[0196] Thus, in an example of the initialisation stage shown in Figure 15, user devices may train parameters of an initial (original) ML model using their local training data. The server may also provide a plurality of input images X to the user device. The user device may process the input images X with the trained ML model to generate outputs Y. The outputs may be, for example, image classifications. Here, the outputs Y may be returned to the server, to enable the server to determine how many different ML models need to be generated to suitably capture the distributions of data across user devices.

[0197] Thus, in an example, prior to receiving the plurality of sets of trainable parameters, the method may comprise: receiving an initial general image generation ML model from the server together with a plurality of input images; updating a first set of trainable parameters of the initial general image generation ML model using the local training data stored on the user device, to obtain a trained initial image generation ML model; processing the plurality of input images with the trained initial image generation ML model to generate a plurality of outputs; adding noise to the plurality of outputs to generate noisy outputs; and transmitting the noisy outputs to the server.

[0198] In an example, adding noise to the first set of trainable parameters or the plurality of outputs may comprise adding noise using local differential privacy to the first set of trainable parameters or the plurality of outputs. In an embodiment, noise may be added to the outputs locally, using any suitable LDP technique. In this way, user data privacy is maintained during the initialisation stage. Here, LDP may be sufficient because during initialisation, noise is being applied to a very small number of parameters. This ensures that the noisy parameters used for initialisation are not too noisy. Note, DDP and secure aggregation may not be used for initialisation since for initialisation it is necessary to collect individual parameter information from each user device, in order to generate the multiple ML models. DDP and secure aggregation (described above) may only allow the server to utilise an aggregated vector.

[0199] The initialisation stage may be repeated for a predefined number of rounds. For example, if a pre-defined number, R, of total federated learning rounds need to be performed, then the initialisation stage may comprise a pre-defined number R0of rounds, where R0<< R. In the first round, the server may send the initial image generation ML model to the user devices. In all subsequent rounds, the server may send the trainable parameters back to the user devices for training using further local training data. This may enable more datapoints (e.g. more trainable parameters) to be obtained for use in the clustering process (described below). Thus, the operations of receiving an initial image generation ML model, training a first set of trainable parameters, and transmitting the first set of noisy trainable parameters or noisy outputs may be repeated for a predefined number of federated learning rounds.

[0200] Figure 16 is a flowchart of example operations of a computer-implemented method, performed by a server, for training a plurality of image generation machine learning, ML, models. The method may comprise: transmitting, to a plurality of user devices, a plurality of sets of trainable parameters of the plurality of image generation ML models that have been pre-trained (S200); receiving, from at least a portion of the plurality of user devices, sets of noisy trainable parameters (S202); for each set of trainable parameters: aggregating the received sets of noisy trainable parameters that correspond to the set of trainable parameters, to generate an updated set of trainable parameters for an image generation ML model corresponding to the set of trainable parameters (S204); and transmitting the updated set of trainable parameters to the at least a portion of the plurality of user devices (S206).

[0201] Many of the features described above with reference to the method performed by a user device also apply to the method performed by the server and are not repeated for the sake of conciseness.

[0202] The operation (S200) of transmitting a plurality of sets of trainable parameters of the plurality of image generation ML models may comprise transmitting a plurality of sets of trainable parameters that have each been pre-trained using local data, as explained above.

[0203] Server side: "Receive all, train one, send one" / "UpdateOne"

[0204] The operation (S202) of receiving sets of noisy trainable parameters may comprise: receiving, from a first user device among the plurality of user devices, a noisy difference vector that indicates a difference between values of a first set of trainable parameters among the plurality of sets of trainable parameters prior to training by the first user device and values of the first set of trainable parameters after training by the first user device.

[0205] In an embodiment, the method may further comprise: receiving, from each of the plurality of user devices, a noised identifier indicating a set of trainable parameters selected by the user device for training; and determining, based on the noised identifiers, a total number of user devices that are training each of the plurality of set of trainable parameters.

[0206] That is, and as explained with reference to Figure 13A, the user devices may transmit ML model identifiers, IDs, to the server so that the server is able to determine the number of user devices that are training each of the C image generation ML models. This may be useful because the server can then provide specific user devices with the data needed for distributed differential privacy. For example, the server may provide the user devices with the maximum number of samples for the specific ML model they are training.

[0207] Server side: "Receive all, train one, send all" / "UpdateAll"

[0208] In an embodiment, the operation (S202) of receiving sets of noisy trainable parameters may comprise: receiving, from a first user device among the plurality of user devices, a noisy modified vector comprising a plurality of segments corresponding to the plurality of sets of trainable parameters, wherein a first segment in the noisy modified vector corresponds to a difference between values of a first set of trainable parameters among the plurality of sets of trainable parameters prior to training by the first user device and values of a first set of trainable parameters after training by the first user device, and segments in the noisy vector other than the first segment are zero values. In an embodiment, each user device may select one model from the C image generation ML models but may send the trainable parameters for all C image generation ML models back to the server. However, since each user device only trains one of the ML models, the server may effectively receive, from each user device, trained values for the trainable parameters of the selected ML model, and zero values for the trainable parameters of the remaining, unselected ML models.

[0209] Initialisation

[0210] As noted above, the plurality of image generation ML models that are received from the server may be pre-trained using different datasets. In an embodiment, the plurality of ML models may be generated in response to determining how distributed user data may be, which impacts how many separate ML models may be needed to ensure a wide variety of user devices are able to participate in federated learning and to prevent data drift. This process may occur during an 'initialisation' stage of the overall method, which is now described.

[0211] In an example of the initialisation stage, user devices may train an initial (original) ML model using their local training data. Parameters of the initial ML model may be returned to the server, to enable the server to determine how many different ML models need to be generated to suitably capture the distributions of data across user devices.

[0212] Thus, in an example as shown in Figure 12, prior to transmitting the sets of trainable parameters, the method may comprise: transmitting an initial image generation ML model to the plurality of user devices; receiving first sets of noisy trainable parameters from at least a portion of the user devices; and generating the plurality of image generation ML models by training the initial image generation ML model using the received first sets of noisy trainable parameters.

[0213] In an embodiment, generating the plurality of image generation ML models may comprise: clustering the received first sets of noisy trainable parameters into a plurality of clusters; and generating for each cluster, an image generation ML model by training the initial image generation ML model using the sets of noisy trainable parameters in the cluster, thereby generating the plurality of image generation ML models. This is shown in Figure 17, which is a schematic illustration of noisy weights projection during an initialization according to an embodiment of the present disclosure. The operations for Figure 17 may follow the operations from Figure 14. In a first operation (1), the number (R0x K0) of noisy client model updates (e.g., noisy trainable parameters) may be (optionally) combined with the frozen parameters and projected to an L-dimensional hyperspace (L: total or only trainable parameters). In a second operation (2), with the help of an unsupervised clustering algorithm (e.g. K-means) and an appropriate distance metric (e.g. Euclidean / L2), the updates may be clustered into groups. In a third operation (3) the centroid of each cluster (denoted by x) may be converted to initial cluster specific models, leading to C models for clustered federated learning (FL) to begin.

[0214] In an example of the initialisation stage, as shown in Figure 15, user devices may train an initial (original) ML model using their local training data. The server may also provide a plurality of input images X to the user devices. The user device may process the input images X with the trained ML model to generate outputs Y. The outputs may be, for example, image classifications. Here, the outputs Y may be returned to the server, to enable the server to determine how many different ML models need to be generated to suitably capture the distributions of data across user devices.

[0215] Thus, in an example, prior to transmitting the plurality of sets of trainable parameters, the method may comprise: transmitting an initial image generation ML model to the plurality of user devices, together with a plurality of input images; receiving noisy outputs corresponding to the plurality of input images from at least a portion of the user devices; and generating the plurality of image generation ML models by training the initial image generation ML model using the plurality of input images and the received noisy outputs.

[0216] In an embodiment, generating the plurality of image generation ML models may comprise: clustering the received noisy outputs into a plurality of clusters; and for each cluster: calculating a centroid corresponding to the cluster; obtaining training data based on the centroid of the cluster, the training data comprising pairs of data items, each pair of data items comprising an image and a classification for the image; and training the initial image generation ML model using the obtained training data corresponding to the cluster. The training data for each cluster may be obtained by user devices or from other data sources, where the data may share similarity traits (e.g. geographic region, age, gender, interests / hobbies). This is shown in Figure 18, which is a schematic illustration of a noisy data projection during an initialization according to an embodiment of the present disclosure. The operations for Figure 18 may follow the operations from Figure 15. In a first operation (1), the received of noisy versions (R0x K0) of the outputs Y may be projected to an L-dimensional hyperspace (L: output dimension). In a second operation (2), with the help of an unsupervised clustering algorithm (e.g. K-means) and an appropriate distance metric (e.g. Euclidean / L2), the noisy versions of the outputsmay be clustered into groups. In a third operation (3), the centroid of each cluster (denoted by Z) may be converted to initial cluster specific models, leading to C versions of Y. In a final operation (4), the server may fine-tune the initial model using sample pairs of (X0, Z) i {1, ..., C} leading to C image generation ML models for clustered FL to begin. During training a combination of original, kl-divergence and contrastive losses could be used.

[0217] In an example, a clustering algorithm may be used to understand the distribution of the data received from the user devices and thereby determine how many image generation ML models need to be generated for the federated learning process to maximise the number of user devices that can participate in the federated learning (while also avoiding data drift problems). In an embodiments, the number of clusters may be equal to the number of image generation ML models C.

[0218] The clustering may comprise using an unsupervised clustering algorithm, such as k-means clustering.

[0219] An embodiment is to construct a manifold from the data received from the user devices. A manifold is a topological space. Thus, generating the plurality of image generation ML models may comprise: constructing a manifold using the received first set of noisy trainable parameters or the received noisy outputs; calculating a plurality of maximally separated points that span the manifold; obtaining training data for each calculated point in the manifold, the training data comprising pairs of data items, each pair of data items comprising an image and a classification for the image; and training the initial image generation ML model using the obtained training data for each calculated point, thereby generating the plurality of image generation ML models. This is shown in Figure 19, which is a schematic diagram illustrating the manifold spanning scheme. The process may begin in operation (1) by obtaining either the first set of noisy trainable parameters or the received noisy outputs (e.g., "updates"), as described above. Operation (2) may comprise constructing a manifold spanning these updates in an embedding space. Operation (3) may comprise generating D maximally separated points spanning the manifold e.g. D=32 points, and converting this to a 1-hot vector index. Finally, operation (4) may comprise creating D initial models based on these points and sending these models to clients.

[0220] As an embodiment to the training schemes described so far, an average model scheme may also be used for the training. This may include having C centroid models on the server, one for each cluster. For any given round, a C_avg model may be created by averaging the centroids either directly on the server or via averaging of all C models on the client. For any given round, C models (+1 if C_avg on the server) may be sent to the clients. Each client may train for b warmup mini-batches (or e epochs) using C_avg as the starting point and compare the distance in hyperspace between C_trained and C_1, C_2, .., C_C models. It is possible to make use of similarity and distance metrics e.g. Euclidian, manhattan, cosine, Canberra. Then the closest model may be chosen as client's cluster. The closest model C_closest from C_1, C_2, ..., C_C may be chosen as starting point and further trained on the client-side as usual. All models are sent back to the server are noised with DDP noise. The disadvantage to DPFedAvg is C* the noise norm because C* the models are sent back to the server, and all but one of these are untrained.

[0221] Some use cases of the present disclosure are now described, to illustrate the advantages of the present disclosure.

[0222] Figure 20 is an example use case of the present disclosure. The left hand side of Figure 20 shows how a single existing image generation ML model is trained using very diverse data and, as a result, is not able to generate an image of a bird that a particular user was expecting. The right hand side shows the present techniques, where multiple image generation ML models have been trained using specific related sets of data, such that each model is able to generate more specific images that are tailored to what a user is interested in.

[0223] Figure 21 is an example use case of the present disclosure, e.g. a drawing assistant. The left hand side of Figure 21 shows how a user who scrolls through their recent images may notice their picture of a garden felt very basic. The user decides to use Drawing Assist to sketch an image of a bird. An on-device image generation ML model generates an image of a dove. However, doves are not found in user's region, and thus the picture looks unrealistic and the user is dissatisfied with the results. In contrast, the right hand side shows an embodiment of the present disclosure, where an on-device image generation ML model that is personalised to user's region using an embodiment generates an image of a parrot. Parrots are found in user's region, and thus the picture looks realistic and the user is happy with the results.

[0224] An example use case of the present disclosure is for a smart fridge, which is able to generate recipe ideas / suggestions for a user based on the food / ingredients in the fridge. When using existing models to do so, the user may be provided with recipe suggestions that are not from the user's region or based on the user's preferences (e.g. vegetarian food). In contrast, an embodiment of the present disclosure provide the user with more personalised recipe suggestions that are based on the user's regions or other preferences. The user is therefore happy with the results.

[0225] An example use case of the present disclosure is for wallpaper or screensaver generation, which may be used on smartphones or smart TVs. When using existing models to generate a user-prompted wallpaper or screensaver (e.g. "let's use trees for wallpaper"), the generated wallpaper or screensaver image may include trees that are not specific to the user's geographic region / location, and the user may be disappointed with the results. In contrast, an embodiment of the present disclosure generate a wallpaper or screensaver image that includes trees that are specific to the user's geographic region / location. The user is therefore happy with the results.

[0226] An example use case is personalized image segmentation for the application named "object eraser". Currently, image segmentation uses a single model for all users around the world, or alternatively, a user manually selects which model they want to use (e.g. they make an explicit choice) or alternatively, the user automatically gets a personalized model but user privacy is not retained. An embodiment of the present disclosure described above mean that based on user data it is possible to determine the best model that will enable local image segmentation and thus may take care of regional specific semantic shifts. In other words, an embodiment of the present disclosur, dynamically determine the best model for user. An embodiment of the present disclosure adapt cluster models based on changing environment. The cluster centroids can act as models that fully adapt to changing environment e.g. newer regions emerging. Full user privacy is maintained in these scenarios, as it is not possible to know which user uses which models for these scenarios and it is not possible to know the demographic of each user individually.

[0227] Object eraser may use an automated boundary selection for object removal. For example, there may be an image which includes a taxi and the user is seeking to remove the taxi from the image. It will be appreciated that a yellow cab of New York is very different from a rickshaw in India but both are taxis. In other words, in a real-world scenario, taxis look different across different regions and the model has been trained to detect yellow cabs but not rickshaws. It will be appreciated that using related techniques, automated boundary selection will work well for images containing yellow cabs but not so well for images with rickshaws.

[0228] A use case is personalised smart reply. In an embodiment, personalized smart reply models may be based on clustering for what language or languages are used by users. As explained above, based on user data it is possible to determine the best model and give user the optimal smart reply experience. By contrast, the related techniques mean using either a single model for all users around the world or the user manually selecting which model they want to use (i.e. an explicit choice). With these techniques, it is possible to dynamically determine best mode for user, adapt cluster models based on changing environment and maintain user privacy in these scenarios.

[0229] A use case is personalised speech to text models. In an embodiment, personalized speech to text models may be separated based on accents, rare words. using user data. Currently, personalised speech to text uses a single model for all users around the world, or alternatively, a user manually selects which model they want to use (e.g. they make an explicit choice) or alternatively, the user automatically gets a personalized model but user privacy is not retained. With these techniques, it is possible to dynamically determine best mode for user, adapt cluster models based on changing environment with cluster centroids acting as models that can fully adapt to changing environment e.g. newer regions emerging and maintain full user privacy in these scenarios, it is not possible to know which user uses which models for these scenarios and it is not possible to know the demographic of each user individually.

[0230] A use case is personalised generative AI models, e.g. to generate wallpaper, or stories. In an embodiment, personalized image and music generation models may be separated based on user data to provide an optimal experience. For example meta data (e.g. styles / subjects) from generated images may be used. Currently, image and music generation models use a single model for all users around the world, or alternatively, a user manually selects which model / LoRA weights they want to use (e.g. they make an explicit choice). With these techniques, it is possible to dynamically determine best mode for user, adapt cluster models based on changing environment with cluster centroids acting as LoRAs that can fully adapt to changing environment e.g. newer regions emerging and maintain full user privacy in these scenarios, it is not possible to tell which user uses which models for these scenarios.

[0231] A use case is personalised image embedding for gallery search. In an embodiment, personalized image-text encoder models (CLIP) for many use cases e.g. Gallery Search, tagging, may be separated based on user data to provide an optimal experience by taking care of region specific semantic shifts. Currently, such models use a single model for all users around the world, or alternatively, a user manually selects which model weights they want to use (i.e. they make an explicit choice). With these techniques, it is possible to dynamically determine best mode for user, adapt cluster models based on changing environment with cluster centroids acting as models that can fully adapt to changing environment e.g. newer regions emerging and maintain full user privacy in these scenarios, it is not possible to know which user uses which models for these scenarios.

[0232] When generating images, specialized models can be used to generate images of a specific content / subject or images in a specific style. Moreover, there are model merging strategies which can be used to combine the models to generate images of specific content / subject in a specific style. For example, when generating wallpaper, e.g. for a phone, a user can generate an image of a personal pet in a particular style (e.g. the style of a particular artist). Similarly, the techniques described herein can be used to generate a personalized screensaver for a digital TV. In the context of generative editing, an embodiment of the present disclosure can be used to provide models which replace selected concepts in existing images with regional specific content and / or styles. Similarly, in a target application such as Portrait Studio or drawing assistants such as sketch to image, starting from a prompt, an embodiment of the present disclosure can be used to provide models which generate images with regional specific content and / or styles.

[0233] An embodiment of the present disclosure can be used to generate music for gallery stories, video editor, ring-tone, alarm or other applications. By using these techniques, the user will enjoy limitless generative AI outputs and customisable options used as background music with images and as ringtones / alarm tones etc. There may be new and unique features - for example generate music with user's everyday images / videos / text. There may be unlimited music generation which avoids problems and cost associated to music licensing and there may be community based personalization for example personalizing based on a group. A user provides text / image inputs, the service generates the music and the user may give preferences and / or feedback for the generated music which indicates whether the music is accepted or should be regenerated based on the feedback.

[0234] Thus, although the present disclosure have been described in relation to generating images, it will be understood that there are many other use cases, such as image / video editing, audio or music generation, and so on. In essence, although the present disclosure are described in relation to image generation, it will be understood that the present disclosure equally apply to other generative tasks such as music generation, recipe suggestions, and so on.

[0235] Figure 22 is a schematic block diagram of a system for implementing the methods described above. The system may comprise a server 1000 which may be a single server or collection of servers (e.g. the cloud). The server 1000 may comprise at least one processor 1002 coupled to memory 1004. The at least one processor 1002 may comprise one or more of: a microprocessor, a microcontroller, and an integrated circuit. The at least one processor 1002 may include one or more central processing units (CPUs) and / or one or more graphics processing units (GPUs). The memory 1004 may comprise volatile memory, such as random access memory (RAM), for use as temporary memory, and / or non-volatile memory such as Flash, read only memory (ROM), or electrically erasable programmable ROM (EEPROM), for storing data, programs, or instructions, for example.

[0236] A first AI model (ML model) 1006 and a second AI model (ML model) 1012 may be stored on the server 1000. It will be appreciated that two models is merely illustrative and many more models may be stored. The server 1000 may also comprise an input / output interface 1008 (or similar communication module) which connects the device to a database 1010. The database 1010 may comprise training dataset(s) for general training of the ML model(s), particularly the second ML model. The server may comprise a clustering module 1010 for clustering as described above. The server 1000 may also be coupled to a plurality of apparatus / user devices 1020 but for ease of reference, only one user device is shown.

[0237] The user device 1020 may also comprise similar standard components to the server 1000. The user device 1020 may comprise at least one processor 1022 coupled to memory 1024. The at least one processor 1022 may comprise one or more of: a microprocessor, a microcontroller, and an integrated circuit. The at least one processor 1022 may include one or more central processing units (CPUs) and / or one or more graphics processing units (GPUs). The memory 1024 may comprise volatile memory, such as random-access memory (RAM), for use as temporary memory, and / or non-volatile memory such as Flash, read only memory (ROM), or electrically erasable programmable ROM (EEPROM), for storing data, programs, or instructions, for example. An ML model 1026 may be stored on the electronic device 1020. There may also be an input / output interface 1028 which connects the user device 1020 to the server 1000.

[0238] Those skilled in the art will appreciate that while the foregoing has described what is considered to be the best mode and where appropriate other modes of performing present disclosure, the present disclosure should not be limited to the specific configurations and methods disclosed in this description of the preferred embodiment. Those skilled in the art will recognise that present techniques have a broad range of applications, and that the embodiments may take a wide range of modifications without departing from any inventive concept as defined in the appended claims.

[0239] In an embodiment, a computer-implemented method, performed by a user device(1020), for training a plurality of image generation machine learning(ML) models may be provided. In an embodiment, the method may include receiving, from a server(1000), a plurality of sets of trainable parameters of the plurality of image generation ML models which have been pre-trained. In an embodiment, the method may include identifying, using local training data stored on the user device(1020), a set of trainable parameters, among the received plurality of sets of trainable parameters, to be updated using the local data of the user device(1020), by: calculating, using the local training data, a loss for each of the received plurality of sets of trainable parameters, by applying the set of trainable parameters to an image generation ML model stored on the user device(1020); and selecting the set of trainable parameters which provide the stored image generation ML model with the lowest calculated loss. In an embodiment, the method may include updating, using the local training data, the selected set of trainable parameters by training the stored image generation ML model configured with the selected set of trainable parameters. In an embodiment, the method may include adding noise to the updated set of trainable parameters to generate a set of noisy trainable parameters. In an embodiment, the method may include transmitting the set of noisy trainable parameters to the server(1000) for aggregation with other sets of noisy trainable parameters, received from other user devices, corresponding to the selected set of trainable parameters.

[0240] In an embodiment, the method may include calculating a difference between values of the updated set of trainable parameters after the training and values of the selected set of trainable parameters before the training. In an embodiment, the method may include generating a difference vector based on the calculated difference. In an embodiment, the method may include adding noise to the generated difference vector to generate a noisy difference vector. In an embodiment, the method may include transmitting the noisy difference vector.

[0241] In an embodiment, the method may include transmitting an identifier indicating the selected set of trainable parameters to the server(1000). In an embodiment, the method may include receiving, from the server(1000), information indicating a number of user devices training the same selected set of trainable parameters for use in adding distributed differential privacy noise to the updated set of trainable parameters.

[0242] In an embodiment, the method may include adding noise to the identifier using local differential privacy. In an embodiment, the method may include transmitting the noised identifier.

[0243] In an embodiment, the method may include calculating a difference between values of the updated set of trainable parameters after the training and values of the selected set of trainable parameters before the training. In an embodiment, the method may include generating a difference vector based on the calculated difference. In an embodiment, the method may include generating a zero vector comprising a plurality of segments, wherein each of the plurality of segments corresponds to each of the plurality of sets of trainable parameters. In an embodiment, the method may include replacing a segment in the zero vector that corresponds to the selected set of trainable parameters with the generated difference vector, to generate a modified vector. In an embodiment, the method may include adding noise to the modified vector to generate a noisy modified vector. In an embodiment, the method may include transmitting the generated noisy modified vector.

[0244] In an embodiment, the method may include adding noise using distributed differential privacy to the difference vector or the modified vector.

[0245] In an embodiment, the method may include receiving an initial general image generation ML model from the server(1000). In an embodiment, the method may include updating a first set of trainable parameters of the initial general image generation ML model using the local training data stored on the user device(1020), to obtain a trained initial image generation ML model. In an embodiment, the method may include adding noise to the updated first set of trainable parameters to generate a first set of noisy trainable parameters. In an embodiment, the method may include transmitting the first set of noisy trainable parameters to the server(1000) for use in generating the plurality of sets of trainable parameters.

[0246] In an embodiment, the method may include identifying a first pre-defined number of largest positive parameters and a second pre-defined number of smallest negative parameters from the updated first set of trainable parameters. In an embodiment, the method may include retaining values of the identified positive parameters and the identified negative parameters and setting all the parameters of the updated first set of trainable parameters other than the identified positive parameters and the identified negative parameters to zero. In an embodiment, the method may include adding noise to the updated first set of trainable parameters.

[0247] In an embodiment, the method may include receiving an initial general image generation ML model from the server(1000) together with a plurality of input images. In an embodiment, the method may include updating a first set of trainable parameters of the initial general image generation ML model using the local training data stored on the user device(1020), to obtain a trained initial image generation ML model. In an embodiment, the method may include processing the plurality of input images with the trained initial image generation ML model to generate a plurality of outputs. In an embodiment, the method may include adding noise to the plurality of outputs to generate noisy outputs. In an embodiment, the method may include transmitting the noisy outputs to the server(1000).

[0248] In an embodiment, the method may include adding noise using local differential privacy to the first set of trainable parameters or the plurality of outputs.

[0249] In an embodiment, a user device(1020) for training a plurality of image generation machine learning(ML) models may be provided. In an embodiment, the user device(1020) may comprise memory(1024) storing instructions, local training data, and an image generation ML model and at least one processor(1022) operatively coupled to the memory(1024) and comprising processing circuitry. In an embodiment, the at least one processor(1022) may individually or collectively execute the instructions to cause the user device(1020) to receive, from a server(1000), a plurality of sets of trainable parameters of the plurality of image generation ML models that have been pre-trained. In an embodiment, the at least one processor(1022) may individually or collectively execute the instructions to cause the user device(1020) to identify, using the local training data, a set of trainable parameters, among the received plurality of sets of trainable parameters, to be updated using the local data of the user device(1020), by: calculating, using the local training data, a loss for each of the received plurality of sets of trainable parameters by applying the set of trainable parameters to the stored image generation ML model; and selecting the set of trainable parameters which provide the stored image generation ML model with the lowest calculated loss. In an embodiment, the at least one processor(1022) may individually or collectively execute the instructions to cause the user device(1020) to update, using the local training data, the selected set of trainable parameters by training the stored image generation ML model configured with the selected set of trainable parameters. In an embodiment, the at least one processor(1022) may individually or collectively execute the instructions to cause the user device(1020) to add noise to the updated set of trainable parameters to generate a set of noisy trainable parameters. In an embodiment, the at least one processor(1022) may individually or collectively execute the instructions to cause the user device(1020) to transmit the set of noisy trainable parameters to the server(1000) for aggregation with other sets of noisy trainable parameters, received from other user devices, corresponding to the selected set of trainable parameters.

[0250] In an embodiment, a computer-implemented method, performed by a server(1000), for training a plurality of image generation machine learning(ML) models may be provided. In an embodiment, the method may include transmitting, to a plurality of user devices, a plurality of sets of trainable parameters of the plurality of image generation ML models that have been pre-trained. In an embodiment, the method may include receiving, from at least a portion of the plurality of user devices, sets of noisy trainable parameters. In an embodiment, the method may include for each set of trainable parameters: aggregating the received sets of noisy trainable parameters that correspond to the set of trainable parameters, to generate an updated set of trainable parameters for an image generation ML model corresponding to the set of trainable parameters; and transmitting the updated set of trainable parameters to the at least a portion of the plurality of user devices.

[0251] In an embodiment, the method may include receiving, from a first user device among the plurality of user devices, a noisy difference vector that indicates a difference between values of a first set of trainable parameters among the plurality of sets of trainable parameters prior to training by the first user device and values of the first set of trainable parameters after training by the first user device.

[0252] In an embodiment, the method may include receiving, from each of the plurality of user devices, a noised identifier indicating a set of trainable parameters selected by the user device for training. In an embodiment, the method may include determining, based on the noised identifiers, a total number of user devices that are training each of the plurality of set of trainable parameters.

[0253] In an embodiment, the method may include receiving, from a first user device among the plurality of user devices, a noisy modified vector comprising a plurality of segments corresponding to the plurality of sets of trainable parameters. In an embodiment, a first segment in the noisy modified vector may correspond to a difference between values of a first set of trainable parameters among the plurality of sets of trainable parameters prior to training by the first user device and values of a first set of trainable parameters after training by the first user device, and segments in the noisy vector other than the first segment are zero values .

[0254] In an embodiment, the method may include transmitting an initial image generation ML model to the plurality of user devices. In an embodiment, the method may include receiving first sets of noisy trainable parameters from at least a portion of the user devices. In an embodiment, the method may include generating the plurality of image generation ML models by training the initial image generation ML model using the received first sets of noisy trainable parameters.

[0255] In an embodiment, the method may include clustering the received first sets of noisy trainable parameters into a plurality of clusters. In an embodiment, the method may include generating, for each cluster, an image generation ML model by training the initial image generation ML model using the sets of noisy trainable parameters in the cluster, thereby generating the plurality of image generation ML models.

[0256] In an embodiment, the method may include transmitting an initial image generation ML model to the plurality of user devices, together with a plurality of input images. In an embodiment, the method may include receiving noisy outputs corresponding to the plurality of input images from at least a portion of the user devices. In an embodiment, the method may include generating the plurality of image generation ML models by training the initial image generation ML model using the plurality of input images and the received noisy outputs.

[0257] In an embodiment, the method may include clustering the received noisy outputs into a plurality of clusters. In an embodiment, the method may include, for each cluster: calculating a centroid corresponding to the cluster; obtaining training data based on the centroid of the cluster, the training data comprising pairs of data items, each pair of data items comprising an image and a classification for the image; and training the initial image generation ML model using the obtained training data corresponding to the cluster.

[0258] In an embodiment, the method may include constructing a manifold using the received first set of noisy trainable parameters or the received noisy outputs. In an embodiment, the method may include calculating a plurality of maximally separated points that span the manifold. In an embodiment, the method may include obtaining training data for each calculated point in the manifold, the training data comprising pairs of data items, each pair of data items comprising an image and a classification for the image. In an embodiment, the method may include training the initial image generation ML model using the obtained training data for each calculated point, thereby generating the plurality of image generation ML models.

Claims

1.A computer-implemented method, performed by a user device(1020), for training a plurality of image generation machine learning(ML) models, the method comprising:receiving, from a server(1000), a plurality of sets of trainable parameters of the plurality of image generation ML models which have been pre-trained;identifying, using local training data stored on the user device(1020), a set of trainable parameters, among the received plurality of sets of trainable parameters, to be updated using the local data of the user device(1020), by:calculating, using the local training data, a loss for each of the received plurality of sets of trainable parameters, by applying the set of trainable parameters to an image generation ML model stored on the user device(1020); andselecting the set of trainable parameters which provide the stored image generation ML model with the lowest calculated loss;updating, using the local training data, the selected set of trainable parameters by training the stored image generation ML model configured with the selected set of trainable parameters;adding noise to the updated set of trainable parameters to generate a set of noisy trainable parameters; andtransmitting the set of noisy trainable parameters to the server(1000) for aggregation with other sets of noisy trainable parameters, received from other user devices, corresponding to the selected set of trainable parameters.2.The method as claimed in claim 1, wherein adding noise to the updated set of trainable parameters comprises:calculating a difference between values of the updated set of trainable parameters after the training and values of the selected set of trainable parameters before the training;generating a difference vector based on the calculated difference; andadding noise to the generated difference vector to generate a noisy difference vector,and wherein transmitting the set of noisy trainable parameters comprises:transmitting the noisy difference vector.3.The method as claimed in claim 1 or 2 wherein prior to updating the selected set of trainable parameters, the method comprises:transmitting an identifier indicating the selected set of trainable parameters to the server(1000); andreceiving, from the server(1000), information indicating a number of user devices training the same selected set of trainable parameters for use in adding distributed differential privacy noise to the updated set of trainable parameters.4.The method as claimed in claim 1 , wherein adding noise to the updated set of trainable parameters comprises:calculating a difference between values of the updated set of trainable parameters after the training and values of the selected set of trainable parameters before the training;generating a difference vector based on the calculated difference;generating a zero vector comprising a plurality of segments, wherein each of the plurality of segments corresponds to each of the plurality of sets of trainable parameters;replacing a segment in the zero vector that corresponds to the selected set of trainable parameters with the generated difference vector, to generate a modified vector; andadding noise to the modified vector to generate a noisy modified vector,and wherein transmitting the set of noisy trainable parameters comprises:transmitting the generated noisy modified vector.5.The method as claimed in claim 2 or 4, wherein adding noise to the difference vector or the modified vector comprises adding noise using distributed differential privacy to the difference vector or the modified vector.6.The method as claimed in any one of claims 1 to 5 wherein, prior to receiving the plurality of sets of trainable parameters, the method comprises:receiving an initial general image generation ML model from the server(1000);updating a first set of trainable parameters of the initial general image generation ML model using the local training data stored on the user device(1020), to obtain a trained initial image generation ML model;adding noise to the updated first set of trainable parameters to generate a first set of noisy trainable parameters; andtransmitting the first set of noisy trainable parameters to the server(1000) for use in generating the plurality of sets of trainable parameters.7.The method as claimed in claim 6 wherein adding noise to the updated first set of noisy trainable parameters comprises:identifying a first pre-defined number of largest positive parameters and a second pre-defined number of smallest negative parameters from the updated first set of trainable parameters;retaining values of the identified positive parameters and the identified negative parameters and setting all the parameters of the updated first set of trainable parameters other than the identified positive parameters and the identified negative parameters to zero; andadding noise to the updated first set of trainable parameters.8.The method as claimed in any one of claims 1 to 5 wherein, prior to receiving the plurality of sets of trainable parameters, the method comprises:receiving an initial general image generation ML model from the server(1000) together with a plurality of input images;updating a first set of trainable parameters of the initial general image generation ML model using the local training data stored on the user device(1020), to obtain a trained initial image generation ML model;processing the plurality of input images with the trained initial image generation ML model to generate a plurality of outputs;adding noise to the plurality of outputs to generate noisy outputs; andtransmitting the noisy outputs to the server(1000).9.A user device(1020) for training a plurality of image generation machine learning(ML) models, the user device(1020) comprising:memory(1024) storing instructions, local training data, and an image generation ML model; andat least one processor(1022) operatively coupled to the memory(1024) and comprising processing circuitry,wherein the at least one processor(1022) individually or collectively executes the instructions to cause the user device(1020) to:receive, from a server(1000), a plurality of sets of trainable parameters of the plurality of image generation ML models that have been pre-trained;identify, using the local training data, a set of trainable parameters, among the received plurality of sets of trainable parameters, to be updated using the local data of the user device(1020), by:calculating, using the local training data, a loss for each of the received plurality of sets of trainable parameters by applying the set of trainable parameters to the stored image generation ML model; andselecting the set of trainable parameters which provide the stored image generation ML model with the lowest calculated loss;update, using the local training data, the selected set of trainable parameters by training the stored image generation ML model configured with the selected set of trainable parameters;add noise to the updated set of trainable parameters to generate a set of noisy trainable parameters; andtransmit the set of noisy trainable parameters to the server(1000) for aggregation with other sets of noisy trainable parameters, received from other user devices, corresponding the selected set of trainable parameters.10.A computer-implemented method, performed by a server(1000), for training a plurality of image generation machine learning(ML) models, the method comprising:transmitting, to a plurality of user devices, a plurality of sets of trainable parameters of the plurality of image generation ML models that have been pre-trained;receiving, from at least a portion of the plurality of user devices, sets of noisy trainable parameters; andfor each set of trainable parameters:aggregating the received sets of noisy trainable parameters that correspond to the set of trainable parameters, to generate an updated set of trainable parameters for an image generation ML model corresponding to the set of trainable parameters; andtransmitting the updated set of trainable parameters to the at least a portion of the plurality of user devices.11.The method as claimed in claim 10, wherein receiving sets of noisy trainable parameters comprises:receiving, from a first user device among the plurality of user devices, a noisy difference vector that indicates a difference between values of a first set of trainable parameters among the plurality of sets of trainable parameters prior to training by the first user device and values of the first set of trainable parameters after training by the first user device.12.The method as claimed in claims 10 or 11 further comprising:receiving, from each of the plurality of user devices, a noised identifier indicating a set of trainable parameters selected by the user device for training; anddetermining, based on the noised identifiers, a total number of user devices that are training each of the plurality of set of trainable parameters.13.The method as claimed in claim 10 wherein receiving sets of noisy trainable parameters comprises:receiving, from a first user device among the plurality of user devices, a noisy modified vector comprising a plurality of segments corresponding to the plurality of sets of trainable parameters,wherein a first segment in the noisy modified vector corresponds to a difference between values of a first set of trainable parameters among the plurality of sets of trainable parameters prior to training by the first user device and values of a first set of trainable parameters after training by the first user device, and segments in the noisy vector other than the first segment are zero values.14.The method as claimed in any one of claims 10 to 13 wherein prior to transmitting the plurality of sets of trainable parameters, the method comprises:transmitting an initial image generation ML model to the plurality of user devices;receiving first sets of noisy trainable parameters from at least a portion of the user devices; andgenerating the plurality of image generation ML models by training the initial image generation ML model using the received first sets of noisy trainable parameters.15.The method as claimed in claim 14 wherein generating the plurality of image generation ML models comprises:clustering the received first sets of noisy trainable parameters into a plurality of clusters; andgenerating, for each cluster, an image generation ML model by training the initial image generation ML model using the sets of noisy trainable parameters in the cluster, thereby generating the plurality of image generation ML models.