Method, system and apparatus for image processing

Federated learning with cluster-specific masks and personalized model weights addresses the challenges of adapting semantic image segmentation to diverse domains and user preferences, enhancing accuracy and efficiency on client devices.

GB2628388BActive Publication Date: 2025-06-11SAMSUNG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
GB2023004180
Authority / Receiving Office
GB · GB
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-03-22
Publication Date
2025-06-11
Estimated Expiration
2043-03-22

AI Technical Summary

Technical Problem

Existing semantic image segmentation techniques are data-hungry and often trained on static source domains that do not adapt well to diverse target domains, leading to inaccuracies due to varying object appearances and labeling differences across regions, and lack of personalization based on user preferences.

Method used

A federated learning approach that clusters client devices based on common characteristics, generates cluster-specific masks to prune neural networks, and personalizes model weights, reducing model size and improving accuracy for specific domains and user preferences.

Benefits of technology

The method effectively personalizes machine learning models for client devices, reducing model size by up to 99.5% while maintaining accuracy, enabling efficient image processing and speech recognition on resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000001_0000
    Figure 00000001_0000
  • Figure 00000002_0000
    Figure 00000002_0000
  • Figure 00000003_0000
    Figure 00000003_0000
Patent Text Reader

Abstract

A method for generating a personalised machine learning model, such as a neural network, for a group or cluster of client devices amongst devices being supplied local models by a federated learning (F
Need to check novelty before this filing date? Find Prior Art

Description

Field

[001] The present application generally relates to methods, systems and apparatuses for processing images. In particular, the present application relates to a computer-implemented method and system for personalising, using federated learning, a machine learning, ML, model for processing images. The method and system for personalising a ML model may be used in other applications, for example in image classification and speech recognition. Background

[002] Over the past few years, advances in deep learning have revolutionised the way people interact with everyday devices. Client devices can be mobile devices such as smartphones, appliances, wearables, or even servers or systems of entities such as hospitals and organisations. Client devices may provide a user with various applications using machine learning, ML, or deep learning, DL, techniques, such as convolutional neural networks (CNNs) and transformers. These applications include image processing, image classification and speech recognition. One example of image processing which is widely used is semantic image segmentation and merely as an example, one paper which describes this technique is “Fully convolutional networks for semantic segmentation” by Long et al published on the Computer Vision and Pattern Recognition (CVPR) Conference (2015).

[003] Semantic image segmentation is a technique which may be used in image editing (e.g. to remove objects). This is a technique that segments an image into different regions based on their semantic meaning by assigning to each pixel in an image, a label representing the class or object to which the pixel belongs. For example, semantic image segmentation of an image of a bowl of fruit may include first classifying the image as a basket of fruit and then identifying each object which in this case is a different type of fruit, e.g. apples, oranges, bananas so that the image can be segmented into the different fruit types. Similarly, an image of a street may be segmented into different segments, including buildings, sky, vegetation, vehicles, people and so on or an image of a room can be segmented into different segments, including ceiling, floor, window, furniture and so on. A segmentation map that indicates the different regions of the image may be created and can then be used to process only specific parts of an image, e.g. to “erase” a specific part of an image.

[004] Many of these semantic segmentation techniques are data-hungry architectures with training typically carried out on a central server. These techniques also often use training data (source domain) which does not change over time and thus the source domain may be very different to the target domain (i.e. the images which are to be processed by the semantic segmentation technique). The same object (e.g. a vehicle in an image of a street) may look very different in different target domains (e.g. different countries or regions) and in some cases, certain objects may only exist in certain target domains (e.g. an Indian rickshaw). It is also possible that the same objects are labelled differently in different target domains (e.g. motorcyclist or rider). The categories of objects with a source domain may also be varied by user preferences and thus source domains may have different training data distributions. The segmentation model which is generated typically has the same capacity (e.g. number of layers, neurons, links) regardless of the complexity of the data which is being processed. Similar issues may be encountered in other use cases such as image classification or speech recognition.

[005] Federated Learning is a relatively new subfield of machine learning, ML, that allows the individual client devices to collaboratively train a ML model by moving the training computation to the client devices, while keeping all the training data private. During each round of the training process, participating client devices download the latest version of a global model and compute an updated model using their local data (i.e. data that is local to or stored on the client devices). These locally trained models are then sent from the participating client devices back to a central server which aggregates all the received locally trained models to generate an updated version of the global model. This is in contrast to centralised training where the training mechanism (e.g. the central server) has access to all the training data.

[006] The present applicant has recognised the need for an improved technique for personalising machine learning models such as image segmentation models. Summary

[007] In a first approach of the present techniques, there is provided a computer-implemented method, performed by a server, for generating personalised machine learning, ML, model data for a plurality of client devices which are connected to a server. The method comprises using the server to: obtain an initial global ML model in the form of a neural network comprising a plurality of neurons each having a global model weight and using at least one round of federated learning to train the obtained global ML model. Each round of federated learning comprises selecting multiple client devices to be used for that round of federated learning; distributing an ML model to each of the selected client devices for local training, wherein in a first round of federated learning the distributed ML model is the initial global ML model and in subsequent rounds of federated learning the distributed ML model is an updated ML model generated in a previous round of federated learning; receiving, from each selected client device, a plurality of local model weights generated by the local training of the distributed ML model and a neural activation score for each neuron; grouping the multiple client devices into a plurality of clusters; generating a cluster-specific mask for each cluster using the neural activation scores received from each client device in the cluster, wherein a mask specifies which neurons within the ML model are to be used; and generating at least one updated ML model based on the received plurality of local model weights. When the federated learning rounds have been completed, the method may further comprise notifying the client devices that personalised model data is available. The personalised model data may be output to the client devices, e.g. when requested by the client devices or automatically pushed to all the client devices.

[008] Each of the plurality ofclusters may comprise one or more of the client devices and each of the client devices in a cluster may have one or more characteristics in common. The cluster may be termed a group or a community of client devices. The client devices may be clustered at a client level. When there are group servers connected between each client device and the serverwhich is coordinating the federated learning forthe global model (i.e. the central server), the client devices may be clustered at a server level. The personalisation of the model data using federated learning as described above is at the level of the cluster rather than the level of each individual client device.

[009] The initial global ML model may be obtained from a data store and / or may be a global ML model which has been trained by the server on data which is accessible by the server. The method described above may generate at least one updated ML model in the form of an updated global ML model. The updated global ML model together with the cluster-specific mask may be the personalised model data for each client device in a cluster. Each cluster may be termed a community and the cluster-specific mask may be termed a community mask and the terms may be used interchangeably. When each cluster-specific mask is applied to an ML model, a pruned ML model may be generated in which the pruned ML model comprises a subset of the neurons within the neural network of the ML model to which the mask is applied. In other words, the cluster-specific mask effectively masks neurons within the neural network. The cluster-specific mask effectively reduces the size (i.e. prunes) of the ML model, e.g. the global ML model, which is to be used forthat cluster. This pruning step may reduce the size of each model to be used by a cluster by a fixed amount, for example by between 80 to 99.5%. In other words, the method may further comprise applying each cluster-specific mask to the updated global ML model to obtain a plurality of pruned models each which is personalised to the respective cluster. Each pruned, personalised model typically requires a smaller amount of storage than and will run more quickly than the global model without compromising the accuracy. The updated model data which is notified to the client devices when the federated learning has been completed may be the pruned model which is generated by the server or may be the mask and the updated global model so that the client device can generate the pruned model. It will be appreciated that for devices with limited storage or processing, sending the pruned model may be preferable.

[010] The mask may be in any suitable format. For example, the plurality of model weights (which may also be termed a set of model weights) are typically represented as a weight matrix or tensor for each layer. For a weight matrix the number of rows corresponding to the input number of channels (e.g., filters in CNNs) or features (e.g., in linear layers) and the number of columns corresponding to the output number of channels or features in a layer. In this example, the mask may be a matrix with the same shape as the weight matrix. Each entry in the mask weight may be one or zero (e.g. a matrix of binary values). Applying each cluster-specific mask to the updated global ML model may comprise multiplying the mask matrix by the weight matrix at each layer.

[011] As set out above, the method may be iterative with several rounds of federated learning. At each round of federated learning, the global ML model may be updated, for example by: selecting multiple client devices for training the global ML model from the previous round; distributing the previous global model data to each selected client device; receiving, from each selected client device, a plurality of local model weights generated by the local training and a neural activation score for each neuron; updating the grouping or partitioning of the multiple client devices into a plurality of clusters; updating the cluster-specific mask for each cluster using the neural activation scores received from each client device in the cluster; updating the updated global ML model based on the plurality of local model weights; and repeating the selecting, distributing, receiving and updating steps until it is determined that no further federated learning is required. Each of these steps is carried out by the server.

[012] During the federated learning optimization stages, a mask may be distributed from the server in the first round of federated learning with the initial global ML model. This may be an initial mask which does not mask any neurons within the neural network. When in the iterative stages of the federated learning, a mask may be generated locally on each client device. The locally generated masks may be sent to the server for storage and the server may aggregate the received masks to form cluster-specific masks based on clusters. When the federated learning optimization stage is completed, the server distributes both the ML model and the maskto the client or, alternatively, the server computes the multiplication ofweights and masks and sends the resulting ML model to the clients.

[013] In addition to personalizing the pruning of the global ML model, the method may personalise the model weights which are used locally by client devices by generating, using the server, a cluster-specific ML model for each cluster using the plurality of local model weights received from each client device in the cluster. In other words, the step of generating the at least one updated ML model in each round of federated learning may comprise generating a cluster-specific ML model for each cluster. Each cluster-specific ML model is in the form of a neural network comprising a plurality of neurons each having a cluster-specific model weight. The number of neurons (and hence weights) in the cluster-specific ML model is the same as the global ML model (initial or updated). As described above, each cluster may be termed a community and the cluster-specific ML model may be termed a community ML model and the terms may be used interchangeably. The personalised model data may comprise the cluster-specific model.

[014] The generation of the cluster-specific ML model may also be iterative and may be done at each of the several rounds of federated learning. In the initial round when no cluster-specific model has been generated or in subsequent rounds when the selected client device has not been assigned to a cluster, the distributed ML model comprises the initial global ML model or the updated global ML model, respectively. When the selected client device has been assigned to a cluster (and a cluster-specific model has been generated in a previous round), the distributed model data may be the cluster-specific ML model.

[015] The cluster-specific model for each cluster may be generated by the server by aggregating the plurality of local model weights received from all client devices in the cluster. An aggregation of the local model weights may be done using the following: within cluster] = 1.....K <**> duster where there are K clients in the cluster, WcUentkis the weight matrix for the kth client within the cluster], Nk is the number of samples at client k in the cluster j.

[016] When a cluster-specific model for a cluster has been generated in a previous round, the cluster-specific model for a subsequent round may be generated by discarding the previous cluster-specific model and generating the updated cluster-specific model by aggregating the plurality of updated local model weights received from all client devices in the (possibly updated) cluster. In other words, the equation above may be used in each round of federated learning. Alternatively, the cluster-specific model generated in a previous round may be updated using any appropriate technique, e.g. a moving average approach. A general methodology to update existing weights in a federated setting is defined in Sashank J. Reddi et. al. “Adaptive Federated Optimization’’, International Conference on Learning Representations (ICLR, 2021).

[017] The updated global model may be generated directly from the local model weights by aggregating the plurality of local model weights received from all selected client devices. Alternatively, the local model weights may be used indirectly. For example, when a clusterspecific model is generated, the updated global model may be generated by aggregating all the weights from the cluster-specific models. The aggregation may be done using the following: Waggregate ^^XPWclusterp Vp = 1,-, P clusters where there are P clusters and Wciuster is the set of weights for the pth cluster. As for the global model, the previous updated global model may be discarded at each round of federated learning and a new global model may be generated at the current round. Alternatively, the global model generated in a previous round may be updated using any appropriate technique, e.g. a moving average approach.

[018] For each selection of multiple client devices for training, the server may select a subset of client devices to which it is connected. The subset selected at each selection may be different. Selecting may comprise determining how many samples are present in a training dataset for each client device and selecting client devices which have more than a minimum threshold of training samples. The minimum threshold may be a fixed numberof samples which ensures an acceptable level of accuracy for the training, e.g. 10. At each iteration of federated learning, selecting may comprise prioritising client devices which have not previously been selected for federated learning.

[019] Grouping the client devices may comprise obtaining a tag (or a label, the terms may be used interchangeably) for each client device, obtaining a cluster for each obtained tag, and assigning each client device to the cluster corresponding to the tag obtained from that client device. The tag may be obtained from storage on the client device. In this case, the tags have been assigned previously to the client devices and may be termed priors. Obtaining a cluster may comprise comparing the obtained tag with a database of paired tags and clusters. Alternatively, obtaining a cluster may comprise defining a cluster, e.g. using any suitable clustering technique. A tag or label may be any identifier which is assigned to a client device, e.g. a geographical location.

[020] Alternatively, the grouping of the client devices may be done based on information received from the client devices during the federated learning process. The client devices may be clustered based on the local model weights generated by the local training on each client device and / or based on the neural activation scores obtained from each client device. Client devices with similar local model weights may be grouped together. Alternatively, client devices with similar neural activation scores may be grouped together. For example, partitioning the client devices may comprise comparing the neural activation scores received from each client device (e.g. determining differences or similarities) between the neural activation scores received from each client device) and grouping client devices into a cluster when the differences between the neural activation scores for the grouped client devices are below an activation difference threshold. Similarly, partitioning the client devices may comprise comparing the local model weights received from each client device (e.g. determining differences or similarities) between the local model weights received from each client device) and grouping client devices into a cluster when the differences between the local model weights for the grouped client devices are below a weight difference threshold. The difference threshold may be set at a fixed amount based on the number of clusters to be defined. For example, a lower threshold will result in a larger number of clusters. When updating the partitioning of the multiple client devices in subsequent rounds of federated learning, this may be done using the same technique as the partitioning used in the first round of federated learning.

[021] The calculation of the neural activation score may be done wholly or partially at each client device. For example, this calculation may be done by obtaining, at each client device, a set of test samples; and for each test sample in the set of test samples applying the ML model which is received from the server to the test sample; and measuring, when applying the ML model to the test sample, an activation value for each neuron in the neural network of the ML model. The activation values for each test sample and for each neuron are then combined to obtain the neural activation score. The test samples may be obtained from a database, e.g. a training dataset, or by newly acquiring the test samples.

[022] In a second approach of the present techniques, there is provided a computer-implemented method, performed by a client device, fortraining a machine learning, ML, model using federated learning, the method comprising: receiving, from a server, a ML model in the form of a neural network comprising a plurality of neurons each having a global model weight; training the received ML model using a plurality of samples in a user database on the client device to generate a plurality of local model weights; selecting a test set of samples from the user database; for each test sample in the set of test samples: applying the received ML model to the test sample and measuring, when applying the ML model to the test sample, an activation value for each neuron in the plurality of neurons; combining the activation values for each test sample and for each neuron to obtain a neural activation score, and outputting the generated plurality of local model weights and the neural activation score to the server.

[023] The ML model which is received from the server will depend on the iteration of the federated learning and / or whether the global model or the cluster-specific model is to be trained. The received ML model may be the initial global model forthe first iteration of federated learning. The received ML model may be an updated or further updated global model forthe subsequent iterations of federated learning, particularly when the client device has not been assigned to a cluster. When the client device has been assigned to cluster in an earlier iteration of the method, the received ML model may be the cluster-specific model orthe updated clusterspecific model.

[024] The neural activation score for each neuron may be calculated by summing, for each neuron, all the measured activation values for each sample. In other words, the neural activation score st may be calculated using: n where a;<n is the activation value of neuron / when passing the n-th test sample.

[025] As an alternative to the sum above, the neural activation score may be normalized. In other words, the neural activation score st for each neuron may be calculated by normalizing the sum of the individual values by the number of samples, i.e. from. where „ is the activation of neuron / when passing the n-th test sample and N is the number of samples. The normalization may be carried out by the client device or the server but by carrying it out at the client device, there is no need to send the number of samples to the server.

[026] As an alternative to the scores which are based on the sum of the values above, the neural activation score may be simplified to a binary output by comparing the normalized sum to an overall threshold. The neural activation score may be 1 when the normalized sum is above an overall threshold T and 0 when the normalized sum is below or equal to the overall threshold T. In this case, the following formula may be used to calculate the neural activation score st for each neuron: Si = K >T], where k(-) is the indicator function which sets all the elements which meet the threshold to 1 and all the elements which do not meet the threshold to 0, aln is the activation for each neuron I for each sample n, T is the threshold, and there are N samples overall.

[027] As an alternative to calculating a sum of all the activation values and comparing to an overall threshold, each activation value may be compared to an activation threshold. In either scenario, the comparisons to a threshold may be carried out by the client device or the server. When comparing each activation value, the neural activation score may be a sum which is incremented by a fixed amount (e.g. one), only for each measured activation value above the activation threshold. In this case, the following formula may be used to calculate the neural activation score st for each neuron: Sl = >7}) n where k(-) is the indicator function, alin is the activation for each neuron / for each sample n, Tt is the activation threshold, and there are A / samples overall. The sum may be normalised by the total number of samples on user devices, for example using the formula below: _ Sn K(al,n>Ti) Sl~ N

[028] In each case, the neural activation scores are used to determine the masks which are applied to prune the models. Each cluster-specific mask may for example be calculated using: Mcluster. .........r-...................:.............~................;..................:—'ZkSdienu >TMWk clients in j-th cluster ciusietj ^clients in clusterj^K CLienLk lvl' J where k(-) is the indicator function, K is the number of clients in the cluster, TM is a threshold on the mask and Sclientk is the activation scores of the k-th client, where Sclientk is a vector, matrix or tensor of neural activation scores (¾) per layer or neuron of the k-th client’s model.

[029] The overall threshold T and / or the activation thresholds Tt may be set in various ways. For example, the overall threshold T and / or the activation thresholds Tt may be set so that the fewer than a fixed percentage (e.g. 90%) of the neurons will have a summed score or an individual activation score which meets the threshold. Only these neurons will then be represented by a one in the mask matrix and all other neurons will be represented by a zero.

[030] The model data which has been trained as described above includes the global ML model which may be personalised and pruned for client devices in a cluster using the clusterspecific mask and / or the cluster-specific ML models which have personalised model weights and which have been optionally pruned using the cluster-specific mask. This personalised model data can be used in a variety of applications, including image processing and speech to text transcription. Once received at the client device, the personalised model data which is personalised at the cluster level may be further personalised to the local client device.

[031] In a third approach of the present techniques there is a computer-implemented method for processing an input image on a client device, the method comprising: receiving, from the server, personalised model data which has been trained as described above and wherein the personalised model data is for image processing; obtaining, at the client device, an input image; processing the input image using the updated model data to generate a processed image in the form of a segmentation map by separating the input image into a plurality of semantically consistent segments. The processed image (i.e. the segmentation map) may then be output. The model data may comprise a ML model in the form of a segmentation model which generates a segmentation map divides the input image into a plurality of segments (or segmented regions). Examples of segmentation models which may be trained as described above are described for example in “Fully Convolutional Networks for Semantic Segmentation” by Long et al. published in the Proceedings of the IEEE / CVF Computer Vision and Pattern Recognition (CVPR) conference (2015) or “Segnet: A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation” by Badrinarayanan et al. published in the Proceedings of the IEEE / CVF Computer Vision and Pattern Recognition (CVPR) conference (2015). Alternatively, segmentation may be done using transformers, e.g. as described in “Segmenter: Transformer for Semantic Segmentation” by Strudel et al. published in the Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV) 2021 or ”Swin Transformer: Hierarchical Vision Transformer using Shifted Windows” by Liu et al. published in the Proceedings of the IEEE / CVF International Conference on Computer Vision (ICCV) 2021. For example, a segmentation model may be used to generate the segmentation map.

[032] The method may further comprise selecting at least one segment from the plurality of semantically consistent segments. In this case, the model data may further comprise a ML model in the form of a recommendation model which automatically selects the at least one segment or the selection may be done by a user. Both the recommendation model and / or the segmentation model may have been trained using federated learning as described above to generate both global ML models and / or cluster-specific ML models where used. The personalisation of the recommendation model should also result in an improved user experience when editing an input image because the recommendation model is more likely to automatically select a segment which a user has previously edited in similar photos.

[033] The method may comprise further processing the input image by applying at least one of the following: erasing the at least one selected segment, applying a filter to the at least one selected segment, and providing a tag for the at least one selected segment to be output with the processed image. Such processing may be done using any suitable application and may be integrated with photo storage applications (e.g. Samsung Photo Gallery, Samsung stories) on a user device. For example, the tag represents a class or label, e.g. dog, cat, people, which shows regions of the image whose pixels belong to the semantic class indicated by the label. The classes may be pre-determined and may be fine-tuned (i.e. increased in number or otherwise updated) by the training.

[034] In a third approach of the present techniques there is a computer-implemented method for processing speech on a client device, the method comprising: receiving, from the server, personalised model data which has been trained as described above, wherein the personalised model data is for converting speech to text; obtaining, at the client device, an input speech segment; and processing the input speech segment using the updated model data to generate a text transcript.

[035] In a related approach of the present techniques, there is provided a client device such as a personal computer or a laptop and may be a more resource constrained device, such as a smartphone or tablet computer and / or other mobile device which is configured to implement the steps carried out by the client device as described above. Similarly, in a related approach of the present techniques, there is provided a server which is configured to implement the steps carried out by the server as described above. Similarly, in a related approach of the present techniques, there is provided a system comprising a server which is connected to multiple client devices and the server and client devices cooperate to implement the steps described above.

[036] In a related approach of the present techniques, there is provided a computer-readable storage medium comprising instructions which, when executed by a processor, causes the processor to carry out any of the methods described herein.

[037] As will be appreciated by one skilled in the art, the present techniques may be embodied as a system, method or computer program product. Accordingly, present techniques may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects.

[038] Furthermore, the present techniques may take the form of a computer program product embodied in a computer readable medium having computer readable program code embodied thereon. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable medium may be, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.

[039] Computer program code for carrying out operations of the present techniques may be written in any combination of one or more programming languages, including object oriented programming languages and conventional procedural programming languages. Code components may be embodied as procedures, methods or the like, and may comprise subcomponents which may take the form of instructions or sequences of instructions at any of the levels of abstraction, from the direct machine instructions of a native instruction set to high-level compiled or interpreted language constructs.

[040] Embodiments of the present techniques also provide a non-transitory data carrier carrying code which, when implemented on a processor, causes the processor to carry out any of the methods described herein.

[041] The techniques further provide processor control code to implement the above-described methods, for example on a general purpose computer system or on a digital signal processor (DSP). The techniques also provide a carrier carrying processor control code to, when running, implement any of the above methods, in particular on a non-transitory data carrier. The code may be provided on a carrier such as a disk, a microprocessor, CD- or DVD-ROM, programmed memory such as non-volatile memory (e.g. Flash) or read-only memory (firmware), or on a data carrier such as an optical or electrical signal carrier. Code (and / or data) to implement embodiments of the techniques described herein may comprise source, object or executable code in a conventional programming language (interpreted or compiled) such as Python, C, or assembly code, code for setting up or controlling an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), or code for a hardware description language such as Verilog (RTM) or VHDL (Very high speed integrated circuit Hardware Description Language). As the skilled person will appreciate, such code and / or data may be distributed between a plurality of coupled components in communication with one another. The techniques may comprise a controller which includes a microprocessor, working memory and program memory coupled to one or more of the components of the system.

[042] It will also be clear to one of skill in the art that all or part of a logical method according to embodiments of the present techniques may suitably be embodied in a logic apparatus comprising logic elements to perform the steps of the above-described methods, and that such logic elements may comprise components such as logic gates in, for example a programmable logic array or application-specific integrated circuit. Such a logic arrangement may further be embodied in enabling elements for temporarily or permanently establishing logic structures in such an array or circuit using, for example, a virtual hardware descriptor language, which may be stored and transmitted using fixed or transmittable carrier media.

[043] In an embodiment, the present techniques may be realised in the form of a data carrier having functional data thereon, said functional data comprising functional computer data structures to, when loaded into a computer system or network and operated upon thereby, enable said computer system to perform all the steps of the above-described method.

[044] The method described above may be wholly or partly performed on an apparatus, i.e. an electronic device, using a machine learning or artificial intelligence model. The model may be processed by an artificial intelligence-dedicated processor designed in a hardware structure specified for artificial intelligence model processing. The artificial intelligence model may be obtained by training. Here, "obtained by training" means that a predefined operation rule or artificial intelligence model configured to perform a desired feature (or purpose) is obtained by training a basic artificial intelligence model with multiple pieces of training data by a training algorithm. The artificial intelligence model may include a plurality of neural network layers. Each of the plurality of neural network layers includes a plurality of weight values and performs neural network computation by computation between a result of computation by a previous layer and the plurality of weight values.

[045] As mentioned above, the present techniques may be implemented using an Al model. A function associated with Al may be performed through the non-volatile memory, the volatile memory, and the processor. The processor may include one or a plurality of processors. At this time, one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an Al-dedicated processor such as a neural processing unit (NPU). The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (Al) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning. Here, being provided through learning means that, by applying a learning algorithm to a plurality of learning data, a predefined operating rule or Al model of a desired characteristic is made. The learning may be performed in a device itself in which Al according to an embodiment is performed, and / o may be implemented through a separate server / system.

[046] The Al model may consist of a plurality of neural network layers. Each layer has a plurality of weight values, and performs a layer operation through calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), deep Q-networks, and transformer or visual transformer networks.

[047] The learning algorithm is a method for training a predetermined target device (for example, a user device or client device) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semisupervised learning, or reinforcement learning. Brief description of the drawings

[048] Implementations of the present techniques will now be described, by way of example only, with reference to the accompanying drawings, in which:

[049] Figure 1a shows a block diagram of a system for performing the methods described herein

[050] Figure 1b shows an expanded version of the system of Figure 1a to illustrate clustering within the system.

[051] Figures 2a to 2c is a flowchart showing a training process for personalising a model to use on a client device within the system of Figure 1 a.

[052] Figure 3 is a Venn diagram illustrating different types of client devices within the system of Figure 1a.

[053] Figure 4a is a schematic drawing of a model which may be used in the process of Figures 2a to 2c.

[054] Figure 4b is a schematic drawing showing how the model of Figure 4a is used to obtain four differently personalised models;

[055] Figure 5a is a schematic illustration of measuring a neuron activation in a model of Figure 3.

[056] Figure 5b is a flowchart for the process of calculating a neural activation score using the measurements shown in Figure 5a.

[057] Figure 6a is a schematic illustration summarising the process in Figures 2a to 2c.

[058] Figure 6b is an alternative illustration of three personalised models generated as shown in Figure 4b.

[059] Figures 7a to 7c are input images of roads taken in three different regions.

[060] Figure 8 is a flowchart for an application using a model obtained from the process of Figures 2a to 2c.

[061] Figure 9 is a segmentation map obtained by applying a segmentation model to the input image shown in Figure 7c.

[062] Figure 10a shows a selection of an object in the input image of Figure 7c.

[063] Figure 10b shows an output image in which the selected object has been removed.

[064] Figure 11 is a flowchart of another application using a model obtained from the process of Figures 2a to 2c.

[065] Figure 12a shows a partial selection of an object in the input image of Figure 7c.

[066] Figure 12b shows an output image in which the partially selected object has been removed.

[067] Figures 13 to 15 are flowcharts of three different applications each using a model obtained from the process of Figures 2a to 2c. Detailed description of the drawings

[068] Broadly speaking, embodiments of the present techniques provide methods, systems and apparatuses fortraining and using ML models, e.g. to process images or speech. Training uses federated learning to personalise model data for all client devices within a cluster. The applications include image processing, image classification and speech recognition. For example, the training may personalise a segmentation model for creating a segmentation map of an input image, and / or personalise a recommendation model for recommending regions of the input image to the user for subsequent processing. The model data may be personalised by applying a cluster-specific mask to prune a global model which is an aggregation of all models generated by client devices and / or by generating cluster-specific models with personalised model weights.

[069] Figure 1 shows a block diagram of a system 100 for performing the training and inference methods described below. The system 100 comprises a plurality of client devices 110 and for ease of reference, a single client device 110 is shown. The client device 110 may be any type of user device (and the terms may be used interchangeably) such as a personal computer or a laptop and may be a more a resource constrained device, such as a smartphone or tablet computer and / or other mobile device. The client device 110 comprises the standard components of such devices, including for example as shown one or more processors 112, memory 114, a communication module 116, a display 122 for displaying information to the user and a user interface 124 such as a mouse, keyboard, voice recognition input device, touch sensitive screen or any similar component for receiving user input. There may be other standard components which are omitted from the Figure for ease of reference.

[070] The display 122 may comprise any suitable display screen, e.g. LCD, LED which may also be touch sensitive to allow user input. The communication module 116 may communicate using any suitable communication, e.g. wireless communication, hypertext transfer protocol (HTTP), message queuing telemetry transport (MQTT), a wireless mobile telecommunication protocol, radio frequency communication (RFID), near field communication (NFC), ZigBee, Thread, Bluetooth, Bluetooth LE, IPv6 over Low Power Wireless Standard (6L0WPAN), Constrained Application Protocol (CoAP) or a wired communication. The memory 114 may be any suitable form of memory, including volatile memory, such as random access memory (RAM), for use as temporary memory, and / or non-volatile memory such as Flash, read only memory (ROM), or electrically erasable programmable ROM (EEPROM), for storing data, programs, or instructions, for example.

[071] The client device 110 optionally comprises an image capture device 120, e.g. camera for taking input images that may then be processed as described below. The client device 110 may also comprise one or more sensors 118, for example a sensor to measure neural activation when using the machine learning module. The sensed data 136 may be stored in storage 130 on the client device which is shown separately from memory 114 but it will be appreciated that the two components may be combined.

[072] The local storage 130 also stores at least one machine learning (ML) model 134 (also termed an artificial intelligence (Al) model) which have been trained to produce an output. Some examples of such models include an image captioning model which produces a text description for an input image, an image classification model which classifies input images into different categories, a speech recognition model which processes input speech and an object eraser model which may comprise one or more sub-models such as an object detection model, a semantic segmentation model and an instance semantic segmentation model.

[073] The storage also locally stores a training dataset 132, e.g. input images, speech data or other user data which may be used to train the ML model(s). By using the user’s personal data to personalise the models, the overall user experience is typically much improved. For example, the user generally needs to provide less supervision when using the models, e.g. to achieve the desired photo editing and hence may generate better results, e.g. higher quality edited images. As shown in Figure 1a, the user’s training dataset is stored locally on the client device and may thus remain on the client device. Thus, the data remains private and there is a reduced risk of data breach or data theft. The personalisation can also then be targeted to users who do not typically share any data.

[074] The client device 110 also comprises a training module 140 which is used as described below for local (i.e. on device) training of the model using the locally stored user data. There is also an inference module 142 which is used to implement the model once training and testing are complete. There may optionally also be a pruning module 144. The pruning module 144 is used to reduce the size of the locally storage model. Merely as an example, uncompressed models such as the image classification and speech recognition models described above typically have between 0.5 and 200 million parameters (including weights). Such large or massive models are difficult to store and use on resource-constrained devices and thus as described below, the method aims to prune the stored models by between 80 to 99.5% (in other words, 90 to 99.5% of the parameters are removed). The modules are shown as separate from the processor or processors 112 but it will be appreciated that the components may be combined as necessary to perform the methods described above.

[075] Using the communication module 116, the client device 110 may communicate with a server 150. The server 150 may also comprise standard components, including for example as shown one or more processors 152, memory 154 and a communication module 156. There may be other standard components which are omitted from the Figure for ease of reference. The server 110 may also comprise a pruning module 160 and the reduction in size of the locally stored models may be done on the server rather than on the client device 110 as described in more detail below.

[076] The central server 150 also communicates with a database 170 which stores the central model(s) 172. The central model 172 may be the aggregate or complete model with all parameters and thus by storing the central model 172 on a separate database, the problem of the size of the model is addressed. The database 170 also stores the mask data 174. A mask is used to tailor the central model to the specific user(s) as explained in more detail below. The database 170 also stores the personalised, local models which have been generated by the client devices 110. The database 170 is shown as a separate component but it will be appreciated that it could be storage which is on the central server 150.

[077] The server 150 also comprises a clustering module 158 which as explained in more detail below generates clusters of client devices. Figure 1 b illustrates that there may be clustering at different hierarchical levels. For example, Figure 1b illustrates a central server 250 (which may also be termed a root server) such as the one shown in Figure 1a. The central server communicates with a plurality of client devices 210a, 210b, 210c, 210m-2, 201 m-1,201m via multiple servers 252a, 252b, 252n. These servers 252a, 252b, 252n may be termed group servers. The client devices may be clustered at the client level, e.g. to form two separate clusters 254a, 254b of client devices. Alternatively, the client devices may be clustered at the group server level, for example by forming a cluster 256a of group servers which means that all the client devices 210a, 210b, 210c are in a single cluster. Each cluster comprises at least one client device and may alternatively be termed a group or a community of client devices.

[078] Clustering may be based on any suitable characteristic of the client devices or group servers. For example, clustering may be based on the geographical coordinates (e.g. GPS) of the client devices and / or group servers. Alternatively, clustering may be based on the local models generated by the client devices, e.g. using one or both of the locally generated model weights or masks. Alternatively, clustering may be on other local client data, e.g. sample statistics uploaded by the users to the servers, including general sample-level statistics such as colour histograms of data samples or the content of data samples, e.g. user preferences for dogs or cats. Any combination of the clusters indicated above can be used. The clustering can be assigned during the process as described in more detail below. Alternatively, the client devices may be labelled or tagged with a particular cluster label, for example if all US users are to be assigned to one cluster and all European users to another cluster, the location of the users is known and thus the clusters can be assigned based on these labels.

[079] Figures 2a to 2c form a flowchart showing an example of the overall process which is applied in the system of Figure 1a. The process comprises a training phase which is a form of federated learning (also known as collaborative learning). Federated learning is a machine learning technique in which an algorithm (i.e. ML model) is trained across multiple decentralized devices or servers each of which hold data samples and which do not share or exchange the samples. Federated learning techniques are described for example in “Learning across domains and devices: Style-driven source-free domain adaptation in clustered federated learning” by Shenaj et al published in WACV 2023, “Federated learning with hierarchical clustering of local updates to improve training on non-IID data” by Briggs et al published in IJCAI2020 and “Federated Learning with Personalization Layers” by Arivazhagan published in 2019.

[080] In a first step S200 of Figure 2a, the central server optionally selects a subset of client devices 110 for federated learning. This optional step may be used when there is a large number of client devices and training on all accessible client devices is not ideal or even possible. This optional step may be used to focus the federated training on client devices with sufficient training data and / or client devices which have not previously been used. This is schematically illustrated in Figure 3 which is a Venn diagram showing a complete set of client devices 310 which may be selected by the central server 150. Within the complete set 310, there are two further sets of client devices. A first set 320 represents the client devices which have not been previously used in the federated learning process, e.g. because they have newly connected to the central server. A second set 330 represents the client devices having a training dataset with sufficient samples, for example more than a sample threshold L. Merely as an example, the sample threshold may be L=10. There are some client devices which sit in the intersection 340 of both sets, i.e. that have sufficient samples and have not been seen previously.

[081] An illustrative process for selecting a subset of client devices for federated learning may comprise, initially reducing the complete set to a first sub-set of client devices having a more manageable number of client devices. Merely as example, the first sub-set may comprise 10% ofthe total number of clients. This initial reduction may be done using any appropriate unbiased sampling technique, e.g. multinomial distribution or uniform distribution. An example of a suitable technique is described in “A General Theory for Client Sampling in Federated Learning” by Yann et al published in FL-IJCAI 2022. From within this first sub-set of client devices, there may be a second step which prioritises selection of certain client devices. For example, the first prioritisation criteria is to ensure that the client device has sufficient samples within the training dataset (i.e. >L). The second prioritisation criteria is to select previously unseen client devices which may represent a small (e.g. 10%) proportion of the first sub-set. Each client device may be assigned a sampling priority based on whether one or both of the prioritisation criteria are met, e.g. • If only one criteria is met - assign a sampling priority of x (e.g. x=0.6) • If both criteria are met - assign a sampling priority of y (e.g. y=0.3) • Otherwise - assign a sampling priority of 1-x-y (e.g. 0.1) The sampling priority thus ensure that the client devices which only meet one criteria are prioritised above those that meet both. There will also still be a small number of client devices which do not meet either criteria.

[082] Returning to Figure 2a, the next step S202 is to send the latest model data to each selected client device. An example of a machine learning or artificial intelligence model is shown in Figure 4a. The model 400 comprises a plurality of neurons (represented by the circles) which are arranged in a plurality of neural network layers 410, 412, 414. A first layer 410 is an input layer which comprises a plurality (p) of inputs (xi, x2,... Xp) and the final layer 414 is an output layer which comprises a plurality (g) of output (yi, y2, ... yg). The number of inputs and outputs may be same or different. Each intermediate layer 412 comprises a plurality of neurons, represented in this example by a(25, a(3), and a(4). Each layer is also associated with a plurality of weights, represented in this example by W(2), W<3), and W(4). The overall model may be considered to have a set of model weights which may be represented as a weight matrix or tensor for each layer. For example, when using a 2D matrix, the number of rows may correspond to the input number of channels (e.g., filters in CNNs) or features (e.g., in linear layers) and the number of columns may correspond to the output number of channels or features in a layer. Each of the plurality of neural network layers includes a plurality of weight values and performs neural network computation by computation between a result of computation by a previous layer and the plurality of weight values. As shown, there are three intermediate layers and hence three weight matrices but it will be appreciated that there may be any suitable number of layers.

[083] Returning to Figure 2a, the model data which is sent at step S202 will depend on whether there have been any previous training iterations or rounds by the central server 150. If the central server 150 has not done any previous federated learning (i.e. this is round 0), the model data which is sent at step S202 comprises the initial or initialised model weights. These initialised model weights may be obtained, for example by pre-training a model such as that shown in Figure 4a, using generic data and the specific task of interest (e.g. image segmentation, speech recognition etc). If the central server 150 has previously performed some federated learning, the model data which is sent at step S202 will depend on whether or not the client device 110 which is receiving the model data has already been assigned to a cluster (as described below). When the client device 110 has not been assigned to a cluster, the client device 110 will receive the latest global aggregate model weights, which may be the mean of all the cluster weights. When the client device 110 has been assigned to a cluster, the client device may receive the cluster-specific model weights but alternatively, the client device may receive the latest global aggregate model weights.

[084] As shown at step S204, each client device 110 which is taking part in the federated learning process receives the latest model data. The next step S206 is to train the model using local data (i.e. the training dataset on the client device). This local data will have been collected previously, e.g. using an image capture device or any other suitable means. The model may be trained using any suitable technique, for example using: Lk = ' IkiWk’Xkj') ^k . J=K;Nk where Nk. number of samples at client k, Wk. model weights at client k, xk7 : j-th input sample at client k, lk: user-specific loss function. Any suitable loss function may be used include for example cross-entropy loss, regression loss or focal loss. The training generates a new set of weights which may be termed local weights. The gradients (i.e. differences between the original weights and new local weights) may also be calculated.

[085] After the model has been locally trained, there is a testing phase in which at step S208, the client device 110 retrieves user data and the locally trained model and then uses the locally trained model on the retrieved user data. In other words, each client device 110 performs a forward pass of the newly trained model on their retrieved data at step S210. While this inference stage is being carried out, a neural activation score (may also be termed a neuron activation score) is determined for each neuron in the model at step S212. The determination of the neural activation score is described in more detail in relation to Figures 5a and 5b. The activation score gives an idea of the average number of activations count, where a higher count indicates important neurons for the user.

[086] During each federated learning training round (which includes the testing phase), the client device 110 may generate a mask, for example using the neural activation scores. Each mask comprises a matrix or tensor for each layer which has the same shape (i.e. same number of rows and columns) as the weight matrix or tensor for each layer and which is multiplied with the weight matrix to personalise the model as illustrated in Figure 4b. The value of each entry in the mask matrix is either one or zero. When the value is zero, the weight is effectively removed and thus the associated neuron is masked (i.e. does not appear). For example, as shown in Figure 4b, four different masks are applied to the central or aggregate model to generate four different models 1 to 4. Each of these models comprises a sub-set of the neurons in the central model and may thus be considered to be a pruned version of the central model because it contains fewer parameters.

[087] Initially, each value in the mask matrix is set to 1 whereby multiplying the weight matrix by the mask matrix makes no change to the weight matrix. In other words, even if there is a personalised mask for a particular cluster which has been generated as explained below, the personalised cluster-specific mask is not sent at step S202 in the training to promote diversity during the learning process. If the cluster-specific mask was distributed during training, the final mask which is generated as a result of training will be a subset of the distributed mask and this could limit the validity of the mask. In effect, during training we mainly considered to leave neurons unpruned and the personalised (i.e. cluster-specific) mask is preferably only used during inference (and testing). Any mask which has been locally generated may also be stored on the client device and / or sent to the server.

[088] As shown at step S214, the local model and / or the neural activation score may be stored in storage on the client device. Each client device 110 then sends to the central server at step S216 its activation score Sciient / cand its newly trained weights Wclient (or gradients based on the original model). The score and updated weights (or gradients) are received at the central sever at step S218.

[089] Figure 2b shows how the central server processes the received score and updated weights to update its central or global model, generate cluster-specific mask(s) and / or generate cluster-specific models. In a first step S220, the central server 150 may store the received scores and / or updated weights. The data may be stored on the central server 150 or in a database which is connected to the central server 150. The data may be stored in a queue and at step S222 there may be a determination to see whether there is sufficient information in the queue to perform the subsequent steps of the method. Thus at optional step S222, it is determined whetherthe size of the current queue exceeds the queue threshold. If the threshold is not exceeded, the process loops back until the threshold is exceeded. There is a further optional step at step S224 in which a subset of the data in the queue may be selected, e.g. using random selection or other suitable technique.

[090] At step S226, the central server groups or partitions the client devices into clusters. There are various options for the clustering. For example, the client devices may be clustered based on the local weights generated by the local training on each client device. In this case, each client device may be assigned to the same cluster for its mask. Alternatively, the client devices may be clustered based on the activation scores (and hence masks) generated by the local testing on each client device. In this case, each client device may be assigned to the same cluster for its weights. Clustering of weights or masks may be done using an unsupervised technique, for example kMeans as in S. Lloyd, "Least squares quantization in PCM," in IEEE Transactions on Information Theory, vol. 28, no. 2, pp. 129-137, (1982) or deepCIuster as in Tian, K., Zhou, et al., “DeepCIuster: A General Clustering Framework Based on Deep Learning”, European Conference on Machine Learning Principles and Practice of Knowledge Discovery in Databases (ECML / PKDD, 2017).

[091] Alternatively, there may be two different clusters for each client device: one cluster for the weights and a different cluster for the mask. In summary, the examples may be: • weights clustering = mask clustering, e.g. {user_1: cluster_4, user_2: cluster_1, user_3: cluster_4, etc.} • mask clustering = weights clustering, e.g. {user_1: cluster_2, user_2: cluster_3, user_3: cluster_1, etc.} • Weights clustering = {user_1: cluster_4, user_2: cluster_1, user_3: cluster_4, etc.} Masks clustering = {user_1: cluster_2, user_2: cluster_3, user_3: cluster_1, etc.}

[092] As explained above, the clustering may alternatively be done based on labels (which may also be termed priors) such as geographical tags or other user related tags. Alternatively, clustering can be based on the domain seen by each client; for example, on the basis of the sample-level statistics of data present in each client (e.g., colour histogram of data samples). Suitable techniques are described for example in “Learning across domains and devices: Style-driven source-free domain adaptation in clustered federated learning” by Shenaj et al published in Winter Applications in Computer Vision (WACV) Conference 2023 or “Personalized image semantic segmentation” by Zhang et al. published in International Conference of Computer Vision (ICCV, 2021). In this case, assignment of clients to groups (also known as clusters) may be done using additional pre-defined rules (which may also be termed priors, e.g., deciding where to separate geographical coordinates) or an unsupervised technique, for example kMeans or deepCIuster. Each cluster may contain the same or a different number of client devices.

[093] Once the clusters are defined, the central server 150 performs aggregation within each cluster. At step S228, for each cluster, the weights of each locally trained model for each client device 110 within the cluster are aggregated to generate a set of cluster weights WcLuster. for that cluster. The aggregation may be done by discarding any previously generated clusterspecific models and generating a new cluster-specific model at each round of federated learning using the following: within cluster] Wrh.„tpr. «- -----------------vfc clients in j-th cluster cluster J #clients in clustercLlentk J where there are K clients in the cluster, WcUentkis the weight matrix for the kth client within the cluster], Nk is the number of samples at client k in the cluster j.

[094] Alternatively, when the training procedure is at round t where t >0 and a previously generated cluster-specific model WPl~^ter. already exists, this model can be updated using any appropriate technique, e.g. a moving average approach. For example, a “pseudo-gradient” -△cluster; may be defined as: —△cluster- —i—7—Sit? m ^cluster■ — ^ciientkvk = clients in j-th cluster ciusieij tfchents in clusterj ciusieij uumt J where there are K clients in the cluster, Wclientkis the weight matrix for the kth client within the cluster], and Nk is the number of samples at client k in the cluster j. The cluster-specific model for the round t, W^luster may then be generated as follows: Cluster; = SERV EROPT where SERVEROPTQ is a general purpose gradient-based optimization algorithm which takes the previous set of weights W^terj, the current pseudo-gradients -Aciustery a learning rate r) and the round index t to produce the set of weights for the current round. This is described in more detail in Sashank J. Reddi et. al. “Adaptive Federated Optimization” , International Conference on Learning Representations (ICLR, 2021).

[095] An aggregated global model may then be obtained from each updated cluster model. For example, to generate the weights for the aggregate global model: ^aggregate #c|ustersSfc ^clusterk Clusters where there are P clusters and Wclusterp is the set of weights for the pth cluster.

[096] At step S230, the activation scores for each client device in the cluster are also aggregated and converted at step S232 to the mask Mcluster. for each cluster. The aggregation and conversion may be done simultaneously, for example using the following: within cluster] MclusterJ K(#clientsinciuster^kSdient,. >T^k clients in j-th cluster where k(-) is the indicator function, k is the number of clients in the cluster, TM is a threshold on the mask and Sclientk is the activation scores of the k-th client. Sclientk is a vector, matrix or tensor of neural activation scores (st) per layer or neuron of the k-th client’s model. Details on how st is calculated are set out below.

[097] When the training procedure is at round t where t >0, a previously generated clusterspecific mask Mcluster. may already exist. When there is such a mask, we can either discard the previous cluster-specific mask and generate a new cluster-specific mask using the equation above or the existing mask(s) can be updated using any appropriate technique, e.g. a moving average approach. As explained above, at each round a different set of client devices is likely to be selected for the federated learning. Accordingly, it may not be ideal to discard the previously generated mask. When updating the previously generated masks, the following method could be used. In a first step, an activation score for each cluster at round t=0 may be stored as = Clients in cluster, Z Sclientk J k When t>0 we compute the “score-difference” -Vciusters from ^cluster, ^kTrjrSciuster, - sciientk vk - 1- - >K clients in j-th cluster where S^ter is the score from the previous round, K is the number of clients in the cluster, and Sciientk is the activation scores of the k-th client. The score Sllllsterj for the round t=0 can be calculated from as follows: ^clusterj — SERV EROPT (yV^sterj ’ ^clusterj’R’ty where SERVEROPTO is a general purpose optimization (weight update) algorithm as described in Sashank J. Reddi et. al. “Adaptive Federated Optimization” , International Conference on Learning Representations (ICLR, 2021). This helps in producing an updated cluster activation scores which is a combinatorial function of the old cluster activation scores and the scores received from the client at the current round.

[098] The next step at step S233 is to determine whether there are any more rounds of federated learning to be completed. If more rounds are required, the method loops back to step S200 in which the next set of client devices are selected for local training data. Steps S200 to S233 are repeated until there is a convergence of accuracy to an acceptable level. This iterative process is schematically illustrated in Figure 6a which shows a global or central server 650 and N client devices 610a, 610b, 61 On. As shown in Figure 6a, at time step t, there is an aggregate global model stored on the central server 650 which is an aggregation of each of the individual cluster models 0-, the aggregate global model may be represented by: 03 = Agg^O^.....0^ Each client device 610a, 610b, 61 On has its own different set of data. At time t, each client device j receives the cluster specific model Of forthat client device and trains the model using its data to generate a new model for time t+1 0-+1. The new models 0j+1 are then returned to the central server 650 at time t+1. The process is repeated until there is no more training to be done.

[099] One outcome of the training process is illustrated in Figure 6b which shows that the global model has a highly complex overall target distribution 670 and is large (perhaps between 0.5 to 200 million parameters). The global model includes a first personalised, pruned model for cluster 1, a second personalised, pruned model for cluster 2 and a third personalised, pruned model for cluster 3. It will be appreciated that three clusters is merely indicative. Each of the personalised models addresses a subdomain of the target domain distribution and thus needs less capacity to tackle it. Each model has its own mean (pn, pi2, and p3) and its own standard deviation (<n, a2, and ct3). There is a relative gain in accuracy of between 1 and 20% due to the cluster specific model weights and typically a faster inference time due to the cluster specific masks. It is typically quicker to implement the pruned model to obtain a result than the aggregate global model. In this example, both a cluster-specific model and a cluster-specific mask is used but it will be appreciated that improvements in performance will also be achieved by using only one of the cluster-specific model or the cluster-specific mask.

[100] Once the central server has finished the federated learning optimisation stage, an output set of masks and weights for each cluster has typically been created. A pruned model for each cluster can then be created at step S234 by applying the mask to the weights, for example by multiplying the mask matrix by the weight matrix. The pruned mask may be represented as ^clusterj ~ Wciusterk ' ^cZuster / ; ■ This pruning step may reduce the size of each cluster model by a fixed amount, for example by between 80 to 99,5%. There is an optional step at step S236 of sending a push notification to all the client devices that the updated (i.e. personalised at a cluster level) model data is ready. Each client device 110 receives the push notification at step S238. There is then an optional check at step S240 to determine whether the client device 110 is idle. If the client device is not idle, there is loop back to wait until the client device is idle. This allows for downloading of data from the central server at convenient times and / or without impacting other applications running on the client device.

[101] When the client device is ready, at step S242, the client device downloads the model data from the central server. The personalised model data may be the cluster-specific mask and the current global ML model or a pruned version in which the cluster-specific mask has been applied to the current global ML model. By pruning at the central server, it is only necessary to send the pruned (i.e. smaller) model to the client devices. When cluster-specific models are also generated, the personalised model data may be the cluster-specific mask and the cluster-specific ML model or a pruned version in which the cluster-specific mask has been applied to the cluster-specific ML model. This may be represented as: ^usert Eclusterk ^clusterk ’ clustery Alternatively, the pruning can be done on the client devices at step S244 and in this case, the client device downloads as model data both the mask and weights for the cluster within which the client device is grouped. This may be represented as: ^user, Ecluster^ ^cluster;,- M-useri Ecluster Mcluster

[102] Once the model data which is personalised at a cluster-specific level has been received, there is a decision at step S246 to see if any further personalisation of the model is to be carried out. There is thus an optional step S248 of fine tuning the model. In other words, step S206 can be repeated to generate a user-level personalization of the model which is for the cluster to which the client device belongs. This can be done by training the cluster specific model on local user data. The fine-tuning can be done without the cluster-specific mask when the model and mask are both received. Such a user-level personalization could lead to further improvements in accuracy and / or inference.

[103] At step S250, the pruned model (whether the cluster specific model or the further personalised user-level model) is then stored. At step S252, the stored model is then applied to a new input. This step could be part of a testing phase or a use (i.e. inference) phase. In either testing or use, the model is used to infer a result whether that is a classification of an image or speech, editing of an image, generating speech etc. The result from the inference of the model is then output.

[104] Figures 5a and 5b provide more detail on the measurement of the neural activation score by the client device. Figure 5a shows an example model comprises a plurality of neurons 502 arranged in a plurality of layers (it will be appreciated that the number of layers and neurons shown is merely illustrative and the model will typically be significantly more complex). Adjacent to each neuron is a bar 504 which is graphically illustrating the sum (or count) of the individual activation scores for the adjacent neuron for each pass of the data. In other words, the count may be calculated using: Si — ’ ^1,n n where ain is the activation of neuron / when passing the n-th input data. The activation may be measured using any standard technique, e.g. after ReLU as introduced in Fukushima, K. “Cognitron: A self-organizing multilayered neural network” on Biological Cybernetics, 20(3), 121-136. or Leaky-ReLU, as introduced in Maas, Andrew L. et al. "Rectifier nonlinearities improve neural network acoustic models." In the International Conference on Machine Learning (ICML). Vol. 30. No. 1.2013 activation functions.

[105] Figure 5b is a flowchart illustrating a method for determining or obtaining the neural activation score. In a first step S510, the current model is applied to each sample in the user data. In other words, a forward pass of the model is done on the samples which are locally stored on the client device. At step S512, the activation value n for each neuron I is measured for each sample n. The individual activation values for each neuron and each sample may be combined in various ways to determine the neural activation score for each neuron. For example, as shown in the first branch, the next step S514 may be calculating the (floating) sum of all activation values for each neuron, e.g. using the equation above. The neural activation score s; for each neuron may then be simply calculated by normalizing at step S516 the sum from the previous calculation by the number of samples. _ Sn al,n Sl “ ~N where the total number of samples is N.

[106] The normalised sum may then be output as the neural activation score at step S518 and optionally stored on the local device at step S540. As an alternative to outputting the normalised sum as the neural activation score, as shown at step S520, the normalised sum may be compared to an overall threshold Tand the neural activation score will be zero (step S522) if the normalised sum is below the threshold and a set value (e.g. one) if the normalised sum is above the threshold (step S524). In this case, the following formula may be used to calculate the neural activation score s{ for each neuron: s; = K[(X„^)>r], where kQ) is the indicator function, aZn is the activation for each neuron / for each sample n and there are N samples overall. Once all neural activation scores are determined, they are output at step S518 and stored at step S540.

[107] As an alternative to calculating a sum of the activation values, after measuring the activation values at step S512, the next step S530 may be to compare each activation value to an activation threshold T (this will be smaller than the overall threshold used in the previous embodiment). If the activation value is above the activation threshold, the count may be incremented, e.g. by one, as shown as step S532. Otherwise, there is no change to the count and the process continues with determining whether there are more samples for each neuron. If yes, steps S530 to S534 are repeated until there are no more samples to consider and the neural activation score can be output at step S518 and stored at step S540. In this case, the following formula may be used to calculate the neural activation score st for each neuron which is a normalized sum of the counts: Sl=----N---- where k(-) is the indicator function which sets the value to zero or one based on the comparison with Tt which is the activation threshold, aln is the activation for each neuron / for each sample n and there are N samples overall.

[108] As explained above, the neural activation scores are used to determine the masks which are applied to prune the models. The overall threshold T and / or the activation thresholds Tt may be set in various ways. For example, the overall threshold T and / or the activation thresholds Tt may be set so that the fewer than a fixed percentage (e.g. 90%) of the neurons will have a summed score or an individual activation score which meets the threshold. Only these neurons will then be represented by a one in the mask matrix and all other neurons will be represented by a zero. The overall threshold T and / or the activation thresholds Tt are set so that the pruned model which is generated is only a fixed percentage of the global model (e.g. 90%). Alternatively, the activation thresholds Tt in a single layer may be set to an arbitrary scalar value x (e.g. 0.2): Ti = x Ml, 0 <I <L where L is the total number of layers.

[109] The pruned, personalised models which are generated using the method of Figures 2a to 2c may be used in a variety of applications, including image processing. Figures 7a to 7c illustrate three different images to which an image processing model, e.g. an image segmentation model and / or an editing model may be implemented. Each of Figures 7a to 7c is an image of a road but they are taken in three different locations to show that the same type of scene in different locations can be appear significantly different. Figure 7a illustrates a countryside, rural road taken for example in Africa. Figure 7b is a street scene in a city such as New York. Figure 7c is an image of a street in a different city, e.g. in Europe. In addition to the same scene looking different, there may be certain objects which only exist in certain regions, e.g. a traffic light of the shape shown in Figure 7c.

[110] As mentioned in the introduction, one method for image processing is image segmentation and it will be appreciated that an image segmentation model which has been personalised and pruned as described above will generate a more accurate segmentation. Figure 8 illustrates some of the steps of using an image segmentation model which has been personalised as described above. As shown in step S800, the first step is to obtain the current local model at the client device. The model may be obtained from a central server by the client device. As explained above, the central server may send a pruned model which is personalised to the cluster to which the client device belongs. Alternatively, the central server may send a model and a mask, both of which are personalised to cluster to which the client device belongs and the model and mask are multiplied together on the client device to obtain the pruned model. The pruned model may optionally be further personalised to the specific client before use.

[111] In a next step S802, an input image is obtained. The obtained input image may be an image which is selected by a user from their photo gallery or an image which has been captured by the user using the camera of the client device. The next step S804 is to use the personalised segmentation model to generate a segmentation map. The segmentation map segments the input image into a plurality of segments. Each segment may be semantically consistent, in other words, each segment may represent a different class, e.g. a different type of object within the image or scenery within the image. Thus, each segment has pixels which belong to the same semantic class. Merely as an example, Figure 9 shows the image in Figure 7c segmented into multiple segments, including for example road 902, background 904, buildings 906, traffic lights 908 and vehicles 910.

[112] Returning to Figure 8, using the segmentation map at least one segment (or segment region and the terms can be used interchangeably) of the image may optionally be selected at step S806. Optionally multiple segmented regions may be selected. This selection may comprise receiving a user input, or the at least one segmented region may be proposed automatically to the user by a recommendation model. The recommendation model may be a pretrained machine learning (ML) model which has been personalised using the method of Figures 2a to 2c. In other words, like the segmentation model, the recommendation model may be a pretrained machine learning (ML) or deep learning (DL) model, such as a convolutional neural network (CNN) or a transformer network, or a combination of both a CNN and a transformer network. The recommendation model and the segmentation model are not typically built using the same specific architecture but both models can be built with convolutional layers, transformer layers and / or fully connected linear layers.

[113] Once the object(s) have been selected, the image may be edited to remove or otherwise adjust the selected object at step S808. The output image may then be output at step S810. The processed image may be stored on the client device and information regarding the processing, e.g. when the editing occurred and what was editing, may be stored with the processed image, e.g. in a log or in metadata. Merely as an example, Figure 10a illustrates the selection of a vehicle in the street scene and its subsequent removal is shown in the example output image of Figure 10b.

[114] Erasing or removing a selected region may be achieved using any suitable technique. For example, erasing or removal may simply be done by masking each of the pixels in the selected region. Masking the pixels leaves a hole in the image which may then be filled using known techniques such as inpainting. These inpainting techniques may use ML or DL models. For example, DL techniques may include CNNs and Transformers. Merely as an example, one inpainting technique is described in “Resolution-robust Large Mask Inpainting with Fourier Convolutions” by Suvorov et al. published in the Winter Conference on Applications of Computer Vision (WACV) in 2022.

[115] Additionally, or alternatively, editing the selected segmented region(s) may involve editing by applying a filter to the selected region(s). The filter may be a semantically aware filter-that is, the filter may be different for different parts of the image depending on what they show. For example, if one selected region contains a person and another selected regions contains greenery, a different filter may be applied to the person’s face in the first selected region than to the greenery in the second selected region. The filter that is applied to the person’s face may, for example, brighten the face if it was shaded and difficult to see, and the filter applied to the greenery may make the greenery look more vivid.

[116] Additionally, or alternatively, the selected region(s) may be processed to improve an image search function on the user’s device. For example, the user may provide a tag or label for the selected region(s) which may be output in step S810 as part of the output image. An image search function may then be used to find this particular tag, for example, object, pet and / or person in other images.

[117] Figure 11 illustrates more detail for a method for erasing an object from within an image. First an input image is received at step S1100 and the image is first processed using a personalised model at step S1102. The personalised model may be a segmentation model which generates a segmentation map of the input image or a boundary detection model which detects boundaries within the image. The next step S1104 is to receive a selection input from the user for the image. Ideally, the selection should be easy to input for the user and may for example be simply a single click on a pixel within the image to indicate a region of the image to be edited.

[118] The next step is identifying the region(s) in the image which corresponds to the user selection at step S1106. The region(s) may be identified using the results of step S1102, e.g. using the segmentation map or based on the detected boundaries. The next step S1108 is to determine whether the user approves of the identification in the previous step. For example, returning to the example in Figure 10a, the user has selected the vehicle (bicycle) and the whole of the vehicle has been correctly identified by the client device (as indicated by the dotted boundary around the vehicle). In this case, the user can approve the selection and the object is removed at step S1114.

[119] Alternatively, as shown in Figure 12a, the rear wheel of the vehicle has been omitted and thus as indicated in Figure 12b, if the object within the dotted boundary is removed, the rear wheel will remain. By using the cluster specific model, it is expected that these errors will be reduced but the method allows for further personalisation of the model on the client device. Returning to Figure 11, when there is such an error, the user indicates a selection by drawing a line around the region to be edited at step S1110. Such a selection is more onerous to the user than the one-click user input suggested above. The user selection is then converted to the boundary of the object (e.g. as shown in Figure 10b). Any appropriate techniques may be used for steps S1110 and S1112. For example, a version of Lasso selection may be used and an example of this technique is described in “ICE-Lasso: An enhanced form of Lasso selection” by Dehmeski et al. published in the IEEE International Conference on Science and Technology for Humanity (2009).

[120] Once the correct object within the image has been identified, the next step S1114 is to remove the object, for example using masking or other techniques as described above. The client device may then check whether the user wishes to make any further edits at step S1116. If there are further edits to be made, the process loops back to the user selection at step S1104. Otherwise, the edited image is output at step S1120. Outputting the edited image may comprise storing the edited image.

[121] Figures 13 to 15 illustrate some alternative uses of a personalised model. Figure 13 is a flowchart for processing an image to generate image captioning. In a first step S1300, an input image is received, for example by selection from a library by a user or by a user taking an image. In this case, the cluster specific model (or further personalised model) which is used is a caption generating model (which may also be termed an image captioning model). As shown at step S1302, the input image is processed using the model to generate a draft caption for the image.

[122] The next step S1304 is to determine whether the user approves of the caption which has been obtained by application of the model. For example, the user reads the caption and can indicate their approval or otherwise using a user interface. When the user can approve the selection, the caption may be output at step S1308. Alternatively, when there is no approval of the caption, the user manually enters the correct caption at step S1304. This caption is output as the result of the process at step S1308. Outputting the caption may comprise outputting an edited image which also comprises the caption. At step S1310, the final caption is stored together with the input image. The stored image and caption form user data which may be used to personalise the cluster specific model.

[123] Figure 14 is a flowchart for processing an image to classify the image, e.g. to generate one or more tags for the image. In a first step S1400, an input image is received, for example by selection from a library by a user or by a user taking an image. In this case, the cluster specific model (or further personalised model) which is used is an image classification model. As shown at step S1402, the input image is processed using the model to generate a draft class or set of classes for the image. A set of classes may be generated, for example, by also using a segmentation model to segment the image and generating a class tag for each segment. Alternatively, a set of classes may be the output for the image as a whole.

[124] The next step S1404 is to determine whether the user approves of the class(es) which has been obtained by application of the model. For example, the user reads the tag(s) for each class and can indicate their approval or otherwise using a user interface. When the user can approve the selection, the class(es) are output at step S1408. Alternatively, when there is no approval, the user manually enters the correct class(es) and / or deletes the incorrect class(es) at step S1404. The manually edited class(es) is output as the result of the process at step S1408. Outputting the class(es) may comprise outputting an edited image which also comprises the class(es). At step S1410, the final class(es) and / or tag(s) is stored together with the input image. The stored image and class(es) and / or tag(s) form user data which may be used to personalise the cluster specific model.

[125] As an alternative or in addition to personalising the cluster specific model, the user data may be used to train a recommendation model which recommends images based on a user input in the form of text and / or an image search model which outputs one or more images in response to a user input in the form of text. The text inputted by the user may correspond to the tag(s) for the class(es) and / or the caption generated using the process of Figure 13. In both cases, the training of these additional models may take place when more than a threshold amount of data has been collected as described above. The training may include generating improved, personalised weights and measuring neuron activation to generate masks as described above. Optionally, the weights and masks can be transferred to the central server in a similar manner to that described in Figures 2a to 2c.

[126] Figure 15 is a flowchart for processing input speech to generate a transcription of the speech. In a first step S1500, input speech is received, for example by recording a voice clip from the user, e.g. using a microphone on the client device or by selection of an audio / video clip from a library of recordings. In this case, the cluster specific model (or further personalised model) which is used is a speech recognition model. As shown at step S1502, the input speech file is processed using the model to generate a draft transcript.

[127] The next step S1504 is to determine whether the user approves of the transcript which has been obtained by application of the model. For example, the user reads the transcript and can indicate their approval or otherwise using a user interface. When the user can approve the selection, the transcript is output at step S1508. Alternatively, when there is no approval, the user manually enters the correct transcript or edits the initial transcript at step S1504. The manual transcript is output as the result of the process at step S1508. Outputting the transcript may comprise outputting a video file which comprises the original speech as audio (either as video or audio file) and the transcript published on screen alongside the original speech. At step S1510, the transcript is stored together with the input speech file. The transcript and input speech file form user data which may be used to personalise the cluster specific model.

[128] As an alternative or in addition to personalising the cluster specific model, the user data may be used to train an on-device offline voice assistant which converts dictated speech to text, such as voice memo. The offline voice assistant can then be used when it is not possible to obtain the cluster specific model from the central server. The training of such an additional machine learning model may take place when more than a threshold amount of user data has been collected as described above. The training may include generating improved, personalised weights for the model itself and measuring neuron activation to generate a mask as described above. Optionally, the weights and mask can be transferred to the central server in a similar manner to that described in Figures 2a to 2c.

[129] At least some of the example embodiments described herein may be constructed, partially or wholly, using dedicated special-purpose hardware. Terms such as ‘component’, ‘module’ or ‘unit’ used herein may include, but are not limited to, a hardware device, such as circuitry in the form of discrete or integrated components, a Field Programmable Gate Array (FPGA) or Application Specific Integrated Circuit (ASIC), which performs certain tasks or provides the associated functionality. In some embodiments, the described elements may be configured to reside on a tangible, persistent, addressable storage medium and may be configured to execute on one or more processors. These functional elements may in some embodiments include, by way of example, components, such as software components, object-oriented software components, class components and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. Although the example embodiments have been described with reference to the components, modules and units discussed herein, such functional elements may be combined into fewer elements or separated into additional elements.

[130] Various combinations of optional features have been described herein, and it will be appreciated that described features may be combined in any suitable combination. In particular, the features of any one example embodiment may be combined with features of any other embodiment, as appropriate, except where such combinations are mutually exclusive. Throughout this specification, the term “comprising” or “comprises” means including the component(s) specified but not to the exclusion of the presence of others.

[131] Attention is directed to all papers and documents which are filed concurrently with or previous to this specification in connection with this application and which are open to public inspection with this specification, and the contents of all such papers and documents are incorporated herein by reference. All of the features disclosed in this specification (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of such features and / or steps are mutually exclusive.

[132] Each feature disclosed in this specification (including any accompanying claims, abstract and drawings) may be replaced by alternative features serving the same, equivalent or similar purpose, unless expressly stated otherwise. Thus, unless expressly stated otherwise, each feature disclosed is one example only of a generic series of equivalent or similar features. The invention is not restricted to the details of the foregoing embodiment(s). The invention extends to any novel one, or any novel combination, of the features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any novel one, or any novel combination, of the steps of any method or process so disclosed. 12 01 24

Claims

1. A computer-implemented method for processing an input image on a client device, the method comprising:receiving, from a server, personalised machine learning, ML, model data, wherein the personalised ML model data is for image processing;obtaining, at the client device, an input image;processing the input image using the personalised ML model data to generate a segmentation map of the input image which comprises a plurality of segments, and outputting the segmentation mapwherein generating the personalised ML model data is done by a server connected to a plurality of client devices and the generating comprises:obtaining, using the server, an initial global ML model in the form of a neural network comprising a plurality of neurons each having a global model weight;using at least one round of federated learning to train the obtained global ML model, wherein each round of federated learning comprises:selecting, using the server, multiple client devices to be used for federated learning;distributing an ML model from the server to each of the selected client devices for local training, wherein in a first round of federated learning the distributed ML model is the initial global ML model and in subsequent rounds of federated learning the distributed ML model is an updated ML model generated in a previous round of federated learning;receiving, at the server from each selected client device, a plurality of local model weights generated by the local training of the distributed ML model and a neural activation score for each neuron;grouping, using the server, the multiple client devices into a plurality of clusters;generating, using the server, personalised ML model data comprising a cluster-specific mask for each cluster using the neural activation scores received from each client device in the cluster, wherein each mask specifies which neurons within the ML model are to be used; andgenerating, using the server, at least one updated ML model based on the received plurality of local model weights; and12 01 24when the federated learning has been completed notifying, using the server, the client devices that personalised ML model data is available.

2. The method as claimed in claim 1, further comprising, when the federated learning has been completed, applying, using the server, each cluster-specific mask to the updated ML model to obtain a plurality of pruned models each of which is personalised to the respective cluster and outputting, to each client device in a particular cluster, the personalised pruned model for the particular cluster.

3. The method as claimed in claim 1 or claim 2 wherein generating the at least one updated ML model comprisesgenerating, using the server, a cluster-specific ML model for each cluster using the plurality of local model weights received from each client device in the cluster.

4. The method as claimed in claim 3, comprisingdetermining, using the server, whether the selected client device has been assigned to a cluster andwhen the selected client device has previously been assigned to a cluster, distributing, from the server to the selected client, the cluster-specific model generated in a previous round of federated learning.

5. The method of claim 3 or claim 4, comprising generating each cluster-specific ML model by aggregating, using the server, the plurality of local model weights received from all client devices in the cluster.

6. The method of claim 5, wherein generating the at least one updated M L model comprises generating, using the server, an updated global ML model by aggregating each cluster-specific model.

7. The method of any one of claims 1 to 5, wherein generating the at least one updated ML model comprises generating, using the server, an updated global ML model by aggregating the plurality of local model weights received from all client devices.

8. The method of any one of the preceding claims, wherein selecting, using the server, multiple client devices to be used for federated learning comprises12 01 24determining how many samples are present in a training dataset for each of the plurality of client devices andselecting client devices which have more than a minimum threshold of training samples.

9. The method of any one of claims 1 to 7, wherein selecting multiple client devices to be used for federated training comprises prioritising, using the server, client devices which have not previously been used for training.

10. The method of any one of the preceding claims, comprising grouping, using the server, the client devices using labels by obtaining a label for each client device, obtaining a cluster for each obtained label, and assigning each client device to the cluster corresponding to the label obtained from that client device.

11. The method of any one of claims 1 to 10, comprising grouping, using the server, the client devices bycomparing the neural activation scores received from each client device, andgrouping client devices into a cluster when the differences between the neural activation scores for the grouped client devices are below an activation difference threshold.

12. The method of any one of claims 1 to 10, comprising grouping, using the server, by comparing the local model weights received from each client device, and grouping client devices into a cluster when the differences between the local model weights for the grouped client devices are below a weight difference threshold.

13. The method of any one of the preceding claims, comprising obtaining the neural activation score at each client device byobtaining, at each client device, a set of test samples;for each test sample in the set of test samplesapplying the ML model which is received from the server to the test sample; and measuring, when applying the ML model to the test sample, an activation value for each neuron in the neural network of the ML model; andcombining the activation values for each test sample and for each neuron to obtain the neural activation score.12 01 2414. The method of claim 13, wherein the neural activation score for each neuron is calculated by summing, for each neuron, all the measured activation values for each sample.

15. The method of claim 13, wherein the neural activation score for each neuron is calculated by summing, for each neuron, all the measured activation values for each sample and normalising the summed activation values.

16. The method of claim 13, wherein the neural activation score for each neuron is calculated bysumming, for each neuron, all the measured activation values for each sample ;normalising the summed activation values;comparing the normalised sum to an overall threshold value, andwhen the normalised sum for a neuron is above the overall threshold value, outputting a neural activation score of one andwhen the normalised sum for a neuron is below the overall threshold value, outputting a neural activation score of zero.

17. The method of claim 13, wherein the neural activation score for each neuron is calculated bycomparing each measured activation value to an activation threshold;for each measured activation value above the activation threshold, incrementing the neural activation score by a fixed amount; andnormalising the incremented neural activation score by the number of test samples.

18. The method of any preceding claim, comprisingselecting at least one segment from the plurality of segments, further processing the input image to generate a processed image by applying at least one of the followingerasing the at least one selected segment, applying a filter to the at least one selected segment, andproviding a tag for the at least one selected segment to be output with the processed image, andoutputting the processed image.

19. A computer-implemented method for processing speech on a client device, the method comprising:12 01 24receiving, from the server, personalised machine learning, ML, model data, wherein the personalised ML model data is for converting speech to text;obtaining, at the client device, an input speech segment; andprocessing the input speech segment using the personalised ML model data to generate a text transcript; andwherein generating the personalised ML model data is done by a server connected to a plurality of client devices and the generating comprises:obtaining, using the server, an initial global ML model in the form of a neural network comprising a plurality of neurons each having a global model weight;using at least one round of federated learning to train the obtained global ML model, wherein each round of federated learning comprises:selecting, using the server, multiple client devices to be used for federated learning;distributing an ML model from the server to each of the selected client devices for local training, wherein in a first round of federated learning the distributed ML model is the initial global ML model and in subsequent rounds of federated learning the distributed ML model is an updated ML model generated in a previous round of federated learning;receiving, at the server from each selected client device, a plurality of local model weights generated by the local training of the distributed ML model and a neural activation score for each neuron;grouping, using the server, the multiple client devices into a plurality of clusters;generating, using the server, personalised ML model data comprising a cluster-specific mask for each cluster using the neural activation scores received from each client device in the cluster, wherein each mask specifies which neurons within the ML model are to be used; andgenerating, using the server, at least one updated ML model based on the received plurality of local model weights; andwhen the federated learning has been completed notifying, using the server, the client devices that personalised ML model data is available.

20. A system comprising a server which is connected to a plurality of client devices, wherein12 01 24the server comprises a processor for generating personalised machine learning, ML, model data by:obtaining, using the server, an initial global ML model in the form of a neural network comprising a plurality of neurons each having a global model weight;using at least one round of federated learning to train the obtained global ML model, wherein each round of federated learning comprises:selecting, using the server, multiple client devices to be used for federated learning;distributing an ML model from the server to each of the selected client devices for local training, wherein in a first round of federated learning the distributed ML model is the initial global ML model and in subsequent rounds of federated learning the distributed ML model is an updated ML model generated in a previous round of federated learning;receiving, at the server from each selected client device, a plurality of local model weights generated by the local training of the distributed ML model and a neural activation score for each neuron;grouping, using the server, the multiple client devices into a plurality of clusters;generating, using the server, personalised ML model data comprising a cluster-specific mask for each cluster using the neural activation scores received from each client device in the cluster, wherein each mask specifies which neurons within the ML model are to be used; andgenerating, using the server, at least one updated ML model based on the received plurality of local model weights; andwhen the federated learning has been completed notifying, using the server, the client devices that personalised ML model data is available;and wherein each client device comprises a processor for processing an input image on the client device, by:receiving, from the server, personalised ML model data, wherein the personalised ML model data is for image processing;obtaining an input image;processing the input image using the personalised ML model data to generate a segmentation map of the input image which comprises a plurality of segments, and outputting the segmentation map.12 01 2421. A system comprising a server which is connected to a plurality of client devices, wherein the server comprises a processor for generating personalised machine learning, ML, model data by:obtaining, using the server, an initial global ML model in the form of a neural network comprising a plurality of neurons each having a global model weight;using at least one round of federated learning to train the obtained global ML model, wherein each round of federated learning comprises:selecting, using the server, multiple client devices to be used for federated learning;distributing an ML model from the server to each of the selected client devices for local training, wherein in a first round of federated learning the distributed ML model is the initial global ML model and in subsequent rounds of federated learning the distributed ML model is an updated ML model generated in a previous round of federated learning;receiving, at the server from each selected client device, a plurality of local model weights generated by the local training of the distributed ML model and a neural activation score for each neuron;grouping, using the server, the multiple client devices into a plurality of clusters;generating, using the server, personalised ML model data comprising a cluster-specific mask for each cluster using the neural activation scores received from each client device in the cluster, wherein each mask specifies which neurons within the ML model are to be used; andgenerating, using the server, at least one updated ML model based on the received plurality of local model weights; andwhen the federated learning has been completed notifying, using the server, the client devices that personalised ML model data is available;and wherein each client device comprises a processor for processing speech on the client device, by:receiving, from the server, personalised ML model data, wherein the personalised ML model data is for converting speech to text;obtaining, at the client device, an input speech segment; andprocessing the input speech segment using the personalised ML model data to generate a text transcript.

Citation Information

Patent Citations

  • Method, system and apparatus for federated learning

    US20220245459A1