Method for federated analytics

The method addresses privacy and efficiency issues in federated analytics by using a machine learning model to remove noise from geospatial heatmaps, ensuring privacy and reducing communication complexity while maintaining data accuracy.

GB2701164APending Publication Date: 2026-04-22SAMSUNG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-05-08
Publication Date
2026-04-22

AI Technical Summary

Technical Problem

Existing federated analytics methods face challenges in efficiently exchanging and analyzing geospatial heatmaps while preserving user privacy and maintaining data accuracy, often requiring high communication complexity and compromising privacy for better results.

Method used

A method using a trained machine learning model to identify and remove excess noise from aggregated geospatial heatmaps, ensuring privacy preservation by learning the underlying data structure without reconstructing original user data, and optimizing communication by sending residuals rather than full resolution images.

Benefits of technology

Enables accurate and privacy-preserving federated analytics by effectively removing noise from geospatial heatmaps, reducing communication overhead, and maintaining data integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for federated analytics, FA. to enable implementation of an FA process using geospatial heatmaps that is both privacy preserving and data efficient using a machine learning model that learns
Need to check novelty before this filing date? Find Prior Art

Description

Field

[001] The present techniques generally relate to a method for federated analytics, FA. In particular, the present techniques provide a communication efficient and privacy preserving method for exchanging and analysing geospatial heatmaps. Background

[002] Federated Analytics, FA, is a principle for collaboratively extracting insights from the data of multiple remote users without sharing the data among users and while preserving the privacy of each user. Thus, FA makes it possible to derive analytic insights from decentralised data without associated privacy risks.

[003] Location data collected on user devices can be particularly sensitive data. However, such data can also be very useful to derive a number of insights for the user, in particular when using applications on a user mobile device that can be tailored depending on the user’s location. Since the data is sensitive, it is important to use privacy-preserving techniques, without compromising the accuracy of geospatial heatmaps that are being used. Additionally, federated analytics can be associated with high-communication complexity, i.e. a large amount of data may need to be shared between a server and a user device.

[004] The present applicant has identified the need for an improved technique for federated analytics that is suitable for obtaining and analysing geospatial heatmaps. Summary

[005] In a first approach of the present techniques, there is provided a computer-implemented method for performing federated analytics, FA at a server, the method comprising: receiving a noisy geospatial heatmap from a plurality of user devices, each noisy geospatial heatmap representing intensity of a specific type of user activity performed by users of the user devices, wherein each noisy geospatial heatmap contains noise to preserve user privacy; generating an aggregated heatmap for the specific type of user activity by aggregating the noisy geospatial heatmaps; removing excess noise from the aggregated heatmap by: identifying structural relationships in the aggregated noisy heatmap using a trained machine learning, ML model; and removing excess noise from the aggregated heatmap using the identified structural relationships to obtain an enhanced aggregated heatmap; and using the enhanced aggregated heatmap to perform federated analytics for the type of user activity.

[006] A geospatial heatmap is a way of collating and visualising user activity data and linking the user activity data to corresponding locations. The geospatial heatmap may be a two-dimensional visualisation of data that uses colour gradients or colour / pixel intensity to represent data points over a geographical area. It allows patterns across locations to be visualised and understood. This means that a geospatial heatmap is an image which represents intensity of a specific type of user activity performed by users of the user devices. For example, each pixel in the image which is the geospatial heatmap may correspond to a certain location, or real-world geographical area. Then, the value of each pixel in the image which is the geospatial heatmap may represent intensity of a user activity. For example, higher pixel values may correspond to higher level of user activity. The geospatial heatmap may be a monochromatic image. That is, there may only be one value associated with each pixel in the heatmap image and this value may represent intensity of a user activity. Alternatively, several values may be associated with each pixel value in the heatmap image. When there are several values associated with each pixel value, each of these values may, for example, correspond to a different user activity. In this way, one heatmap may encode information about the intensity of several user activities at the same time.

[007] Advantageously, the present techniques enable a server to receive heatmaps from users that preserve data privacy for users, but also permit accurate data analysis to be performed by the server. This is because the server is able to remove some of the noise added into the heatmaps to preserver data privacy. For example, if a user device adds noise to a generated heatmap, such noise may be equally distributed across the heatmap, even at locations where it is unlikely or not possible for a user to be present. For example, it is much less likely that a user will have spent time in, for example, a river and thus, location data at this point is likely to be noise and can be filtered out by the trained machine learning model. However, it is important to note that the machine learning model has no information about the underlying geography of a heatmap, i.e. the machine learning does not know that a data point is located in a river. Rather, the ML model has learned an underlying data structure that enables it to recognise noisy data and differentiate noisy data from data that follows a pattern consistent with the underlying structure of the data.

[008] Since the ML model only learns the underlying data structure, any user data will still remain private. That is, the ML model is trained to remove noise that does not conform with the underlying data structure, but it is not trained to remove all noise. Noise which correlates to where data would be expected will remain, thus ensuring the privacy of the original user data.

[009] Additionally, using an ML model which has learned the expected structure of data means that the ML model does not need to be trained on non-private user data. For example, the ML model can be trained by adding additional noise to already anonymised heatmaps. The original noisy heatmap then is the ground truth, which the ML model tries to reach by learning structural information and removing at least some of the additional noise.

[010] Processing the aggregated heatmap using an ML model may comprise using a differentially private ML model that protects privacy of individual data points. That is, the ML model is such that it is unable to reconstruct the original user data. This is consistent with an ML model which learns the underlying structure of data. Suitable models include deep learning, DL, models which comprise a trainable set of parameters. For example, one type of suitable DL model may be a U-net model (Ronneberger O, Fischer P, Brox T (2015). "U-Net: Convolutional Networks for Biomedical Image Segmentation").

[011] Prior to receiving a noisy geospatial heatmap from a plurality of user devices, the method may comprise: sending, to each user device, a number of communication rounds to use when sending their noisy geospatial maps to the server.

[012] In one example, the number of communication rounds to use is one. In this case, receiving a noisy geospatial heatmap from a plurality of user devices may comprise: receiving, during a single communication round, a full resolution noisy geospatial heatmap from each user device of the plurality of user devices. That is, each user device uses a single communication round to send a single noisy geospatial heatmap. The heatmap sent in this example is a full resolution heatmap. This reduces the number of communication rounds, and ensures that the server receives a high resolution heatmap which improves the accuracy of the federated analytics. However, this comes at the cost of having to send a very high resolution heatmap, which may reduce the possibility of some user devices from being able to send their heatmaps.

[013] Full resolution means that the heatmap includes all detail available on the user device and is required at a resolution to show all of this detail. Alternatively, full resolution may mean a maximum resolution determined by the server at which an aggregated heatmap is to be constructed. Resolution may refer to number of pixels per real-world geographical area, wherein each pixel in the geospatial heatmap covers a certain real-world geographical area. All location data available for this pixel may then be pooled or averaged together to obtain a pooled heatmap value for the real-world geographical area.

[014] Thus, in another example, the number of communication rounds to use is at least two. In this case, receiving a noisy geospatial heatmap from a plurality of user devices may comprise: receiving, during a first communication round, a first low resolution noisy heatmap from each user device; and receiving, during at least one further communication round, a residual for at least one higher resolution noisy heatmap from each user device, wherein the residual represents a difference between a lower resolution noisy heatmap (corresponding to the previous communication round) and a higher resolution noisy heatmap (corresponding to the current communication round).

[015] That is, instead of sending a full resolution heatmap (e.g. 1024 x 1024 pixels), the user device sends a first low resolution version of the heatmap (e.g. 32 x 32 pixels) during a first communication round. However, the server still needs as much data as possible to perform federated analytics. Thus, during the at least one further communication round, the user device generates a higher resolution heatmap, which has a higher resolution than the heatmap associated with a previous communication round. However, instead of sending each higher resolution heatmap, the user device sends a residual. In this way, although multiple communication rounds are used to send the data, since only the residual is sent for each further communication round after the first round, fewer bytes are needed overall.

[016] When multiple communication rounds are used, receiving a noisy geospatial heatmap from a plurality of user devices may further comprise: generating, using the ML model, a full resolution noisy geospatial heatmap for each user device using the first low resolution noisy heatmap and the at least one residual. That is, the ML model generates a full resolution noisy geospatial heatmap for each user device, using the first low resolution noisy heatmap and each residual received from the user device. This ensures the noise removal step is performed on the full resolution noisy geospatial heatmap.

[017] Receiving at least one residual may comprise: receiving, during each further communication round, a residual between a higher resolution noisy heatmap and a lower resolution noisy heatmap.

[018] In one example, during each further communication round, the higher resolution noisy heatmap may have a resolution which is four times greater than the resolution of the lower resolution noisy heatmap. In a standard multiple communication round approach, the resolution between a first and second communication round may be doubled. For example, if the resolution of a first image is 32x32 pixels, the resolution of the second image may be 64x64 pixels, the resolution of the third image may be 128x128 pixels and so on. In an approach to the present techniques, the second resolution may be skipped and only the first image and a residual between the first and third image may be sent, i.e. a residual between the 32x32 pixels image (generated and sent in the previous communication round) and the 128x128 pixels image (generated for this current communication round). In the conventional approach, the residual between the 32x32 pixels image and the 64x64 pixels image would also be sent. Thus, the present techniques may be used to reduce the number of communication rounds and the amount of data sent from user device to server. This is possible because the ML model may be used to enhance to received image, even with the reduced amount of data available. Thus, the ML model may also act as a super resolution model and effectively able to upscale an image from reduced data.

[019] Prior to receiving a noisy geospatial heatmap from the plurality of user devices, the method may comprise: requesting, from each user device, a geospatial heatmap according to at least one heatmap parameter. For example, the at least one heatmap parameter may comprise any of: a region of interest, a time interval, and / or a target application for creating the geospatial heatmap. Requesting heatmaps according to at least one heatmap parameter may ensure that only relevant heatmaps are received. For example, the server may determine that for a certain geographical region or for a certain user activity, data is missing or that there is not sufficient data. In this case, the server may request data for these areas and / or activities only. Alternatively, the server may request heatmap data for a certain timeframe only. For example, the server may only want data from the last week, or month, or another time frame. In this case, the server may request all location data the user devices have for this timeframe.

[020] Generating an aggregated heatmap may comprise aggregating the received noisy geospatial heatmaps according to the at least one heatmap parameter. That is, heatmaps with the same heatmap parameters may be aggregated for only. For example, only heatmaps from the same timeframe, location and / or activity may be aggregated together. In this way, there may be several aggregated heatmaps stored on the server, each for a different set of parameters. For example, there may be an aggregated heatmap for the month of May, covering London and only coffee shops, i.e. the user activity may be visiting coffee shops. Similarly, there may be corresponding heatmaps for other timeframes, areas and activities.

[021] Once the enhanced aggregated heatmap has been generated, it can be used to perform federated analytics. For example, the federated analytics performed may reveal trends in user activity in different locations. This can then be used to provide recommendations to users when they are in a particular location. For example, when a user is in west London, England and is searching, via their user device, for a coffee shop to visit, the server may be able to provide a response (e.g. “try Java Junction” on Main Road”) that is based on the enhanced aggregated heatmap for that activity in that location. In another example, when a user is in Gangnam, South Korea and is applying a filter to an image they are taking / have taken, the server may be able to recommend a filter that is based on popular filters used by others in the same location.

[022] Thus, the method may further comprise: receiving a request for a recommendation from a user of a user device in relation to a specific user activity, together with the location of the user device; and providing a recommendation to the user of the user device based on the enhanced aggregated heatmap for the received location and the received specific user activity.

[023] In a second approach of the present techniques, there is provided a computer-implemented method fortraining a machine learning, ML, model to enhance a noisy geospatial heatmap, the method comprising: obtaining a plurality of noisy geospatial heatmaps, the heatmaps representing intensity of at least one user activity; adding noise to each of the noisy geospatial heatmaps; and training the ML model to remove the added noise from the noisy geospatial heatmaps.

[024] Advantageously, using such a machine learning model means that user data can be kept completely private. While the ML model is trained to remove data, it is never trained on non-anonymised user data, meaning that it is not trained to reconstruct original nonanonymised user data. This is a further safeguard to ensure that privacy is maintained at all times. Additionally, this ML model is suitable for enhancing several types of noisy heatmaps. For example, it can be used as in the first approach above to enhance an aggregated heatmap and remove any superfluous noise from this. The ML model can also be used, as described with reference to the first approach, to effectively perform a super resolution task and reconstruct a noisy heatmap from less data, for example, when a residual between a higher and lower resolution noisy heatmap is transmitted to a server, wherein the difference in resolution between the higher and lower resolution heatmaps is great. For example, the difference in resolution may be fourfold.

[025] Training the ML model to remove the added noise from the noisy geospatial heatmaps may comprise training the ML model to recognise structural relationships in the noisy heatmaps and to remove the added noise using the recognised structural relationships.

[026] In particular, the ML model is in this way trained to learn structural relationships in noisy heatmaps, which are independent of added noise. However, the ML model is not able to reconstruct user data, as a more conventional super resolution model may be taught to. That is, the ML observes differential privacy principles.

[027] Training the ML model to recognise structural relationships in the noisy heatmaps may comprise training the model to learn spatial and / or temporal relationships in the noisy heatmaps. That is, the ML model may learn patterns in both geography and timing in the noisy heatmaps and use these patterns to filter out noisy data points that don’t fit these patterns.

[028] Training the ML model may comprise training the ML model for a fixed number of iterations to generalise the ML model. Early stopping techniques may be used to determine the optimal number of iterations to arrive at a ML model which has learned the required structural information but is not overfitting, i.e. is still applicable to finding a general solution.

[029] Adding noise to each of the noisy geospatial heatmaps may comprise adding noise that is drawn from a different probability distribution than the noise in the noisy geospatial heatmaps. Advantageously, it is therefore not necessary to know which probability distribution the noise in the noisy heatmaps is drawn from. The ML learns structural relationships, independent of the noise used in training.

[030] Alternatively, adding noise to each of the noisy geospatial heatmaps may comprise adding noise that is drawn from the same probability distribution than the noise in the noisy geospatial heatmaps.

[031] In a third approach of the present techniques, there is provided a computer-implemented method for performing federated analytics, FA, at a user device, the method comprising: receiving, from a server, a request for a geospatial heatmap; generating a geospatial heatmap, using data items stored on the user device, the heatmap representing intensity of a specific type of user activity performed by a user of the user device; adding noise to the generated geospatial heatmap to anonymise the heatmap; and sending the noisy geospatial heatmap to the server for use in performing federated analytics.

[032] Adding noise may comprise: sampling random noise from a Laplacian Distribution; and adding the sampled random noise to the generated heatmap.

[033] Sampling random noise from a Laplacian Distribution may comprise obtaining a privacy parameter e which determines a variance of the Laplacian Distribution. The privacy parameter may also be termed a privacy budget. The privacy parameter may be the inverse of the variance of the Laplacian Distribution.

[034] Sampling random noise from a Laplacian Distribution may comprise sampling random noise from a Laplacian Distribution according to a Discrete Laplacian Mechanism, or according to a Geo-Indistinguishability method.

[035] The step of receiving, from a server, a request for a geospatial heatmap may comprise: receiving, from the server, a number of communication rounds to use when sending the geospatial heatmap. The server may also specify the resolution for each communication round. Alternatively, the user device may infer the resolutions from the number of communication rounds. For example, a full resolution of the noisy geospatial heatmap may be standardised and saved to the user device. The same may be true for a low resolution version of the noisy heatmap. That is the user device may know which resolution to start from and which resolution to end at, and may infer the resolutions at each communication round from these two resolutions.

[036] When the number of communication rounds is one, generating a geospatial heatmap may comprise generating a full resolution geospatial heatmap.

[037] When the number of communication rounds is at least two, generating a geospatial heatmap may comprise: generating a full resolution geospatial heatmap; generating, from the full resolution geospatial heatmap, a first low resolution heatmap; and generating, using the full resolution geospatial heatmap, a higher resolution geospatial heatmap for each further communication round, wherein the resolution of the higher resolution geospatial heatmap created for each further communication round is higher than the first low resolution heatmap and increases for each successive communication round.

[038] Adding noise to the generated geospatial heatmap may comprise adding noise to the first low resolution heatmap and each higher resolution geospatial heatmap generated for each further communication round.

[039] Sending the noisy heatmap to a server may comprise: sending, during a first communication round, the first low resolution noisy heatmap; calculating a residual between each higher resolution noisy geospatial heatmap and the noisy heatmap corresponding to the previous communication round; and sending, during each further communication round, the calculated residual.

[040] In a fourth approach of the present techniques, there is provided a server for federated analytics, FA, the server comprising at least one processor and memory storing instructions that, when executed by the at least one processor individually or collectively, cause the server to: receive a noisy geospatial heatmap from a plurality of user devices, each noisy geospatial heatmap representing intensity of a specific type of user activity performed by users of the user devices, wherein each noisy geospatial heatmap contains noise to preserve user privacy; generate an aggregated heatmap for the specific type of user activity by aggregating the noisy geospatial heatmaps; remove excess noise from the aggregated heatmap by: identify structural relationships in the aggregated noisy heatmap using a trained machine learning, ML model; and remove excess noise from the aggregated heatmap using the identified structural relationships to obtain an enhanced aggregated heatmap; and using the enhanced aggregated heatmap to perform federated analytics for the type of user activity.

[041] The features described above with respect to the first approach apply equally to the fourth approach and therefore, for the sake of conciseness, are not repeated.

[042] In a fifth approach of the present techniques, there is provided a server for training a machine learning, ML, model to enhance a noisy geospatial heatmap, the server comprising at least one processor and memory storing instructions that, when executed by the at least one processor individually or collectively, cause the server to: obtain a plurality of noisy geospatial heatmaps, the heatmaps representing intensity of at least one user activity; add noise to each of the noisy geospatial heatmaps; and train the ML model to the remove the added noise from the noisy geospatial heatmaps.

[043] The features described above with respect to the second approach apply equally to the fifth approach and therefore, for the sake of conciseness, are not repeated.

[044] In a sixth approach of the present techniques, there is provided a user device for federated analytics, FA, the user device comprising at least one processor and memory storing instructions that, when executed by the at least one processor individually or collectively, cause the user device to: receive, from a server, a request for a geospatial heatmap; generate a geospatial heatmap, using data items stored on the user device, the heatmap representing intensity of a specific type of user activity performed by a user of the user device; add noise to the generated geospatial heatmap to anonymise the heatmap; and send the noisy geospatial heatmap to the server for use in performing federated analytics.

[045] The features described above with respect to the third approach apply equally to the sixth approach and therefore, for the sake of conciseness, are not repeated.

[046] The user device may be a smart device. The user device may be a smartphone. A smartphone is an example of a smart device. The user device may be a smart appliance. A smart appliance is another example of a smart device. An example of a smart appliance is a smart television (TV), a smart fridge, a smart oven, a smart vacuum cleaner, a smart robotic device, a smart lawn mower, and so on. More generally, the user device may be a constrained-resource device, but which has the minimum hardware capabilities to perform the federated analytics task described above. The user device may be any one of: a smartphone, tablet, laptop, computer or computing device, virtual assistant device, a vehicle, an autonomous vehicle, a robot or robotic device, a robotic assistant, image capture system or device, an augmented reality system or device, a virtual reality system or device, a gaming system, an Internet of Things device, or a smart consumer device (such as a smart fridge, smart vacuum cleaner, smart lawn mower, smart oven, etc). It will be understood that this is a non-exhaustive and non-limiting list of example devices.

[047] In a related approach of the present techniques, there is provided a computer-readable storage medium comprising instructions which, when executed by at least one processor, causes the processor to carry out any of the methods described herein.

[048] In the cases where the present techniques are implemented or executed on a device comprising multiple processors, the present techniques may be implemented by one or more of the multiple processors. That is, the present techniques may be implemented by or executed by the processors individually or collectively.

[049] As will be appreciated by one skilled in the art, the present techniques may be embodied as a system, method or computer program product. Accordingly, present techniques may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects.

[050] Furthermore, the present techniques may take the form of a computer program product embodied in a computer readable medium having computer readable program code embodied thereon. The computer readable medium may be a computer readable signal medium or a computer readable storage medium. A computer readable medium may be, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing.

[051] Computer program code for carrying out operations of the present techniques may be written in any combination of one or more programming languages, including object oriented programming languages and conventional procedural programming languages. Code components may be embodied as procedures, methods or the like, and may comprise subcomponents which may take the form of instructions or sequences of instructions at any of the levels of abstraction, from the direct machine instructions of a native instruction set to high-level compiled or interpreted language constructs.

[052] Embodiments of the present techniques also provide a non-transitory data carrier carrying code which, when implemented on a processor, causes the processor to carry out any of the methods described herein.

[053] The techniques further provide processor control code to implement the abovedescribed methods, for example on a general purpose computer system or on a digital signal processor (DSP). The techniques also provide a carrier carrying processor control code to, when running, implement any of the above methods, in particular on a non-transitory data carrier. The code may be provided on a carrier such as a disk, a microprocessor, CD- or DVD-ROM, programmed memory such as non-volatile memory (e.g. Flash) or read-only memory (firmware), or on a data carrier such as an optical or electrical signal carrier. Code (and / or data) to implement embodiments of the techniques described herein may comprise source, object or executable code in a conventional programming language (interpreted or compiled) such as Python, C, or assembly code, code for setting up or controlling an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), or code for a hardware description language such as Verilog (RTM) or VHDL (Very high speed integrated circuit Hardware Description Language). As the skilled person will appreciate, such code and / or data may be distributed between a plurality of coupled components in communication with one another. The techniques may comprise a controller which includes a microprocessor, working memory and program memory coupled to one or more of the components of the system.

[054] It will also be clear to one of skill in the art that all or part of a logical method according to embodiments of the present techniques may suitably be embodied in a logic apparatus comprising logic elements to perform the steps of the above-described methods, and that such logic elements may comprise components such as logic gates in, for example a programmable logic array or application-specific integrated circuit. Such a logic arrangement may further be embodied in enabling elements for temporarily or permanently establishing logic structures in such an array or circuit using, for example, a virtual hardware descriptor language, which may be stored and transmitted using fixed or transmittable carrier media.

[055] In an embodiment, the present techniques may be realised in the form of a data carrier having functional data thereon, said functional data comprising functional computer data structures to, when loaded into a computer system or network and operated upon thereby, enable said computer system to perform all the steps of the above-described method.

[056] The method described above may be wholly or partly performed on an apparatus, i.e. an electronic device, using a machine learning or artificial intelligence model. The model may be processed by an artificial intelligence-dedicated processor designed in a hardware structure specified for artificial intelligence model processing. The artificial intelligence model may be obtained by training. Here, "obtained by training" means that a predefined operation rule or artificial intelligence model configured to perform a desired feature (or purpose) is obtained by training a basic artificial intelligence model with multiple pieces of training data by a training algorithm. The artificial intelligence model may include a plurality of neural network layers. Each of the plurality of neural network layers includes a plurality of weight values and performs neural network computation by computation between a result of computation by a previous layer and the plurality of weight values.

[057] As mentioned above, the present techniques may be implemented using an Al model. A function associated with Al may be performed through the non-volatile memory, the volatile memory, and the processor. The processor may include one or a plurality of processors. At this time, one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and / or an Al-dedicated processor such as a neural processing unit (NPU). The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (Al) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning. Here, being provided through learning means that, by applying a learning algorithm to a plurality of learning data, a predefined operating rule or Al model of a desired characteristic is made. The learning may be performed in a device itself in which Al according to an embodiment is performed, and / o may be implemented through a separate server / system.

[058] The Al model may consist of a plurality of neural network layers. Each layer has a plurality of weight values, and performs a layer operation through calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.

[059] The learning algorithm is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. Brief description of the drawings

[060] Implementations of the present techniques will now be described, by way of example only, with reference to the accompanying drawings, in which:

[061] Figure 1A is an example geospatial heatmap;

[062] Figure 1B shows at what point in the federated analytics pipeline different issues might arise;

[063] Figure 2 shows an overview of the federated analytics process according to the present techniques;

[064] Figure 3 is a flowchart for a method performed by a server at inference time according to the present techniques;

[065] Figure 4 illustrates how conventional communication rounds, and communication rounds in the present techniques differ;

[066] Figures 5A and 5B show algorithms for implementing the machine learning, ML, model of the present techniques;

[067] Figure 6 is a flowchart for a method performed by a server at training time according to the present techniques;

[068] Figure 7 is a flowchart for a method performed by a user device according to the present techniques;

[069] Figures 8 to 11 show results obtained using the present techniques; and

[070] Figure 12 is a block diagram illustrating a server / user device set-up according to the present techniques. Detailed description of the drawings

[071] Broadly speaking, the present techniques generally relate to a method for federated analytics (FA). Advantageously, the present techniques enable implementation of an FA process for geospatial heatmaps that is both privacy preserving and data efficient using a machine learning model that learns the underlying structure of noisy geospatial heatmaps. In particular, the present techniques provide a method for using federated analytics that is able to extract insights from aggregated geospatial data without impacting accuracy and while preserving privacy.

[072] Figure 1A is an example heatmap. In particular, a geospatial heatmap encodes an intensity of user activities at a certain location on a two-dimensional map. Areas that are shown in orange and brighter colours indicate higher intensity of activity at these locations. I.e. the more time a user spends at a certain location, the more likely the location is to be shown in increasingly brighter orange colour. Additionally, the more often a user visits a certain location, the more likely the location is to be shown in increasingly brighter colour.

[073] A geospatial heatmap is a way of collating and visualising user activity data and linking the user activity data to corresponding locations. This means that a geospatial heatmap is an image which represents intensity of a specific type of user activity performed by users of the user devices. For example, each pixel in the image which is the geospatial heatmap may correspond to a certain location, or real-world geographical area. Then, the value of each pixel in the image which is the geospatial heatmap may represent intensity of a user activity. For example, higher pixel values may correspond to higher level of user activity. The geospatial heatmap may be a monochromatic image. That is, there may only be one value associated with each pixel in the heatmap image and this value may represent intensity of a user activity. Alternatively, several values may be associated with each pixel value in the heatmap image. When there are several values associated with each pixel value, each of these values may, for example, correspond to a different user activity. In this way, one heatmap may encode information about the intensity of several user activities at the same time.

[074] Figure 1B shows at what point in the federated analytics pipeline different issues might arise. For example, when user devices, i.e. clients, transmit data to a server, there may be issues with communication complexity. That is, data transfer may not be as efficient as possible because data may be transmitted from a user device to the server over several FA rounds. This means that more data than necessary may be transmitted which may be an obstacle for an efficient FA algorithm.

[075] Additionally, because privacy-preserving FA algorithms naturally require some modification of collected user data, accuracy at the server side may become an issue. That is, perturbation of data which happens at the user device may lead to less accurate results when FA insights are used for resulting tasks. Known methods do not make use of spatial prediction and processing at a server and therefore require more data to be passed between user device and server, while still leading to lower accuracy results.

[076] Figure 2 shows an overview of the federated analytics process according to the present techniques. By using a ML model to reconstruct noisy heatmaps taking into account their structure, the present techniques allow for an accurate, privacy-preserving and communication efficient method for federated analytics.

[077] In particular, Figure 2 shows a plurality of user devices (client 1...k) which send partitioned noisy heatmaps to a server. Partitioning may happen according to a QuadTree structure. QuadTrees are used to recursively subdivide a two-dimensional space into four quadrants or regions. When using a QuadTree structure on the heatmaps of the present techniques, this may, for example, mean dividing a heatmap into four quadrants according to the four cardinal directions, i.e. North, South, East, West. That is, a heatmap may be split into four quadrants along the North-South and East- West lines of the heatmap. Using a QuadTree structure may ensure that communication between user devices and server happens as efficiently as possible.

[078] The server then aggregates any received noisy heatmaps from the user devices and processes these to generate an aggregated heatmap. The aggregated heatmap may be used to perform analytics. For example, the aggregated heatmap may be used to output recommendations to a user. For example, the recommendation may be for a specific place, such as a specific restaurant or cafe. The recommendation may also be a for a specific image filter. The server may determine, by analysing where a specific user is located, and which image filters other users, or the same user, have previously used at this location, which image filter is most appropriate to suggest. The same may be true for other use cases. That is, recommendations for, for example, specific music playlists, neighbourhoods for specific types of cuisines, or recommendations based on other user activities, may all be based on analysing location history obtained from the aggregated heatmap. This is the final step shown in Figure 2, in which downstream tasks are based on insights inferred from processing the aggregated heatmap.

[079] Figure 3 is a flowchart for a method performed by a server at inference time according to the present techniques. The method comprises: receiving a noisy geospatial heatmap from a plurality of user devices, each noisy geospatial heatmap representing intensity of a specific type of user activity performed by users of the user devices, wherein each noisy geospatial heatmap contains noise to preserve user privacy; generating an aggregated heatmap for the specific type of user activity by aggregating the noisy geospatial heatmaps; removing excess noise from the aggregated heatmap by: identifying structural relationships in the aggregated noisy heatmap using a trained machine learning, ML model; and removing excess noise from the aggregated heatmap using the identified structural relationships to obtain an enhanced aggregated heatmap; and using the enhanced aggregated heatmap to perform federated analytics for the type of user activity.

[080] Advantageously, the present techniques enable the analysis of more accurate heatmaps because unhelpful noise can be removed from noisy heatmaps at the server side. For example, if a user device adds noise to a generated heatmap, such noise may be equally distributed across the heatmap, even at locations where it is unlikely or not possible for a user to be present. For example, it is much less likely that a user will have spent time in, for example, a river and thus, location data at this point is likely to be noise and will be filtered out by the machine learning model. However, it is important to note that the machine learning model has no information about the underlying geography of a heatmap, i.e. the machine learning does not know that a data point is located in a river. Rather, the ML model learns an underlying data structure that enables it to recognise noisy data and differentiate noisy data from data that follows a pattern consistent with the underlying structure of the data.

[081] Thus, receiving a noisy heatmap S300 may comprise location history data from a user device, with the location history data being aggregated in the form of a heatmap, i.e. in some cases an image representing the location history data. This may be for a specific time frame or a specific location, i.e. the location information may only be for a specific city and / or for a time period, such as, for example a specific day, week, month etc. The time period may, for example, also be the same day, e.g. Thursday, over a longer time period. For example, location information may mean all location data for every Thursday in the month of January. This location data is then used to generate a geospatial heatmap, i.e. a map that links intensity of a user activity to a geospatial location. Intensity of a user activity can, for example, mean how often a user visits a specific location and / or how much time a user spends at a specific location. A geospatial heatmap is a way of collating and visualising this information.

[082] Importantly, all heatmaps received by the server have been anonymised by the user device using differential privacy techniques. That is, the user device has already added random noise to the heatmap generated on the user device before this is sent to the server. For example, noise added to the heatmap on the user device may be drawn from a Laplacian Distribution.

[083] Prior to receiving a noisy geospatial heatmap from a plurality of user devices, the method may comprise: sending, to each user device, a number of communication rounds to use when sending their noisy geospatial maps to the server.

[084] Additionally, prior to receiving a noisy geospatial heatmap from the plurality of user devices, the method may also comprise: requesting, from each user device, a geospatial heatmap according to at least one heatmap parameter. For example, the at least one heatmap parameter may comprise any of: a region of interest, a time interval, and / or a target application for creating the geospatial heatmap. Requesting heatmaps according to at least one heatmap parameter may ensure that only relevant heatmaps are received. For example, the server may determine that for a certain geographical region or for a certain user activity, data is missing or that there is not sufficient data. In this case, the server may request data for these areas and / or activities only. Alternatively, the server may request heatmap data for a certain timeframe only. For example, the server may only want data from the last week, or month, or another time frame. In this case, the server may request all location data the user devices have for this timeframe.

[085] In one example, the number of communication rounds to use is one. In this case, receiving a noisy geospatial heatmap from a plurality of user devices may comprise: receiving, during a single communication round, a full resolution noisy geospatial heatmap from each user device of the plurality of user devices. That is, each user device uses a single communication round to send a single noisy geospatial heatmap. The heatmap sent in this example is a full resolution heatmap. This reduces the number of communication rounds, and ensures that the server receives a high resolution heatmap which improves the accuracy of the federated analytics. However, this comes at the cost of having to send a very high resolution heatmap, which may reduce the possibility of some user devices from being able to send their heatmaps.

[086] Full resolution means that the heatmap includes all detail available on the user device and is required at a resolution to show all of this detail. Alternatively, full resolution may mean a maximum resolution determined by the server at which an aggregated heatmap is to be constructed. Resolution may refer to number of pixels per real-world geographical area, wherein each pixel in the geospatial heatmap covers a certain real-world geographical area. All location data available for this pixel may then be pooled or averaged together to obtain a pooled heatmap value for the real-world geographical area. Alternatively, noisy heatmaps may be received over several communication rounds, as described in more detail below.

[087] Next the received geospatial heatmaps are aggregated S302. Aggregating may simply mean normalising and adding all corresponding heatmaps. That is, all heatmaps for the same geographical area, time interval and / or user activity may be added together. Thus, generating an aggregated heatmap may comprise aggregating the received noisy geospatial heatmaps according to the at least one heatmap parameter. Normalising means ensuring that the values of each individual user heatmaps have the same mean and standard deviation. That is, some users may move around in an area very frequently, and therefore, their heatmaps may show a lot of activity for a whole area. Relevant activity, i.e. staying at a specific place for a prolonged time, may then have to be measured relative to the other activity. This means that for another user, who may, for example, not move around as much in the same area, the same intensity values may have different meaning. Therefore, normalising heatmaps ensures that all user data is comparable.

[088] Then, the ML model removes excess noise from the aggregated heatmap to enhance the aggregated heatmap S304. This ML model has been trained to detect the underlying structure of heatmap data and is capable of filtering out superfluous noise from the user heatmaps by using this underlying structure. Because the ML model only learns the underlying data structure, any user data will still remain private. That is, the ML model is trained to remove noise that does not conform with the underlying data structure, but it is not trained to remove all noise. Noise which correlates where data would be expected will remain, thus ensuring the privacy of the original user data.

[089] Additionally, using an ML model which learns the expected structure of data means that the ML model does not need to be trained on non-private user data. For example, the ML model can be trained by adding additional noise to already anonymised user data. The original noisy user data then is the ground truth for the user data to which additional noise has been added.

[090] This also means that the ML model obeys differential privacy principles. That is, the ML model is such that it is unable to reconstruct the original user data. This is consistent with an ML model which learns the underlying structure of data. Suitable models include deep learning, DL, models, such as U-nets, which comprise a trainable set of parameters.

[091] Noisy heatmaps received from the user devices are are temporally aggregated with 7(.) for t + 1 rounds on centralized server as: / / ^ = 7(^..... where Ht+1 refers to the aggregated heatmap, 7(.) is a function used to aggregate each instance of the plurality of noisy heatmaps received from the user devices / , / 2, and refers to the aggregated heatmap that was generated at a previous communication round on the server.

[092] Server side processing can be described by the following equation: MIU.....^,^)) = / ( / / ^) where / (.) refers to the ML model. This method is data independent and therefore £-differentially private as M(.,£). In theory, to guarantee this, this function should behave nearly identically for two databases D± and D2, similar to a differential mechanism M(., £) as follows: Pr(M( / )1, £)) <exp(f) Py(M(D2, £)) Hence, / (.) should be independent from any specific data statistic that is previously learned / trained with other datasets: Pr( / (M(D1,£))) = Pr(M(D1,£)) < exp(£)Pr(M( / )2, £)) = exp(£)Pr( / (M(D2, £)))

[093] Then, the composition f ° m can be also ^differentially private. So, we utilize a blind processing method (independent from data) that is optimized with the input data instantly to preserve the privacy. That is, the ML model described above is used as the blind processing method here.

[094] Because an ML model is used to process the aggregated heatmap, it is also possible for the present techniques to deliver accurate results, with less heatmap data received. In conventional communication methods, using, for example, QuadTrees, information is sent in several rounds, each round containing more detailed information than the previous round.

[095] Figure 4 illustrates how conventional communication rounds, and communication rounds in the present techniques differ. In conventional communication techniques (shown in blue), resolution of an image sent from a user device to the server is increased incrementally with each communication round. For example, the resolution of each image in each communication round may be doubled from communication round to communication round. However, bandwidth is saved by not simply sending the full resolution image in one go. Instead, at each communication round, a residual of the higher resolution image of the current communication round and the previous communication round are sent to the server. Sending the residual instead of the actual image is more efficient, because it requires less information to be included at each communication round. That is, the number of bits required to represent the information that is sent at each communication round is reduced. The first, lowest resolution image (H°) is sent to the server as normal, but each subsequent image is only sent as a residual with respect to the previously sent image. Having the previously sent image and the residual, allows for each image in the communication rounds to be reconstructed at the server. That is: Client Client ~ 0-:. OH ............................Server, pate On Server: w* Data On Sewer: H-v •• + ) Ma On Server -- / spOO") -J- where fup is a function to upscale the previous image in the communication rounds to be able to calculate a residual for each communication round. Similarly, fup\s used on the server to reconstruct the higher resolution image from the residual and the image received in the previous communication round.

[096] In Figure 4, the brown arrows illustrate how bandwidth can be saved by instead using the ML model of the present techniques. Because the ML model as described above is trained to learn the structure of noisy heatmaps, less information is needed to reconstruct noisy heatmaps sent from user devices. In particular, the ML model is used in place of function f above on the server to reconstruct the higher resolution image. This effectively means that communication rounds which would be necessary in the conventional approach can be skipped when using the present techniques, i.e.: Client Gate On Server: atom .................................'.Swesr,, Date On Server: Client ..................................Server Data On Server#2’ - --#2 - where fdip is the ML model above which is now used to upscale the received noisy heatmap on the server. Because the ML model is much more efficient at upscaling, while at the same time preserving privacy, communication rounds that would be necessary in the conventional approach can be skipped, and instead, residuals calculated with a larger difference in resolution between higher and lower resolution images can be sent from user device to server.

[097] Thus, in this example, the number of communication rounds to use is at least two. In this case, receiving a noisy geospatial heatmap from a plurality of user devices may comprise: receiving, during a first communication round, a first low resolution noisy heatmap from each user device; and receiving, during at least one further communication round, a residual for at least one higher resolution noisy heatmap from each user device, wherein the residual represents a difference between a lower resolution noisy heatmap and a higher resolution noisy heatmap.

[098] In the example above and shown in Figure 4, the first low resolution noisy heatmap is H°, and the residual of the second higher resolution noisy heatmap and the first low resolution noisy heatmap is H2 - fdiP(H°y In this example, there is one intermediate communication step (at H2) before a residual of the full resolution noisy heatmap and the reconstructed heatmap at H2 is received by the server.

[099] Receiving at least one residual may comprise: specifying a number of intermediate communication steps; receiving, at each intermediate communication step, a residual of a higher resolution noisy heatmap and a lower resolution noisy heatmap, wherein the residual represents the difference in information between the higher resolution noisy heatmap and the lower resolution noisy heatmap, and wherein the lower resolution noisy heatmap was received at an intermediate communication step immediately preceding the intermediate communication step at which the higher resolution noisy heatmap is received.

[100] That is, instead of sending a full resolution heatmap (e.g. 1024 x 1024 pixels), the user device sends a first low resolution version of the heatmap (e.g. 32 x 32 pixels) during a first communication round. However, the server still needs as much data as possible to perform federated analytics. Thus, during the at least one further communication round, the user device generates a higher resolution heatmap, which has a higher resolution than the heatmap associated with a previous communication round. However, instead of sending each higher resolution heatmap, the user device sends a residual. In this way, although multiple communication rounds are used to send the data, since only the residual is sent for each further communication round after the first round, fewer bytes are needed overall.

[101] The higher resolution noisy heatmap may have a resolution which is four times greater than the resolution of the lower resolution noisy heatmap. In a standard communication approach, the resolution between a first and second communication round may be doubled. For example, if the resolution of a first image is 32x32 pixels, the resolution of the second image may be 64x64 pixels, the resolution of the third image may be 128x128 pixels and so on. In an approach to the present techniques, the second resolution may be skipped and only the first image and a residual between the first and third image may be sent, i.e. a residual between the 32x32 pixels and the 128x128 pixels image. In the conventional approach, the residual between the 32x32 pixels image and the 64x64 pixels image would also be sent. Thus, the present techniques may be used to reduce the number of communication rounds and the amount of data sent from user device to server. This is possible because the ML model may be used to enhance the received image, even with the reduced amount of data available. Thus, the ML model may also act as a super resolution model and is effectively able to upscale an image from reduced data.

[102] Thus, when multiple communication rounds are used, receiving a noisy geospatial heatmap from a plurality of user devices may further comprise: generating, using the ML model, a full resolution noisy geospatial heatmap for each user device using the first low resolution noisy heatmap and the at least one residual. That is, the ML model generates a full resolution noisy geospatial heatmap for each user device, using the first low resolution noisy heatmap and each residual received from the user device. This ensures the noise removal step is performed on the full resolution noisy geospatial heatmap. Receiving at least one residual may comprise: receiving, during each further communication round, a residual between a higher resolution noisy heatmap and a lower resolution noisy heatmap.

[103] In the example shown in Figure 4, the resolution of each noisy heatmap doubles for each communication round, as described above. Thus, using a resolution that is four times the resolution of the previous communication round is equivalent to skipping a communication round, as illustrated in Figure 4. Thus, bandwidth savings are possible compared to conventional communication techniques.

[104] Figures 5A and 5B show algorithms for implementing the machine learning, ML, model of the present techniques. In particular, both Figures illustrate how the ML model is implemented on the server in practice.

[105] Figure 6 is a flowchart for a method performed by a server at training time according to the present techniques. There is provided a computer-implemented method for training a machine learning, ML, model to enhance a noisy geospatial heatmap, the method comprising: obtaining a plurality of noisy geospatial heatmaps, the heatmaps representing intensity of at least one user activity; adding noise to each of the noisy geospatial heatmaps; and training the ML model to remove the added noise from the noisy geospatial heatmaps.

[106] Receiving a plurality of noisy geospatial heatmaps S500 may comprise receiving a plurality of noisy geospatial heatmaps that are stored on the server. Alternatively, this may comprise receiving a plurality of noisy geospatial heatmaps from a plurality of user devices.

[107] Advantageously, using such a machine learning model means that user data can be kept completely private. While the ML model is trained to remove data, it is never trained on non-anonymised user data, meaning that it is not trained to reconstruct original nonanonymised user data. This is a further safeguard to ensure that privacy is maintained at all times. Additionally, this ML model is suitable for enhancing several types of noisy heatmaps. For example, it can be used as in the first approach above to enhance an aggregated heatmap and remove any superfluous noise from this. The ML model can also be used, as described with reference to the first approach, to effectively perform a super resolution task and reconstruct a noisy heatmap from less data, for example, when a residual between a higher and lower resolution noisy heatmap is transmitted to a server, wherein the difference in resolution between the higher and lower resolution heatmaps is great. For example, the difference in resolution may be fourfold.

[108] Adding noise S502 to the noisy heatmaps may comprise adding noise that is drawn from a different probability distribution than the noise in the noisy geospatial heatmaps. Advantageously, it is therefore not necessary to know which probability distribution the noise in the noisy heatmaps is drawn from. The ML learns structural relationships, independent of the noise used in training. Alternatively, adding noise to each of the noisy geospatial heatmaps may comprise adding noise that is drawn from the same probability distribution than the noise in the noisy geospatial heatmaps.

[109] Training the machine learning, ML, model S504 may comprise using a loss function to adapt the parameters of the ML model. Our function f (. ) = #(.; 0) is a Deep Neural Network with a trainable parameter set 6. The aim is to iteratively update the parameter set 0 by learning structural (i.e., spatial or temporal) relations in the input data by minimizing the following loss function: min / ( / , g(Z; 0)) + AR(g(Z; 0)) 0 where / ( / , g(Z-, 0)) is the data fitting loss and AR(g(Z; 6)) is a regularization term.

[110] Here, Z is a random noise and iteratively overfits to the output of aggregation step H. With early stopping, the structure of data can be captured by g(.; 0) that filters out distorted privacy components from the data which affects the accuracy and communications complexity. This ultimately will help to enhance the accuracy and communication complexity. That is, training the ML model may comprise training the ML model for a fixed number of iterations to avoid overfitting and generalise the ML model. Early stopping techniques may be used to determine the optimal number of iterations to arrive at a ML model which has learned the required structural information but is not overfitting.

[111] Figure 7 is a flowchart for a method performed by a user device according to the present techniques. There is provided a computer-implemented method for performing federated analytics, FA, performed by a user device, the method comprising: receiving, from a server, a request for a geospatial heatmap (step S700); generating a geospatial heatmap, using data items stored on the user device, the heatmap representing intensity of a specific type of user activity performed by a user of the user device (step S702); adding noise to the generated geospatial heatmap to anonymise the heatmap (step S704); and sending the noisy geospatial heatmap to the server for use in federated analytics (step S706).

[112] Receiving, from a server, a request for a geospatial heatmap S700 may comprise: receiving, from the server, a number of communication rounds to use when sending the geospatial heatmap. The server may also specify the resolution for each communication round. Alternatively, the user device may infer the resolutions from the number of communication rounds. For example, a full resolution of the noisy geospatial heatmap may be standardised and saved to the user device. The same may be true for a low resolution version of the noisy heatmap. That is the user device may know which resolution to start from and which resolution to end at, and may infer the resolutions at each communication round from these two resolutions.

[113] Generating a heatmap S702 may comprise generating a heatmap according to certain heatmap parameters, such as timeframe, location and / or user activity. Generating the heatmap may comprise retrieving location data from storage and generating a heatmap as described above with reference to Figure 3. The heatmap parameters may be received from the server, or the user device may determine the heatmap parameters.

[114] The heatmap may also be generated according to the number of communication rounds received from the server. When the number of communication rounds is one, generating a geospatial heatmap may comprise generating a full resolution geospatial heatmap. A full resolution heatmap is as described above.

[115] When the number of communication rounds is at least two, generating a geospatial heatmap may comprise: generating a full resolution geospatial heatmap; generating, from the full resolution geospatial heatmap, a first low resolution heatmap; and generating, using the full resolution geospatial heatmap, a higher resolution geospatial heatmap for each further communication round, wherein the resolution of the higher resolution geospatial heatmap created for each further communication round is higher than the first low resolution heatmap and increases for each successive communication round. That is, when there are several communication rounds, the user device creates geospatial heatmaps with resolutions adapted for each of the communication rounds.

[116] Adding noise to the generated heatmap S704 may comprise: sampling random noise from a Laplacian Distribution; and adding the sampled random noise to the generated heatmap. This is done to anonymise the heatmap, i.e. to ensure that it is not possible to trace an individual user using their geospatial heatmap. Therefore, the noise is added on the user device, and thus, the original location data never leaves the user device, ensuing that the present techniques are privacy-preserving.

[117] Adding noise to the generated geospatial heatmap may comprise adding noise to the first low resolution heatmap and each higher resolution geospatial heatmap generated for each further communication round. I.e. noise is added separately to each generated geospatial heatmap. Alternatively, noise may be added to the full resolution heatmap and is then automatically incorporated into each of the other heatmaps generated for each communication round.

[118] In particular, the present techniques use differential privacy to ensure privacy of user data. A number of different techniques can be used to ensure this privacy, but generally speaking, differential privacy techniques are implemented by adding Laplacian noise to generated heatmaps.

[119] That is, this random noise may be sampled from a Laplacian distribution Lap(b, pX.) as follows: Lap(b, p)(Z) = — exp I---—) Zb \ b / Here, p e R^s the expectation of the Laplace distribution and b is the scale parameter. Thus, adding noise to a geospatial heatmap comprises using a Laplace Mechanism M(.) which takes a sample X ERD(i.e., the geospatial heatmap) and noise Y e RD drawn from the Laplacian distribution above. Both the heatmap and the noise are combined to obtain a noisy heatmap for the kth client / k: M(Xk,£)=Xk + Yk = Ik

[120] Here, the noise Y is a random variable independently sampled from Lap(0, 1 / e), i.e. from a Laplace distribution with mean 0 and variance l / £. £ g R1 denotes the privacy budget, also referred to as privacy parameter, i.e. this term decides how much privacy will be added. Because e is linked to the variance of the Laplacian distribution in this way, e determines the shape of the Laplacian distribution and therefore the noise values that are drawn from the distribution. Therefore, e also determines how much noise will be added to a data sample, and therefore, how much “privacy” is added to a data sample. The function M(., e) is ^-differentially private.

[121] Thus, sampling random noise from a Laplacian Distribution may comprise obtaining a privacy parameter e which determines a variance of the Laplacian Distribution. The privacy parameter may also be termed a privacy budget. The privacy parameter may be the inverse of the variance of the Laplacian Distribution.

[122] Sampling random noise from a Laplacian Distribution may comprise sampling random noise from a Laplacian Distribution according to a Discrete Laplacian Mechanism (“Towards Sparse Federated Analytics: Location Heatmaps under Distributed Differential Privacy with Secure Aggregation", Arxiv 2022), or according to a Geo-Indistinguishability (“Geo-Indistinguishability: Differential Privacy for Location-Based Systems”, Arxiv 2012.) method.

[123] Sending the noisy heatmap to a server S706 may comprise: sending, during a first communication round, the first low resolution noisy heatmap; calculating a residual between each higher resolution noisy geospatial heatmap and the noisy heatmap corresponding to the previous communication round; and sending, during each further communication round, the calculated residual.

[124] The residual may be calculated using the ML model that is also used on the server, as described above, to reconstruct a higher resolution image from a residual and a lower resolution image. However, the user device may use an ML model that is trained on-device, using user data only. This ensures that there are no privacy issues, i.e. no transfer of data from server to user device or vice versa. The model may be trained on the user device in the same way as it is trained on the server. That is, noise may be added to already noisy images on the user device and the model is trained to remove this noise from the images. Again, the ML model may be a deep learning, DL, model, such as a U-net, as described above.

[125] The present techniques may be used in a number of different scenarios. For example, the user activity may be using an image filter or image effect. An image filter or image effect is used to modify or enhance an image by altering its pixels. For example, the shades or colours may be altered, or the image may be sharpened or blurred, or an artistic effect or overlay may be added. Then, performing federated analytics means using the geospatial heatmaps received from the plurality of user devices to determine which image filters are particularly popular at which locations. Other factors may be taken into account in the analysis, for example, time of day, month etc. For example, different image filters may be recommended for the same location at different times of the year.

[126] In this example use case, the server may receive a request for a recommendation from a user of a user device in relation to image filters, together with the geographical location of the user device. The request may be automatically or autonomously sent to the server when the user launches an app to use image filters. In response, the server may provide a recommendation to the user of the user device based on the enhanced aggregated heatmap for the received location and the specific user activity (image filters).

[127] Similarly, a user device may transmit a request for a recommendation to a server in relation to image filters, together with the geographical location of the user device (e.g. as determined via GPS or otherwise). The request may be automatically or autonomously sent to the server when the user launches an app to use image filters. In response, the user device may receive, from the server, a recommendation based on the enhanced aggregated heatmap for the received location and the specific user activity (image filters). The user device may apply the recommended filter to one or more images (e.g. images they have captured in that geographical location), and thereby generate a modified image.

[128] Similarly, local music playlists can be suggested using the above methods. For example, the user activity may be playing a particular song, or particular songs, or one or more particular music genres. Then, performing federated analytics means using the geospatial heatmaps received from the plurality of user devices to determine a music playlist that may be particularly suited to a user at a specific location. For example, if the user is at a location which is popular for running, playlists that are popular for running may be suggested because other users will also have been running and playing similar music at the same location.

[129] In this example use case, the server may receive a request for a recommendation from a user of a user device in relation to music (song, playlist, etc.), together with the geographical location of the user device. The request may be automatically or autonomously sent to the server when the user launches an app to listen to music. In response, the server may provide a recommendation to the user of the user device based on the enhanced aggregated heatmap for the received location and the specific user activity (music).

[130] Similarly, a user device may transmit a request for a recommendation to a server in relation to music, together with the geographical location of the user device (e.g. as determined via GPS or otherwise). The request may be automatically or autonomously sent to the server when the user launches an app to listen to music. In response, the user device may receive, from the server, a recommendation based on the enhanced aggregated heatmap for the received location and the specific user activity (music). The user device may listen to the recommended music, or search for and then listen to music based on the recommended music genres, artists, and so on.

[131] In yet another example, replies to user queries may be optimised. For example, a user device may transmit a geospatial location of interest and a query to the server. In response, the user may expect a visual answer to their query, i.e. that the answer is shown on a map. For example, the user may currently be in a specific neighbourhood (i.e. the geospatial location of interest) and may be searching for the best place to eat Korean food. Taking into account the data collected from other user devices, may then help generate a map for the user showing which area may be particularly suitable for finding a Korean restaurant (e.g. Chinatown and / or SoHo in London). Alternatively, the server may also use the enhanced aggregated heatmap to find the most popular Korean restaurant at the current location of the user and suggest this restaurant. The suggestion may, for example, involve generating a map showing the recommended restaurant.

[132] In this example use case, the server may receive a request for a recommendation from a user of a user device in relation to a specific query (user activity type), together with the geographical location of the user device. The request may be automatically or autonomously sent to the server when the user makes the query. For example, the user may input a text query in a web browser or search engine, or make a voice query to an Al assistant. In response, the server may provide a recommendation to the user of the user device based on the enhanced aggregated heatmap for the received location and the specific query.

[133] Similarly, a user device may transmit a request for a recommendation to a server in relation to a specific query (user activity type), together with the geographical location of the user device (e.g. as determined via GPS or otherwise). The request may be automatically or autonomously sent to the server when the user makes the query, via text input, voice input, or otherwise. In response, the user device may receive, from the server, a recommendation based on the enhanced aggregated heatmap for the received location and the specific user activity or query. The user device may output the recommendation to the user either via a display of the user device or an audio output device (e.g. speaker) of the user device, or otherwise. If the user query is for a restaurant, for example, the user device may receive or generate a map showing the location of the recommended restaurant, and optionally generate direction information for the user based on where the user device is currently located.

[134] Results

[135] Figures 8 to 11 show results obtained using the present techniques. We illustrate inference performance before and after our enhancement step is applied. Since we don’t have GT labels, we cluster NoDP prediction in an unsupervised manner. Cluster centers of NoDP prediction is accepted as GT. A chamfer distance formula is used to find pairs between GT and estimations by minimizing the distances. CH(A, B) = -^-^gA minbeB dX(a, b)

[136] Recognition Accuracy: A threshold is set and this metric calculates the percentage of pairs whose distance is smaller than this threshold. Confidence Score: Variance of Euclidean distances between GT and estimated points (cluster centers). Then, this variance value is projected to [0, 1] with an exponential function.

[137] Figure 8 shows (left to right) a heatmap and corresponding distribution of points on the heatmap for an original heatmap with no differential privacy applied, for a heatmap to which noise has been applied, and for a heatmap to which noise has been applied after the heatmap has been processed by the ML model of the present techniques. The differential privacy mechanism used is a discrete Laplacian mechanism (as described above) with privacy parameter e = 0.1.

[138] Figure 9 shows (left to right) a heatmap and corresponding distribution of points on the heatmap for an original heatmap with no differential privacy applied, for a heatmap to which noise has been applied, and for a heatmap to which noise has been applied after the heatmap has been processed by the ML model of the present techniques. The differential privacy mechanism used is a discrete Laplacian mechanism (as described above) with privacy parameter e = 1.0.

[139] Figure 8 shows (left to right) an original heatmap with no differential privacy applied, a heatmap to which noise has been applied, and for a heatmap to which noise has been applied after the heatmap has been processed by the ML model of the present techniques. The differential privacy mechanism used is a Geo indistinguishability mechanism (as described above) with privacy parameter e = 0.1 and R = 1000.

[140] Figure 11 illustrates (left to right) a heatmap and corresponding distribution of points on the heatmap for an original heatmap with no differential privacy applied, for a heatmap to which noise has been applied, and for a heatmap to which noise has been applied after the heatmap has been processed by the ML model of the present techniques. The differential privacy mechanism used is a discrete Laplacian mechanism (as described above) with privacy parameter 6 = 0.1. Figure 11 also illustrates for each heatmap recognition accuracy and confidence score.

[141] Figure 12 is a block diagram illustrating a server / user device set-up 200 according to the present techniques. The user device 100 may comprise at least one processor 102 coupled to memory 104. The user device may communicate with the server 112 via a communication module 110 in the user device and a communication module 120 in the server. The server may have at least one processor 114 coupled to memory 116. The machine learning, ML, model 118 is stored on the server. The ML model is executed by the at least one processor coupled to memory.

[142] Those skilled in the art will appreciate that while the foregoing has described what is considered to be the best mode and where appropriate other modes of performing present techniques, the present techniques should not be limited to the specific configurations and methods disclosed in this description of the preferred embodiment. Those skilled in the art will recognise that present techniques have a broad range of applications, and that the embodiments may take a wide range of modifications without departing from any inventive concept as defined in the appended claims.

Claims

1. A computer-implemented method for performing federated analytics, FA at a server, the method comprising:receiving a noisy geospatial heatmap from a plurality of user devices, each noisy geospatial heatmap representing intensity of a specific type of user activity performed by users of the user devices, wherein each noisy geospatial heatmap contains noise to preserve user privacy;generating an aggregated heatmap for the specific type of user activity by aggregating the noisy geospatial heatmaps;removing excess noise from the aggregated heatmap by:identifying structural relationships in the aggregated noisy heatmap using a trained machine learning, ML model; andremoving excess noise from the aggregated heatmap using the identified structural relationships to obtain an enhanced aggregated heatmap; andusing the enhanced aggregated heatmap to perform federated analytics for the type of user activity.

2. The method as claimed in claim 1 wherein processing the aggregated heatmap using an ML model comprises using a differentially private ML model that protects privacy of individual data points.

3. The method as claimed in claim 1 or 2 wherein prior to receiving a noisy geospatial heatmap from a plurality of user devices, the method comprises:sending, to each user device, a number of communication rounds to use when sending their noisy geospatial maps to the server.

4. The method as claimed in claim 3 wherein the number of communication rounds to use is one, and wherein receiving a noisy geospatial heatmap from a plurality of user devices comprises:receiving, during a single communication round, a full resolution noisy geospatial heatmap from each user device of the plurality of user devices.

5. The method as claimed in claim 3 wherein the number of communication rounds to use is at least two, and wherein receiving a noisy geospatial heatmap from a plurality of user devices comprises:receiving, during a first communication round, a first low resolution noisy heatmap from each user device; andreceiving, during at least one further communication round, a residual for at least one higher resolution noisy heatmap from each user device, wherein the residual represents a difference between a lower resolution noisy heatmap corresponding to the previous communication round and a higher resolution noisy heatmap corresponding to the current communication round.

6. The method as claimed in claim 5 wherein receiving a noisy geospatial heatmap from a plurality of user devices further comprises:generating, using the ML model, a full resolution noisy geospatial heatmap for each user device using the first low resolution noisy heatmap and the at least one residual.

7. The method as claimed in 5 or 6 receiving at least one residual comprises:receiving, during each further communication round, a residual between a higher resolution noisy heatmap and a lower resolution noisy heatmap.

8. The method as claimed in claim 7 wherein, during each further communication round, the higher resolution noisy heatmap has a resolution which is four times greater than the resolution of the lower resolution noisy heatmap.

9. The method as claimed in any preceding claim wherein prior to receiving a noisy geospatial heatmap from the plurality of user devices, the method comprises:requesting, from each user device, a geospatial heatmap according to at least one heatmap parameter.

10. The method as claimed in claim 9 wherein the at least one heatmap parameter comprises any one or more of: a region of interest; a time interval; and a target application for creating the geospatial heatmap.

11. The method as claimed in claim 9 or 10 wherein generating an aggregated heatmap comprises aggregating the received noisy geospatial heatmaps using the at least one heatmap parameter.

12. A computer-implemented method for training a machine learning, ML, model to enhance a noisy geospatial heatmap, the method comprising:obtaining a plurality of noisy geospatial heatmaps, the heatmaps representing intensity of at least one user activity;adding noise to each of the noisy geospatial heatmaps; andtraining the ML model to the remove the added noise from the noisy geospatial heatmaps.

13. The method as claimed in claim 11 wherein training the ML model to remove the added noise from the noisy geospatial heatmaps comprises training the ML model to recognise structural relationships in the noisy heatmaps and to remove the added noise using the recognised structural relationships.

14. The method as claimed in claim 13 wherein training the ML model to recognise structural relationships in the noisy heatmaps comprises training the model to learn spatial and / or temporal relationships in the noisy heatmaps.

15. The method as claimed in any of claims 11 to 14 wherein training the ML model comprises training the ML model for a fixed number of iterations to generalise the ML model.

16. The method as claimed in any of claims 11 to 15 wherein adding noise to each of the noisy geospatial heatmaps comprises adding noise that is obtained from a different probability distribution than the noise in the noisy geospatial heatmaps.

17. The method as claimed in any of claims 11 to 15 wherein adding noise to each of the noisy geospatial heatmaps comprises adding noise that is obtained from the same probability distribution than the noise in the noisy geospatial heatmaps.

18. A computer-implemented method for performing federated analytics, FA, at a user device, the method comprising:receiving, from a server, a request for a geospatial heatmap;generating a geospatial heatmap, using data items stored on the user device, the heatmap representing intensity of a specific type of user activity performed by a user of the user device;adding noise to the generated geospatial heatmap to anonymise the heatmap; andsending the noisy geospatial heatmap to the server for use in performing federated analytics.

19. The method as claimed in claim 18 wherein adding noise comprises:sampling random noise from a Laplacian Distribution; andadding the sampled random noise to the generated geospatial heatmap.

20. The method as claimed in claim 19 wherein sampling random noise from a Laplacian Distribution comprises obtaining a privacy parameter e which determines a variance of the Laplacian Distribution.

21. The method as claimed in claim 19 or 20 wherein sampling random noise from a Laplacian Distribution comprises sampling random noise from a Laplacian Distribution according to a Discrete Laplacian Mechanism, or according to a Geo-Indistinguishability method.

22. The method as claimed in any of claims 18 to 21 wherein receiving, from a server, a request for a geospatial heatmap comprises:receiving, from the server, a number of communication rounds to use when sending the geospatial heatmap.

23. The method as claimed in claim 22 wherein the number of communication rounds to use is at least two, and wherein generating a geospatial heatmap comprises:generating a full resolution geospatial heatmap;generating, from the full resolution geospatial heatmap, a first low resolution heatmap; andgenerating, using the full resolution geospatial heatmap, a higher resolution geospatial heatmap for each further communication round, wherein the resolution of the higher resolution geospatial heatmap created for each further communication round is higher than the first low resolution heatmap and increases for each successive communication round.

24. The method as claimed in claim 23 wherein adding noise to the generated geospatial heatmap comprises adding noise to the first low resolution heatmap and each higher resolution geospatial heatmap generated for each further communication round.

25. The method as claimed in claim 24 wherein sending the noisy geospatial heatmap to the server comprises:sending, during a first communication round, the first low resolution noisy heatmap;calculating a residual between each higher resolution noisy geospatial heatmap and the corresponding noisy heatmap of the previous communication round; andsending, during each further communication round, the calculated residual.IntellectualPropertyOfficeApplication GB2507086.3Search report under Section 17 of the Patents Act 1977Date search completed: 14 November 2025Claims searched: 1-11International classificationSubclass and subgroup Valid from G06F21 / 62 01 / 01 / 2013 G06T7 / 20 01 / 01 / 2017 G06V10 / 46 01 / 01 / 2022Field of searchWorldwide search of patent documents classified in the following areas of the IPC:G06T, G06F, G06VDatabases used in the preparation of this search report:SEARCH-NPL; SEARCH-PATENTDocuments considered to be relevantIntellectual Property Office is an operating name of the Patent Officewww.gov.uk / ipoNon-patent literatureRef. Category Relevant claims Document of relevance D1 X 1 at least Towards Sparse Federated Analytics: Location Heatmaps under Distributed Differential Privacy with Secure Aggregation, Eugene Bagdasaryan, Peter Kairouz, Stefan Mellem, Adria Gascon, Kallista Bonawitz, Deborah Estrin, Marco Gruteser, 11 / 11 / 2021, Accession No. XP091094827, See whole document, describes Federated analytics from privately generated geospatial heatmaps.Patent literatureRef. Category Relevant claims Document of relevance D2 A - US 2023 / 0032705 A1 NAVALPAKKAM et al., See whole document, a method of creating a noisy geospatial heatmap from a plurality of user devices. Generating an aggregated heatmap and the removal of noise.CategoriesLetter or symbol Description X Document indicating lack of novelty or inventive step. Y Document indicating lack of inventive step, if combined with another document of the same category. & Member of the same patent family. A Document indicating technological background. P Document published on or after the priority date but before the fling date of the present application.Letter or symbol Description E Earlier application published on or after the filing date of the present application.

Citation Information

Patent Citations

  • Differentially Private Heatmaps

    US20230032705A1