Distributed privacy budget service

By managing privacy budgets in the server and allocating and decreasing privacy budgets on a per-set data basis, the risks of differential privacy attacks in the prior art are solved, and a secure analysis of user behavior and interactive data is achieved.

CN119968632APending Publication Date: 2025-05-09GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380063925.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-31
Filing Date
2023-12-29
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively prevent differential privacy attacks, especially when analyzing user behavior and interactive data, it is easy to disclose user's private information.

Method used

Verify that there is an adequate privacy budget to analyze the dataset by managing the privacy budget in the server, receiving requests to analyze the dataset, and sending requests to multiple servers. The method includes sorting the data set into groups and allocating and decreasing privacy budgets on a per-group basis to prevent reuse of records.

Benefits of technology

It effectively mitigates the risk of differential privacy attacks, and by limiting the number of analysis times of each set of data, preventing malicious actors from tracing the data of individual users, enhancing the differential privacy of the data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119968632A_ABST
    Figure CN119968632A_ABST
Patent Text Reader

Abstract

A server can implement a method for managing privacy budget. The method includes receiving a request to analyze a dataset, the dataset associated with a privacy budget representing a number of times the dataset can be analyzed. The method further includes sending a first request to a first server implementing a first privacy budget service to verify whether a sufficient privacy budget exists to analyze the data set, and sending a second request to a second server implementing a second privacy budget service to verify whether a sufficient privacy budget exists to analyze the data set, the second privacy budget service is independent of the first privacy budget service. The method further includes receiving, from the first server, a first response indicating whether there is a sufficient privacy budget; receiving, from the second server, a second response indicating whether there is a sufficient privacy budget; and processing the data set based on the first response and the second response.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to and the benefit of provisional U.S. patent application No. 63 / 478,140, ​​entitled “Distributed Privacy Budget Service,” filed on December 31, 2022. The entire contents of the provisional application are hereby expressly incorporated herein by reference. Technical Field

[0003] The present disclosure relates to enforcing privacy budgets, and more particularly, to techniques for distributing privacy budget monitoring to improve security and enhance differential privacy. Background Art

[0004] The background description provided herein is for the purpose of generally presenting the background of the present disclosure. The work of the presently named inventors (to the extent that it is described in this background section) and aspects of this description that may not be identified as prior art at the time of filing are neither explicitly nor implicitly admitted to be prior art to the present disclosure.

[0005] Some services collect data related to user behavior and user interactions with electronic resources. This data may contain personally identifiable information (PII), and the service may analyze the data to generate reports and metrics. For example, a service may aggregate records containing data from multiple users and analyze the aggregated records to produce a combined report. Even if a service or operator edits the data to remove PII, this type of analysis may still be susceptible to security risks.

[0006] For example, if a service that operates on user input data does not maintain differential privacy, a malicious actor may be able to retrieve the user input data based on the output of the service. Differential privacy refers to a mathematically provable guarantee that an individual user's information is protected. A set of analysis results can be considered differentially private if it is not possible to trace a single user's data (i.e., input data) from the analysis results (i.e., output data).

[0007] There is a class of differential privacy attacks in which batches of records may be crafted, where each new batch is generated by adding or removing a small number of records and passing these batches to a service for processing. By knowing some metadata about the records (e.g., the Internet Protocol (IP) address from which the records originated) and understanding the aggregation algorithm, analysis of reports generated from these crafted batches may provide private information contained in some records. This class of differential privacy attacks relies on reusing the same records over and over again to generate multiple insights that can be compared and analyzed later.

[0008] Therefore, there is a need for improved techniques for analyzing records in a manner that enforces differential privacy of the outputs. Summary of the invention

[0009] An example implementation of these techniques is a method for managing a privacy budget in one or more servers. The method may include: receiving a request to analyze a data set, the data set being associated with a privacy budget representing a number of times the data set can be analyzed; sending a first request to a first server implementing a first privacy budget service to verify whether there is sufficient privacy budget to analyze the data set, the first privacy budget service maintaining a first instance of the privacy budget associated with the data set; sending a second request to a second server implementing a second privacy budget service to verify whether there is sufficient privacy budget to analyze the data set, the second privacy budget service being independent of the first privacy budget service, and the second privacy budget service maintaining a second instance of the privacy budget associated with the data set; receiving a first response from the first server indicating whether there is sufficient privacy budget according to the first privacy budget service; receiving a second response from the second server indicating whether there is sufficient privacy budget according to the second privacy budget service; and processing the data set based on the first response and the second response.

[0010] Another example embodiment is a method for managing a privacy budget in one or more servers. The method may include: receiving a plurality of data sets, each of the plurality of data sets including encrypted data and metadata for the data set; classifying the plurality of data sets into one or more groups based on a corresponding plurality of metadata included in the plurality of data sets; and for one of the one or more groups, querying whether there is a sufficient privacy budget for the group to store the results of an analysis of the group, the privacy budget for the group indicating the number of times the group can be analyzed.

[0011] Yet another example embodiment is a computing system comprising one or more servers and a non-transitory computer-readable medium storing instructions thereon. When executed by the one or more processors, the computing system is caused to implement any of the methods described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a block diagram of an example computing system in which the techniques of this disclosure may be implemented;

[0013] Figure 2A is a block diagram illustrating an example computing architecture including a security control plane and a data plane of the present disclosure;

[0014] Figure 2B is an example similar to Figure 2A A block diagram of another example computing architecture is shown in FIG. 1 , except that Figure 2B Additional infrastructure for managing encryption keys and privacy budgets is illustrated;

[0015] Figure 3 is an example FIG. 2A to FIG. 2B A messaging diagram for an example scenario in which a control plane receives a request to perform a computation using logic executed on a data plane;

[0016] Figure 4 is a block diagram illustrating an example technique for grouping records and allocating a privacy budget on a per-group basis;

[0017] Figure 5 is a messaging diagram illustrating an example scenario in which records are grouped, privacy budgets are allocated on a per-group basis, and privacy budgets are verified by a distributed privacy budget service;

[0018] Figure 6 is a flow chart illustrating an example method for managing a privacy budget on a per-group basis; and

[0019] Figure 7 is a flow chart illustrating an example method for managing a privacy budget using a distributed privacy budget service. DETAILED DESCRIPTION

[0020] The service of the present disclosure implements techniques for mitigating differential privacy attacks. According to an example technique, the service prevents records from being reused in multiple batches. The service accomplishes this task by monitoring which records have been used and allocating a budget (called a privacy budget) to each of these records. The service decrements the privacy budget each time a record is used in a processed batch. For example, if the privacy budget is '2' for record N, record N can be used in a batch that is only processed twice. After a record is used in a batch and processed by the service, the service decrements the budget for the record by one. The service will then refuse to process any batches containing record N.

[0021] In some implementations, the service further optimizes the technique by tracking the privacy budget on a per-batch basis rather than on a per-record basis. For example, a service may process trillions of events per day. Therefore, tracking the privacy budget on a per-record basis may be inefficient. In order to reduce the computational resources required to monitor the privacy budget for a set of records, the service may divide the set of records into "buckets" or groups and assign a privacy budget to each group. Such groups may be defined, for example, by a time range and a transmitter. An example group may include "all records transmitted by transmitter X from 9 am to 10 am". The service may then assign a privacy budget to the example group and decrement the budget whenever the example group is processed (including whenever a batch of records including the group is processed). This optimization allows a large number of records to be processed efficiently with minimal overhead while maintaining differential privacy.

[0022] The second example technique is to distribute privacy budget operations. A service that analyzes records may call an external privacy budget service operated by a trusted party to verify whether there is a privacy budget available for processing a record or a group of records. However, if a single trusted party operates the privacy budget service, such a trusted party has the opportunity to tamper with the privacy budget. To address this potential vulnerability, trust can be decentralized by distributing the privacy budget service over more than one operator. Instead of a single trusted party, there may be N independent co-owners of the privacy budget (i.e., multiple trusted parties). Each trusted party may implement an instance of the privacy budget service, where each of these instances has a copy of the entire privacy budget. In order to consume the budget, the service that analyzes the record may be required to perform atomic transactions on all of these instances, where "atomic" refers to the condition that all instances must verify the privacy budget or the process cannot continue. The transaction can only proceed if all privacy budgets are consistent on all of these instances. Therefore, if a trusted party tampers with its privacy budget, the transaction will fail. Therefore, the distributed privacy budget service reduces the likelihood that the privacy budget will be tampered with, because all trusted parties that implement instances of the privacy budget must collude to modify the privacy budget.

[0023] Sample computing environment

[0024] In some embodiments, the techniques discussed above may be implemented using a security control plane (SCP), which in turn provides an isolated secure execution environment for a data plane (DP). Any arbitrary business logic (such as any logic for analyzing batches of records) may be executed within the DP, and all sensitive data that traverses the SCP and enters the DP is encrypted. It should be understood that while the examples provided in this disclosure primarily describe techniques for managing privacy budgets performed using an SCP architecture, the techniques described herein may be implemented using any suitable computing environment.

[0025] The SCP described herein provides an unobservable secure execution environment in which services can be deployed. Specifically, any business logic (e.g., code for an application) that provides a service can be executed in the secure execution environment to provide the security and privacy guarantees required by the workflow, while no party can observe the computation at runtime. The state of the environment is opaque even to the administrator of the service, and the service can be deployed on any supported cloud.

[0026] As an example, two clients generating data (Client 1 and Client 2) may wish to combine the data streams they receive from their respective clients so that the clients can generate quantitative metrics related to these clients that cannot be derived from their individual data sets. As a more specific example, for example, Client 1 may be a retailer having data indicative of customer transactions, and Client 2 may be an analytics engine capable of measuring the effectiveness of an advertising campaign for products offered by the retailer.

[0027] Client 2 may offer a service with an algorithm that Client 2 claims will securely perform data analysis. However, Client 1 may not wish to expose that client's customer data to Client 2 in a manner that could allow the data to be leaked or used in a manner that does not comply with Client 1's privacy and security guarantees. Therefore, Client 1 wants to ensure that (1) its customer data cannot be leaked by Client 2 or any other party, and (2) the logic used to analyze the customer data complies with Client 2's security requirements. The technology disclosed herein provides a secure execution environment for executing business logic, such that sensitive data analyzed by the business logic remains encrypted everywhere except within the secure execution environment, and provides proof so that any party can ensure that logic running within the secure execution environment executed as guaranteed.

[0028] In general, services that perform calculations (i.e., use business logic to process events or requests) are divided between the data plane (DP) and the security control plane (SCP). The business logic specific to the calculation is hosted in the DP, where the DP is located in a trusted execution environment (TEE), which is also referred to as a secure area (enclave) in this article. The business logic can be provided to the DP as a container, where the container is a software package containing all the necessary elements to run the business logic in any environment. The container can be provided to the SCP, for example, by the business logic owner. Functionally, the SCP provides a secure execution environment and facilities to deploy and operate the DP on a large scale, including managing encryption keys, buffering requests, tracking privacy budgets, accessing storage, coordinating policy-based horizontal automatic expansion, etc. The SCP execution environment isolates the DP from the details of the cloud environment, thereby allowing the service to be deployed on any supported cloud provider without making changes to the DP. Both the DP and the SCP work together by communicating via an input / output (I / O) application programming interface (API), which is also referred to as a control plane I / O API or CPIO API in this article.

[0029] In an example implementation, all data traversing the SCP is always encrypted, and only the DP has access to the decryption keys. For example, the SCP may facilitate trusted data exchange in which data from multiple parties that may not trust each other may be joined, but none of these parties have access to the keys used to decrypt the data. Additionally, the decryption keys may be bit-sliced ​​when located outside of the DP so that only the DP can assemble the decryption keys within the TEE. Depending on the desired application, the outputs from the DPs may be edited or aggregated in a way that the outputs may be shared and no individual user's data may be identified or leaked.

[0030] SCP provides several privacy, trust and security guarantees. Regarding privacy, services using SCP can guarantee that no relevant party (e.g., a device operated by a client, a cloud platform, a third party) can act alone to access or leak plaintext (i.e., non-encrypted) sensitive information, including administrators deployed by SCP. In addition, regarding trust, DP runs in a secure execution environment in a trusted state when it is booted in a secure area. For example, SCP can be implemented using a trusted platform module (TPM) or a virtual trusted platform module (vTPM) according to secure boot standards, and / or using a trusted and / or certified operating system (OS), using technologies to ensure process isolation in hardware (including memory encryption and / or memory address space segmentation, and trust chains from boot). Starting from an audited code base and reproducible builds, cryptographic proofs are used to prove the DP binary identity and origin to a key management service (KMS) at runtime, which is configured to release encryption keys only to verified secure areas. Therefore, any tampering with the DP image will result in the system being unable to decrypt any data. Given that cloud providers have a strong incentive to guarantee their terms of service (ToS) guarantees, cloud providers are undoubtedly trustworthy. Regarding security, the Secure Execution Environment is unobservable. The Secure Execution Environment's memory is encrypted or otherwise hardware-protected to prevent access from other processes. In the example implementation, core dumps are not possible. All data is encrypted in transit and at rest, and all I / O to and from the DP is encrypted. No one can access the private keys in plain text (e.g., the KMS is locked, the keys are sharded, and the keys are only available within the DP, which is within the Secure Execution Environment.

[0031] SCP distributes trust in a way that three parties need to cooperate to leak plaintext user event data. SCP also uses a distributed trust model to ensure that two parties need to cooperate to tamper with the privacy budget service (see Figure 2B and Figure 5 Described in more detail). Distributed trust is used for both event decryption and privacy budget services. With respect to event decryption, the private key required to decrypt events received at the SCP is generated in a secure environment and bit-sharded between at least two KMSs, each of which is under the control of an independent trusted party. For example, each trusted party may further encrypt the corresponding key shard of each trusted party using a KMS key owned by a trusted party in the cloud provider's KMS. The KMS is configured to release key material only to DPs that match a specific hash. If the DP is tampered with, the key shard will not be released. In this scenario, the service can be started, but no events will be able to be decrypted. Similarly, the privacy budget service can be distributed between two independent trusted parties, and transaction semantics can be used to ensure that the budgets of the two trusted parties match, which allows detection of budget tampering.

[0032] If you will refer to Figure 2B As discussed, SCP also provides a mechanism for proving that any business logic running on DP corresponds to publicly released code, allowing other parties to verify the business logic used to analyze sensitive data. In general, the complete code base for business logic is available for inspection and audit by all relevant parties. The build is reproducible, and any relevant party can build a DP container. Building a DP container generates a set of encrypted hashes (e.g., platform configuration registers (PCRs)). Therefore, all parties can verify whether the deployed product matches the released code base by comparing hashes. The trusted party publishes the hash to the parties requesting verification of the build logic. For example, KMS is configured to release key material only to images that match the hash generated from building the released logic. This ensures that the private key used to decrypt sensitive information is only available to images corresponding to specific submissions to a specific repository.

[0033] Turning to an example computing system in which the SCP of the present disclosure may be implemented, Figure 1 An example computing system 100 is illustrated. Computing system 100 includes a client computing device 102 (also referred to herein as client device 102) coupled to a cloud platform 122 (also referred to herein as cloud 122) via a network 120. Network 120 may typically include one or more wired and / or wireless communication links, and may include, for example, a wide area network (WAN) such as the Internet, a local area network (LAN), a cellular telephone network, or other suitable type of network or combination of networks. Although examples of the present disclosure are primarily directed to cloud-implemented architectures, it should be understood that the techniques disclosed herein (including techniques for providing a secure execution environment in which to process sensitive data, for generating, sharding, and distributing keys, for managing privacy budgets, and for providing mechanisms by which proprietary business logic can be authenticated) may also be applied in non-cloud systems.

[0034] For example, client device 102 may be a portable device such as a smartphone or tablet computer. Client device 102 may also be a laptop computer, a desktop computer, a personal digital assistant (PDA), a wearable device such as smart glasses, or other suitable computing devices. Client device 102 may include memory 106, one or more processors (CPU) 104, a network interface 114, a user interface 116, and an input / output (I / O) interface 118. Client device 102 may also include Figure 1Components not shown, such as a graphics processing unit (GPU). The client device 102 may be associated with a service user, which is an end user of the services provided by the SCP, as discussed below. The end user operates the client device 102 (or more specifically, a browser or application on the client device 102) that sends requests / events to the service. In order to transmit the request or event to the service, the client device 102 encrypts the request / event using a public key, which the client device 102 may retrieve from a public key repository (e.g., a public key repository server 178). The client device 102 is merely exemplary. As discussed below, the cloud platform 122 may receive incoming events and / or requests from the client device 102, from a browser / application / client process executing on the client device 102, or from another computing device that issues a request on behalf of the client device 102 or forwards a request from the client device 102. In addition, although in Figure 1 Only one client device is illustrated in FIG. 1 , but the computing system 100 may include multiple client devices capable of communicating with the cloud platform 122 .

[0035] The network interface 114 may include one or more communication interfaces, such as hardware, software, and / or firmware for enabling communication via a cellular network, a WiFi network, or any other suitable network, such as the network 120. The user interface 116 may be configured to provide information to a user, such as a response to a request / event received from the cloud platform 122. The I / O interface 118 may include various I / O components (e.g., ports, capacitive or resistive touch-sensitive input panels, keys, buttons, lights, LEDs). For example, the I / O interface 118 may be a touch screen.

[0036] The memory 106 may be a non-transitory memory and may include one or more suitable memory modules, such as random access memory (RAM), read-only memory (ROM), flash memory, other types of permanent memory, etc. The memory 106 may store machine-readable instructions that can be executed on one or more processors 104 and / or special processing units of the client device 102. The memory 106 also stores an operating system (OS) 110, which may be any suitable mobile or general-purpose OS. In addition, the memory 106 may store one or more applications that communicate data with the cloud platform 122 via the network 120. Communicating data may include sending data, receiving data, or both. For example, the memory 106 may store instructions for implementing a browser, an online service, or an application that requests data from / sends data to an application (i.e., business logic) implemented on a DP of a secure execution environment on the cloud platform 122, as discussed below.

[0037] The cloud platform 122 may include multiple servers associated with a cloud provider to provide cloud services via the network 120. The cloud provider is the owner of the cloud platform 122 on which the SCP 126 is deployed. Figure 1 Only one cloud platform is illustrated in FIG. 1 , but SCP 126 may be deployed on multiple cloud platforms, even if the cloud platforms are operated by different cloud providers. The servers providing cloud platform 122 may be distributed across multiple sites to improve reliability and reduce latency. Individual servers or groups of servers within cloud platform 122 may communicate with client devices 102 and with each other via network 120. Example servers that may be included in cloud platform 122 are discussed in further detail below. Although not specifically described herein, the server configuration of cloud platform 122 may be described in detail below. Figure 1 106. Each server in the cloud platform 122 is exemplified, but each server included in the cloud platform 122 may include one or more processors similar to processor 104, which are suitable and configured to execute various software stored in one or more memories similar to memory 106. The server may further include databases, which may be local databases stored in the memory of a particular server or network databases stored in a network-connected memory (e.g., in a storage area network). The server may also include network interfaces and I / O interfaces similar to interfaces 114 and 118, respectively. In addition, it should be understood that although certain components are described as individual servers, in general, the term "server" may refer to one or more servers. In addition, although functions are generally described as being performed by separate servers, some of the functions described herein may be performed by the same server.

[0038] The cloud platform 122 includes an SCP 126, which includes a TEE 124. TEE 124 is a secure execution environment in which DP 128 is isolated. A TEE such as TEE 124 is an environment that provides execution isolation and provides a higher level of security than a conventional system. TEE 124 can use hardware to enforce execution isolation (known as confidential computing). The cloud provider is considered to be the root of trust for SCP 126 and complies with the Terms of Service (ToS) agreement of the cloud platform 122. The hardware manufacturer of the server that provides TEE 124 also has a ToS guarantee, thereby also providing an additional layer of trust. SCP 126 also uses technology to ensure that the state at startup is secure, including the use of a minimalist OS image recommended by the cloud provider, and a TPM / vTPM-based secure boot sequence used in the OS image.

[0039] One or more servers of the cloud platform 122 perform control plane (CP) functions (i.e., to support SCP 126), and one or more servers perform data plane (DP) functions. For example, CP functions including key management and privacy budget services may be distributed on more than one trusted party. All functions of DP 128 are performed by processes within TEE 124. Depending on the specific implementation, there may be more than one TEE per DP server. TEE 124 can be deployed and operated by an administrator. The administrator can audit the logic to be implemented on DP128 and verify against the hash of the binary image to deploy logic 142. On the CP, there may be a front-end server or process 134 that receives external request / event indications (e.g., from client device 102), buffers requests / events until they can be processed by DP 128, and forwards the received request to DP 128. In general, as used herein, a request may also refer to an event or record, or may include one or more events or records, unless otherwise specified. The terms "record," "request," and "event" are used interchangeably herein unless otherwise noted. A request may include encrypted data (e.g., encrypted data representing an interaction of a client device 102 with a website, application, or other online resource), as well as metadata for the request. The metadata may include a timestamp indicating the date and / or time the request was generated. Depending on the nature of the data, the metadata may also include other information, such as an indication of the source of the request or an indication of where the results of processing the data should be output. For example, if the request is generated based on an interaction of a client device 102 with an advertiser's online advertisement placed on an online resource published by a publisher, the metadata may include the advertiser's domain name and / or the publisher's domain name, or other indications of the advertiser and / or publisher.

[0040] In some implementations, there is a third-party server 136 between the client device 102 and the SCP 126. The third-party server 136 (which may include one or more servers and may or may not be hosted on the cloud platform 122) may be responsible for receiving requests from the client device 102 (these requests are encrypted by the client device 102) and later dispatching the encrypted requests to the SCP 126. In some cases, the third party is the administrator of the service. The third-party server 136 does not have the key used to decrypt the requests. The third-party server 136 may, for example, aggregate the requests into batches and store these batches (e.g., on the cloud storage device 160). The third-party server 136 or the cloud storage server 160 may notify the front-end server 134 that it is ready to process the request, and / or the front-end server 134 may subscribe to notifications that are pushed to the front-end server 134 when a batch is added to the cloud storage device 160.

[0041] DP 128 includes a server (which may include one or more servers) including one or more processors 138 (similar to processor 104) and one or more memories 140 (similar to memory 106). Memory 140 includes business logic 142 (also referred to as logic 142) that can be executed by processor 138. Business logic 142 is used to implement any application or service deployed on TEE 124. Memory 140 may also store a key cache 146 that stores encryption keys for encrypting and decrypting communications. In addition, memory 140 includes a CPIO API 144 that includes a function library for communicating with other elements of cloud platform 122 (including components on the CP of SCP 126). CPIO API 144 can be configured to dock with any cloud platform provided by a cloud provider. For example, in a first deployment, SCP 126 can be deployed to a first cloud platform provided by a first cloud provider. DP 128 hosts specific business logic 142, and CPIO API 144 facilitates communication between logic 142 and the first cloud platform. In a second deployment, SCP 126 may be deployed to a second cloud platform provided by a second cloud provider. DP 128 may host the same business logic 142 as the first deployment, and CPIO API 144 is configured to facilitate communication between logic 142 and the second cloud platform. Thus, SCP 126 may be deployed to a different cloud platform without editing the underlying business logic 142, and only CPIO API 144 is configured to interface with a specific cloud platform.

[0042] There may be additional CP-level services provided by servers of the cloud platform 122 that support the SCP 126. For example, the validator server 148 may provide a validator module that can verify whether the business logic 142 complies with the security policy. Figure 1 Although not explicitly illustrated in FIG. 1 , the verifier module may operate within TEE 124. As another example, privacy budget service server 152 may implement a privacy budget service that verifies whether a privacy budget for a user or device has been exhausted. Additionally or alternatively, one or more privacy budget services may be implemented by a trusted party, such as reference 100. Figure 2B discussed.

[0043] Additionally, the cloud platform 122 may include other servers and databases that communicate with the SCP 126, as described in the following paragraphs. These servers may facilitate the CP functions of the SCP 126. Specifically, the CP functions may be distributed across several servers, as will be discussed below. However, the processes of the DP 128 remain within the TEE 124 and are not distributed outside of the TEE 124.

[0044] As mentioned above, the cloud storage 160 may store encrypted batches of requests before they are received by the front-end server 134. The cloud storage 160 may also be used to store responses after the DP 128 has processed the received requests, or to perform storage functions for other components of the cloud platform 122. The queue 162 may be used by the front-end server 134 to store pending requests before they can be analyzed by the DP 128. For example, after receiving a request from the client device 102, the front-end server 134 may receive the request and temporarily store the pending request in the queue 162 until the DP 128 is ready to process the request. As another example, after receiving notification from the third-party server 136 that a batch of requests is stored in the cloud storage 160, the front-end 134 may retrieve the batch and place the batch in the queue 162, where the batch awaits analysis by the DP 128.

[0045] KMS service 164 provides a KMS that generates, deletes, distributes, replaces, rotates, and otherwise manages encryption keys. The functions of KMS 164 may be performed by one or more servers. Thus, KMS 164 may be a cloud KMS. Trusted party 1 server 166 and trusted party 2 server 172 are servers associated with trusted party 1 and trusted party 2, respectively, that provide the functions of each trusted party. Although Figure 1 Only two trusted parties are illustrated, but the cloud platform 122 may include multiple trusted parties. Each trusted party may manage an instance of a privacy budget (such as reference Figure 2B and Figure 5 Detailed description further), and the logic 142 implemented on DP128 can also be audited to verify the built product against the hash of the published logic. The trusted party owns the creation and management of asymmetric keys for encrypting and decrypting user data. The trusted party can securely generate keys and publish public keys to the world. The private key can be bit-sharded into two parts (one shard under the control of each trusted party, but any number of N shards can also be supported, such as when there are N trusted parties). Envelope encryption technology can be used, in which each trusted party encrypts the shard for each key of the trusted party using the symmetric key of the KMS, and the encrypted shard is stored in the repository of the trusted party. Envelope encryption allows the envelope to be rotated without having to rotate the keys within the envelope. The public keys can be stored and managed by the public key repository server 178. Additionally or alternatively, the KMS server 164 can manage the public keys.

[0046] The computing system 100 may also include a public security policy storage 180, which may be located on or off the cloud platform 122. The public security policy storage 180 stores security policies so that the security policies are accessible to the public (e.g., by the client device 102, by components of the cloud platform 122). A security policy (also referred to herein as a policy) describes what actions or fields are allowed in order to constitute the output of a service. A policy may also be described as a machine-readable and machine-enforceable privacy design document (PDD).

[0047] Next reference Figure 2A , example architecture 200A illustrates connections between components and software elements of computing system 100. Client device 102 may retrieve a public key (e.g., from public key repository server 178) in order to address a request to a service being implemented on DP 128 (i.e., by business logic 142). For example, client device 102 may initiate a request to access content provided by a service, or may emit an event including user behavior data.

[0048] The encrypted request from the client device 102 is first received by the front-end module 234 (i.e., a module implemented by the front-end server 134) of the SCP 126. In some specific implementations, the request is first received by a third party that batches the request before notifying the front-end 234 (or causing the front-end 234 to be notified). The notification to the front-end 234 may include the location within the cloud storage device 160 where the encrypted request resides (e.g., the location of the cloud storage bucket), and may include (e.g., by including metadata indicating such information) an indication of where the output from the DP 128 should be output. In such a case, the front-end 234 may retrieve the encrypted request from the cloud storage device 160. In any event, the front-end 234 passes the encrypted request to the DP 128 using the functions defined by the CPIO API 144. The front-end 234 may store the encrypted request in the queue 162 until the DP 128 is ready to process the requests and retrieve the requests from the queue 162. DP 128 decrypts the requests and processes the requests according to business logic 142. Decrypting the requests may include communicating with KMS 164 (e.g., a cloud KMS implemented by a distributed server) to retrieve and assemble a private key for decrypting the requests and / or may include communicating with a trusted party, such as in Figure 2B These are examples of cloud native services integrated with SCP 126, but the concept extends to other cloud infrastructure and services, with SCP 126 mediating between these services and business logic 142 by using CPIO API 144 to translate semantics, making business logic 142 agnostic to the specific cloud environment.

[0049] Processing these requests may include using CPIO API 144 functionality to communicate with a privacy budget service 252 (e.g., implemented by a privacy budget service server 152) to check the privacy budget and ensure compliance with the privacy budget. The privacy budget tracks processed requests and events. For example, there may be a maximum number of requests from a particular user that can be processed during a particular computation or cycle. Ensuring compliance with the privacy budget prevents parties analyzing the output from DP 128 from extracting information about a particular user. By checking compliance with the privacy budget, DP 128 provides a differentially private output. As will be described with reference to Figures 4 to 5 As discussed, a privacy budget may be allocated to a request group such that a request group may only be processed a maximum number of times.

[0050] The results from processing the request may be encrypted by DP 128 and may be edited and / or aggregated so that the output does not reveal information about a particular user. DP 128 may store the results, for example, in cloud storage 160, where the results may be retrieved by parties having a decryption key for the results. As an example, if the results are processed for a third-party server 136, DP 128 may encrypt the results using a key that the third-party server 136 may decrypt.

[0051] Steering Figure 2B , architecture 200B is similar to architecture 200A, except that additional details regarding key management and privacy budget are illustrated. Figure 2A compared to, Figure 2B Also illustrated are trusted party 1 server 166 (referred to herein as trusted party 1 166 for the sake of brevity), trusted party 2 server 172 (referred to herein as trusted party 2 172 for the sake of brevity), and public key distribution service 278. Public key distribution service 278 provides public keys to client devices 102, which can use these public keys to address requests to DP 128, front end 234, or third-party server 136 (which aggregates requests) Figure 2B164). The public key distribution service 278 may be operated by the public key repository server 178 or by the KMS server 164. Trusted party 1 166 includes a key cache 268 containing an encrypted shard-1 key (i.e., an encrypted first portion of a private key), and trusted party 2 172 includes a key cache 274 containing an encrypted shard-2 key (i.e., an encrypted second portion of a private key). Each of the trusted parties 166, 172 may also provide a privacy budget service 270, 276, and may each manage an instance of a privacy budget. Distributing the management of the privacy budget to two trusted parties helps ensure that no one trusted party can tamper with the privacy budget. Both privacy budget services 270, 276 should enforce the same privacy budget; therefore, if the two services return different outputs, the SCP 126 may recognize that one of the trusted parties 166, 172 has tampered with the privacy budget. Figure 2B The architecture exemplified in prevents any one trusted party from having full control over the private decryption keys or privacy budget. A single trusted party cannot individually provide an unlimited budget to any user, and therefore a single trusted party cannot repeatedly aggregate the same batch of data.

[0052] The elements illustrated in architecture 200B may implement Figure 3 and Figure 5 The actions illustrated in .

[0053] Example scenario including shard keys

[0054] exist Figure 3 During the example scenario 300 illustrated in , the client device 102 retrieves 302 a public key from a public key distribution service 278 (which may be implemented by a public key repository server 178, or may be provided by a KMS 164). The client device 102 encrypts 304 a request for a service implemented on a TEE 124 using a public key associated with the service. The client device 102 then transmits 306 the encrypted request to a front-end server 134. As previously explained, in some implementations, a third-party server 136 receives the encrypted request before the request reaches the front-end server 134, and stores the encrypted request in a cloud storage device 160, where the encrypted request can be retrieved by the front-end server 134. The front-end server 134 passes 308 the encrypted request to the DP 128. The business logic 142 retrieves the encrypted request for processing, and attempts to retrieve a decryption key from the key cache 146. If business logic 142 retrieves the decryption key, the scenario continues from event 308 to event 320 .

[0055] As applicable to Figure 3 and Figure 5, the service implemented on the TEE 124 may be an aggregation service. In this case, the encrypted request may include events detected by an API implemented on the client device 102. Each event may be a transformation, where the transformation may correspond to the client device 102 interacting with an advertisement provided by a publisher. The aggregation service may be a service configured to provide an aggregate report that provides aggregate statistics of events from client software (e.g., a browser, an application, or a process implemented by the client device 102). Thus, in such an example, the business logic 142 may include performing event aggregation and outputting an output report.

[0056] Return to reference Figure 3 , if the decryption key does not exist in the key cache 146, the business logic 142 uses the CPIOAPI 144 to request the decryption key. More specifically, the DP 128 transmits 310 a request for key shard 1 to the trusted party 1 166, and transmits 312 a request for key shard 2 to the trusted party 2 172. In response, the trusted party 1 166 transmits 314 key shard 1 to the DP 128, and the trusted party 2 172 transmits 316 key shard 2 to the DP 128. When the key shards 1 and 2 are received 314, 316 by the DP 128, the key shards may be encrypted using the symmetric keys generated by the trusted party 1 166 and the trusted party 2 172, respectively. Therefore, the DP 128 may also need to request the key shards to be decrypted by, for example, the KMS 164 (i.e., the cloud KMS), which may store the symmetric keys operated by the trusted parties 166, 172.

[0057] Prior to transmitting 314 and 316 the key shards to DP 128, the trusted parties 166, 172 may first verify that the business logic 142 corresponds to code that is publicly released on a commit to the code repository. This may be accomplished by proof. The code base of the business logic 142 is available for inspection and audit by all interested parties (client devices 102, trusted parties 166, 172, cloud platforms 122, administrators, third parties, etc.). As discussed above, any interested party may build a DP container that includes the business logic 142 and generate PCRs for the published logic. Thus, any party may verify that the business logic 142 built and deployed on DP 128 matches the published code base by comparing the PCRs of the deployed business logic 142 with the PCRs of the published code base. CPIO API 144 can communicate the PCRs of deployed business logic 142 to other parties (ie, to client device 102, to cloud platform 122, or other components of computing system 100) to prove that deployed business logic 142 corresponds to a released code base and has not been modified.

[0058] Thus, in the request 310, 312 sent by the DP 128, using the CPIO API 144, the business logic 142 may include the PCRs of the deployed business logic 142. Alternatively or additionally, the trusted party 166, 172 may request the PCRs. The trusted party 166, 172 may then confirm that the deployed image matches the PCRs of the published code base. After performing this verification process, the trusted party 166, 172 may release 314, 316 the key shards to the DP 128. Likewise, the KMS 164 may also verify that the binary image deployed on the DP 128 is certified, and only release symmetric decryption keys to certified DP binaries running in the TEE 124.

[0059] The business logic 142 then assembles 318 a private decryption key from the key shards and may store the assembled private key in the key cache 146. The assembly of the private key occurs only in the TEE 124, and the CPIO API 144 utilizes secure channels and authentication when communicating with the trusted parties 166, 172. Since each trusted party 166, 172 contains only a portion of each key, after the key shards are combined, the entire private key exists only within the secure TEE 124. Therefore, this prevents any party from leaking plaintext data or private keys. The business logic 142 retrieves 320 the private key from the key cache 146 and uses the private key to decrypt the request. Before processing the request, the business logic 142 may verify 322 that there is a privacy budget available to process the request. According to the CPIO API 144, the business logic 142 may perform such verification by communicating with the privacy budget service 270, 272, and / or 154. Verifying that there is sufficient privacy budget available to process the request may include sending a request to trusted party 1 166 and trusted party 2 172 (as will be referred to in the examples below) that each manages a privacy budget. Figure 5 discussed) both transmit requests.

[0060] Assuming that there is sufficient privacy budget, business logic 142 may process 324 the request. Before storing 326 the results of the processing, in some cases business logic 142 again checks whether the privacy budget is still available. Depending on the specific implementation, the business logic 142 may store 326 the results of the processing before processing the request. Figure 3), after processing the request and before storing the processed result, or before processing the request and before storing the processed result, check the privacy budget. The business logic 142 may then encrypt 326 the result and store the result for later retrieval. For encryption, the business logic 142 may use a system-provided or customer-managed encryption key (CMEK) for static encryption. The result is then ready to be retrieved and consumed by any party (e.g., an administrator, client device 102, third-party server 136, etc.) that possesses the CMEK key.

[0061] Example scenarios include privacy budgets and distributed trust on a per-group basis

[0062] Steering Figure 4 , which is a block diagram illustrating an example technique for grouping records and allocating a privacy budget on a per-group basis. TEE 124 may receive several data sets 404A-404E individually or in one or more batches. TEE 124 may receive data sets (such as Figure 3 ), or receiving a data set from a third-party server 136. For example, each data set may include an encrypted request discussed above with reference to event 306. Thus, each data set may include encrypted data (e.g., representing an interaction of the client device 102 with an online resource) and metadata associated with the encrypted data. Metadata may also be included in the encrypted data. As used herein, a "data set" may correspond to the "requests," "reports," and "events" discussed above. In addition, as described above, the metadata may be a timestamp indicating the date and / or time when the data set was generated. Depending on the nature of the data, the metadata may also include other information, such as an indication of the source of the data set or an indication of where the results of processing the data should be output. For example, if the request is generated based on the client device 102 interacting with an advertiser's online advertisement placed on an online resource published by a publisher, the metadata may include the advertiser's domain name and / or the publisher's domain name, or other indications of the advertiser and / or publisher. The metadata or a portion of the metadata may be included in plain text (i.e., unencrypted).

[0063] In order to reduce the computing resources required to monitor the privacy budget of each data set 404A-E, the TEE 124 may classify the data sets 404A-404E into groups and allocate the privacy budget to each group rather than to each individual data set 404A-E. The metric used to classify the data sets 404A-E may vary depending on the specific implementation. In general, the classification may be based on metadata included in each data set 404A-E. In one example, the classification may be based on a timestamp included in the metadata. The example timestamp may indicate both time and data, such as 2021-01-01 12:29:33AM. In order to group the data sets, the timestamp may be rounded down to the hour (e.g., by rounding down the example timestamp 2021-01-01 12:29:33AM to 2021-01-01 12:00:00AM). As another example, the classification may be based on a publisher domain name and / or an advertiser domain name included in the metadata. In yet another example, the classification may be based on both timestamp and publisher domain name and / or advertiser domain name. In such an example, each group may include data sets from the same advertiser domain name included in the same rounded-down time window.

[0064] exist Figure 4 In the example of , five data sets 404A-404E are classified into two groups: group 408A and 408B based on metadata included in each of the data sets 404A-E. For example, if the data sets 404A-E are classified based on timestamps and advertiser domain names, data sets 404A-404C may be generated from the same advertiser domain name during the same time window, and data sets 404D and 404E may be generated from the same advertiser domain name during a later time window. Then, a first privacy budget 411A is assigned to group 408A including data sets 404A-404C, and a second privacy budget 411B is assigned to group 408B including data sets 404D and 404E. The sizes of the first privacy budget 411A and the second privacy budget 411B may be equal. As will be described with reference to Figure 5 Further discussing, when group 408A is analyzed or included in a batch being processed, privacy budget 411A is decremented by the privacy budget service, which may be a distributed privacy budget service. Similarly, when group 408B is analyzed, privacy budget 411B is decremented. Figure 3Before processing the group using business logic 142 at event 324 in , and / or before storing the analysis of group 408A, TEE 124 can verify that there is sufficient privacy budget to analyze group 408A and / or store the analysis of the group. This verification can be performed on a per-group basis rather than a per-record basis, which reduces the computational resources required to maintain differential privacy for multiple records.

[0065] In order to track the privacy budget for each group, the TEE 124 may generate a privacy budget key for each group. The privacy budget key may be generated based on information included in the metadata (such as by hash information included in the metadata). More specifically, the privacy budget key may be generated based on information used to classify the data set into groups, so that each group is assigned a unique privacy budget key. For example, an example privacy budget key may be generated by hashing an advertiser domain name, a publisher domain name, or both an advertiser domain name and a publisher domain name. An example privacy budget key may also be generated based on a rounded-down time window corresponding to the group (e.g., an hourly window on a particular date), or a privacy budget key and an indication of a rounded-down time window may be used to track the privacy budget for the group. In some implementations, the TEE 124 generates a privacy budget key for the group. In other implementations, the privacy budget key may be generated by the originator of the data set and included in the data set. For example, if a data set is generated based on an interaction between a client device 102 and a publisher, when the publisher generates a data set for sending to the TEE 124 for analysis, the publisher may hash information such as the advertiser domain name and / or the publisher domain name to generate a privacy budget key, and include the privacy budget key in the data set. When the TEE 124 queries a privacy budget service (e.g., trusted party 1 166 and trusted party 2 172) to determine whether there is sufficient privacy budget to process the data set, the TEE 124 may include the privacy budget key in the request and, if not indicated by the privacy budget key itself, the time window.

[0066] Steering Figure 5 , scenario 500 depicts two techniques of the present disclosure: distributing a privacy budget on a per-group basis, and verifying a privacy budget using a distributed privacy budget service. It should be understood that these techniques can be implemented independently of each other. For example, in one implementation, TEE 124 can allocate a privacy budget on a per-group basis, but rely on only a single trusted party to verify the privacy budget. In another implementation, TEE 124 can allocate a privacy budget on a per-record basis, but rely on multiple trusted parties to verify the privacy budget. In yet another implementation, Figure 4 As illustrated in , TEE 124 can either allocate a privacy budget on a per-group basis or rely on multiple trusted parties to verify the privacy budget.

[0067] Although Figure 5 Not shown, but scenario 500 may begin similarly to scenario 300 (i.e., by including events similar to events 302, 304, 306, and 308. DP 128 may receive 508A-508N multiple encrypted requests (i.e., data sets, such as Figure 4 DP 128 may then analyze metadata included in each of the encrypted requests to classify the encrypted requests into groups (e.g., similar to groups 408A, 408B). DP 128 allocates 530 a privacy budget to each group. The privacy budgets for different groups may be the same so that each group has the same initial privacy budget. The privacy budget for each group may be a number N, where N may be an integer equal to 1 or greater. If the privacy budget is equal to N=1, the group may be analyzed only once before the privacy budget for the group is completely consumed.

[0068] Although Figure 5 Although not shown, scene 500 may also include events 310, 312, 314, 316, 318, and 320, such as Figure 3 exemplified in .

[0069] Next, DP 128 verifies 522 that there is sufficient privacy budget to process the group. Each of the trusted parties 166, 172 may implement a privacy budget service that maintains a respective instance of a privacy budget for the group. The first instance of the privacy budget maintained by trusted party 1 166 should be the same as the second instance of the privacy budget maintained by trusted party 2 172, because each of trusted parties 166, 172 should enforce the same privacy budget for a given group. Therefore, trusted party 1 166 and trusted party 2 172 may be referred to as implementing a distributed privacy budget service. Each of the trusted parties 166, 172 is independent of each other (i.e., implemented on separate servers). Furthermore, it is not required that the trusted parties 166, 172 be implemented on the same cloud platform; each may be implemented on a different cloud platform.

[0070] Verifying 522 the privacy budget may include sending 532 to trusted party 1 166 a verification that (according to trusted party 1 166) sufficient privacy budget exists to process the group's request, and sending 534 to trusted party 2 172 a verification that (according to trusted party 2 172) sufficient privacy budget exists to process the group's request. Although (for clarity) in Figure 5166, 172 as occurring sequentially, but DP 128 may send 532, 434 the first request and the second request simultaneously. Each request may include a privacy budget key (and a time window rounded down if not indicated by the privacy budget key) so that each trusted party can track which group the request belongs to. Each of the trusted parties 166, 172 may then check whether there is enough privacy budget to process the group (i.e., check whether the privacy budget is at least 1) according to each of the respective instances of the privacy budget of each. If there is enough privacy budget according to trusted party 1 166, trusted party 1 166 decrements the instance of the privacy budget of the trusted party. Similarly, if there is enough privacy budget according to trusted party 2 172, trusted party 2 172 decrements the instance of the privacy budget of the trusted party. According to trusted party 1 166 , trusted party 1 166 transmits 536 a response to DP 128 indicating whether there is a sufficient privacy budget, and according to trusted party 2 172 , trusted party 2 172 transmits 538 a response to DP 128 indicating whether there is a sufficient privacy budget.

[0071] Based on the two responses, DP 128 verifies 540 whether there is sufficient privacy budget to process the group. The two responses must match and indicate that there is sufficient privacy budget for the process to continue. If the responses do not match, or one or both of the trusted parties 166, 172 indicate that there is insufficient privacy budget to process the group, the process is aborted and the group is not analyzed (i.e., the process does not continue to event 524 or 526). Verifying 522 that there is sufficient privacy budget to process the group may include events 532, 534, 536, 538, and 540. DP 128 may communicate 532, 534, 536, 538 with the trusted parties 166, 172 according to the CPIO API 144.

[0072] In scenario 500, after verifying 522 that there is a privacy budget available for processing the group, DP 128 may then process 524 the group (e.g., by using business logic 142 to analyze the group, similar to event 324). DP 128 may analyze the group within a larger batch of requests, but enforces that each request included in the group is analyzed within the same group and, if applicable, within the same batch. DP 128 may then encrypt 526 the results of the analysis at event 524 and store the results for later retrieval, similar to event 326. Although in scenario 500, DP 128 verifies 522 the privacy budget before analyzing 524 the group, in other scenarios, DP 128 may analyze 524 the group and verify 522 the privacy budget before storing 526 the results. If there is not enough privacy budget, DP 128 discards the results of the analysis at event 524 and does not store 526 the results.

[0073] DP 128 and trusted parties 166 and 172 may execute Figure 5 Additional operations not shown are performed to ensure that consumption of the privacy budget for the group occurs in an atomic manner, i.e., the privacy budget on all instances is either consumed or not consumed. For example, after verifying 522 that there is a privacy budget available for processing the group, DP 128 may transmit a commit message to each of the trusted parties 166, 172 (e.g., simultaneously) so that each of the trusted parties 166, 172 simultaneously decrements the respective instances of the privacy budget. The trusted parties 166, 172 may then transmit a notification to DP 128 that the privacy budget instances for each trusted party 166, 172 have been successfully consumed. DP 128 may then continue processing 524 the group or store 526 the results of any analysis after confirming that all instances of the privacy budget have been successfully consumed. If such confirmation is not received from either of the trusted parties 166, 172, the process may be aborted.

[0074] Example methods for privacy budgeting and distributed trust on a per-group basis

[0075] Figure 6 6 is a flow chart illustrating an example method 600 for managing a privacy budget on a per-group basis. The method 600 may be implemented by one or more servers (e.g., servers supporting SCP 126). The method 600 may begin at block 602, where the server receives a plurality of data sets (e.g., events 508A-508N). Each data set may include encrypted data and metadata for the data set. For example, the encrypted data may represent an interaction of a user with an online resource (e.g., an advertisement, or other interactive element of a browser or application) via a client device 102.

[0076] At box 604, the server classifies the plurality of data sets into one or more groups (e.g., event 528) based on metadata included in each data set. For each group, the server may assign a privacy budget, indicating the number of times the group may be analyzed, such that the privacy budget for the one or more groups is performed on a per-group basis rather than on a per-request basis. The privacy budget for each group may be the same initial privacy budget. Classifying the data sets may include classifying based on a timestamp included in the metadata for the data sets. For example, the timestamp may be rounded down to the hour so that each group corresponds to data received or generated during a specific hour window on a date.

[0077] At block 606, for a group in the one or more groups, the server queries whether there is sufficient privacy budget for the group to store the results of the analysis of the group. In some implementations, before the query, the server analyzes the data set included in the group to produce an output (e.g., event 524). If, based on the query, there is sufficient privacy budget for the group, the server stores the output (e.g., event 526). If there is not sufficient privacy budget for the group, the server suppresses storing the output and instead discards the output. In other implementations, the server queries whether there is sufficient privacy budget before analyzing the data set in the group. If there is sufficient privacy budget, the server analyzes the data set included in the group and stores the output of the analysis. If there is not sufficient privacy budget, the server suppresses analyzing the data set in the group. While analyzing the data set or after analyzing the data set, the server decrements the privacy budget for the group, or decrements the privacy budget by notifying the privacy budget service (e.g., trusted party 1 166 and / or trusted party 2 172) to decrement the privacy budget. In some implementations, the privacy budget service decrements the privacy budget in response to receiving a request to verify whether there is an available privacy budget for the group.

[0078] Querying whether there is sufficient privacy budget for the group may include (e.g., via an API call) sending a request to consume the privacy budget for the group to a privacy budget service (e.g., event 322, event 522, 532, 534) (e.g., in a scenario involving a distributed privacy budget service, to trusted party 1 166, trusted party 2 172, or to both trusted parties 166, 172). The server may then receive a response from the privacy budget service indicating whether there is sufficient privacy budget to consume the privacy budget for the group. In a scenario involving a distributed privacy budget service, the query may include sending a first request to consume the privacy budget for the group to a first privacy budget service (e.g., trusted party 1 166), sending a second request to consume the privacy budget for the group to a second privacy budget service (e.g., trusted party 2 172), and receiving a first response and a second response from the first privacy budget service and the second privacy budget service, respectively (e.g., events 536, 538). The server may determine whether there is sufficient privacy budget based on the two responses. The two responses must match and indicate that sufficient privacy budget exists for the server to continue analyzing the datasets included in the group or to store the results of any analysis.

[0079] In addition, the query may include sending a request to the privacy budget service, wherein the request includes a privacy budget key representing the group. The privacy budget key may be generated by the server based on metadata common to the group (e.g., advertiser domain name, publisher domain name, timestamp, which depends on how the data sets are classified into groups) (such as by hashing at least a portion of the metadata). For example, the key may be generated based on hashing an indication in the metadata of the domain name of the collection from which the data set was received (e.g., a publisher domain name that published an advertisement and / or an advertiser domain name of the advertisement).

[0080] Figure 7 7 is a flow chart illustrating an example method 700 for managing a privacy budget using a distributed privacy budget service. The method 700 may be implemented by one or more servers (e.g., servers supporting SCP 126). The method 700 may begin at block 702, where a server receives a request to analyze a data set (e.g., events 308, 508A-508N). The data set may include encrypted data and metadata for the data set. For example, the encrypted data may represent an interaction of a user with an online resource (e.g., an advertisement, or other interactive element of a browser or application) via a client device 102. Depending on the specific implementation, the server may receive multiple data sets, and the request may be to analyze multiple data sets. The data set may be associated with a privacy budget representing the number of times the data set may be analyzed.

[0081] At block 704, the server sends a first request (e.g., event 532) to a first server that implements a first privacy budget service (e.g., trusted party 1 166) to verify whether sufficient privacy budget exists to analyze the data set. The first privacy budget service maintains a first instance of a privacy budget associated with the data set. Similarly, at block 706, the server sends a second request (e.g., event 534) to a second server that implements a second privacy budget service (e.g., trusted party 2 172) to verify whether sufficient privacy budget exists to analyze the data set. The second privacy budget service is independent of the first privacy budget service (i.e., implemented on a server independent of the server that implements the first privacy budget service) and maintains a second instance of a privacy budget associated with the data set.

[0082] At block 708, the server receives a first response from the first server according to the first privacy budget service indicating whether there is sufficient privacy budget (e.g., event 536). Similarly, at block 710, the server receives a second response from the second server according to the second privacy budget service indicating whether there is sufficient privacy budget (e.g., event 538). Based on the first response and the second response, at block 712, the server processes the data set. If both responses indicate that there is sufficient privacy budget, the processing at block 712 includes analyzing the data set (e.g., event 524) and storing the results of the analysis (e.g., event 526). In some implementations, the server may first analyze the data set and, if both responses indicate that there is sufficient privacy budget, store the results of the analysis. If there is sufficient privacy budget to continue the analysis and store the results of the analysis, the server causes the first privacy budget service and the second privacy budget service to each decrement their respective instances of the privacy budget (e.g., as described with reference to block 606). The first privacy budget service and the second privacy budget service decrement their budgets atomically (i.e., either both services will decrement their respective instances of the privacy budget because there is sufficient privacy budget to consume, or neither service will decrement their respective instances of the privacy budget). If one or both responses indicate that there is not sufficient privacy budget, the server refrains from analyzing the data set, or if the data set was analyzed before determining that there is not sufficient privacy budget, refrains from storing the results of the analysis.

[0083] In some implementations, method 700 can be combined with aspects of method 600 such that the data set can be part of a data set group to which a privacy budget is allocated. In such implementations, the first request and the second request to the privacy budget service are requests to verify a privacy budget for the group.

[0084] Example

[0085] The following example list reflects the various embodiments that are clearly expected by the present disclosure. It will be readily understood by those skilled in the art that the following examples are neither limitations on the embodiments disclosed herein nor exhaustive of all embodiments that can be conceived from the above disclosure, but are meant to be exemplary in nature.

[0086] Example 1. A method for managing a privacy budget in one or more servers, the method comprising: receiving a request to analyze a data set, the data set being associated with a privacy budget representing the number of times the data set can be analyzed; sending a first request to a first server implementing a first privacy budget service to verify whether there is sufficient privacy budget to analyze the data set, the first privacy budget service maintaining a first instance of the privacy budget associated with the data set; sending a second request to a second server implementing a second privacy budget service to verify whether there is sufficient privacy budget to analyze the data set, the second privacy budget service being independent of the first privacy budget service, and the second privacy budget service maintaining a second instance of the privacy budget associated with the data set; receiving a first response from the first server indicating whether there is sufficient privacy budget according to the first privacy budget service; receiving a second response from the second server indicating whether there is sufficient privacy budget according to the second privacy budget service; and processing the data set based on the first response and the second response.

[0087] Example 2. The method of Example 1, wherein the processing comprises: determining that both the first response and the second response indicate that there is a sufficient privacy budget; and in response to the determination, analyzing the data set and storing results of the analysis.

[0088] Example 3. The method of Example 1, wherein the processing comprises: analyzing the data set; determining that both the first response and the second response indicate that there is a sufficient privacy budget; and in response to the determination, storing a result of the analysis.

[0089] Example 4. The method of example 2 or 3, further comprising causing the first privacy budget service and the second privacy budget service to atomically decrement the first instance of the privacy budget and the second instance of the privacy budget.

[0090] Example 5. The method of example 1, wherein the processing comprises: determining that the first response indicates that sufficient privacy budget exists and the second response indicates that sufficient privacy budget does not exist; and responsive to the determination, refraining from analyzing the data set.

[0091] Example 6. The method of Example 1, wherein the processing comprises: analyzing the data set; determining that the first response indicates that sufficient privacy budget exists and the second response indicates that sufficient privacy budget does not exist; and responsive to the determination, suppressing storage of results of the analysis.

[0092] Example 7. The method of example 1, wherein the processing comprises: determining that both the first response and the second response indicate that there is not sufficient privacy budget; and responsive to the determination, refraining from analyzing the data set.

[0093] Example 8. The method of Example 1, wherein the processing comprises: analyzing the data set; determining that both the first response and the second response indicate that there is not sufficient privacy budget; and responsive to the determination, suppressing storage of results of the analysis.

[0094] Example 9. A method according to any of the preceding examples, wherein: the dataset is included in a dataset group, the privacy budget is defined for the dataset group, and receiving the request to process the dataset includes receiving the request to process the dataset group.

[0095] Example 10. The method of any of the preceding examples, wherein sending the first request comprises sending the first request via an application programming interface (API) call to the first privacy budget service.

[0096] Example 11. A method for managing a privacy budget in one or more servers, the method comprising: receiving a plurality of data sets, each of the plurality of data sets comprising encrypted data and metadata for the data set; classifying the plurality of data sets into one or more groups based on a corresponding plurality of metadata included in the plurality of data sets; and for one of the one or more groups, querying whether there is sufficient privacy budget for the group to store results of an analysis of the group, the privacy budget for the group representing the number of times the group can be analyzed.

[0097] Example 12. The method according to Example 11 further includes: before querying whether there is a sufficient privacy budget for the group, analyzing the data set included in the group to produce an output; and in response to determining that there is a sufficient privacy budget for the group based on the query, storing the output.

[0098] Example 13. The method according to Example 11 further includes: before querying whether there is a sufficient privacy budget for the group, analyzing a data set included in the group to produce an output; and in response to determining that there is a sufficient privacy budget for the group based on the query, suppressing storage of the output.

[0099] Example 14. According to the method of Example 11, the method further includes: before analyzing the data set included in the group, querying whether there is a sufficient privacy budget for the group; and in response to determining that there is a sufficient privacy budget for the group based on the query: analyzing the data set included in the group to generate an output, and storing the output.

[0100] Example 15. The method according to Example 11 further includes: before analyzing the data set included in the group, querying whether there is a sufficient privacy budget for the group; and in response to determining that there is no sufficient privacy budget for the group based on the query, suppressing analysis of the data set included in the group.

[0101] Example 16. The method of Example 12 or 14, further comprising: decrementing the privacy budget for the group.

[0102] Example 17. A method according to any one of Examples 11 to 16, wherein: for each data set, the metadata for the data set includes a timestamp; and classifying the multiple data sets into the one or more groups includes classifying the multiple data sets based on the corresponding multiple timestamps corresponding to the multiple data sets.

[0103] Example 18. The method of Example 17, wherein classifying the plurality of data sets based on the respective plurality of timestamps comprises classifying the plurality of data sets into a plurality of groups such that each of the one or more groups corresponds to an hour window on a date.

[0104] Example 19. A method according to any one of Examples 11 to 18, wherein each of the one or more groups has an equal initial privacy budget.

[0105] Example 20. A method according to any one of Examples 11 to 19, wherein querying whether there is sufficient privacy budget for the group includes: sending a request to a privacy budget service to consume the privacy budget for the group; and receiving a response from the privacy budget service indicating whether there is sufficient privacy budget to consume the privacy budget for the group.

[0106] Example 21. A method according to Example 20, wherein the privacy budget service is a first privacy budget service, the request is a first request, and the response is a first response, and wherein the query further includes: sending a second request to a second privacy budget service to consume the privacy budget for the group; receiving a second response from the second privacy budget service indicating that there is sufficient privacy budget to consume the privacy budget for the group; and determining whether there is sufficient privacy budget based on whether both the first response and the second response indicate that there is sufficient privacy budget.

[0107] Example 22. A method according to any one of Examples 11 to 21, wherein the group comprises a set of data sets, the method further comprising: using metadata included in the set of data sets to generate a key representing the group, wherein the query comprises sending a request including the key.

[0108] Example 23. The method of Example 22, wherein generating the key comprises generating the key by applying a hash operation to at least a portion of the metadata included in the set of data sets.

[0109] Example 24. The method of Example 23, wherein the at least a portion of the metadata included in the set of data sets indicates a domain name from which the set of data sets was received.

[0110] Example 25. The method of any one of Examples 11 to 24, wherein for a dataset in the plurality of datasets, the encrypted data represents an interaction of a user with an online resource.

[0111] Example 26. A computing system for managing a privacy budget, the computing system comprising: one or more servers; and a non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by the one or more processors, causing the computing system to implement a method according to any of the preceding examples.

[0112] Additional considerations

[0113] The following additional considerations apply to the foregoing discussion.

[0114] The client device (e.g., client device 102) in which the techniques of the present disclosure may be implemented may be any suitable device capable of wireless communication, such as a smartphone, a tablet computer, a laptop computer, a desktop computer, a mobile game console, a point of sale (POS) terminal, a health monitoring device, a drone, a camera, a media streaming dongle or another personal media device, a wearable device such as a smart watch, a wireless hotspot, a femtocell, or a broadband router. In addition, in some cases, the client device may be embedded in an electronic system such as a head unit of a vehicle or an advanced driver assistance system (ADAS). Further, the client device may operate as an Internet of Things (IoT) device or a mobile Internet device (MID). Depending on the type, the client device may include one or more general purpose processors, a computer readable memory, a user interface, one or more network interfaces, one or more sensors, and the like.

[0115] Certain embodiments are described in the present disclosure as including logic or multiple components or modules. A module may be a software module (e.g., a code stored on a non-transitory machine-readable medium) or a hardware module. A hardware module is a tangible unit that can perform certain operations and can be configured or arranged in a particular manner. A hardware module may include a dedicated circuit system (circuitry) or logic that is permanently configured to perform certain operations (e.g., as a dedicated processor, such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)). A hardware module may also include a programmable logic or circuit system (e.g., as contained in a general-purpose processor or other programmable processor) that is temporarily configured by software to perform certain operations. The decision to implement a hardware module with a dedicated and permanently configured circuit system or with a temporarily configured circuit system (e.g., configured by software) may be driven by cost and time considerations.

[0116] When implemented in software, the techniques may be provided as part of an operating system, a library used by multiple applications, a specific software application, etc. The software may be executed by one or more general-purpose processors or one or more special-purpose processors.

Claims

1. A method for managing a privacy budget in one or more servers, the method comprising: receiving a request to analyze a data set, the data set being associated with a privacy budget corresponding to a number of times the data set can be analyzed; sending a first request to a first server implementing a first privacy budget service for maintaining a first instance of the privacy budget to determine whether there is sufficient privacy budget to analyze the data set; sending a second request to a second server implementing a second privacy budget service to maintain a second instance of the privacy budget to determine whether sufficient privacy budget exists to analyze the data set, wherein the second privacy budget service is independent of the first privacy budget service; receiving, according to the first privacy budget service, a first response from the first server indicating whether there is a sufficient privacy budget; receiving, according to the second privacy budget service, a second response from the second server indicating whether there is sufficient privacy budget; as well as The data set is processed based on the first response and the second response.

2. The method according to claim 1, wherein the processing comprises: determining that both the first response and the second response indicate that a sufficient privacy budget exists; as well as Responsive to the determination, the data set is analyzed and results of the analysis are stored.

3. The method of claim 1, wherein the processing comprises: analyzing the data set; determining that both the first response and the second response indicate that a sufficient privacy budget exists; as well as Responsive to the determining, results of the analyzing are stored.

4. The method according to claim 2 or 3, further comprising: The first privacy budget service and the second privacy budget service are caused to atomically decrement the first instance of the privacy budget and the second instance of the privacy budget.

5. The method of claim 1, wherein the processing comprises: determining that the first response indicates that sufficient privacy budget exists and the second response indicates that sufficient privacy budget does not exist; as well as Responsive to the determination, analyzing the data set is suppressed.

6. The method of claim 1, wherein the processing comprises: analyzing the data set; determining that the first response indicates that sufficient privacy budget exists and the second response indicates that sufficient privacy budget does not exist; as well as Responsive to the determination, storage of results of the analysis is suppressed.

7. The method of claim 1, wherein the processing comprises: determining that both the first response and the second response indicate that there is not enough privacy budget; as well as Responsive to the determination, analyzing the data set is suppressed.

8. The method of claim 1, wherein the processing comprises: analyzing the data set; determining that both the first response and the second response indicate that there is not enough privacy budget; as well as Responsive to the determination, storage of results of the analysis is suppressed.

9. A method according to any one of the preceding claims, wherein: The data set is included in a data set group, The privacy budget is defined for the set of datasets, and Receiving the request to process the data set includes receiving a request to process the group of data sets.

10. The method of any preceding claim, wherein sending the first request comprises: The first request is sent via an application programming interface (API) call to the first privacy budget service.

11. A computing system for managing a privacy budget, the computing system comprising: One or more servers; as well as a non-transitory computer-readable medium having stored thereon instructions that, when executed by the one or more processors, cause the computing system to: receiving a request to analyze a data set, the data set being associated with a privacy budget corresponding to a number of times the data set can be analyzed; sending a first request to a first server implementing a first privacy budget service for maintaining a first instance of the privacy budget to determine whether there is sufficient privacy budget to analyze the data set; sending a second request to a second server implementing a second privacy budget service to maintain a second instance of the privacy budget to determine whether sufficient privacy budget exists to analyze the data set, wherein the second privacy budget service is independent of the first privacy budget service; receiving, according to the first privacy budget service, a first response from the first server indicating whether there is a sufficient privacy budget; receiving, according to the second privacy budget service, a second response from the second server indicating whether there is a sufficient privacy budget; and The data set is processed based on the first response and the second response.

12. The computing system of claim 11, wherein to process the data set, the instructions cause the computing system to: determining that both the first response and the second response indicate that a sufficient privacy budget exists; and Responsive to the determination, the data set is analyzed and results of the analysis are stored.

13. The computing system of claim 11, wherein to process the data set, the instructions cause the computing system to: determining that at least one of the first response or the second response indicates that there is not enough privacy budget; and Responsive to the determination, analyzing the data set is suppressed.

14. A computing system according to any one of claims 11 to 13, wherein: The data set is included in a data set group, The privacy budget is defined for the set of datasets, and To receive the request to process the data set, the instructions cause the computing system to: A request to process the set of data sets is received.

15. The computing system of any one of claims 11 to 14, wherein to send the first request, the instructions cause the computing system to: The first request is sent via an application programming interface (API) call to the first privacy budget service.