Verifiably secure data set operations using private binding keys
The secure computing environment with a secure connector and PII match module within a TEE addresses the challenge of securely combining datasets by encrypting PII and controlling access to encryption keys, achieving efficient and secure data processing.
Patent Information
- Application Number
- JP2024562335
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-24
- Filing Date
- 2023-07-24
- Publication Date
- 2025-06-05
- Estimated Expiration
- 2043-07-24
AI Technical Summary
Current techniques for combining datasets from multiple parties face challenges in securely and efficiently handling sensitive data, such as personally identifiable information (PII), without exposing it to unauthorized parties or requiring computationally expensive hashing and encryption processes.
The proposed solution involves a secure computing environment where a secure connector and a PII match module operate within a Trusted Execution Environment (TEE), enabling secure dataset joining by encrypting PII fields and ensuring that only attested secure code can access the encryption keys, thus protecting sensitive information from inspection and modification.
This approach ensures the secure and efficient combination of datasets by preventing unauthorized access to PII, reducing the computational burden of data obfuscation, and maintaining end-to-end privacy and integrity of data processing.
Smart Images

Figure 2025517282000001_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and claims the benefit of the filing date of U.S. Provisional Patent Application No. 63 / 391,794, entitled "VERIFIABLE SECURE DATASET JOINING WITH PRIVATE JOIN KEYS," filed July 24, 2022, the entire contents of which are expressly incorporated herein by reference.
[0002] The present disclosure relates to secure computing environments, and more particularly to techniques implemented in a cloud or other suitable environment for improving data security and computational efficiency when performing operations such as combining data sets from multiple parties. [Background technology]
[0003] The background art provided herein is described for purposes of generally presenting the context of the present disclosure. The work of the inventors named herein is not admitted, expressly or impliedly, as prior art to the present disclosure to the extent described in this Background section, including aspects of the description that may not qualify as prior art at the time of filing.
[0004] Currently, certain services and applications may attempt to combine datasets from different independent parties. The datasets often contain data that one party does not want and / or is not allowed to share with the other party, which may be referred to as "restricted data" for simplicity. An example of such restricted data is personally identifiable information (PII). Because this data can act as a combine key, i.e., data that logically links records from separate datasets, this data may not be easily removed before performing the combine operation.
[0005] For example, a data service DS 1 The device identifier ID at different times d1 , ID d2 , , , ID N A set of devices S identified by 1 The temperature sensor readings can be stored in a separate data service called DS 2 Set S 1 A set S that at least partially overlaps with 2 It is possible to maintain a set S of pressure sensor readings without disclosing the identity of the device that corresponds to a particular sensor reading. 1 and S. 2 It may be desirable to combine crossover temperature and pressure readings.
[0006] It is desirable to provide a computing environment in which join operations of data sets from multiple sources can be performed securely and efficiently. Summary of the Invention
[0007] The techniques of this disclosure support data set join operations that eliminate the need for first-party data (1PD) sources to expose sensitive data, such as PII, to other parties, without the 1PD source (or “customer data source”) having to perform computationally expensive hashing and / or encryption locally or hand over the data to other parties for these operations.
[0008] Using the techniques of this disclosure, the system can ensure that customer data sources only need to connect to secure connectors running in a trusted execution environment (TEE) to provide data, and that the secure connectors do not provide access to the customer data to any other party. These techniques further enable customer data sources not to share credentials with modules other than the secure connectors.
[0009] The secure connector receives the 1PD and can at least partially encrypt the 1PD, e.g., the PII fields. The encrypted data then flows securely through an extract-transform-load (ETL) pipeline to a PII match module, also implemented in a TEE. Only attested secure code can access the encryption key(s) necessary to decrypt the encrypted fields, no party can extract sensitive information from the encrypted PII, and no party can modify the functionality of the secure connector or the PII match module. [Brief description of the drawings]
[0010] [Figure 1] FIG. 1 is a block diagram of an example computing environment in which at least some of the techniques of this disclosure can be implemented. [Figure 2A] FIG. 1B is a block diagram illustrating an example computing architecture including a secure control plane and a data plane that can be utilized in the computing environment of FIG. 1A. [Figure 2B] FIG. 2B is a block diagram illustrating another example of a computing architecture similar to FIG. 2A, except that here the environment includes additional infrastructure for managing cryptographic keys and privacy budgets. [Figure 3A] FIG. 1B is a block diagram of an example pipeline for performing a secure combine operation of 1PD with other datasets, which may be implemented in the computing environment of FIG. 1A or in another suitable environment. [Figure 3B] A pipeline block diagram generally similar to FIG. 3A, but with the secure connector and PII match module combined into a single entity. [Figure 4A] FIG. 4 is a flow diagram of an example method in a secure connector for taking in plaintext 1PD, pre-processing 1PD, and re-encrypting 1PD that can be implemented in the environment of FIG. 3A or FIG. 3B. [Figure 4B]FIG. 4B is a flow diagram of an example method generally similar to FIG. 4A, but in which at least a portion of the 1PD arrives at the secure connector in an encrypted format. [Figure 4C] 4B is a flow diagram of an example method generally similar to FIG. 4A, except that the secure connector hashes the PII in the 1PD before sending the 1PD to the PII match module. [Figure 5A] FIG. 3B is a flow diagram of an exemplary method for a PII match module to receive a PD containing pre-processed and encrypted PII from a secure connector, decrypt the PII, and match the PPD with other data sets using the PII fields, which can be implemented in the environment shown in FIG. 3A or FIG. 3B. [Figure 5B] FIG. 5B is a flow diagram of an example method generally similar to FIG. 5A, but in which a PII match module that may be implemented in the environment shown in FIG. 3B performs matching using hashed PII fields or values. [Figure 5C] 3C is a flow diagram of an example method in a PII match module for generating a combined data set for a data service that may be implemented in the environment of FIG. 3A or FIG. 3B. [Figure 6A] FIG. 13 is a flow diagram of an example method at a customer data source for providing 1PD in clear text to a secure connector. [Figure 6B] FIG. 13 is a flow diagram of an example method at a customer data source for encrypting PII fields of a 1PD and providing the 1PD to a secure connector. [Figure 7A] FIG. 5 is a block diagram illustrating the conversion of PII and non-PII data of one PD as one PD passes through the environment of FIG. 3A or FIG. 3B according to the methods of FIG. 4A, FIG. 5A, and FIG. 6A. [Figure 7B] FIG. 5 is a block diagram illustrating the conversion of PII and non-PII data of one PD as one PD passes through the environment of FIG. 3A or FIG. 3B according to the methods of FIGS. 4A, 5A, and 6B. [Figure 7C]FIG. 5 is a block diagram illustrating the conversion of PII and non-PII data of one PD as one PD passes through the environment of FIG. 3A or FIG. 3B according to the methods of FIG. 4B, FIG. 5B, and FIG. 6A. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0011] As described in more detail below, the secure connector and PII match module may execute in the TEE to securely and efficiently perform join operations of data sets from different parties. The secure connector of some embodiments also performs pre-processing of the PII so that the PII from different data sets is in the same format to enable efficient matching operations. The secure connector, PII match module, and components of the ETL pipeline may be implemented in a cloud computing environment, or simply "the cloud."
[0012] These components enable the burden of data obfuscation, which may include hashing and / or encryption, to be shifted from Customer Data sources to the cloud while protecting the PII from inspection by other parties. Implementing these modules in the TEE allows the PII match module to perform matching and / or joining in the clear, while ensuring end-to-end privacy of PII and integrity of data processing.
[0013] These techniques address technical issues associated with traditional approaches, such as 1P data owners not being able to retain control over their datasets and preventing other parties from accessing individual-level PII, or 1P data owners having to share PII with various intermediate parties (e.g., services that apply analytics to 1PD). Even when data owners hash PII fields to obfuscate certain information and certain platforms then use the hashed fields as a join key to correlate or combine datasets, these approaches are computationally intensive as data often needs to conform to a specific format for correct ingestion. Because there are often many sources in 1PD, hashing and comparison in multiple different formats introduces inefficiencies and even errors.
[0014] As a more specific example, a phone number may be in a format such as "555.555.5555," "555-555-5555," or "(555)555-5555," each of which corresponds to a different hash. Mailing addresses may be in even more diverse formats. Furthermore, while hashing provides obfuscation, hashed data has security vulnerabilities, e.g., exposed to dictionary attacks.
[0015] These techniques can be applied in a wide variety of applications, such as in the adtech industry, where advertisers measure the effectiveness of their advertising campaigns by determining, for example, which consumer segments or audiences purchase a particular type of product, or which advertisements result in the highest number of sales. To this end, the system can combine 1PD (e.g., sales data in a customer relationship management (CRM) system) and advertising campaign data (e.g., information about people who interacted with an advertisement). Because both 1PD and campaign data contain PII, such as phone numbers, IP addresses, email addresses, physical addresses, etc., the PII can act as the combining key(s).
[0016] An exemplary environment suitable for implementing such techniques is first described with reference to Figures 1, 2A, and 2B. An exemplary pipeline for matching / merge operations is then discussed with reference to Figures 3A and 3B. This is followed by a description of exemplary methods in a secure connector, a PII match module, and a 1P data source.
[0017] Computing environment for secure multi-party computation and joint operations - Patents.com The Secure Control Plane (sometimes referred to herein as "SCP") described herein provides an unobservable secure execution environment in which services can be deployed. In particular, any business logic (e.g., application code) that provides the service can be executed within the secure execution environment, and no party can observe the computations at runtime, to provide the security and privacy guarantees required for the workflow. The state of the environment is opaque even to administrators of the service, and the service can be deployed on any supported cloud.
[0018] As an example, two data generating clients, Client 1 and Client 2, may wish to combine the data streams they receive from each of their customers so that the clients can generate quantitative metrics related to these customers that cannot be derived from their individual data sets. As a more specific example, Client 1 can be a retailer that has data indicative of customer transactions, and Client 2 can be an analytics engine that can measure, for example, the effectiveness of an advertising campaign for products offered by the retailer.
[0019] Client 2 may provide a service with an algorithm that requires Client 2 to perform data analysis securely. However, Client 1 may not want to expose its customer data to Client 2 in a way that potentially allows the data to be exfiltrated or used in a way that does not provide Client 1's privacy and security guarantees. Thus, Client 1 wants to ensure that (1) its customer data cannot be exfiltrated by Client 2 or any other party, and (2) the logic used to analyze the customer data complies with Client 2's security requirements. The techniques disclosed herein provide a secure execution environment in which the business logic executes, such that sensitive data analyzed by the business logic remains encrypted everywhere except within the secure execution environment, and provide proofs such that any party can ensure that the logic operating within the secure execution environment executes as guaranteed.
[0020] Generally speaking, services that perform computations (i.e., process events or requests using business logic) are split between a Data Plane (DP) and a Secure Control Plane (SCP). The business logic specific to the computation is hosted in the DP, which resides in a TEE (also referred to herein as an enclave). The business logic may be provided to the DP as a container, which is a software package that includes all of the elements required to run the business logic in any environment. The container may be provided to the SCP, for example, by the owner of the business logic. Functionally, the SCP provides a secure execution environment and capabilities for the DP to be deployed and run at scale, including managing encryption keys, buffering requests, monitoring privacy budgets, accessing storage, orchestrating policy-based horizontal auto-scaling, etc. The SCP execution environment allows the service to be deployed on any supported cloud vendor without modification of the DP by isolating the DP from the specifics of the cloud environment. Both the DP and the SCP work together by communicating through an Input / Output (I / O) Application Programming Interface (API) (also referred to herein as the Control Plane I / O API, or CPIO API).
[0021] In an exemplary implementation, all data traversing the SCP is always encrypted and only the DP has access to the decryption key. For example, for a particular service, the business logic may include performing event aggregation and outputting an aggregate summary report. In such an example, the SCP sends encrypted requests from one or more event sources to the DP, which in time decrypts the request, processes the request, checks the privacy budget, and generates and sends the encrypted report. Furthermore, the decryption key may be bit-split when outside the DP such that only the DP can assemble the decryption key in the TEE. Depending on the desired application, the output from the DP may be edited or aggregated in such a way that the output can be shared and individual users' data cannot be identified or exfiltrated.
[0022] SCP provides several privacy, authenticity, and security guarantees. With regard to privacy, a service using SCP can provide assurance that no stakeholder (e.g., the device on which the client runs, the cloud platform, a third party), acting alone, including the administrator of the SCP deployment, can access or exfiltrate plaintext (i.e., unencrypted) sensitive information. Furthermore, with regard to authenticity, the DP is running in a secure execution environment that is in a trusted state at the time the enclave is started. For example, SCP may be implemented in a Trusted Platform Module (TPM) or Virtual Trusted Platform Module (vTPM) according to the Secure Boot standard and / or using a trusted and / or certified operating system (OS). Starting from an audited code base and a repeatable build, cryptographic attestations are used to prove DP binary identity and provenance at runtime (as described in more detail below). Furthermore, a Key Management Service (KMS) releases cryptographic keys only to validated enclaves. As a result, if the DP image is tampered with, the system cannot decrypt any data. Cloud providers are implicitly trusted, given the strong incentives that cloud providers have to guarantee their Terms of Service (ToS) guarantees. With respect to security, the Secure Execution Environment is non-observable. The memory of the Secure Execution Environment is encrypted or otherwise hardware protected from access by other processes. Core dumps are not possible in the exemplary embodiment. All data is encrypted in transit and at rest, and all I / O to / from the DP is encrypted. No human has access to the plaintext private key (e.g., the KMS is locked down, the keys are split, and the keys are only valid within the DP that is within the Secure Execution Environment).
[0023] SCP distributes trust in such a way that three stakeholders need to cooperate to exfiltrate plaintext user event data. SCP also uses a distributed trust model to ensure that two stakeholders need to cooperate to tamper with the privacy budget service. Distributed trust works using both event decryption and privacy budget services. For event decryption, the private keys required to decrypt events received at the SCP are generated in a secure environment and bit-split between at least two KMSs, each under the control of an independent trusted party. The KMSs are configured to release keying material only to DPs that match specific hashes. If the DPs have been tampered with, their keys will not be released. In such a scenario, the service can be started, but none of the events can be decrypted. Similarly, privacy budget services may be distributed between two independent trusted parties, and transaction semantics may be used to ensure that the budgets of both trusted parties match. This allows for detection of budget tampering.
[0024] SCP also provides a mechanism to prove that any business logic running in the DP corresponds to publicly released code, as described with reference to Figure 2B, allowing other parties to verify the business logic being used to analyze sensitive data. The entire code base of the business logic (except for the scenario described with reference to Figure 5 for proprietary business logic) is available to all stakeholders for inspection and auditing. Builds are reproducible, and any stakeholder can build a DP container. Building a deployable image generates a set of cryptographic hashes (e.g., Platform Configuration Registers (PCRs)). Thus, all parties can verify that the deployed product matches the published code base by comparing the PCRs. After building the logic, the DP provides the PCRs (e.g., via the CPIO API) to parties that request verification of the built logic. For example, the KMS is configured to release keying material only to images that match the PCRs generated from building the published logic. This ensures that the private key for decrypting sensitive information is only valid for images that correspond to a specific commit in a specific repository.
[0025] Turning to an exemplary computing system in which the SCP of the present disclosure can be implemented, FIG. 1 illustrates an exemplary computing system 100. The computing system 100 includes a client computing device 102 (also referred to herein as a client device 102) coupled to a cloud platform 122 (also referred to herein as a cloud 122) via a network 120. The network 120 can generally include one or more wired and / or wireless communication links, and may include, for example, a wide area network (WAN), such as the Internet, a local area network (LAN), a cellular network, or any other suitable type or combination of networks. Although the examples of the present disclosure primarily refer to cloud implementation architectures, it should be understood that the techniques disclosed herein, such as techniques for providing a secure execution environment for processing sensitive data, techniques for generating, splitting, and distributing keys, and techniques for providing mechanisms for verifying unique business logic, can also be applied to non-cloud systems.
[0026] The client device 102 may be, for example, a portable device such as a smartphone or tablet computer. The client device 102 may also be a laptop computer, a desktop computer, a personal digital assistant (PDA), a wearable device such as smart glasses, or other suitable computing device. The client device 102 may include memory 106, one or more processors (CPUs) 104, a network interface 114, a user interface 116, and an input / output (I / O) interface 118. The client device 102 may also include components not shown in FIG. 1, such as a graphics processing unit (GPU). The client device 102 may be associated with a service user, who is an end user of the services provided by the SCP, as described below. The end user operates the client device 102 (or, more specifically, a browser or application on the client device 102) that sends requests / events to the service. To send a request or event to the service, the client device 102 encrypts the request / event using a public key, which the client device 102 may obtain from a public key repository (e.g., a public key repository server 178). Client device 102 is exemplary only. As described below, cloud platform 122 may receive events and / or requests from client device 102, from a browser / application / client process executing on client device 102, or from other computing devices that issue requests on behalf of or forward requests from client device 102. Additionally, although only one client device is shown in FIG. 1, computing system 100 may include multiple client devices that can communicate with cloud platform 122.
[0027] Network interface 114 may include one or more communications interfaces, such as hardware, software, and / or firmware, to enable communication over a cellular network, a Wi-Fi network, or any other suitable network, such as network 120. User interface 116 may be configured to provide information to a user, such as responses to requests / events received from cloud platform 122. I / O interface 118 may include various I / O components (e.g., ports, capacitive or resistive touch-sensitive input panels, keys, buttons, lights, LEDs). For example, I / O interface 118 may be a touch screen.
[0028] The memory 106 may be a non-transitory memory and may include one or more suitable memory modules, such as random access memory (RAM), read only memory (ROM), flash memory, other types of persistent memory, etc. The memory 106 may store machine-readable instructions executable by one or more processors 104 and / or specialized processing units of the client device 102. The memory 106 also stores an operating system (OS) 110, which may be any suitable mobile or general-purpose OS. Additionally, the memory 106 may store one or more applications that communicate data with the cloud platform 122 over the network 120. Communicating data may include transmitting data, receiving data, or both. For example, the memory 106 may store instructions for implementing a browser, online service, or application that requests / sends data from / to an application (i.e., business logic) implemented in a DP of the secure execution environment of the cloud platform 122, described below.
[0029] Cloud platform 122 may include multiple servers associated with a cloud provider to provide cloud services over network 120. The cloud provider is the owner of cloud platform 122 on which SCP 126 is deployed. Although only one cloud platform is shown in FIG. 1, SCP 126 may be deployed on multiple cloud platforms, even if the cloud platforms are operated by different cloud providers. The servers providing cloud platform 122 may be distributed across multiple sites for improved reliability and reduced latency. Individual servers or servers within cloud platform 122 may communicate with client device 102 and with each other over network 120. Exemplary servers that may be included in cloud platform 122 are described in more detail below. Although not shown in FIG. 1, each server included in cloud platform 122 may include one or more processors (similar to processor(s) 104) adapted and configured to execute various software stored in one or more memories (similar to memory 106). The servers may further include a database, which may be a local database stored in a particular server's memory or a network database stored in a network-attached memory (e.g., in a storage area network). The servers may also include network interfaces and I / O interfaces, as well as interfaces 114 and 118, respectively. Additionally, while certain components are described as separate servers, it should be understood that generally speaking, the term "server" may refer to one or more servers. Additionally, although functions are generally described as being performed by separate servers, some of the functions described herein may be performed by the same server.
[0030] The cloud platform 122 includes an SCP 126 that includes a TEE 124. The TEE 124 is a secure execution environment in which the DP 128 is isolated. A TEE, such as the TEE 124, is an environment that provides execution isolation and offers a higher level of security than a typical system. The TEE 124 may utilize hardware to enforce the isolation (referred to as confidential computing). The cloud provider is considered the root of trust for the SCP 126, adhering to the terms of service (ToS) agreement of the cloud platform 122. The hardware manufacturer of the server that provides the TEE 124 also has ToS assurance, thus providing an additional layer of trust. The SCP 126 also utilizes techniques to ensure that the boot-time state is secure, including using a minimal OS image recommended by the cloud provider and using a TPM / vTPM-based secure boot sequence for that OS image.
[0031] One or more servers of the cloud platform 122 perform control plane (CP) functions (i.e., support the SCP 126), and one or more servers perform data plane (DP) functions. All functions of the DP 128 are performed by the servers in the TEE 124. The TEE 124 may be deployed and operated by an administrator. The administrator may audit the logic implemented in the DP 128 and verify against a hash of a binary image to deploy the logic 142. On the CP, there may be a front-end server 134 that receives external request / event indications (e.g., from the client device 102), buffers the request / event until the DP 128 can process it, and forwards the received request to the DP 128. Generally speaking, as used herein, a request may also refer to an event or may include one or more events, unless otherwise noted. In some implementations, there is a third-party server 136 between the client device 102 and the SCP 126. The third party server 136 (which may include one or more servers and may or may not be hosted on the cloud platform 122) may be responsible for receiving requests (encrypted by the client device 102) from the client device 102 and later dispatching the encrypted requests to the SCP 126. In some cases, the third party is an administrator of the service. The third party server 136 does not have the key to decrypt the requests. The third party server 136 may, for example, aggregate the requests into batches and store the batches (e.g., in the cloud storage 160). The third party server 136 or the cloud storage server 160 may notify the front-end server 134 that the request is ready to be processed and / or the front-end server 134 may subscribe to notifications that are pushed to the front-end server 134 when batches are added to the cloud storage 160.
[0032] The DP 128 includes a server (which may include one or more servers) including one or more processors 138 (similar to the processor(s) 104) and one or more memories 140 (similar to the memory 106). The memory 140 includes business logic 142 (also referred to as logic 142) that may be executed by the processor 138. The business logic 142 is for implementing applications or services that are deployed on the TEE 124. The memory 140 may also store a key cache 146, which stores cryptographic keys for encrypting and decrypting communications. Additionally, the memory 140 includes a CPIO API 144 that includes a library of functions for communicating with other elements of the cloud platform 122, such as components on the CP of the SCP 126. The CPIO API 144 may be configured to interface with any cloud platform provided by a cloud provider. For example, in a first deployment, the SCP 126 may be deployed to a first cloud platform provided by a first cloud provider. The DP 128 hosts a particular business logic 142 and the CPIO API 144 facilitates communication between the logic 142 and a first cloud platform. In a second deployment, the SCP 126 may be deployed to a second cloud platform provided by a second cloud provider. The DP 128 may host the same business logic 142 as in the first deployment and the CPIO API 144 is configured to facilitate communication between the logic 142 and the second cloud platform. Thus, the SCP 126 may be deployed to a different cloud platform without editing the underlying business logic 142 and only by configuring the CPIO API 144 to interface with the particular cloud platform.
[0033] There may be additional CP-level services provided by the servers of cloud platform 122 that support SCP 126. For example, verifier server 148 may implement a verifier module that can verify whether business logic 142 complies with security policies, as described below with reference to FIG. 5. As another example, privacy budget service server 152 may implement a privacy budget service that verifies whether a user's or device's privacy budget has been exhausted. One or more privacy budget services may additionally or alternatively be implemented by a trusted party, as described with reference to FIG. 2B.
[0034] Additionally, the cloud platform 122 may include other servers and databases in communication with the SCP 126, as described in the following paragraphs. These servers may facilitate the CP functionality of the SCP 126. In particular, the CP functionality may be distributed across several servers, as described below. However, the DP 128 remains within the TEE 124 and is not distributed outside the TEE 124.
[0035] Cloud storage 160 may store the encrypted batch of requests, as described above, before the encrypted batch is received by front-end server 134. Cloud storage 160 may also be used to store responses after DP 128 processes received requests, or to perform storage functions of other components of cloud platform 122. Queue 162 may be used by front-end server 134 to store pending requests before they can be analyzed by DP 128. For example, after receiving a request from client device 102, front-end server 134 may receive the request and temporarily store the pending request in queue 162 until DP 128 is ready to process the request. As another example, after receiving a notification from third-party server 136 that a batch of requests is stored in cloud storage 160, front-end 134 may retrieve the batch and place the batch in queue 162 where it awaits analysis by DP 128.
[0036] Key Management Server (KMS) 164 provides the KMS that generates, deletes, distributes, replaces, rotates, and otherwise manages encryption keys. Trusted Party 1 Server 166 and Trusted Party 2 Server 172 are servers associated with Trusted Party 1 and Trusted Party 2, respectively, and provide the functionality of each trusted party. Although FIG. 1 shows only two trusted parties, cloud platform 122 may include multiple trusted parties. Each trusted party may manage a privacy budget and may audit logic 142 implemented in DP 128 to verify build products against a hash of published logic. Trusted parties own the creation and management of asymmetric keys used to encrypt and decrypt user data. Trusted parties may securely generate keys and publish public keys to the world. The private key may be bit-split into two parts (one split under the control of each trusted party, although any number of N splits may be supported if there are N trusted parties), as described in more detail with reference to FIG. 4. Each trusted party may use an envelope encryption technique, encrypting its splits per key with a KMS symmetric key and storing the encrypted splits in its repository. Envelope encryption allows the envelope to be rotated without necessarily rotating the key within the envelope. The public keys may be stored and managed by a public key repository server 178. Additionally or alternatively, the KMS server 164 may manage the public keys.
[0037] The computing system 100 may also include a public security policy storage 180, which may be located on the cloud platform 122 or outside the cloud platform 122. The public security policy storage 180 stores security policies such that the security policies are publicly accessible (e.g., by the client devices 102, by components of the cloud platform 122). A security policy (also referred to herein as a policy) describes what actions or fields are allowed to configure the output of a service. A policy may also be described as a machine-readable and machine-executable privacy design document (PDD). Policies are further described with reference to FIG. 5.
[0038] 2A, an example architecture 200A illustrates connections between components and software elements of computing system 100. Client device 102 may obtain a public key (e.g., from public key repository server 178) to address a request to a service implemented in DP 128 (i.e., by business logic 142). For example, client device 102 may initiate a request to access content provided by a service or may issue an event that includes user behavior data.
[0039] The encrypted request from the client device 102 is first received by the front-end module 234 of the SCP 126 (i.e., a module implemented by the front-end server 134). In some implementations, the request is first received by a third party that batches the requests before notifying (or having the front-end 234 notify). In such a case, the front-end 234 may retrieve the encrypted request from the cloud storage 160. In any event, the front-end 234 passes the encrypted request to the DP 128 using functions defined by the CPIO API 144. The front-end 234 may store the encrypted request in the queue 162 until the DP 128 is ready to process the request and retrieves the request from the queue 162. The DP 128 decrypts the request and processes the request according to the business logic 142. Decrypting the request may include communicating with KMS 264 (i.e., cloud KMS implemented by KMS server 164) and / or communicating with a trusted party to obtain and assemble a private key to decrypt the request, as in FIG. 2B.
[0040] Processing the request may include using CPIO API 144 functions to communicate with a privacy budget service 252 (e.g., implemented by a privacy budget service server 152) to check the privacy budget and ensure compliance with the privacy budget. The privacy budget monitors the requests and events processed. For example, there may be a maximum number of requests originating from a particular user that can be processed during a particular calculation or period. Ensuring compliance with the privacy budget prevents a party analyzing the output from DP 128 from exposing information about a particular user. By checking compliance with the privacy budget, DP 128 provides differential private outputs.
[0041] The results of processing the requests may be encrypted by DP 128 and may be edited and / or aggregated so that the output does not reveal information about a particular user. DP 128 may store the results, for example, in cloud storage 160, where a party with a decryption key for the results may retrieve the results. As an example, when processing results from a third-party server 136, DP 128 may encrypt the results using a key that the third-party server 136 can decrypt.
[0042] Turning to FIG. 2B, architecture 200B is similar to architecture 200A, except that additional details regarding key management and privacy budgets are shown. In comparison to FIG. 2A, FIG. 2B also shows trusted party 1 server 166 (for brevity, referred to herein as trusted party 1 166), trusted party 2 server 172 (for brevity, referred to herein as trusted party 2 172), and public key distribution service 278. Public key distribution service 278 provides public keys to client devices 102, which can be used by client devices 102 to address requests to DP 128, front end 234, or third party server 136, which aggregates the requests (not shown in FIG. 2B). Public key distribution service 278 can be operated by public key repository server 178 or KMS server 164. Trusted party 1 166 includes a key cache 268 that contains the encrypted split-1 key (i.e., the encrypted first part of the private key), while trusted party 2 172 includes a key cache 274 that contains the encrypted split-2 key (i.e., the encrypted second part of the private key). Each of the trusted parties 166, 172 may also provide a privacy budget service 270, 276 and may manage an instance of the privacy budget, respectively. Distributing the management of the privacy budget to the two trusted parties helps ensure that no trusted party can tamper with the privacy budget. Both privacy budget services 270, 276 should enforce the same privacy budget. Thus, if the two services return different outputs, the SCP 126 can recognize that one of the trusted parties 166, 172 is tampering with the privacy budget. The architecture shown in Figure 2B prevents any one trusted party from gaining complete control over the private decryption key or the privacy budget: no single trusted party acting alone can provide an unlimited budget to any user.Thus, a single trusted party cannot repeatedly aggregate the same batches of data.
[0043] Example Pipeline for Performing Secure Match / Merge Operations 3A illustrates a pipeline 300A that may be implemented at least in part within the above environment. The pipeline 300A receives a data set from a 1P data source 302 and provides the results of the matching / joining to a data service 304. The parties controlling the systems 302 and 304 are separate and independent, and the data service 304 desirably performs operations (e.g., analysis) using the data set from the 1P data source 302 without relying on or having access to the data contained in the data set, particularly the PII. The 1P data source 302 may be any suitable external source of 1PD keyed by plaintext PII or any other suitable data. The 1P data source 302 may be, for example, a CRM, a proprietary system, a file available on the Internet, etc.
[0044] The 1P data source 302 provides a data set to a secure connector 320 implemented in the cloud 310 via an encrypted link 303. The link 303 can be, for example, an SSL / TLS connection established over the Internet. The secure connector 320 can operate with an audited and certified TEE. As described in more detail below, the secure connector 320 in operation can hash and / or encrypt some or all of the received data set. The secure connector 320 provides the hashed / encrypted data set to an ETL pipeline 324 via an encrypted link 322. The ETL pipeline 324 can move the data set to either a data repository 330 or a secure join module 328 via an encrypted link 326. The ETL pipeline 324 can generally perform data transformations and field mapping to conform to a particular schema and format non-encrypted fields. The repository 330 can be a data storage service that enables time-deferred consumption of data ingested from the 1P data source 302.
[0045] The PII match module 328, like the secure connector 320, can operate in an audited and certified TEE. In operation, the PII match module 328 can match and combine the 1P data set with other data sets that may be from other 1P data sources or may reside internal to the data service 304, for example. The PII match module 328 then provides a privacy secure output to the data service 304, which may operate on the cloud platform 312 or any other suitable platform.
[0046] 3A, the ETL pipeline 324 can send data either to a secure binding module 328 for immediate consumption by the data services 304, or to a repository 330. The repository 330 supports a "take once, use many" workflow. The repository 330 always stores sensitive PII hashed or encrypted. In some implementations, an additional layer, such as encryption of data at rest, further ensures the security of the data stored in the repository 330.
[0047] 3B, pipeline 300B is similar to pipeline 300A, except that here a single component 325 running in the TEE implements the functionality of both secure connector 320 and PII match module 328. However, this simplified architecture does not support storing encrypted PII in repository 330 for later consumption.
[0048] 3A and 3B in general, one or more TEEs supporting secure connectors and PII match modules are services with provable properties of security and privacy. More specifically, these services ensure that a party can verify that it is connected to the correct server. That is, a party can verify what the TEE box does by inspecting the code repository, and a party can verify that the repository code corresponds exactly to the image running on the server. Furthermore, the cloud provider 310's attestation infrastructure ensures that the requested decryption key is only available in a TEE with a specific signature.
[0049] Example Workflow for Performing Secure Match / Merge Operations Some example workflows that the pipelines of Figures 3A and 3B can support are now described with reference to Figures 4A-7C. The methods of Figures 4A-6B can be implemented using suitable processing hardware, for example as a set of software instructions stored on a non-transitory computer-readable medium and executable by one or more processors.
[0050] Referring first to FIG. 4A, method 400A can be implemented in secure connector 320 or 325. Method 400A includes plaintext PII matching, server-side encryption, and the use of a customer-generated key. Method 400A begins at block 403 where the secure connector performs authentication with a 1PD source. More specifically, the secure connector can initiate a connection between a 1P data source (e.g., 1P data source 302) and the secure connector. The customer can first provide encrypted credentials to ensure that the secure connector is the only entity that can connect to the 1P data source. KMS 164 (see FIGS. 1, 2A, and 2B) can use the decryption key of the credentials with an account owned by the customer, and KMS 164 ensures that only secure connector 320 can perform decryption operations using these credentials.
[0051] The secure connector can decrypt the credentials and use the decrypted credentials to authenticate to the 1P data source. Data transfer occurs via SSL / TLS or a similar protocol that allows for authentication of the endpoint(s). The secure connector and the 1P data source can potentially use mutual authentication (mTLS) to ensure both ends of the connection that data flows from and to the intended endpoint. Certain 1P data sources require repeated use of credentials, while other 1P data sources rely on tokens, certificates, or other techniques to fetch data over the secure connection. According to other embodiments, the secure connector and the 1P data source use certificates and encryption schemes to provide access to data on behalf of credentials. The certificates required for the connection are encrypted and used in such a way that only the secure connector has a valid certificate to establish a successful connection.
[0052] In either case, a customer associated with the 1P data source uses the cloud KMS described above to locally generate a data encryption key DEK and a key encryption key (KEK). The customer's computing system can use the cloud KMS's APIs to encrypt the DEK with the KEK. The customer also configures the KMS to enable the secure connector and the PII match module to decrypt the KEK. In block 404, the secure connector receives the encrypted DEK associated with the 1P data source. In block 405, the secure connector provides the encrypted DEK to the PII match module.
[0053] In block 410, the secure connector ingests a data set in clear from a 1P data source. As shown for further clarity in FIG. 7A, according to this workflow, the data set in stage 702A includes both non-PII and PII fields in clear. In block 420, the secure connector pre-processes the PII to match a specific standard format. This transformation improves the matching rate in the PII match module, potentially reducing error rates and improving efficiency.
[0054] At block 422, the secure connector decrypts the DEK using the KMS and encrypts at least the PII field of the ingested data set with the DEK (see FIG. 7A, stage 704A). At block 430, the secure connector sends the data through the pipeline to a PII match module for matching with other data sets based on PII. As described with reference to FIG. 5A, the PII match module can perform matching and combining in the clear. FIG. 7A illustrates this clear text comparison at stage 706A.
[0055] Next, FIG. 4B illustrates method 400B. Like blocks are labeled with like reference numbers, and only the differences between methods 400A and 400B are discussed below. Method 400B includes clear text PII matching and client side encryption. In block 411, the secure connector ingests (e.g., fetches) a data set from a 1P data source, where the data set includes encrypted PII, as also shown in FIG. 7B, stage 702A.
[0056] FIG. 4C illustrates method 400C, which includes using hashed PII matching. Like blocks are labeled with like reference numbers, and only the differences between methods 400A and 400B are discussed below. Method 400B includes clear PII matching and client-side encryption. At block 421, the secure connector hashes the PII fields, and at block 432, the secure connector provides a data set comprising the hashed PII fields to a PII match module for hash-based comparison. FIG. 7C illustrates that at stage 702C, the data set includes clear PII data and non-PII data. At stage 704C, the PII is pre-processed and hashed. At stage 70BC, the comparison is based on the formatted / processed and hashed PII.
[0057] 5A is a flow diagram of an example method 500A of a PII match module, such as PII match module 325. Method 500A may correspond to method 400A or 400B of the secure connector.
[0058] At block 501, the PII match module receives a data set comprising pre-processed and encrypted PII from the secure connector over an encrypted link (see FIG. 7A, stage 704A). At block 510, the PII match module uses the KMS to decrypt the data using the DEK, and then at block 520, decrypts the encrypted PII fields using the DEK.
[0059] At block 530, the PII match module matches the 1PD dataset with other datasets, such as an internal dataset, based on the PII fields (see FIG. 7A, stage 706A). The PII match module may also discard all non-matching rows. At block 540, the PII match module may provide the matched dataset to a data service, such as data service 304.
[0060] 5B is a flow diagram of another example method 500B of a PII match module, such as PII match module 328 or 325. Method 500B may correspond to secure connector method 400C. At block 502, the PII match module receives a data set comprising pre-processed and hashed PII from the secure connector. At block 531, the PII match module may match the data set with other data sets based on the hashed PII fields and discard non-matching rows. At block 540, the PII match module may provide the matched data set to a data service, such as data service 304.
[0061] 5C is a flow diagram of an example method 500C in a PII match module for generating a combined data set for a data service. At block 550, the PII match module can determine matches between data sets using PII in unhashed or hashed format according to methods 500A and 500B, respectively.
[0062] At block 560, the PII match module can map external identifiers to internal identifiers for matched rows in the dataset. At block 570, the PII match module can also augment each row of the output dataset with metadata indicating the type of match that occurred (e.g., based on email, phone, address) for post-processing (e.g., conflict and duplicate resolution). At block 572, the PII match module can remove all PII from the output dataset.
[0063] Additionally or alternatively to block 560, at block 562, the PII match module may generate a list of matched internal identifiers between the datasets. Flow may also proceed to block 570 where the PII match module augments each row with metadata as described above. Furthermore, additionally or alternatively to blocks 560 and 562, at block 564, the PII match module may include any combination of fields from both the datasets and / or the metadata, but may not include any PII fields.
[0064] FIG. 6A is a flow diagram of an example method 600A that may be implemented in a customer data source (e.g., 1P data source 302) to provide 1PD in clear to a secure connector. In block 601, the customer data source generates a DEK and a KEK locally using the cloud KMS. The customer data source encrypts the DEK with the KEK using an API of the cloud KMSI. In block 602, the customer data source performs authentication with the secure connector. The secure credential service can then configure the cloud KMS to enable decryption of the KEK at the secure connector and the PII match module. In block 620, the customer data source provides the encrypted DEK to the secure connector, and in block 630, provides the data in clear to the secure connector over a secured link (see FIG. 7A, stage 702A or FIG. C, stage 702C).
[0065] Figure 6B is a flow diagram of an example method 600B that is generally similar to Figure 6A, except here the customer data source encrypts the PII fields at block 622 (see Figure 7B, stage 702B) and provides the data set to the secure connector over an encrypted link.
[0066] Other considerations The following additional considerations apply to the above discussion:
[0067] A client device (e.g., client device 102) capable of implementing the techniques of this disclosure can be any suitable device capable of wireless communication, such as a smartphone, tablet computer, laptop computer, desktop computer, mobile game console, point of sale (POS) terminal, health management device, drone, camera, media streaming dongle or other personal media device, wearable device such as a smart watch, wireless hotspot, femtocell, or broadband router. Additionally, the client device may be embedded in an electronic system such as a vehicle's head unit or advanced driver assistance system (ADAS) in some cases. Additionally, the client device may operate as an Internet of Things (IoT) device or a Mobile Internet Device (MID). Depending on the type, the client device may include one or more general-purpose processors, computer-readable memory, a user interface, one or more network interfaces, one or more sensors, and the like.
[0068] Certain embodiments are described in this disclosure as including logic or several components or modules. The modules can be software modules (e.g., code stored on a non-transitory machine-readable medium) or hardware modules. A hardware module is a tangible unit that can perform certain operations and can be configured or arranged in a certain manner. A hardware module can include dedicated circuitry or logic that is permanently configured (e.g., as a special-purpose processor such as a field programmable gate array (FPGA) or an application-specific integrated circuit (ASIC)) to perform certain operations. A hardware module can also include programmable logic or circuitry that is temporarily configured by software (e.g., contained within a general-purpose processor or other programmable processor) to perform certain operations. The decision to implement a hardware module in dedicated and permanently configured circuitry or in temporarily configured circuitry (e.g., configured by software) can be subject to cost and time considerations.
[0069] If implemented in software, the techniques may be provided as part of an operating system, a library used by multiple applications, a specific software application, etc. The software may be executable by one or more general-purpose processors or one or more special-purpose processors.
Claims
1. 1. A method in one or more servers for performing a join operation, comprising: receiving a first data set from a first party (1P) data source, the first data set including personally identifiable information (PII) data and non-PII data ... pre-processing the PII data to generate first formatted PII data, the first formatted PII data conforming to a predetermined format; matching, in the TEE, the first formatted PII data with second formatted PII data contained in a second data set; performing a join operation between the first data set and the second data set based on the matching to generate a joined data set; providing the combined data set to a data service that operates independently of the 1P data source; The method includes:
2. The method of claim 1 , further comprising performing, by the module, authentication with the 1P data source prior to the reception of the first data set.
3. The method of claim 2 , wherein the performing the authentication includes performing a decryption operation using credentials associated with the 1P data source.
4. the module implements a secure connector configured to use credentials associated with the 1P data source; The method of claim 1 or 2, wherein the matching is implemented in a secure binding module that is prevented from accessing the credentials associated with the 1P data source.
5. The method further comprising:
5. The method of claim 4, further comprising providing the first formatted PII data from the secure connector to the secure coupling module via an extract-transform-load (ETL) pipeline.
6. 6. The method of claim 5, wherein the ETL pipeline is configured to provide the first formatted PII data to (i) the data service and (ii) a repository for time-deferred consumption of the first formatted PII data.
7. receiving, at the secure connector, the encrypted DEK associated with the 1P data source; The method of any one of claims 4 to 6, further comprising providing the encrypted DEK from the secure connector to the protected binding module.
8. decrypting the encrypted DEK using a key management service (KMS) to generate a DEK; the PII data received from the 1P data source along with the first data set is encrypted; The method of claim 7 , wherein the pre-processing of the PII data includes decrypting the received PII data prior to generating the first formatted PII data.
9. 9. The method of claim 8, further comprising encrypting the first formatted PII data using the DEK in the secure connector prior to providing the first formatted PII data to the secure coupling module.
10. 10. The method of claim 8, further comprising hashing the first formatted PII data at the secure connector prior to providing the first formatted PII data to the secure coupling module.
11. The method of claim 10 , wherein the matching of the first formatted PII data with the second formatted PII data is implemented in the secure combination module and is based on hashed PII data.
12. The method of claim 11 , further comprising discarding non-matching rows in the first and second data sets.
13. both the receiving and the matching are implemented in the module configured to operate as a secure connector and a PII match component; 4. The method of claim 1, wherein providing the combined data set comprises providing the combined data set from the secure connector and the PII match component to the data service via an ETL pipeline.
14. 14. The method of any one of claims 1 to 13, wherein said receiving the first data set comprises receiving the first data set in plaintext over an encrypted link.
15. A system including one or more servers equipped with processing hardware configured to implement the method of any one of claims 1 to 14.
Citation Information
Patent Citations
Establishing links between identifiers without disclosing specific identifying information
JP2020507826A
Customized Trusted Computer For Secure Data Processing and Storage
US20160306995A1
Providing and obtaining one or more data sets via a digital communication network
WO2021116046A1