Data Clean Room
By establishing a data clean room in each account, using security functions and anonymization technology, the problems of loss of data control and privacy risks in cross-account data analysis are solved, and secure and low-cost data sharing and analysis are achieved.
Patent Information
- Application Number
- CN202180005136.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-07-31
- Filing Date
- 2021-06-30
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-06-30
AI Technical Summary
In the prior art, companies need to rely on third parties when conducting data analysis, resulting in loss of data control, high analysis costs and possible violation of privacy regulations, and at the same time, they cannot achieve secure data sharing and analysis across accounts.
By establishing a data clean room in each account, using security functions and anonymization technology, limiting access and output of other account data, implementing secure data analysis across accounts.
It realizes data analysis across accounts without leaking sensitive information, maintains control over its own data by each account, avoids the loss of data control rights and privacy risks, and reduces analysis costs and time delays.
Smart Images

Figure CN114303146B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of priority to U.S. patent application serial number 16 / 944,929, filed on July 31, 2020, the contents of which are hereby incorporated by reference in their entirety. Technical Field
[0003] The present disclosure generally relates to using a data clean room to securely analyze data across different accounts.
[0004] background
[0005] Currently, most digital advertising is carried out using third-party cookies. Cookies are small pieces of data generated and sent by a web server and stored on a user's computer through their web browser. They are used to collect data about a user's website browsing habits based on their website browsing history. Due to privacy concerns, the use of cookies is restricted.
[0006] Companies may want to create target groups to target advertising or marketing efforts at specific audience segments. To do this, they may want to compare their customer information with that of other companies to see if their customer lists overlap, creating these target groups. Consequently, companies may want to perform data analysis, such as overlap analysis, on their customer or other data. To perform this type of data analysis, companies can use "trusted" third parties that have access to data from each company and perform the analysis. However, this third-party approach has significant drawbacks. First, companies cede control of their customer data to these third parties, which can lead to unforeseen and harmful consequences, as this data may contain sensitive information, such as personally identifiable information. Second, the analysis is performed by the third party, not the company itself. Consequently, companies must return to the third party for more detailed or different analysis. This can increase the costs associated with the analysis and introduce time delays. Furthermore, providing this information to third parties for this purpose may conflict with evolving data privacy regulations and general industry policies. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The various drawings depict only example embodiments of the disclosure and should not be considered limiting of its scope.
[0009] Figure 1 An example computing environment is shown in which a network-based data warehouse system may implement streams on shared database objects, according to some example embodiments.
[0010] Figure 2 is a block diagram illustrating components of a computing service manager according to some example embodiments.
[0011] Figure 3 is a block diagram illustrating components of an execution platform according to some example embodiments.
[0012] Figure 4 is a block diagram illustrating accounts in a data warehouse system according to some example embodiments.
[0013] Figure 5 is a block diagram illustrating a data clean room according to some example embodiments.
[0014] Figure 6 is a block diagram illustrating a double-blind data clean room, according to some example embodiments.
[0015] Figure 7 A flow chart for loading data into a data warehouse system is shown, according to some example embodiments.
[0016] Figures 8A-8C is a block diagram illustrating secure querying using a data clean room, according to some example embodiments.
[0017] Figure 9 A diagrammatic representation of a machine in the form of a computer system within which a set of instructions, for causing the machine to perform any one or more of the methodologies discussed herein, may be executed is shown according to some embodiments of the present disclosure.
[0018] Detailed description
[0019] The following description includes systems, methods, techniques, instruction sequences, and computing machine program products that embody illustrative embodiments of the present disclosure. In the following description, for the purpose of explanation, many specific details are set forth to provide an understanding of the various embodiments of the subject matter of the present invention. However, it will be apparent to those skilled in the art that embodiments of the subject matter of the present invention can be implemented without these specific details. In general, well-known instruction instances, protocols, structures, and techniques are not necessarily shown in detail.
[0020] Embodiments of the present disclosure can provide a data clean room that allows secure data analysis across multiple accounts without the use of a third party. Each account can be associated with a different company or party. The data clean room can provide security functions to protect sensitive information. For example, the data clean room can limit access to data in other accounts. The data clean room can also limit which data can be used for analysis and can limit output. For example, the output can be limited based on a minimum threshold of overlapping data (e.g., elements of each output data row). Thus, each account (e.g., company, party) can maintain control of the data in its own account while being able to perform data analysis using its own data and data from other accounts. Each account can set policies for which types of data and which types of analysis it is willing to allow other accounts to perform. Overlapping data can be anonymized to prevent sensitive information from being leaked.
[0021] Figure 1 An example shared data processing platform 100 is shown that implements secure messaging between deployments according to some embodiments of the present disclosure. To avoid obscuring the present invention with unnecessary detail, various functional components that are not closely related to conveying an understanding of the present invention have been omitted from the figure. However, those skilled in the art will readily appreciate that various additional functional components can be included as part of shared data processing platform 100 to facilitate additional functionality not specifically described herein.
[0022] As shown in the figure, the shared data processing platform 100 includes a network-based data warehouse system 102, a cloud computing storage platform 104 (eg, a storage platform, Services, Microsoft or Google Cloud ) and remote computing devices 106. The network-based data warehouse system 102 is a network-based system for storing and accessing data in an integrated manner (e.g., storing data internally and accessing data located externally remotely), and reporting and analyzing integrated data from one or more different sources (e.g., cloud computing storage platform 104). The cloud computing storage platform 104 includes multiple computing machines and provides computer system resources, such as data storage and computing power, to the network-based data warehouse system 102 on demand. Although Figure 1 A data warehouse is depicted in the illustrated embodiment, but other embodiments may include other types of databases or other data processing systems.
[0023] The remote computing device 106 (e.g., a user device such as a laptop computer) includes one or more computing machines (e.g., a user device such as a laptop computer) that execute remote software components 108 (e.g., a browser-accessed cloud service) to provide additional functionality to users of the network-based data warehouse system 102. The remote software components 108 include a collection of machine-readable instructions (e.g., code) that, when executed by the remote computing device 106, cause the remote computing device 106 to provide certain functionality. The remote software components 108 can operate on input data and generate result data based on processing, analyzing, or otherwise transforming the input data. By way of example, as discussed in further detail below, the remote software components 108 can be data providers or data consumers that enable database tracking processes (e.g., flows on shared tables and views).
[0024] The network-based data warehouse system 102 includes an access management system 110, a computing service manager 112, an execution platform 114, and a database 116. The access management system 110 enables administrative users to manage access to resources and services provided by the network-based data warehouse system 102. Administrative users can create and manage users, roles, and groups, and use permissions to allow or deny access to resources and services. As discussed in further detail below, the access management system 110 can store shared data that securely manages shared access to storage resources of the cloud computing storage platform 104 between different users of the network-based data warehouse system 102.
[0025] The computing service manager 112 coordinates and manages the operation of the network-based data warehouse system 102. The computing service manager 112 also performs query optimization and compilation, and manages clusters of computing services that provide computing resources (e.g., virtual warehouses, virtual machines, EC2 clusters). The computing service manager 112 can support any number of client accounts, such as end users providing data storage and retrieval requests, system administrators who manage the systems and methods described herein, and other components / devices that interact with the computing service manager 112.
[0026] Computing service manager 112 is also coupled to database 116, which is associated with all data stored on shared data processing platform 100. Database 116 stores data related to various functions and aspects associated with network-based data warehouse system 102 and its users.
[0027] In some embodiments, database 116 includes a summary of data stored in remote data storage systems and data available from one or more local caches. In addition, database 116 may include information about how data is organized in remote data storage systems and local caches. Database 116 allows systems and services to determine whether a piece of data needs to be accessed without having to load or access the actual data from a storage device. Computing service manager 112 is also coupled to execution platform 114, which provides multiple computing resources (e.g., virtual warehouses) that perform various data storage and data retrieval tasks, as discussed in more detail below.
[0028] The execution platform 114 is coupled to a plurality of data storage devices 124-1 to 124-N that are part of the cloud computing storage platform 104. In some embodiments, the data storage devices 124-1 to 124-N are cloud-based storage devices located in one or more geographic locations. For example, the data storage devices 124-1 to 124-N can be part of a public cloud infrastructure or a private cloud infrastructure. The data storage devices 124-1 to 124-N can be hard disk drives (HDDs), solid-state drives (SSDs), a storage device cluster, the Amazon S3 storage system, or any other data storage technology. In addition, the cloud computing storage platform 104 can include a distributed file system (e.g., the Hadoop Distributed File System (HDFS)), an object storage system, and the like.
[0029] The execution platform 114 includes multiple compute nodes (e.g., virtual warehouses). A set of processes on the compute nodes executes a query plan compiled by the compute service manager 112. The set of processes may include: a first process that executes the query plan; a second process that uses a least recently used (LRU) strategy to monitor and delete micro-partition files and implement out-of-memory (OOM) error mitigation; a third process that extracts health information from process logs and status information and sends it back to the compute service manager 112; a fourth process that establishes communication with the compute service manager 112 after the system boots; and a fifth process that handles all communications with the compute cluster for a given job provided by the compute service manager 112 and transmits information back to the compute service manager 112 and other compute nodes of the execution platform 114.
[0030] The cloud computing storage platform 104 also includes an access management system 118 and a web proxy 120. Like the access management system 110, the access management system 118 allows users to create and manage users, roles, and groups, and use permissions to allow or deny access to cloud services and resources. The access management system 110 of the network-based data warehouse system 102 and the access management system 118 of the cloud computing storage platform 104 can communicate and share information to enable access and management of resources and services shared by users of both the network-based data warehouse system 102 and the cloud computing storage platform 104. The web proxy 120 handles the tasks involved in accepting and processing concurrent API calls, including traffic management, authorization and access control, monitoring, and API version management. The web proxy 120 provides HTTP proxy services for creating, publishing, maintaining, protecting, and monitoring APIs (e.g., REST APIs).
[0031] In some embodiments, the communication links between the elements of the shared data processing platform 100 are implemented via one or more data communication networks. These data communication networks can utilize any communication protocol and any type of communication medium. In some embodiments, the data communication network is a combination of two or more data communication networks (or subnetworks) coupled to each other. In alternative embodiments, these communication links are implemented using any type of communication medium and any communication protocol.
[0032] like Figure 1As shown, data storage devices 124-1 through 124-N are decoupled from the computing resources associated with execution platform 114. That is, new virtual warehouses can be created and terminated within execution platform 114, and additional data storage devices can be created and terminated independently on cloud computing storage platform 104. This architecture supports dynamic changes to network-based data warehouse system 102 based on changing data storage / retrieval requirements and the changing needs of users and systems accessing shared data processing platform 100. Support for dynamic changes allows network-based data warehouse system 102 to scale rapidly in response to the ever-changing demands on systems and components within network-based data warehouse system 102. The decoupling of computing resources from data storage devices 124-1 through 124-N supports the storage of large amounts of data without requiring a correspondingly large amount of computing resources. Similarly, this decoupling of resources supports a significant increase in the computing resources used at a given time without requiring a corresponding increase in available data storage resources. Furthermore, the decoupling of resources enables different accounts to create additional computing resources to process data shared by other users without impacting the other users' systems. For example, a data provider may have three computing resources and share data with a data consumer, and the data consumer may generate a new computing resource to perform queries on the shared data, where the new computing resource is managed by the data consumer without affecting or interacting with the data provider's computing resources.
[0033] The computing service manager 112, the database 116, the execution platform 114, the cloud computing storage platform 104 and the remote computing device 106 are Figure 1 104. However, each of the computing service manager 112, database 116, execution platform 114, cloud computing storage platform 104, and remote computing environment can be implemented as a distributed system (e.g., distributed across multiple systems / platforms in multiple geographic locations) connected via an API and access information (e.g., tokens, login data). In addition, each of the computing service manager 112, database 116, execution platform 114, and cloud computing storage platform 104 can be scaled up or down (independently of each other) based on changes in received requests and changing needs of the shared data processing platform 100. Therefore, in the described embodiment, the network-based data warehouse system 102 is dynamic and supports regular changes to meet current data processing needs.
[0034] During typical operation, the network-based data warehouse system 102 processes multiple jobs (e.g., queries) determined by the computing service manager 112. These jobs are scheduled and managed by the computing service manager 112 to determine when and how the jobs are executed. For example, the computing service manager 112 may divide the job into multiple discrete tasks and determine what data is required to execute each of the multiple discrete tasks. The computing service manager 112 may assign each of the multiple discrete tasks to one or more nodes of the execution platform 114 to process the task. The computing service manager 112 may determine what data is required to process the task and further determine which nodes within the execution platform 114 are best suited to process the task. Some nodes may already have cached the data required to process the task (because they recently downloaded data from the cloud computing storage platform 104 for a previous job), making them good candidates for processing the task. Metadata stored in the database 116 helps the computing service manager 112 determine which nodes within the execution platform 114 have cached at least a portion of the data required to process the task. One or more nodes in the execution platform 114 process the task using data cached by those nodes and, when necessary, data retrieved from the cloud computing storage platform 104. It is desirable to retrieve as much data as possible from the cache within the execution platform 114 because the retrieval speed is typically much faster than retrieving data from the cloud computing storage platform 104.
[0035] like Figure 1 As shown, the shared data processing platform 100 separates the execution platform 114 from the cloud computing storage platform 104. In this arrangement, the processing resources and cache resources in the execution platform 114 operate independently of the data storage devices 124-1 through 124-N in the cloud computing storage platform 104. Therefore, the computing resources and cache resources are not limited to specific data storage devices 124-1 through 124-N. Instead, all computing resources and all cache resources can retrieve data from and store data to any data storage resource in the cloud computing storage platform 104.
[0036] Figure 2 1 is a block diagram illustrating components of the computing service manager 112 according to some embodiments of the present disclosure. Figure 2As shown, the request processing service 202 manages received data storage requests and data retrieval requests (e.g., jobs to be performed on database data). For example, the request processing service 202 can determine the data required to process a received query (e.g., a data storage request or a data retrieval request). The data may be stored in a cache within the execution platform 114 or in a data storage device in the cloud computing storage platform 104. The management console service 204 supports administrators and other system managers' access to various systems and processes. In addition, the management console service 204 can receive requests to execute jobs and monitor the workload on the system. According to some example embodiments, and as discussed in further detail below, the stream sharing engine 225 manages change tracking of database objects (e.g., data shares (e.g., shared tables) or shared views).
[0037] The computing service manager 112 also includes a job compiler 206, a job optimizer 208, and a job executor 210. The job compiler 206 parses a job into multiple discrete tasks and generates execution code for each of the multiple discrete tasks. The job optimizer 208 determines the optimal method for executing the multiple discrete tasks based on the data to be processed. The job optimizer 208 also handles various data pruning operations and other data optimization techniques to improve the speed and efficiency of job execution. The job executor 210 executes the execution code of a job received from a queue or determined by the computing service manager 112.
[0038] The job scheduler and coordinator 212 sends the received jobs to the appropriate service or system for compilation, optimization and dispatch to the execution platform 114. For example, the jobs can be prioritized and processed according to the priority order. In an embodiment, the job scheduler and coordinator 212 prioritizes internal jobs scheduled by the computing service manager 112 and other "external" jobs (such as user queries) that can be scheduled by other systems in the database but can utilize the same processing resources in the execution platform 114. In some embodiments, the job scheduler and coordinator 212 identifies or assigns specific nodes in the execution platform 114 to process specific tasks. The virtual warehouse manager 214 manages the operation of multiple virtual warehouses implemented in the execution platform 114. As discussed below, each virtual warehouse includes multiple execution nodes, each execution node including a cache and a processor (e.g., a virtual machine, an operating system-level container execution environment).
[0039] Additionally, the compute service manager 112 includes a configuration and metadata manager 216, which manages information related to data stored in remote data storage devices and local caches (i.e., caches within the execution platform 114). The configuration and metadata manager 216 uses metadata to determine which data micro-partitions need to be accessed to retrieve data for processing a particular task or job. A monitor and workload analyzer 218 oversees the processes executed by the compute service manager 112 and manages the distribution of tasks (e.g., workloads) across virtual repositories and execution nodes within the execution platform 114. The monitor and workload analyzer 218 also reallocates tasks as needed based on the changing workload across the network-based data warehouse system 102 and can also reallocate tasks based on user (e.g., "external") query workloads that can also be processed by the execution platform 114. The configuration and metadata manager 216 and the monitor and workload analyzer 218 are coupled to a data storage device 220. Figure 2 The data storage device 220 in represents any data storage device within the network-based data warehouse system 102. For example, the data storage device 220 may represent a cache in the execution platform 114, a storage device in the cloud computing storage platform 104, or any other storage device.
[0040] Figure 3 is a block diagram illustrating components of the execution platform 114 according to some embodiments of the present disclosure. Figure 3 As shown, the execution platform 114 includes multiple virtual warehouses, which are elastic clusters of computing instances such as virtual machines. In the example shown, the virtual warehouses include virtual warehouse 1, virtual warehouse 2, and virtual warehouse N. Each virtual warehouse (e.g., an EC2 cluster) includes multiple execution nodes (e.g., virtual machines), each of which includes a data cache and a processor. Virtual warehouses can execute multiple tasks in parallel by using multiple execution nodes. As discussed herein, the execution platform 114 can add new virtual warehouses and discard existing virtual warehouses in real time based on the current processing needs of the system and users. This flexibility allows the execution platform 114 to quickly deploy large amounts of computing resources when needed, without being forced to continue paying for them when they are no longer needed. All virtual warehouses can access data from any data storage device (e.g., any storage device in the cloud computing storage platform 104).
[0041] although Figure 3 Each virtual warehouse shown in FIG includes three execution nodes, but a particular virtual warehouse may include any number of execution nodes. Furthermore, the number of execution nodes in a virtual warehouse is dynamic, such that new execution nodes are created when there is additional demand, and existing execution nodes are deleted when they are no longer needed (e.g., when a query or job is completed).
[0042] Each virtual warehouse can access Figure 1 Thus, the virtual warehouse does not need to be assigned to a specific data storage device 124-1 to 124-N, but can access data from any of the data storage devices 124-1 to 124-N within the cloud computing storage platform 104. Similarly, Figure 3 Each execution node shown in can access data from any one of the data storage devices 124-1 to 124-N. For example, the storage device 124-1 of a first user (e.g., a provider account user) can be shared with a worker node in a virtual warehouse of another user (e.g., a consumer account user), so that the other user can create a database (e.g., a read-only database) and use the data in the storage device 124-1 directly without copying the data (e.g., copying it to a new disk managed by the consumer account user). In some embodiments, a particular virtual warehouse or a particular execution node can be temporarily assigned to a particular data storage device, but the virtual warehouse or execution node can later access data from any other data storage device.
[0043] exist Figure 3 In the example shown in FIG, virtual warehouse 1 includes three execution nodes 302-1, 302-2, and 302-N. Execution node 302-1 includes a cache 304-1 and a processor 306-1. Execution node 302-2 includes a cache 304-2 and a processor 306-2. Execution node 302-N includes a cache 304-N and a processor 306-N. Each execution node 302-1, 302-2, and 302-N is associated with processing one or more data storage and / or data retrieval tasks. For example, a virtual warehouse may process data storage and data retrieval tasks associated with an internal service (e.g., a clustering service, a materialized view refresh service, a file compression service, a stored procedure service, or a file upgrade service). In other embodiments, a specific virtual warehouse may process data storage and data retrieval tasks associated with a specific data storage system or a specific category of data.
[0044] Similar to virtual warehouse 1 discussed above, virtual warehouse 2 includes three execution nodes 312-1, 312-2, and 312-N. Execution node 312-1 includes a cache 314-1 and a processor 316-1. Execution node 312-2 includes a cache 314-2 and a processor 316-2. Execution node 312-N includes a cache 314-N and a processor 316-N. Additionally, virtual warehouse 3 includes three execution nodes 322-1, 322-2, and 322-N. Execution node 322-1 includes a cache 324-1 and a processor 326-1. Execution node 322-2 includes a cache 324-2 and a processor 326-2. Execution node 322-N includes a cache 324-N and a processor 326-N.
[0045] In some embodiments, relative to the data being cached by the execution node, Figure 3 The execution nodes shown are stateless. For example, they do not store or otherwise maintain state information about the execution nodes or data cached by a particular execution node. Therefore, in the event of an execution node failure, the failed node can be transparently replaced with another node. Because there is no state information associated with the failed execution node, a new (replacement) execution node can easily replace the failed node without having to worry about recreating specific state.
[0046] although Figure 3 The execution nodes shown each include a data cache and a processor, but alternative embodiments may include execution nodes that include any number of processors and any number of caches. Additionally, the size of the caches may vary between different execution nodes. Figure 3 The illustrated cache stores data retrieved from one or more data storage devices in the cloud computing storage platform 104 (e.g., the S3 objects most recently accessed by a given node) locally on the execution node (e.g., on a local disk). In some example embodiments, the cache stores file headers and individual columns of a file when a query downloads only the columns required for the query.
[0047] To improve cache hits and avoid overlapping, redundant data being stored in node caches, job optimizer 208 distributes input file sets to nodes using a consistent hashing scheme to hash the table file names of the accessed data (e.g., data in database 116 or database 122). According to some example embodiments, subsequent or concurrent queries that access the same table file will therefore be executed on the same node.
[0048] As discussed, nodes and virtual warehouses can change dynamically in response to environmental conditions (e.g., disaster scenarios), hardware / software issues (e.g., failures), or management changes (e.g., changing from a large cluster to a smaller cluster to reduce costs). In some example embodiments, when the set of nodes changes, no data is immediately reshuffled. Instead, a least recently used replacement strategy is implemented to eventually replace cached content that is lost across multiple jobs. Thus, the cache reduces or eliminates the bottleneck problem that occurs in platforms that continually retrieve data from remote storage systems. Instead of repeatedly accessing data from remote storage devices, the systems and methods described herein access data from a cache in an execution node, which is significantly faster and avoids the bottleneck problem discussed above. In some embodiments, the cache is implemented using a high-speed memory device that provides fast access to cached data. Each cache can store data from any storage device in the cloud computing storage platform 104.
[0049] In addition, cache resources and computing resources can vary between different execution nodes. For example, one execution node may contain a large amount of computing resources and minimal cache resources, making it useful for tasks that require a large amount of computing resources. Another execution node may contain a large amount of cache resources and minimal computing resources, making it useful for tasks that require caching a large amount of data. Yet another execution node may contain cache resources that provide faster input-output operations, which is useful for tasks that require rapid scanning of large amounts of data. In some embodiments, the execution platform 114 implements skew handling to allocate work between cache resources and computing resources associated with a particular execution, wherein the allocation can be further based on the expected tasks to be performed by the execution node. For example, if the tasks performed by the execution node become more processor-intensive, more processing resources can be allocated to the execution node. Similarly, if the tasks performed by the execution node require a larger cache capacity, more cache resources can be allocated to the execution node. In addition, due to various issues (e.g., virtualization issues, network overhead), some nodes may execute much slower than other nodes. In some example embodiments, a file stealing solution is used to resolve the imbalance at the scan level. In particular, each time a node process completes a scan of its input file set, it requests additional files from other nodes. If one of the other nodes receives such a request, that node analyzes its own set (e.g., how many files are left in the input file set when the request is received) and then transfers ownership of one or more remaining files for the duration of the current job (e.g., query). The requesting node (e.g., a file stealing node) then receives the data (e.g., header data) and downloads the file from the cloud computing storage platform 104 (e.g., from the data storage device 124-1) without downloading the file from the transfer node. In this way, a lagging node can transfer files via file stealing in a manner that does not exacerbate the load on the lagging node.
[0050] Although virtual warehouses 1, 2, and N are associated with the same execution platform 114, the virtual warehouses may be implemented using multiple computing systems in multiple geographic locations. For example, virtual warehouse 1 may be implemented by a computing system in a first geographic location, while virtual warehouse 2 and virtual warehouse N may be implemented by another computing system in a second geographic location. In some embodiments, these different computing systems are cloud-based computing systems maintained by one or more different entities.
[0051] In addition, each virtual warehouse Figure 31 and 2 as having multiple execution nodes. Multiple computing systems at multiple geographic locations can be used to implement the multiple execution nodes associated with each virtual warehouse. For example, an instance of virtual warehouse 1 implements execution nodes 302-1 and 302-2 on one computing platform at one geographic location, while implementing execution node 302-N on a different computing platform at another geographic location. The selection of a particular computing system to implement an execution node can depend on various factors, such as the level of resources required for a particular execution node (e.g., processing resource requirements and cache requirements), the resources available at a particular computing system, the communication capabilities of networks within or between geographic locations, and which computing systems have already implemented other execution nodes in the virtual warehouse.
[0052] The execution platform 114 is also fault-tolerant. For example, if one virtual warehouse fails, the virtual warehouse will be quickly replaced by a different virtual warehouse located in a different geographical location.
[0053] A particular execution platform 114 may include any number of virtual warehouses. Furthermore, the number of virtual warehouses in a particular execution platform may be dynamic, such that new virtual warehouses are created when additional processing and / or caching resources are needed. Similarly, existing virtual warehouses may be deleted when the resources associated with the virtual warehouse are no longer necessary.
[0054] In some embodiments, virtual warehouses can operate on the same data in the cloud computing storage platform 104, but each virtual warehouse has its own execution node with independent processing and cache resources. This configuration allows requests on different virtual warehouses to be processed independently without interference between requests. This independent processing, combined with the ability to dynamically add and remove virtual warehouses, supports adding new processing capacity for new users without affecting the performance observed by existing users.
[0055] Figure 4An example of two independent accounts in a data warehouse system according to some example embodiments is shown. Here, Company A can operate Account A 402 using the network-based data warehouse system described herein. In Account A 402, Company A data 404 can be stored. Company A data 404 may include, for example, customer data 406 related to Company A's customers. Customer data 406 may be stored in a table or other format that stores customer information and other related information. Other related information may include identifying information (e.g., email) and other known characteristics of the customer (e.g., gender, geographic location, purchasing habits, etc.). For example, if Company A is a consumer goods company, purchasing characteristics (e.g., whether the customer is single, married, a member of a suburban household or a member of an urban household, etc.) may be stored. If Company A is a streaming service company, information about the customer's viewing habits (e.g., whether the customer likes science fiction, nature, reality, action, etc.) may be stored.
[0056] Similarly, Company B may operate Account B 412 using the network-based data warehouse system described herein. Within Account B 412, Company B data 414 may be stored. Company B data 414 may include, for example, customer data related to Company B's customers. Customer data 416 may be stored in a table or other format that stores customer information and other related information. As described above, other related information may include identifying information (e.g., email address) and other known characteristics of the customer (e.g., gender, geographic location, purchasing habits, etc.).
[0057] For security reasons, Company B might not be able to access Company A's data, and vice versa. However, Company A and Company B might want to share at least some data with each other without revealing sensitive information (such as personally identifiable information about their customers). For example, Company A and Company B might want to explore cross-marketing or advertising opportunities and might want to see how much overlap there is between their customers and filter based on certain characteristics of those customers to identify relationships and patterns.
[0058] To this end, a network-based data warehouse system as described herein can provide a data clean room. Figure 5 is a block diagram illustrating a method for operating a data clean room according to some example embodiments. The data clean room can enable Company A and Company B to perform overlapping analysis on their company data without sharing sensitive data and without losing control of the data. The data clean room can establish connections between the data of each account and can include a set of blind cross-reference tables.
[0059] Next, an example operation of creating a data clean room is described. Account B may include customer data 502 for company B, and account A may include customer data 504 for company A. In this example, account B may initiate the creation of a data clean room; however, either account may initiate the creation of a data clean room. Account B may create a security function 506. The security function 506 may look up specific identifier information in the customer data 502 of account B. The security function 506 may anonymize the information (e.g., generate a first result set) by creating an identifier for each customer data. The security function 506 may be a secure user-defined function (UDF) and may be implemented using the technology described in U.S. patent application No. 16 / 814,875, entitled “System and Method for Global Data Sharing,” filed on March 10, 2020, which is incorporated herein by reference in its entirety, including but not limited to those portions specifically appearing below, and the incorporation by reference is made with the following exception: if any portion of the above-mentioned application is inconsistent with this application, this application replaces the above-mentioned application.
[0060] The secure function 506 can be implemented as a SQL UDF. The secure function 506 can be defined to protect the underlying data used to process the function. In this way, the secure function 506 can limit the direct and indirect exposure of the underlying data.
[0061] Secure function 506 can then be shared with Account A using secure share 508. Secure share 508 can allow Account A to execute secure function 506 while limiting Account A's access to Account B's underlying data used by the function and restricting Account A from viewing the function's code. Secure share 508 can also limit Account A's access to the code of secure function 506. Furthermore, secure share 508 can limit Account A from viewing any logs or other information regarding Account B's use of secure function 506, or the parameters to secure function 506 provided by Account B when secure function 506 was called.
[0062] Account A can execute a secure function 506 (e.g., to generate a second result set) using its customer data 504. The result of executing the secure function 506 can be transmitted to account B. For example, a cross-reference table 510 can be created in account B, which can include anonymized customer information 512 (e.g., anonymized identity information). Similarly, a cross-reference table 514 can be created in account A, which can include anonymized customer information 516 for matching overlapping customers of the two companies, as well as dummy identifiers for unmatched records. The data from the two companies can be securely connected so that neither account can access the underlying data or other identifiable information. For example, data can be securely joined using the technology described in U.S. patent application Ser. No. 16 / 368,339, filed on May 28, 2019, entitled “Secure Data Joins in a Multiple Tenant Database System,” which is incorporated herein by reference in its entirety, including but not limited to those portions specifically appearing below, with the following exception: this application supersedes the above-referenced application if any portion is inconsistent with the present application.
[0063] For example, cross-reference table 510 (and anonymized customer information 512) may include fields: "my_cust_id," which may correspond to a customer ID in Account B's data; "my_link_id," which may correspond to an anonymous link to the identified customer information; and "their_link_id," which may correspond to an anonymized matching customer in Company A. "their_link_id" may be anonymized so that Company B cannot discern the identity of the matching customer. Anonymization may be performed using hashing, encryption, tokenization, or other suitable techniques.
[0064] Furthermore, to further anonymize identities, all listed customers of Company B in the cross-reference table 510 (and anonymized customer information 512) may have a unique matching customer listed from Company A, regardless of whether there is an actual match. A false "their_link_id" may be created for unmatched customers. Thus, neither company may be able to determine the identity information of a matching customer. Neither company may be able to discern where there is an actual match, rather than a false identifier returned (no match). Therefore, the cross-reference table 510 may include anonymized key-value pairs. A summary report may be created noting the total number of matches, but other details about the matched customers may not be provided to protect the customers' identities.
[0065] Data cleanrooms can operate in one or two directions, meaning double-blind cleanrooms can be provided. Figure 6 A block diagram of a method for operating a double-blind cleanroom, according to some example embodiments, is shown. The double-blind cleanroom can enable Company A to perform overlap analysis using its corporate data with Company B's corporate data, and vice versa, without sharing sensitive data and without losing control of their own data. The double-blind cleanroom can establish connections between each account's data and can include a set of double-blind cross-reference tables.
[0066] Here, Account A may include its customer data 602, and Account B may include its customer data 604. As described above, Account A may create a secure function 606 ("Get_Link_ID"). As described above, secure function 606 may be shared with Account B using secure sharing 608. Additionally, a stored-procedure function 610 may detect changes in the data in the corresponding customer data and may update and refresh the link accordingly.
[0067] The same or similar process can be applied from Account B to Account A using secure functions 612, secure shares 614, and stored procedures 616. Thus, a cross-reference table 618 in Account A can include information about customer overlap between the two companies. For example, the cross-reference table 618 includes fields: "my_cust_id," which can correspond to a customer ID in Account A's data; "my_link_id," which can correspond to an anonymized link to Company A's identified customer information; and "their_link_id," which can correspond to an anonymized matching customer in Company B. Anonymization can be performed using hashing, encryption, tokenization, and the like.
[0068] Similarly, the cross-reference table 620 in Account B may include information about customer overlap between the two companies. For example, the cross-reference table 620 includes fields: "my_cust_id," which may correspond to a customer ID in Account B's data; "my_link_id," which may correspond to an anonymized link to Company B's identified customer information; and "their_link_id," which may correspond to an anonymized matching customer in Company A. Anonymization may be performed using hashing, encryption, tokenization, or other suitable techniques.
[0069] In the example above, both Company A and Company B have accounts with data warehouse systems. However, the blind cleanroom technique described in this article can also find application when one or both companies do not have accounts with data warehouse systems. Figure 7Techniques for loading data into a data warehouse system, according to some example embodiments, are illustrated. Here, Company C may not have an account with a data warehouse system but may still want to employ the data cleanroom techniques described herein. Company C can load its company data into a load file 702 (e.g., a .csv file). Then, using a browser 704 or an app, Company C can use a file uploader 706 to upload the load file 702 to a secure cloud storage location 708 (also known as an "enclave bucket"). Data can be moved from the enclave bucket 708 into an enclave account 710. Data can be moved using batch data ingestion commands, copy commands, and the like. A customized web GUI 712 can also be used to send control information to the enclave account 710. For example, the control information can set access restrictions for the data in the load file 702. As described above, the enclave account 710 can then function and operate as a regular account for the purposes of creating and using a data cleanroom.
[0070] After creating a data clean room, each party can run secure queries on the secure data to gather more detailed information. In the example of securely matching customers of two companies, either company can send a query request to the other company to determine the number of matches based on selected criteria.
[0071] Figures 8A-8C is a block diagram illustrating a method of processing security queries using a data clean room, according to some example embodiments. Figures 8A-8C The examples in build on Figure 6 Here, Account A may include its customer data 602 and Account B may include its customer data 604. As explained above, cross-reference tables 618 and 620 may be created and included in Account A and Account B, respectively. Additionally, sales data 802 ( Figure 8C ) may be included in Account A, and viewing data 804, which may include information about the customer's viewing habits, may be included in Account B.
[0072] In addition, account B can use secure view 806 to allow company A to access selected data called secure query available data 808. Account B can notify account A of the secure query available data 808 in various ways. Account B can publish this information to account A in advance. It can share the structure and lookup key of the secure query available data 808 with account A. It can also use private data exchange using the technology described in U.S. patent application No. 16 / 746,673, entitled "Private Data Exchange", filed on January 17, 2020, which is incorporated herein by reference in its entirety, including but not limited to those parts specifically appearing below, and the incorporation by reference is made with the following exception: if any part of the above-mentioned application is inconsistent with this application, then this application replaces the above-mentioned application.
[0073] A user in Account A can run a secure query 810 in the data clean room. The query can request data from both Account A and Account B, but can restrict the user's access to sensitive data from Account B. For example, a user can run a query requesting information: How many of my (Company A's) customers watched a show from Company B, grouped by Company B's shows and my (Company A's) segment also bought my "Top Paper Towels" product, who I know who lives in the United States or Canada and who is also anime fan from Company B, and where there are at least two customers in each result group?
[0074] Using Company A's customer data 602, cross-reference table 618, and sales data 802, the query can generate a temporary table 812 ("MyData1"). In this example, temporary table 812 can include anonymized Company A customer information for customers who purchased "Top Paper Towels" and reside in the United States or Canada, securely linked to matching anonymized Company B customers (which may or may not include "fake" matching accounts as described above). A secure query request 814 can also be generated and sent to Company B. Secure query request 814 can request Company B to run the remainder of the query and return the final results. For example, secure query request 814 can include a request ID, filters for Company B to apply ("select_c" and "where_c"), and the output format of the final results ("result_table"). Secure query request 814 can be provided in the form of a request table (as shown) or can be another type of remote procedure call (e.g., a SQL statement). Temporary table 812 and secure query request 814 can be shared with Account B using secure query request sharing 816.
[0075] Next, account B can complete the remainder of the original query. At account B, a copy of the temporary table 818 can be stored. The secure query request 814 can be received by the stream 820 on the request and the task 822 function on the new request. The secure query can then be executed by the restricted query process function 824. The function can use data from various sources (such as the copy of the temporary table 818, the cross-reference table 620, the viewing data 804, and the secure query available data 808) to perform the query. In this example, the restricted query process function 824 can filter the customers identified in the temporary table 818 to those customers who have watched one of company B's shows and are animation fans, and the function can group the results by the show and company A's segment group, as shown in the query state 826. The results can be generated and output to the temporary result table 828. Next, the last part of the query can be executed: filtering out results that have less than two customers in each result group (e.g., a minimum threshold). Figures 8A-8C In the example shown, there is only one result for the grouping of "Movie B" and the specific Company A segment, so it is removed. The final results can then be shared with Account A using the secure query result sharing function 830. The results in Account A can be provided as a final results table 832 ("res1"), which shows matches for queries grouped by both programs.
[0076] Figure 9 A diagrammatic representation of a machine 900 in the form of a computer system within which a set of instructions for causing the machine 900 to perform any one or more of the methodologies discussed herein may be executed is shown according to an example embodiment. Figure 9 A diagrammatic representation of a machine 900 is shown in the form of an example computer system within which instructions 916 (e.g., software, programs, applications, applet, apps, or other executable code) for causing the machine 900 to perform any one or more of the methods discussed herein can be executed. For example, the instructions 916 can cause the machine 900 to perform any one or more operations of any one or more of the methods described herein. As another example, the instructions 916 can cause the machine 900 to implement portions of the data flows described herein. In this manner, the instructions 916 transform a general-purpose, unprogrammed machine into a specific machine 900 (e.g., a remote computing device 106, an access management system 110, a computing service manager 112, an execution platform 114, an access management system 118, a web agent 120, a remote computing device 106) that is specifically configured to perform any of the functions described and illustrated in the manner described herein.
[0077] In alternative embodiments, the machine 900 operates as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 900 may operate as a server or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 900 may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a smartphone, a mobile device, a network router, a network switch, a network bridge, or any machine capable of executing instructions 916, sequentially or otherwise, to specify actions to be taken by the machine 900. Further, while a single machine 900 is illustrated, the term "machine" shall also be taken to include any collection of machines 900 that individually or jointly execute instructions 916 to perform any one or more of the methodologies discussed herein.
[0078] The machine 900 includes a processor 910, a memory 930, and an input / output (I / O) component 950 that are configured to communicate with each other, for example, via a bus 902. In an example embodiment, the processor 910 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, a processor 912 and a processor 914 that may execute instructions 916. The term "processor" is intended to include a multi-core processor 910 that may include two or more independent processors (sometimes referred to as "cores") that may execute instructions 916 concurrently. Although Figure 9 Multiple processors 910 are shown, but the machine 900 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0079] The memory 930 may include a main memory 932, a static memory 934, and a storage unit 936, all of which may be accessed by the processor 910, for example, via the bus 902. The main memory 932, the static memory 934, and the storage unit 936 store instructions 916, which embody any one or more of the methodologies or functionality described herein. During execution by the machine 900, the instructions 916 may reside, in whole or in part, within the main memory 932, the static memory 934, the storage unit 936, within at least one processor 910 (e.g., within a cache memory of a processor), or any suitable combination thereof.
[0080] The I / O components 950 include components for receiving input, providing output, generating output, transmitting information, exchanging information, capturing measurements, and the like. The specific I / O components 950 included in a particular machine 900 will depend on the type of machine. For example, a portable machine such as a mobile phone will likely include a touch input device or other such input mechanism, while a headless server machine will be less likely to include such a touch input device. It will be appreciated that the I / O components 950 may include Figure 9 Many other components are not shown in the figure. The I / O components 950 are grouped according to function only to simplify the following discussion, and this grouping is in no way limiting. In various example embodiments, the I / O components 950 may include output components 952 and input components 954. The output components 952 may include visual components (e.g., displays such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tubes (CRTs)), acoustic components (e.g., speakers), other signal generators, etc. The input components 954 may include alphanumeric input components (e.g., keyboards, touch screens configured to receive alphanumeric input, optical keyboards, or other alphanumeric input components), point-based input components (e.g., mice, trackpads, trackballs, joysticks, motion sensors, or other pointing instruments), tactile input components (e.g., physical buttons, touch screens or other tactile input components that provide location and / or force for touch or touch gestures), audio input components (e.g., microphones), etc.
[0081] Communication can be implemented using a variety of technologies. The I / O component 950 may include a communication component 964 that is operable to couple the machine 900 to a network 980 or a device 970 via coupling 982 and coupling 972, respectively. For example, the communication component 964 may include a network interface component or another suitable device that interfaces with the network 980. In further examples, the communication component 964 may include a wired communication component, a wireless communication component, a cellular communication component, and other communication components that provide communication via other modalities. The device 970 may be another machine or any of a variety of peripheral devices (e.g., a peripheral device coupled via a universal serial bus (USB)). For example, as described above, the machine 900 may correspond to any of the remote computing device 106, the access management system 110, the computing service manager 112, the execution platform 114, the access management system 118, and the web agent 120, and the device 970 may include any of these systems and devices.
[0082] Various memories (e.g., memories of 930, 932, 934 and / or processor 910 and / or storage unit 936) may store one or more sets of instructions 916 and data structures (e.g., software) that embody or are utilized by any one or more of the methods or functions described herein. When executed by processor 910, these instructions 916 cause various operations to implement the disclosed embodiments.
[0083] As used herein, the terms “machine storage medium,” “device storage medium,” and “computer storage medium” mean the same and may be used interchangeably in this disclosure. These terms refer to a single or multiple storage devices and / or media (e.g., a centralized or distributed database and / or associated caches and servers) that store executable instructions and / or data. Accordingly, these terms should be considered to include, but are not limited to, solid-state memory and optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage medium, computer storage medium, and / or device storage medium include non-volatile memory, including, for example: semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), field programmable gate arrays (FPGAs), and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM optical disks. The terms “machine storage medium,” “computer storage medium,” and “device storage medium” specifically exclude carrier waves, modulated data signals, and other such media (at least some of which are included in the term “signal media” discussed below).
[0084] In various example embodiments, one or more portions of network 980 may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of a public switched telephone network (PSTN), a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, 980 may include a wireless or cellular network, another type of network, or a combination of two or more such networks. For example, network 980 or a portion of network 980 may include a wireless or cellular network, and coupling 982 may be a code division multiple access (CDMA) connection, a global system for mobile communications (GSM) connection, or another type of cellular or wireless coupling. In this example, coupling 982 may implement any of a variety of types of data transmission technologies, such as single carrier radio transmission technology (1xRTT), evolution data optimized (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rates for GSM evolution (EDGE) technology, third generation partnership project (3GPP) standards including 3G, fourth generation wireless (4G) networks, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), world wide interoperability for microwave access (WiMAX), long term evolution (LTE), other technologies defined by various standards setting organizations, other long range protocols, or other data transmission technologies.
[0085] Instructions 916 can be transmitted or received over network 980 using a transmission medium via a network interface device (e.g., a network interface component included in communication component 964) and utilizing any of a variety of well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instructions 916 can be transmitted or received to device 970 using a transmission medium via coupling 972 (e.g., a peer-to-peer coupling). The terms "transmission medium" and "signal medium" are synonymous and are used interchangeably in this disclosure. The terms "transmission medium" and "signal medium" should be understood to include any intangible medium capable of storing, encoding, or carrying instructions 916 for execution by machine 900, and include digital or analog communication signals or other intangible media that facilitate the communication of such software. Accordingly, the terms "transmission medium" and "signal medium" should be understood to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal.
[0086] The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" are synonymous and may be used interchangeably in this disclosure. These terms are defined to include both machine storage media and transmission media. Thus, these terms include both storage devices / medium and carrier / modulated data signals.
[0087] The various operations of the example methods described herein can be performed at least in part by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the related operations. Similarly, the method described herein can be processor-implemented at least in part. For example, at least some operations of the method described herein can be performed by one or more processors. The execution of some operations can be distributed between one or more processors, and the one or more processors not only reside in a single machine, but also are deployed across multiple machines. In some example embodiments, one or more processors can be located in a single location (e.g., in a home environment, an office environment, or a server farm), and in other embodiments, the processor can be distributed across multiple locations.
[0088] Although embodiments of the present disclosure have been described with reference to specific example embodiments, it will be apparent that various modifications and changes may be made to these embodiments without departing from the broader scope of the subject matter of the present invention. Accordingly, the description and drawings are to be regarded as illustrative and not restrictive. The drawings forming a part of this application show, by way of illustration and not limitation, specific embodiments in which the subject matter may be implemented. The illustrated embodiments are described in sufficient detail to enable those skilled in the art to implement the teachings disclosed herein. Other embodiments and embodiments derived therefrom may be used so that structural or logical substitutions and changes may be made without departing from the scope of the present disclosure. Therefore, this detailed description should not be understood in a limiting sense, and the scope of the various embodiments is limited solely by the appended claims, together with the full range of equivalents to which such claims are entitled.
[0089] Such embodiments of the subject matter of the present invention may be referred to herein individually and / or collectively by the term "invention", which is merely for convenience and is not intended to voluntarily limit the scope of this application to any single invention or inventive concept (if more than one invention or inventive concept is actually disclosed). Therefore, although specific embodiments have been illustrated and described herein, it should be understood that the specific embodiments shown can be replaced with any arrangement calculated to achieve the same purpose. This disclosure is intended to cover any and all adaptations or variations of the various embodiments. Combinations of the above embodiments, as well as other embodiments not specifically described herein, will be apparent to those skilled in the art upon reading the above description.
[0090] In this document, the terms "a" or "an," as is common in patent documents, are used to include one or more than one, independent of any other instance or usage of "at least one" or "one or more." In this document, the term "or" is used to refer to a non-exclusive or, so that "A or B" includes "A but not B," "B but not A," and "A and B," unless otherwise stated. In the appended claims, the terms "including" and "in which" are used as the plain-English equivalents of the respective terms "comprising" and "wherein." Furthermore, in the appended claims, the terms "including" and "comprising" are open-ended; that is, systems, apparatus, articles, or processes that include elements in addition to those listed after such terms in a claim are still considered to fall within the scope of that claim.
[0091] The following numbered examples are embodiments:
[0092] Example 1. A method comprising: providing first-party data in a first account; providing second-party data in a second account; performing, by a processor, a security function using the first-party data to generate a first result, including creating a link to the first-party data and anonymizing identifying information in the first-party data; sharing the security function with a second account; performing the security function using the second-party data to generate a second result, and restricting the second account from accessing the first-party data; and generating a cross-reference table having the first result and the second result, the cross-reference table providing an anonymized match of the first result and the second result.
[0093] Example 2. The method according to Example 1 further includes: restricting the second account from accessing the code of the security function.
[0094] Example 3. The method according to any one of Examples 1-2 further includes: restricting the second account from viewing logs related to the execution of the first part of the security function.
[0095] Example 4. The method according to any one of Examples 1-3 further includes: generating false matching information in the second result in the case of no match.
[0096] Example 5. The method of any one of Examples 1-4, further comprising: generating a summary report of the anonymized matches.
[0097] Example 6. The method of any of Examples 1-5, further comprising: when the number of matches is below a minimum threshold, limiting access to the number of matches.
[0098] Example 7. The method of any of Examples 1-6, wherein providing the first-party data comprises: uploading a load file to a secure cloud storage location; storing data from the load file in the enclave account; and setting access restrictions on the data from the load file based on the control information.
[0099] Example 8. The method according to any one of Examples 1-7 further includes: receiving a query request; executing a first part of the query request based on at least the first-party data and the cross-reference table; generating a temporary table based on executing the first part of the query request; generating a secure query request, the secure query request including instructions related to executing the second part of the query request; and sharing the secure query request and the temporary table with a second account.
[0100] Example 9. The method of any one of Examples 1-8 further includes: executing a secure query request at the second account and concatenating a result of the secure query request with information from the temporary table to generate a final result of the query request.
[0101] Example 10. A system comprising: one or more processors of a machine; and a memory storing instructions that, when executed by the one or more processors, cause the machine to perform operations implementing any of Example Methods 1 to 9.
[0102] Example 11. A machine-readable storage device embodying instructions that, when executed by a machine, cause the machine to perform operations implementing any one of Example Methods 1 to 9.
Claims
1. A method for operating a data clean room, comprising: providing first-party data in a first account in a web-based data warehouse system; providing second-party data in a second account in the network-based data warehouse system; performing, by a first account, a secure function using the first-party data to generate a first result, wherein performing the secure function using the first-party data includes creating a link to the first-party data and anonymizing identifying information in the first-party data; sharing the secure function by the first account with the second account using secure sharing, the secure sharing allowing the second account to execute the secure function while restricting the second account from accessing the first-party data; executing, by the second account, the secure function using the second party data and thereby generating a second result, wherein the second account is restricted from accessing the first party data; as well as A cross-reference table is generated having the first result and the second result, the cross-reference table providing an anonymized match of the first result and the second result.
2. The method according to claim 1, wherein The cross-reference table is accessible via the network-based data warehouse system to perform analysis of overlapping first-party data and second-party data.
3. The method according to claim 1 or 2, further comprising: The code restricts the second account from accessing the secure function.
4. The method according to claim 3, further comprising: The second account is restricted from viewing logs related to execution of the first portion of the secure function.
5. The method according to claim 1, further comprising: In the case of no match, false match information in the second result is generated.
6. The method according to claim 1, further comprising: Generates a summary report of anonymized matches.
7. The method according to claim 6, further comprising: When the number of anonymized matches falls below a minimum threshold, access is restricted to that number of anonymized matches.
8. The method according to claim 1, wherein Providing the first-party data includes: Upload the load file to a secure cloud storage location; storing data from the loaded file into an enclave account; and Based on the control information, access restrictions are set on the data from the loaded file.
9. The method according to claim 1, further comprising: receiving query requests; executing a first portion of the query request based on at least the first-party data and the cross-reference table; generating a temporary table based on executing the first part of the query request; generating a secure query request, the secure query request including instructions related to executing the second portion of the query request; as well as The secure query request and the temporary table are shared with the second account.
10. The method according to claim 9, further comprising: At the second account, the secure query request is executed and the result of the secure query request is connected with the information from the temporary table to generate a final result of the query request.
11. A computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 10.
12. A system for operating a data clean room, comprising: one or more processors of the machine; as well as A memory storing instructions which, when executed by the one or more processors, cause the machine to perform a method according to any one of claims 1-10.
Citation Information
Patent Citations
System and method for global data sharing
US10999355B1
Secure Data Joins In A Multiple Tenant Database System
US20200311297A1
Private data exchange
US20210081439A1
Secure data joins in a multiple tenant database system
US10713380B1