Secure and scalable private set intersection for large datasets

CN116261721BActive Publication Date: 2026-09-29VISA INTERNATIONAL SERVICE ASSOCIATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202180065778.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-10-07
Filing Date
2021-10-06
Publication Date
2026-09-29
Estimated Expiration
2041-10-06

AI Technical Summary

Technical Problem

许多即使不是全部实施的PSI协议(例如,那些基于加密布尔电路、布隆过滤器或布谷鸟散列的协议)很快超过了主存储器,从而需要更多的工程工作

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116261721B_ABST
    Figure CN116261721B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to methods and systems for determining private set intersection (PSI) and performing private database join (PDJ). Some embodiments feature a grouping technique that enables the PSI and PDJ methods to be executed in parallel by worker nodes in a computing cluster, thereby reducing execution time. A first-party computing system and a second-party computing system can each tokenize their respective datasets and then distribute the datasets to chunk groups. The chunk groups can each be filled with dummy tokens. The first-party computing system and the second-party computing system can then perform several parallel PSIs on pairs of corresponding chunk groups. The results can then be combined to produce a set of tokenized intersection sets, which can then be de-tokenized to produce a set intersection.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-referencing related applications

[0002] This application is an international application claiming priority to U.S. Provisional Application No. 63 / 088,863, filed on October 7, 2020, the disclosures of which are hereby incorporated herein by reference in their entirety for all purposes. Background Technology

[0003] The intersection of two sets includes elements common to both sets. Determining the intersection of sets is a practice associated with the use of databases and digital data and has many applications. For example, a nonprofit organization might have a dataset corresponding to a list of people who previously volunteered for the organization, and another dataset including members of the organization residing in a specific city. The nonprofit organization can determine the intersection of these datasets to identify a list of previous volunteers residing in that city. The nonprofit organization can then send mass communications (e.g., text messages, emails, etc.) to inform these volunteers of new volunteering opportunities in the city.

[0004] Typically, determining the intersection of two sets involves comparing the individual elements of those sets. This means that full access to both sets is usually required to determine the intersection. This may not be a problem if the datasets are owned or controlled by one party. However, if the sets are owned or held by different parties, each party will need to disclose their sets to each other to determine the intersection. This can be problematic when the sets contain private or sensitive data, such as personally identifiable information, medical records, etc.

[0005] Fortunately, privacy-preserving methods exist for determining the intersection of sets. Private Set Intersection (PSI) allows parties each holding a private set of elements to compute the intersection of two sets while only displaying the intersection itself. PSI has applications in a variety of settings. For example, PSI can be used to measure the effectiveness of online advertising

[39] , perform private contact discovery [12,21,62], perform privacy-preserving location sharing [50,31], perform privacy-preserving remote diagnostics

[10] , and detect botnets

[49] . Several recent studies, most notably [11,54,55], have investigated the balance between computation and communication. Some have even optimized PSI protocols based on the cost of operating these protocols in the cloud.

[0006] While progress has been made in improving the efficiency of PSI protocols, almost all literature on balanced PSI (e.g., where each party has a private set containing roughly the same number of elements) focuses on sets with a maximum size of 2. 24On a setting of approximately 16 million elements. A notable exception is the study

[67] , which demonstrated the feasibility of using a non-standard “server-assisted” PSI on sets of size one billion elements. In this work, mutually trusted third-party servers helped determine the intersection. Another notable exception is the recent study [59,60]. In this study, two servers (each with more than 16 GB of memory) determined the PSI of two sets of one billion elements in 34.2 hours. This result leaves room for improvement.

[0007] Furthermore, there are many issues associated with extending existing PSI protocols to large datasets, such as memory consumption. Broadly speaking, memory consumption is a problem when implementing cryptographic schemes that operate on large amounts of data. Many, if not all, PSI protocols (e.g., those based on cryptographic Boolean circuits, Bloom filters, or Cuckoo hashes) quickly exceed main memory limits, requiring more engineering work. Even computing the plaintext intersection of billions of elements is a significant challenge.

[0008] The implementation plan addressed these and other issues individually and collectively. Summary of the Invention

[0009] Embodiments of this disclosure relate to improved methods for determining private set intersections (PSIs) in parallel. These methods are fast and efficient, particularly for determining PSIs of large sets (e.g., sets of billions of elements). For example, these methods have been used to determine the PSI of two sets of billions of elements, each containing 128-bit elements, in 83 minutes; this is 25 times faster than the current state-of-the-art solution described in

[60] , which determined the same PSI in 34.2 hours. Furthermore, the embodiments of this disclosure are relatively flexible and easy to implement because they can be used to parallelize most existing methods for determining PSIs.

[0010] Embodiments of this disclosure also relate to an improved method for performing a Private Database Join (PDJ), primarily based on the improved method for performing PSI mentioned above. Broadly speaking, a “join query” (e.g., an SQL statement or a JSON-style request) can be reinterpreted as a PSI operation between sets of “join keys.” The method described herein can be used to determine the intersection set of join keys, which can then be used to generate a join table, thereby completing the PDJ.

[0011] Embodiments of this disclosure also relate to systems, computers, and other apparatuses that can be used to perform the methods described above. These systems may include, for example, an orchestrator computer that interprets requests from clients (corresponding to PSI or PDJ) and sends the requests to first-party servers and second-party servers, each associated with a corresponding database and a corresponding computing cluster. The first-party servers and second-party servers may communicate with each other and with their respective computing clusters to compute results for the PSI or PDJ and return the results to the orchestrator, which may then return the results to the client computer.

[0012] More specifically, one embodiment relates to a method performed by a first-party computing system. The first-party computing system can tokenize a first-party set to generate a tokenized first-party set. The tokenized first-party set may include multiple first-party tokens. The first-party computing system then generates multiple tokenized first-party subsets by assigning each of the multiple first-party tokens to a subset of tokenized first-party subsets using an allocation function. Then, for each tokenized first-party subset of the multiple tokenized first-party subsets, the first-party computing system can perform a private set intersection protocol with the second-party computing system using the tokenized first-party subset and a tokenized second-party subset corresponding to the second-party computing system. In this way, the first-party computing system can perform multiple private set intersection protocols and generate multiple intersecting token subsets. Subsequently, the first-party computing system can combine the multiple intersecting token subsets to generate an intersecting token set, and then detoxify the intersecting token set to generate an intersecting set.

[0013] Another embodiment relates to a different method performed by a first-party computing system. The first-party computing system may receive a private database table join query identifying one or more first database tables and one or more attributes. The first-party computing system may retrieve one or more first database tables from a first-party database and then determine a plurality of first-party join keys based on the one or more first database tables and one or more attributes. Next, the first-party computing system may tokenize the plurality of first-party join keys to generate a set of tokenized first-party join keys, wherein the set of tokenized first-party join keys includes a plurality of first-party tokens. The first-party computing system may generate the plurality of tokenized first-party subsets by using an allocation function to assign each of the plurality of first-party tokens to a subset of tokenized first-party subsets of a plurality of tokenized first-party subsets. Then, for each tokenized first-party subset of the plurality of tokenized first-party subsets, the first-party computing system may use the tokenized first-party subset and a tokenized second-party subset corresponding to a second-party computing system to perform a private set intersection protocol with the second-party computing system, thereby performing multiple private set intersection protocols and generating multiple intersecting token subsets. The first-party computing system can combine multiple intersecting token subsets to generate an intersecting connection key set, and then detoxify the intersecting token set to generate another intersecting connection key set. Subsequently, the first-party computing system can use the intersecting connection key set to filter the one or more first database tables, thereby generating one or more filtered first database tables. The first-party computing system can receive one or more filtered second database tables from the second-party computing system, and then combine the one or more filtered first database tables with the one or more filtered second database tables to generate a join database table.

[0014] These and other embodiments of this disclosure are described in detail below. For example, other embodiments relate to systems, apparatuses, and computer-readable media associated with the methods described herein.

[0015] Before discussing specific embodiments of this disclosure, some terms may be described in detail.

[0016] the term

[0017] A "server computer" can include a powerful computer or cluster of computers. For example, a server computer can include a mainframe, a small cluster of computers, or a group of servers operating as a single unit. In one example, a server computer can include a database server coupled to a web server. A server computer can include one or more computing devices and can use any of a variety of computing architectures, arrangements, and compilations to serve requests from one or more client computers.

[0018] An "edge server" can refer to a server located at the "edge" of a computing domain or network. Edge servers can communicate with computers both within and outside the computing network. Edge servers can allow external computers (e.g., client computers) to access resources or services provided by the computing domain or network.

[0019] "Client computer" can refer to a computer that uses services from other computers or devices, such as server computers. Client computers can connect to these other computers or devices via a network, such as the Internet. As an example, a client computer may include a laptop computer that connects to an image hosting server to view images stored on the image hosting server.

[0020] "Memory" can refer to any suitable one or more devices capable of storing electronic data. Suitable memory can include non-transient computer-readable media storing instructions executable by a processor to perform desired methods. Examples of memory can include one or more memory chips, disk drives, etc. Such memory can be operated using any suitable electrical, optical, and / or magnetic modes of operation.

[0021] "Processor" can refer to any suitable one or more data computing devices. A processor can include one or more microprocessors that work together to perform the desired function. A processor can include a CPU, which includes at least one high-speed data processor sufficient to execute program components for performing user and / or system-generated requests. The CPU can be a microprocessor such as AMD's Athlon, Duron, and / or Opteron; IBM and / or Motorola's PowerPC; IBM and Sony's Cell processor; Intel's Celeron, Itanium, Pentium, Xeon, and / or XScale; and / or similar processors.

[0022] A "hash function" can refer to any function that can be used to map data of arbitrary length or size to data of fixed length or size. Hash functions can also be used to obfuscate data by replacing it with its corresponding "hash value." The hash value can be used as a token.

[0023] A "token" can refer to data used as a substitute for other data. A token can include a sequence of numbers or alphanumeric characters. Tokens can be used to hide confidential or sensitive data. The process of converting data into a token is called "tokenization." Tokenization can be implemented using a hash function. The process of converting a token into its replacement data is called "detokenization." Detokenization can be achieved through a mapping (such as a lookup table) that associates a token with the data it replaces. "Reverse lookup" can refer to a technique that can be used to determine the replacement data based on the token using a mapping.

[0024] A "virtual value" can refer to a value that has no meaning or significance. Virtual values ​​can be generated using a random or pseudo-random number generator. Virtual values ​​can include "virtual tokens," which are tokens that do not correspond to any substitute data.

[0025] The term "multi-party computation (MPC)" can refer to a computation performed by multiple parties. Each party, such as a computer, server, or cryptographic device, may have some computational inputs. Each party can use the inputs to collectively compute the output of the computation.

[0026] The term "secure multi-party computation (secure MPC)" can refer to secure multi-party computation. In some cases, "secure multi-party computation" refers to multi-party computation in which the parties do not share information or other inputs with each other. It is determined that PSI can be implemented using secure MPC.

[0027] An "Oblivious Transmission (OT) protocol" refers to a process where one party can send a message (or other data) to another party without knowing what message was sent. An OT protocol can be n=1, meaning one party can send one of n potential messages to another without knowing which of n messages was sent. OT protocols can be used to implement various forms of secure MPC, including the PSI protocol.

[0028] A "pseudo-random function" can refer to a deterministic function that produces seemingly random outputs. Pseudo-random functions can include collision-resistant hash functions, elliptic curve groups, etc. A pseudo-random function approximates a "random oracle," and is an ideal cryptographic primitive that maps an input to a random output in its output domain. Pseudo-random functions can be constructed using a pseudo-random number generator.

[0029] An "unintentional pseudo-random function" (OPRF) can refer to a function that uses a pseudo-random function and input provided by a second party to deliver a pseudo-random output to a first party. The first party may not learn the input, and the second party may not learn the pseudo-random output. OPRFs can be used to implement many forms of secure MPC, including the PSI protocol.

[0030] A “message” can refer to any data that can be sent between two entities. A message can include plaintext or ciphertext data. A message can include alphanumeric sequences (e.g., “hello123”) or any other data (e.g., an image or video file). Messages can be sent between computers or other entities.

[0031] A "log file" or "audit log" can be a data file that stores recorded information. For example, a log file can include records of usage of a specific service, such as a private database connection service. Log files may contain additional information, such as the time associated with service usage, identifiers associated with clients using the service, and the nature of service usage. Attached Figure Description

[0032] Figure 1 Exemplary use cases of PDJ according to some embodiments are shown.

[0033] Figure 2 Exemplary PPSI and PPDJ systems according to some embodiments are shown.

[0034] Figure 3 A high-level description of PSI parallelization according to some embodiments is shown.

[0035] Figure 4 An exemplary PPSI method including grouping is shown according to some embodiments.

[0036] Figure 5 A flowchart corresponding to an exemplary PPSI method, including grouping, is shown according to some embodiments.

[0037] Figure 6 A flowchart corresponding to an exemplary PPDJ method according to some embodiments is shown.

[0038] Figure 7 A system block diagram of an exemplary SPARK-PSI system according to some embodiments is shown.

[0039] Figure 8 A diagram illustrating the detailed SPARK-PSI workflow according to some embodiments is shown.

[0040] Figure 9 The graph shows a summary of the results from the SPARK-PSI benchmark experiments.

[0041] Figure 10 An exemplary computer system according to some embodiments is shown. Detailed Implementation

[0042] Embodiments of this disclosure relate to an improved implementation of an overtly transmitted (OT) PSI protocol that can be used to rapidly determine the PSI of a set comprising a large number (e.g., billions) of elements. Benchmark experiments performed using these protocols (see Section VII) were used to determine the PSI of two sets, each comprising one billion 128-bit elements. The PSI was determined in approximately 83 minutes.

[0043] In comparison, the original hash protocol for standard set intersection takes 74 minutes to complete, with 19 minutes (26%) for hashing and transmitting data and 55 minutes (74%) for computing the plaintext intersection. Therefore, in terms of execution time, the Parallel PSI (PSI) protocol according to the embodiment is only slightly slower than the insecure set intersection protocol.

[0044] As an additional comparison, the study in

[60] used a solid-state drive to determine the PSI of two billion-element sets containing 128-bit elements in 34.2 hours. Simple hashing took 30.0 hours (88%). OT computation took 3 hours (9%), and plaintext intersection took 1.2 hours (4%). Therefore, the PPSI protocol according to the embodiment can determine the PSI approximately 25 times faster than current state-of-the-art solutions.

[0045] The embodiments achieve these results by using novel techniques that enable the parallelization of the PSI protocol. In this way, parties (e.g., computer systems storing private collections) can distribute their computational workload across multiple worker nodes in a computing cluster (e.g., using a large-scale data processing engine such as Apache Spark), thereby reducing the total amount of time required to compute PSI. Furthermore, these parallelization techniques can be used with many different existing PSI protocols (e.g., KKRT

[41] , PSSZ15

[56] , etc.) without additional modification to those protocols. Therefore, the embodiments can include “plug-and-play” solutions that entities and organizations can more easily integrate into existing PSI systems or infrastructures.

[0046] This disclosure describes the following aspects. In aspect (1), “grouping” can be used to securely generate tokenized subsets based on an input dataset. In aspect (2), the first and second parties can use “Parallel Private Set Intersection (PPSI) technology” (sometimes referred to as the PPSI protocol or Π). PPSIThe private set intersection of the first set and the second set is determined. PPSI technology may involve the use of the grouping technology described above. In aspect (3), the first party and the second party may use the "Parallel Private Database Join (PPDJ)" technology (sometimes referred to as the PPDJ protocol) to perform a private join of one or more first database tables and one or more second database tables. In aspect (4), a "PPSI or PPDJ system" including a computer and other devices may be used to perform PPSI or PPDJ technology. In aspect (5), an implementation of a PPSI or PPDJ system called "SPARK-PSI" may use Apache open-source software, in particular Apache Spark. In aspect (6), benchmark experiments performed on SPARK-PSI demonstrate its speed and efficiency. In aspect (7), cryptographic threat modeling, analysis, and simulation may be used to demonstrate the security of grouping technology, PPSI technology, etc. In aspect (8), various related studies, theories, and additional concepts related to the PSI field are provided, which may help in understanding the embodiments of this disclosure.

[0047] In a broad sense, grouping can involve tokenizing elements of two sets (e.g., a first-party set and a second-party set), distributing tokens into subsets (or “blocks”) of roughly the same size, and filling the subsets with random dummy tokens, thereby masking the actual number of tokens in the subsets. Grouping can securely divide the elements of a set into these subsets without revealing any information about the number or distribution of elements in the set. In this way, grouping enables the parallelization of PSI protocols. Instead of using a single PSI protocol with (large) first and (large) second sets, first-party and second-party computational systems can use pairs of tokenized subsets to execute multiple PSI protocols.

[0048] The PPSI technology described above can involve the application of the grouping techniques. First-party and second-party computing systems can use grouping techniques to partition the first-party set and the second-party set into multiple tokenized first-party subsets and multiple tokenized second-party subsets. Then, the first-party and second-party computing systems can use the corresponding tokenized subset pairs to execute an m-PSI protocol (where m is the number of tokenized subsets corresponding to each party). These PSI protocols can be executed in parallel using a computing cluster comprising multiple worker nodes. The results of these m-PSI protocols can include m intersecting token subsets. One or both of the first-party and second-party computing systems can combine the m intersecting token subsets to generate an intersecting token set. The first-party and second-party computing systems can detoxify the intersecting token set to generate the intersecting set. In this way, the PSI of the first-party set and the second-party set can be determined.

[0049] PPDJ technology can involve the application of the aforementioned PPSI technology. The client computer can send a join query (also known as a "join request," "private database table join query," and other similar terms) to the orchestrator computer. The join query can identify the database tables corresponding to both parties and a set of "attributes" that form the basis of the join operation. Generally, the orchestrator computer can reinterpret this join query as one or more PSI operations on a set of "join keys." The reinterpreted join query can be sent to both the first-party and second-party computing systems. The first-party and second-party computing systems can use the first-party join key set and the second-party join key set to perform PPSI technology (using the corresponding computing clusters of the first-party and second-party computing systems). This may produce an intersection set of join keys. Using the intersection set of join keys, each party can filter its corresponding database tables and then send the filtered database tables to each other. The filtered database tables can then be combined (joined) to complete the PDJ.

[0050] Figure 1 An exemplary use case of a private database connection is illustrated. A first party and a second party may each possess a first-party dataset 102 and a second-party dataset 104, respectively. Both parties may wish to use their datasets to generate a machine learning model 112. For example, these datasets may include advertising data tables, and the machine learning model 112 may include a model for predicting the effectiveness of advertising campaigns. Both parties benefit from using data from the other party to train the machine learning model 112. However, the parties may not wish to freely share data with each other.

[0051] Alternatively, each party can use its respective dataset as input to perform a private database join 106. Thus, each party can enrich its dataset without revealing any additional information to the other. During the training phase, the joined dataset can be used as input to a machine learning algorithm 110. The machine learning algorithm 110 can generate a machine learning model 112. Afterward, either party can use the machine learning model 112 to make any number of inferences 114 on the data, such as whether a particular advertising campaign is effective.

[0052] PPSI or PPDJ systems, which can be used to implement PPSI and / or PPDJ technologies, are described in more detail in Chapter I. Generally, this system includes a client computer, an orchestrator computer, a "first-party domain," and a "second-party domain." The first-party domain may include a first-party computing system, which may include a first-party server and a first-party computing cluster. The second-party domain may include a second-party computing system, which may include a second-party server and a second-party computing cluster. The first-party server and the second-party server may be referred to as "edge servers." Each party can use its respective computing system to perform PPSI or PPDJ technologies together with the other party.

[0053] It should be understood that various methods exist for implementing the PPSI or PPDJ systems described above. Such implementations can utilize a variety of hardware systems, software packages, frameworks, libraries, etc. However, for illustrative purposes, a specific implementation using Apache open-source software, including Apache Spark (SPARK-PSI), is described. At the time of writing, Apache open-source software is widely used in academia, research, and industry for big data applications. Therefore, SPARK-PSI demonstrates practical implementations of some examples.

[0054] In addition, the SPARK-PSI implementation was used in a series of benchmark experiments described in Chapter VII. These benchmark experiments demonstrate the speed and power of the PPSI technique described herein, particularly when compared with existing state-of-the-art PSI protocols. For example, the SPARK-PSI implementation performed a parallel private set intersection between two sets, each containing one billion 128-bit elements, in 83 minutes. The current state-of-the-art PSI protocol described in

[60] achieved the same result in 34.2 hours. Therefore, the method according to the embodiments can be used to perform PSI (on large datasets) at approximately 25 times faster than the current state-of-the-art PSI protocol.

[0055] I.PPMI and PPDJ systems

[0056] The PCSI technique for determining the intersection of two sets and the PPDJ technique for generating join database tables can be performed by PPSI and PPDJ systems, computer networks, databases, and other devices that enable both parties to perform secure private set intersections or secure private database joins.

[0057] A. System Block Diagram

[0058] Figure 2 A system block diagram of an exemplary PPSI and PPDJ system according to some embodiments is shown. Figure 2The system includes: client computer 202, orchestrator (also known as orchestrator computer) 204, first domain 206 and second domain 208.

[0059] First domain 206 and second domain 208 broadly include computing resources corresponding to the first party and the second party, respectively. First domain 206 may include a first-party server 210, a first-party database 222, and a first-party computing cluster 226. Second domain 208 may include a second-party server 212, a second-party database 224, and a second-party computing cluster 228. The combination of first-party server 210 and first-party computing cluster 226 can be referred to as a "first-party computing system." Similarly, the combination of second-party server 212 and second-party computing cluster 228 can be referred to as a "second-party computing system."

[0060] In some embodiments, the first-party computing system and the second-party computing system may comprise a single computer entity, rather than a combination of computer entities as described above. Therefore, it should be understood that in these embodiments, messages sent or received by, for example, the first-party server 210 may instead be sent or received by a single computer entity comprising the first-party computing system, and the same applies to the second-party computing system.

[0061] Figure 2 Computers and devices can communicate with each other via a communication network, which can take any suitable form and may include any one and / or a combination of the following: direct interconnection; the Internet; a local area network (LAN); a metropolitan area network (MAN); an Operational Mission as a Node on the Internet (OMNI); a secure custom connection; a wide area network (WAN); a wireless network (e.g., employing protocols such as, but not limited to, Wireless Application Protocol (WAP), I-mode, etc.). Messages between computers and devices can be sent using secure communication protocols, such as, but not limited to, File Transfer Protocol (FTP); Hypertext Transfer Protocol (HTTP); Secure Hypertext Transfer Protocol (HTTPS), Secure Sockets Layer (SSL), ISO (e.g., ISO 8583), etc.

[0062] 1. Client computer

[0063] Client computer 202 may include a computer system associated with the client. The client may request the output of a PPSI or a PPDJ (join table) on two datasets (intersecting sets). The client can use client computer 202 to request this output by sending a request message to orchestrator 204. When using client computer 202 to request the output of a PPDJ, the request message may include a database query, such as an SQL-style query. Alternatively, the request message may include a JSON request. Client computer 202 may be a computer system associated with a first or second party. After determining that the PSI or PPDJ operation is complete, client computer 202 may receive the results from orchestrator 204. Client computer 202 may communicate with orchestrator 204 via an interface exposed by the orchestrator (e.g., a UI application, portal, Jupyter Labs interface, etc.).

[0064] 2. Arranger

[0065] The orchestrator computer 204 may include a computer system that manages or otherwise directs PPSI and PPDJ operations. The orchestrator 204 may receive request messages from client computers, interpret those request messages, and communicate with first-party server 210 and second-party server 212 to fulfill those requests. For example, if a request message includes a PDJ query, the orchestrator may verify the correctness of the PDJ query, reinterpret the query as a PPSI operation, and then send a request message detailing those operations to first-party server 210 and second-party server 212. Messages from the orchestrator to first-party server 210 and second-party server 212 may, for example, identify a specific dataset on which first-party server 210 and second-party server 212 should perform PPSI or PPDJ operations.

[0066] These messages may also include metadata or data schemas that can be used to perform PPSI or PPDJ operations. The orchestrator computer 204 may acquire this metadata and schemas during an initialization phase that occurs between the orchestrator 204, the first-party server 210, and the second-party server 212. During this initialization phase, the first-party server 210 and the second-party server 212 may send their respective metadata and schemas to the orchestrator 204.

[0067] The orchestrator 204 can connect to the first-party server 210 and the second-party server 212 via their respective cluster interfaces 214 and 220. Once the first-party computing system and the second-party computing system complete the PPSI or PPDJ operation, they can return the results (e.g., intersection sets or join database tables) to the orchestrator 204 through their respective cluster interfaces. The orchestrator 204 can then return the results to the client computer 202.

[0068] Additionally, although the arranger 204 is shown outside the first domain 206 and the second domain 208, in some embodiments the arranger 204 may be included in either of these domains and thus may be operated by either the first or the second party.

[0069] 3. First and second-party servers

[0070] First-party server 210 and second-party server 212 may include edge servers located at the "edge" of first-party domain 206 and second-party domain 208, respectively. First-party server 210 and second-party server 212 can manage PPSI and PPDJ operations performed by their respective compute clusters. First-party server 210 and second-party server 212 can communicate with their respective compute clusters using their respective cluster interfaces 214 and 220. In some embodiments, cluster interfaces 214 and 220 may be implemented using Apache Livy. First-party server 210 and second-party server 212 can communicate with each other via their respective data stream processors 216 and 218. In some embodiments, data stream processors 216 and 218 may be implemented using Apache Kafka. These data stream processors 216 and 218 can also be used to communicate with worker nodes 238-244.

[0071] First-party server 210 can connect to first-party database 222 to retrieve any relevant sets or database tables used to perform PPSI or PPDJ operations. First-party server 210 can perform a grouping technique (described in more detail below) to generate tokenized subsets, which it can send to first-party computing cluster 226 (via cluster interface 214 and driver node 230). First-party computing cluster 226 can then perform PPSI on these tokenized subsets, returning a tokenized intersection set. First-party server 210 can then detoxify the tokenized intersection set to generate an intersection set, which can be returned to client computer 202 via orchestrator 204. Alternatively, if first-party server 210 is performing a PDJ operation, it can use the intersection subset to generate a join database table, which can then be returned to client computer 202 via orchestrator 204.

[0072] Similarly, the second-party server 212 can connect to the second-party database 224 to retrieve any relevant sets or database tables used to perform PPSI or PPDJ operations. The second-party server 212 can perform a grouping technique (described in more detail below) to produce tokenized subsets, which it can send to the second-party computing cluster 228 (via cluster interface 220 and driver node 232). The second-party computing cluster 228 can then perform PPSI on these tokenized subsets, returning a tokenized intersection set. The second-party server 212 can then detoxify the tokenized intersection set to produce an intersection set, which can be returned to the client computer 202 via orchestrator 204. Alternatively, if the second-party server 212 is performing a PDJ operation, it can use the intersection subset to generate a join database table, which can then be returned to the client computer 202 via orchestrator 204.

[0073] 4. First and second-party databases

[0074] First-party database 222 and second-party database 224 may include databases (sometimes referred to as "first-party sets" and "second-party sets") and database tables (sometimes referred to as "first-party database tables" and "second-party database tables") storing datasets. The first-party computing system and the second-party computing system can access their respective databases to retrieve these datasets and database tables in order to perform PPSI and PPDJ operations. Notably, first-party database 222 may be isolated from second-party domain 208. Similarly, second-party database 224 may be isolated from first-party domain 206. This prevents either party from accessing private data belonging to the other party.

[0075] 5. First and second-party computing clusters

[0076] The first-party computing cluster 226 and the second-party computing cluster 228 may include computer nodes that can execute the PSI protocol in parallel to perform PPSI technology according to an embodiment. These may include driver nodes 230 and 232 (also referred to as master nodes) and worker nodes 238-244. Each node may store code that enables it to perform its corresponding function. For example, driver nodes 230 and 232 may each store corresponding PSI driver libraries 234 and 236. Similarly, worker nodes 238-244 may store PSI working libraries 246-252. Worker nodes 238-244 can use these PSI working libraries to execute multiple private set intersection protocols to produce multiple intersecting subsets, which can then be combined to produce an intersecting set.

[0077] In a broader sense, driver nodes 230 and 232 can distribute computational workload among worker nodes in their respective computing clusters. This can include workload related to determining the PSI of tokenized subsets. For example, driver node 230 can assign a specific tokenized subset i to worker node 238 and can identify the corresponding worker node in the second-party computing cluster 228. Thus, the task of worker node 238 is to perform the PSI protocol with the corresponding worker node using tokenized subset i. When it completes its task, worker node 238 can return the result to driver node 230, and driver node 230 can assign a new tokenized subset j to worker nodes. This process can be repeated until the intersection of each tokenized subset is determined. Driver node 230 can then combine these tokenized intersection subsets to produce a tokenized intersection set, and then send the tokenized intersection set to first-party server 210. Alternatively, driver node 230 can send the tokenized intersection subsets to first-party server 210, which can then perform the combination process itself.

[0078] B. Typical PPSI and PPDJ system data flow

[0079] Figure 3 An exemplary data flow corresponding to the PPSI and PPDJ systems is shown. Figure 3 This also largely corresponds to some methods according to the embodiments. The first-party computing system within the first domain 302 and the second-party computing system within the second domain 304 can each tokenize their respective datasets (first-party dataset 306 and second-party dataset 308), thereby generating a tokenized first-party dataset 310 and a tokenized second-party dataset 312. The first-party computing system and the second-party computing system can then map their respective tokens to token block groups 312-322. These token block groups can each be assigned to different worker nodes among multiple worker nodes. Worker nodes 312-322 can execute multiple PSI protocol instances 324-334 on the first domain 302 and the second domain 304. When data exchange is required to execute the PSI protocol, the worker nodes can communicate via a data stream processor (e.g., Figure 2 The data stream processors 216 and 218 in the PSI exchange data. These PSI instances 324-334 can generate multiple intersecting token subsets, which can then be combined and detoxified to produce an intersecting data set. This process is described in more detail in sections II and III below.

[0080] II. Grouping

[0081] As described above, grouping techniques can be used to tokenize a first-party set and a second-party set, each containing n elements. The tokenized first-party set and the second-party set are then separated into m tokenized subsets or "blocks." Each party can then populate each tokenized subset with virtual tokens. In some cases, each party can populate each tokenized subset with virtual tokens to ensure that each subset contains (1+δ0)n / m tokens with some parameter δ0. A subset can also be referred to as a "partition."

[0082] After performing the grouping technique, the first and second parties can execute a series of PSI protocols (e.g., the KKRT protocol) on each corresponding tokenized subset pair. As a result, multiple intersecting token subsets can be combined to produce an intersecting token set. This intersecting token set can then be detoxified, resulting in an intersecting set.

[0083] You can refer to this. Figure 4 and 5 To better understand the PPSI technology involving grouping applications, the diagram illustrates the process for determining a first set 402 and a second set 404, each set including a list of animals. One or more steps in this process may be optional. The first set 402 and the second set 404 may include data records stored in a first database and a second database, respectively. Each record may include additional data fields corresponding to the corresponding animal (e.g., weight, place of origin, etc.).

[0084] refer to Figure 4 and 5 In step 406, the orchestrator computer may receive a request from the client computer. The request may indicate that the client computer wishes to receive the intersection of the first set 402 and the second set 404.

[0085] In step 408, the first-party computing system and the second-party computing system may receive a request message from the orchestrator computer. The request message may correspond to a request received by the orchestrator from the client computer. The request may indicate a first-party set 402 and a second-party set 404. In this way, the first-party computing system and the second-party computing system know which sets are used to perform PPSI.

[0086] In step 410, the first-party computing system can obtain data from a first-party database (e.g., ...). Figure 2 The first-party database 222 retrieves a first-party set 402. Similarly, the second-party computing system can retrieve a second-party set 404 from the second-party database. In some embodiments, the first-party set 402 and the second-party set 404 may include an equal number of elements (referred to as "first-party elements" and "second-party elements"). This number of elements may be represented as n.

[0087] A. Tokenization

[0088] In step 412, the first-party computing system and the second-party computing system can respectively tokenize the first-party set 402 and the second-party set 404, thereby generating a tokenized first-party set and a tokenized second-party set. The tokenized first-party set may include multiple "first-party tokens". Similarly, the tokenized second-party set may include multiple "second-party tokens".

[0089] The first-party computing system and the second-party computing system may use any appropriate means to tokenize the first-party set 402 and the second-party set 404, provided that the means are consistent, i.e., when both parties tokenize the same data element (e.g., “CAMEL”), they produce the same token.

[0090] For example, a first-party computation system and a second-party computation system can use a collision-resistant hash function to tokenize their respective sets. The first-party computation system can generate multiple hash values ​​by hashing each first-party element using the hash function. The tokenized first-party set can include these multiple hash values. Similarly, the second-party computation system can generate a second set of multiple hash values ​​by hashing each second-party element, and the tokenized second-party set can include these second set of multiple hash values. The computation system can then generate a mapping that associates its tokens with elements of the original set. This mapping can include, for example, value pairs corresponding to the token and its original set elements. For example, in step 422, this mapping can later be used to perform detoxification via a reverse lookup.

[0091] B. Subset allocation

[0092] Subsequently, in step 414, the first-party computing system and the second-party computing system can partition their respective tokenized sets into multiple tokenized subsets. For example, the first-party computing system can generate multiple tokenized first-party subsets by using an allocation function to assign each of a plurality of first-party tokens to a tokenized first-party subset of a plurality of tokenized first-party subsets. Similarly, the second-party computing system can generate multiple second-party tokenized subsets by using an allocation function to assign each of a plurality of second-party tokens to a tokenized second-party subset of a plurality of tokenized second-party subsets. In some embodiments, each party can generate an equal number of tokenized subsets. This may include a predetermined number of subsets, which may represent m.

[0093] As mentioned above, a computational system can use an allocation function to perform subset allocation. Many potential allocation functions exist. However, ideally, the allocation function always matches tokens to subsets. That is, if a first-party computational system maps tokens to a specific subset, a second-party computational system should map the same tokens to the corresponding subset.

[0094] For example, in some embodiments, the allocation function may include a token-based dictionary ordering mapping to tokenized subsets T1, ..., T m The dictionary sorting function is used by a first-party computation system to assign each of a plurality of first-party tokens to a corresponding tokenized first-party subset based on the dictionary sorting of the first-party tokens. For example, one subset might include numeric tokens starting with the number "1", another subset might include numeric tokens starting with the number "2", and so on. The same process can be performed by a second-party computation system. Assuming the tokens are generated using a hash function with approximately uniform pseudo-randomness, each subset might include approximately n / m elements.

[0095] Another example technique is hash-based allocation. The first-party computation system and the second-party computation system can each locally assign values ​​to the random hash function h: {0, 1}. * → Sample from {1, ..., m}. Note that this hash function h should be different from any hash function used to generate the tokenized set. The hash function h can take any value (e.g., a token) and returns values ​​from 1 to m containing end values. Each party can tokenize its set S = {s1, ..., sm}. n} Transform into subsets T1, ..., T m Such that for all s∈S, it considers s∈T h(s) In other words, each tokenized first (and second) subset can be associated with a numeric identifier between a predetermined number m of subsets containing end values. The allocation function may include a hash function h that produces a hash value between a predetermined number m of subsets containing end values. The first-party computation system can assign each of a plurality of first-party tokens to a tokenized first-party subset by generating hash values ​​using first-party tokens as input to the hash function h and assigning the first-party tokens to tokenized first-party subsets with numeric identifiers having hash values ​​equal to the hash values. The second-party computation system can perform a similar process. Modeling h as a random function ensures that all elements {h(s)|s∈S} are uniformly distributed. This means...

[0096] C. Subset filling

[0097] Subsequently, in step 416, the first-party computing system and the second-party computing system can populate each of their respective plurality of tokenized subsets using virtual tokens. For illustrative purposes, two virtual tokens, 428 and 430, are shown. Subset population prevents either party from determining any information about the other party's set based on the number of tokens in each subset. For example, if the first-party subset T is tokenized... iThe absence of any tokens means that the first-party set S does not contain any elements that will be assigned to the subset after tokenization. However, for the padded subset, neither party can determine the distribution of the other party's set.

[0098] The first-party computing system and the second-party computing system can fill each of their tokenized subsets with uniformly random virtual tokens. In some embodiments, the computing system can fill each tokenized subset with virtual tokens such that each tokenized subset is of equal size. In some embodiments, the computing system can fill each tokenized subset with virtual tokens such that the size of each tokenized subset is equal to (1+δ0)n / m tokens of some parameter δ0.

[0099] In some embodiments, the first-party computing system may determine a padding value for each of a plurality of tokenized first-party subsets. This padding value may include the difference between the size of the tokenized first-party subset and a target value. This target value may include, for example, the value (1+δ0)n / m described above. The padding value then includes the number of virtual tokens that can be added to the particular tokenized subset to achieve the target value. The first-party computing system may generate a plurality of random virtual tokens (e.g., using a random number generator), wherein the plurality of random virtual tokens includes a number of random virtual tokens equal to the padding value. The first-party computing system may then assign the plurality of random virtual tokens to the tokenized first-party subsets. The first-party computing system may repeat this process for each tokenized first-party subset. A second-party computing system may perform a similar procedure.

[0100] Even for relatively small tokens (e.g., length κ = 128 bits), there are a large number (2 κ The possible virtual tokens are such that the probability of any virtual token in the tokenized intersection set is negligible. Alternatively, if κ is large enough, the first-party computation system and the second-party computation system can be represented by {0, 1}. κ Virtual tokens s′ sampled from non-overlapping subsets are used to fill their j-th subsets such that h(s′)=j′≠j. This ensures that there are no virtual tokens in the tokenized intersection set.

[0101] The following calculates the value of parameter δ0 to ensure that the subset assignment step will not fail, unless the probability is negligible. For fixed i∈[n] and j∈{1,...,m}, assume X i,j An indicator variable is equal to 1 if and only if the i-th element s i In T j End in, and assume X j =∑ i∈[n] X i,j Indicates a size of T j For a fixed j, since X i,jSince the variables are independent of each other (because h is modeled as a random function), Chernof's bounding is generated. in And 0≤δ≤1 (for a single block group T) j By using union bounding, the probability that any block has more than (1+δ)μ elements is... Therefore, if the failure probability is set to The result is In other words, the above grouping technique only requires that the maximum block size is at most (1+δ0)n / m and the probability is very high. More specifically, assume the set size is n=10. 9 Given that the statistical parameter σ = 80, and choosing parameter m = 64, it can be seen that the maximum block size of any one of the 64 block groups is at most n′ ≈ 15.68 × 10⁻⁶. 6 (where δ0 = 0.0034) and the probability is (1-2) -80 ).

[0102] III.PPSI

[0103] After assigning and populating the tokenized elements to subsets, in step 418, the first-party computing system and the second-party computing system can participate in m parallel instances of the PSI protocol, where in the i-th instance π i In this approach, the first-party computing system and the second-party computing system input their respective i-th padded tokenized subsets. That is, for each of the multiple tokenized first-party subsets, the first-party computing system can use the tokenized first-party subset and the corresponding tokenized second-party subset of the second-party computing system to execute a private set intersection protocol with the second-party computing system. In this way, the first-party computing system and the second-party computing system can execute multiple private set intersection protocols and generate multiple intersecting token subsets.

[0104] First-party computing systems and second-party computing systems can use any suitable PSI protocol. One noteworthy PSI protocol is KKRT

[41] , which, at the time of writing, is one of the fastest and most efficient PSI protocols. However, any underlying PSI protocol can be used to implement embodiments, such as PSSZ15

[56] , PSWW18

[57] , etc.

[0105] In step 420, the first-party computing system and the second-party computing system can each combine multiple tokenized intersection subsets to generate a tokenized intersection set. This tokenized intersection set may include the union of multiple tokenized intersection subsets. Therefore, combining multiple intersecting token subsets may include determining the union of multiple intersecting token subsets.

[0106] Subsequently, in step 422, the first-party computing system and the second-party computing system can detoxify the tokenized intersection set to generate an intersection set. The intersection set may include elements common to both the first-party set 402 and the second-party set 404. Figure 4 In the example, this could include the set {CAMEL, BEAR, ANT, BAT}. In this way, both parties can learn the intersection of their respective sets without learning the other elements in those sets.

[0107] In optional step 424, if the intersection of sets is requested by either the orchestrator computer or the client computer communicating with the orchestrator computer, the first-party computing system (and optionally the second-party computing system) can send the intersection set to the orchestrator computer. Subsequently, in optional step 426, the orchestrator computer can send the intersection set to the client computer.

[0108] IV. Threat Modeling and Security Simulation

[0109] A. Threat Modeling

[0110] This section considers semi-honest adversaries and details their capabilities in deploying PPSI technology on computing clusters (such as Spark clusters) and big data frameworks (such as the Spark framework).

[0111] In standard cryptographic terminology, it is assumed that the underlying PSI protocol is secure against “semi-honest” (originally called honest but curious) adversaries. That is, both parties and their respective computational systems should faithfully follow the instructions of the PSI protocol. However, parties may attempt to learn as much as possible from the PSI protocol messages. This assumption applies to many common use cases where parties may have already been honestly engaged according to some protocol. Furthermore, it is assumed that all cryptographic primitives are secure. Finally, it should be noted that the PSI protocol does reveal the size of the set to both parties and the final output in plaintext (see

[29] for an example of size-hidden PSI, and [47, 57] for an example of protected output).

[0112] It is assumed that each computing cluster (e.g., a Spark cluster) has enabled built-in security features, and that no big data framework implementation (e.g., a Spark implementation) is vulnerable. These features may include data at rest encryption, access management, quota management, queue management, etc. It is also assumed that these features guarantee a secure local computing environment for each local cluster, preventing attackers from gaining access to the computing cluster unless authorized.

[0113] Furthermore, it is assumed that only authorized users can issue commands to the orchestrator. It should also be noted that the orchestrator can be operated by some (semi-honest) third party without compromising security.

[0114] In this threat model, the adversary can observe network communications between different parties during protocol execution. It can also control some of the parties' observations of the data present in their cluster's storage devices and memories, as well as the order of memory access. The semi-honest adversary model implies that participants need to provide correct input to the PSI protocol.

[0115] B. Safety Simulation

[0116] This chapter provides a security proof in the so-called "simulation paradigm," a standard in cryptography. In short, it proves that all attacks an adversary can carry in a designed protocol can be simulated in an ideal world where the parties interact only with a hypothetical trusted third party F who receives input from each party. PSI The interaction computes the intersection locally and returns only the intersection to each party. Since the above grouping techniques self-reduce PSI, they acquire the security properties of the underlying PSI protocol Π (e.g., protocols like KKRT operating on tokenized subsets). For reduction, assuming the hash function h is statistically close to a random function (or, a non-programmable random oracle), this proves that PSI self-reduction is statistically secure. Protocol Π PPSI (where the underlying PSI instance Π is instantiated using the real PSI protocol) is computationally safe. Assuming that the underlying PSI protocol Π depends on DDH, then assuming that DDH holds, the protocol ΠPPSI remains safe. For example, this is the case when the underlying PSI protocol is

[41] and OT is instantiated via DDH

[48] .

[0117] A brief summary of Π PPSI Agreement in F PSI Simulation in an ideal world. Note that the protocol operates in semi-honest mode, so the simulator can access the input tape of the corrupted party. Also, for F... PSI - Protocols in a hybrid model, wherein the protocol can enable PSI functionality F PSI The calls are subroutines. However, it should be noted that these calls will be made on a subset of the total data. For readability, this function is represented as F′. PSI .

[0118] Since the protocol is effectively symmetric, without loss of generality, the first party, P1, is assumed to be the corrupt party. The simulator begins by feeding the input S1 to the ideal PSI function F. PSI This is done to obtain the PSI output I′ = s1 ∩ s2. Next, it partitions I′ into m block groups specified by h, i.e., I... j ={i∈I′|h(i)=j} denotes block group j. If any block group I jIf there are more than (1+δ0)n / m elements, the simulator will abort. Then, for each block group j, the emulation of F′ in the hybrid world... PSI The simulator receives a filled set of size (1+δ0)n / m from P1 and returns I. j As for F′ PSI The output of the call. Finally, the simulator outputs I′. This completes the description of the simulation.

[0119] Please note that the simulation fails if (1) the simulator encounters a grouping failure (i.e., the block group size exceeds (1+δ0)n / m), or (2) a virtual item added by one side matches an item added by the other side. Therefore, based on the analysis described in the Subset Assignment subsection, it can be concluded that the ideal world simulation is statistically indistinguishable from the hybrid world protocol.

[0120] V. Parallel PDJ based on parallel PSI

[0121] This section describes how to perform SQL-style join queries using the PPSI technique described above. As mentioned above, in a Private Database Join (PDJ), both parties may wish to perform a join operation on their private data. These parties can obtain assistance from an orchestrator, and the computer system can expose metadata, such as a dataset schema that can be used to perform the join operation.

[0122] Figure 6 A flowchart illustrating an exemplary PPDJ method according to some embodiments is shown. This PPDJ method can be used to perform private database joins based on such queries using some of the grouping techniques and PPSI techniques described in Sections II and III above. Figure 6 The various steps are optional.

[0123] In step 602, the orchestrator computer may receive a request from the client computer. This request may include a Private Database Table Join Query (PDTJQ). The client computer may include a computer system associated with any party or any other suitable client (e.g., a client authorized by any party to receive the output of the PDJ). The query executing the Private Database Table Join may be submitted to the orchestrator computer using an orchestrator API such as the JupyterLab interface.

[0124] In step 604, the orchestrator computer can verify the correctness of the query. This can include verifying the query syntax and verifying that the PPSI and PPDJ systems can execute the PDJ based on the received query. Implementations can support any query that can be divided into: a "select" clause specifying one or more columns (sometimes called attributes) in two tables; a "join" clause comparing one or more columns for equality between a first set and a second set; and a "where" clause that can be broken down into combination clauses, where each join is a function of a single table. Therefore, verifying the correctness of the query can include verifying whether the query contains one or more of these supported clauses.

[0125] For example, for illustrative purposes, the embodiments may support the following queries:

[0126] SELECT P2.table0.col4,P1.table0.col3

[0127] FROM P1.table0

[0128] JOIN P2.table0

[0129] ON P1.table0.col1=P2.table0.col2

[0130] AND P1.table0.col2=P2.table0.col6

[0131] WHERE P1.table0.col3>23.

[0132] In this example, columns from both sides are selected and joined based on the equality of the join keys:

[0133] P1.table0.col1=P2.table0.col2

[0134] P1.table0.col2=P2.table0.col6

[0135] And the added constraints:

[0136] P1.table0.col3>23.

[0137] After verifying the correctness of the private database table join query, the orchestrator can reinterpret the query if necessary so that both the first-party and second-party computing systems can understand it. This reinterpretation may involve reconstructing the query as a PSI, Spark code, one or more Spark jobs, etc.

[0138] In step 606, the first-party computing system and the second-party computing system may receive a private database table join query from the orchestrator (if reinterpretation is required). The private database table join query may identify one or more first database tables and one or more second database tables (e.g., tables that can be joined), as well as one or more attributes. These attributes may correspond to columns in the identified tables on which join operations can be performed. In some embodiments, the first-party computing system and the second-party computing system may review the reinterpreted private database table join query and approve or reject the query before executing the remainder of the PDJ.

[0139] In step 608, the first-party computing system and the second-party computing system can retrieve one or more first database tables and one or more second database tables from the first-party database and the second-party database, respectively.

[0140] In optional step 610, in some embodiments, the private database table join query may include an "wherein" clause. In these embodiments, the first-party computing system and the second-party computing system may pre-filter one or more first database tables and one or more second database tables based on the "wherein" clause. For example, this might include deleting rows from database tables whose corresponding columns do not conform to the "wherein" clause. If a more complex underlying PSI protocol is used, such as a PSI protocol that can maintain the output set in a secret shared form, it is possible to implement the "wherein" clause as a function of multiple tables.

[0141] In step 612, the first-party computing system and the second-party computing system may each determine a set of join keys (alternatively referred to as multiple first-party or second-party join keys, or a first-party set and a second-party set) corresponding to a private database table join query. This set of join keys may include data entries corresponding to one or more columns in one or more first and second database tables. These columns themselves may correspond to attributes identified by the private database table join query. Therefore, the first-party computing system and the second-party computing system may determine multiple first-party join keys and multiple second-party join keys based on one or more first or second database tables and one or more attributes.

[0142] Once the input tables are filtered using the local "where" clause, both the first-party and second-party computational systems can treat the join key columns as first-party and second-party sets, and then perform the grouping techniques described in Chapter II. The join key columns can refer to the columns appearing in the "join" clause. In the example above, for the first party, these are P1.table0.col1 and P1.table0.col2, and for the second party, these are P2.table0.col2 and P2.table0.col6.

[0143] In step 614, the first-party computing system and the second-party computing system can tokenize the multiple first-party connection keys and the multiple second-party connection keys respectively, thereby generating a set of tokenized first-party connection keys (including multiple first-party tokens) and a set of tokenized second-party connection keys (including multiple second-party tokens). Step 614 can be similar to step 412, as referred to in Section II.A above. Figure 4 and 5 As described.

[0144] In some embodiments where multiple attributes exist, the first-party computing system can concatenate each first-party connection key corresponding to said attribute to generate multiple concatenated first-party connection keys. The first-party computing system can then hash the multiple concatenated first-party connection keys to generate multiple hash values, which may include a set of tokenized connection keys. Using the example private database table join query described above, the first party can generate its set of tokenized connection keys P1 as follows:

[0145] P1={H(P1.table0.col1[i],P1.table0.col2[i])|i∈{1,…,n}}

[0146] Let P2 denote a similar set of tokens from the second party. In other words, the first-party computation system can combine these concatenation key sets via concatenation before hashing, instead of hashing each concatenation key set separately (e.g., hashing P1.table0.col1[i] and P1.table0.col2[i] separately). This concatenation operation reduces the number of PPSI operations that need to be performed and thus improves performance. Note that rows with the same concatenation key will have the same tokens, so the tokenized sets P1 and P2 can contain only a single copy of the tokens.

[0147] In step 616, the first-party computing system and the second-party computing system can each generate a mapping that associates a corresponding token with an original data value (e.g., a connection key). In some embodiments, this can be achieved by appending a “token” column to one or more first-party data tables and one or more second-party database tables. The first-party computing system can generate a token column comprising a set of tokenized first-party connection keys and append it to one or more first database tables. The second-party computing system can perform a similar process. That is, for the example above:

[0148] P1.table0.token=H(P1.table0.col1[i],P1.table0.col2[i])

[0149] In step 618, the first-party computing system and the second-party computing system can allocate their respective tokenized connection key sets to tokenized first and second-party subsets. The first-party computing system can generate multiple tokenized first-party subsets by using an allocation function to allocate each of a plurality of first-party tokens to a tokenized first-party subset of multiple tokenized first-party subsets, such as the one referenced above. Figure 4 and 5 Step 414 describes an assignment function based on dictionary sorting or hashing. A second-party computing system can perform a similar process.

[0150] At step 620, if necessary, the first-party computing system and the second-party computing system may populate each of their tokenized subsets with virtual tokens, for example, as referenced in Section II.C. Figure 4 and 5 As described in step 416.

[0151] In step 622, for each of the multiple tokenized subsets, the first-party computing system and the second-party computing system can execute a private set intersection protocol, thereby executing multiple private set intersection protocols and generating multiple intersecting token subsets, for example, as referenced in Chapter III. Figure 4 and 5 As described in step 418. Both the first-party and second-party computing systems can use any suitable PSI protocol, such as KKRT, PSSZ15, PSSW18, etc.

[0152] In step 624, the first-party computing system and the second-party computing system can combine multiple intersecting subsets of tokens to generate, for example, the reference in Chapter III. Figure 4 and 5 The set of intersecting tokens described in step 420. Both the first-party and second-party computational systems can use a union operation to combine multiple subsets of intersecting tokens.

[0153] In step 626, the first-party computing system and the second-party computing system can detoxify the intersecting token set to generate, for example, the reference in Chapter III. Figure 4 and 5 The corresponding interconnection key set described in step 422. The first-party computing system and the second-party computing system can use, for example, a "token" column, which is generated and appended to the database table at step 616 above.

[0154] In step 628, the first-party computing system and the second-party computing system can filter their respective database tables using a corresponding set of connection keys, thereby generating one or more filtered first-party database tables and one or more filtered second-party database tables. This may involve, for example, removing one or more rows from one or more first database tables based on a token column, the one or more rows corresponding to one or more tokenized first-party connection keys that are not in the corresponding set of connection keys, and the same applies to one or more second database tables.

[0155] In step 630, the first-party computing system may send one or more filtered first database tables to the second-party computing system. Similarly, the second-party computing system may send one or more filtered first database tables to the first-party computing system. This sending enables both parties to construct join database tables. It is worth noting that because both tables are filtered using the same set of join keys, they do not leak any additional information.

[0156] In step 632, the first-party computing system and the second-party computing system can combine one or more filtered first database tables and one or more filtered second database tables to generate a join table. This can be achieved using a standard (e.g., non-private) join operation with a corresponding set of join keys between one or more filtered first database tables and one or more filtered second database tables.

[0157] In step 634, the first-party computing system and the second-party computing system may each send their join database tables to the orchestrator computer. Optionally, the orchestrator computer may verify that the two join database tables are identical in order to confirm that both the first-party computing system and the second-party computing system are acting semi-honestly.

[0158] In step 636, the orchestrator computing system can send the connection database table to the client computer via, for example, the orchestrator API described above.

[0159] In summary, PDJ operations can be performed using a series of stages. In one stage, the PSI and PDJ systems can reinterpret the PDJ query as a set intersection operation. In subsequent stages, the tables corresponding to the PDJ query can be retrieved by their respective parties (e.g., from a first-party database and a second-party database). Next, the parties can determine a set of join keys based on the reinterpreted PDJ query. Using the grouping techniques described above, the parties can generate tokenized subsets of join keys, the intersection of which can be determined using PPSI techniques. The intersection can then be used to perform a "reverse token lookup," enabling the parties to filter their respective data tables. The parties can send their filtered data tables to each other and then use the two filtered data tables to construct the join table.

[0160] VI.SPARK-PSI

[0161] This section describes a specific implementation of embodiments of this disclosure using existing open-source Apache software, including Apache Spark. This implementation is referred to as "SPARK-PSI" and uses a C++ implementation of the KKRT protocol as the underlying PSI protocol used in the PPSI technology described above. This section demonstrates how to implement embodiments of this disclosure in practice using industry-standard software.

[0162] A. Apache Spark

[0163] Apache Spark is an open-source distributed computing framework for large-scale data workloads. It utilizes in-memory caching and optimizes query execution for data of any size. On top of Spark, there are libraries for running distributed computing, such as SQL queries, machine learning algorithms, graph analysis, and data streaming. Spark applications consist of “drivers” (operated by “drivers” or “master” nodes) that transform user-provided data processing pipelines into individual tasks and distribute these tasks to “worker nodes”. The basic abstractions available in Spark are built on a distributed data structure called “Resilient Distributed Datasets (RDDs)”

[73] , and these abstractions provide distributed data processing operators such as mapping, filtering, reduction, broadcasting, etc. Higher-level abstractions reveal popular APIs such as SQL, streaming, and graph processing.

[0164] Due to the capabilities demonstrated by Spark and similar data platforms, implementing PSI using Apache Spark can potentially achieve performance gains. However, Spark lacks any multi-tenancy concept, and all applications and tasks run within a single security domain. This is incompatible with the fundamental setup of the PSI protocol, which involves two or more untrusted parties that require multiple security domains and strong isolation between them. The implementation addresses this issue by assigning each party to a Spark cluster, thereby achieving isolation through physical separation of each party's computation. Furthermore, the implementation introduces an orchestrator computer that coordinates multiple independent Spark clusters in different data centers to jointly execute PSI tasks.

[0165] The second security issue with Apache (addressed by Spark-PSI) is the default data partitioning scheme, which can reveal information about the datasets of various parties. For example, if data is partitioned to worker nodes based on the first byte of each data element, a malicious user could know how many data elements begin with a specific byte (e.g., 0x00, 0x01, etc.). This could leak information about the distribution of data within the dataset and compromise the security associated with the PSI protocol. This problem is addressed using the secure grouping technique described in Chapter II.

[0166] Another potential issue (addressed by Spark-PSI) is that adding an orchestrator outside the Spark cluster can result in suboptimal execution plans. Specifically, local scheduling optimizations for each cluster can degrade the performance of collaborative computations across multiple clusters with varying data sizes and hardware configurations. However, Spark-PSI can leverage Spark's lazy evaluation capability, which can be used to defer task execution until an action is triggered. In this way, wireless evaluation can be used to efficiently coordinate operations across clusters.

[0167] B.SPARK-PSI System

[0168] Figure 7 The overall system architecture of the SPARK-PSI system is illustrated. The first and second parties can use the SPARK-PSI system to determine the PSI of their respective sets (e.g., two private datasets). Alternatively or additionally, the first and second parties can use the SPARK-PSI system to perform PDJ operations. In this way, the first and second parties can generate join tables.

[0169] See Chapter I for reference Figure 2 As described, each party may possess its own corresponding domain (first party domain 706 and second party domain 708), which contain the party's data (e.g., collections, database tables, etc.) and computing resources. These may include a first party (edge) server 710, a first party database 722, and a first party Spark cluster 726, as well as corresponding second party (edge) servers 712, second party databases 724, and second party Spark clusters 728.

[0170] The orchestrator computer 704 can coordinate the computing resources of the first domain 706 and the second domain 708 so that both parties can determine the PSI or complete the PDJ. The orchestrator 704 may expose an interface (e.g., a UI application, portal, Jupyter Lab interface, etc.) that enables the client computer 702 to send private database join queries and receive the results of the queries (e.g., join database tables). The orchestrator 704 may connect to the first-party server 710 and the second-party server 712 via its corresponding Apache Livy

[45] cluster interfaces 714 and 720. Although the orchestrator computer 704 is shown outside the first domain 706 and the second domain 708, in practice the orchestrator 704 may be included in either of these domains.

[0171] As mentioned above, refer to Figure 2 The orchestrator 704 can store and manage various metadata, including schemas of any dataset stored by the first and second parties (e.g., stored in first-party database 722 and second-party database 724). The orchestrator computer 704 can acquire this metadata and schemas during an initialization phase performed between the orchestrator 704, the first-party server 710, and the second-party server 712. During this initialization phase, the first-party server 710 and the second-party server 712 can send their respective metadata and schemas to the orchestrator 704.

[0172] During PDJ operation, in step 754, the client computer 702 may first authenticate itself with the orchestrator 704 using any suitable authentication technology. The client computer may include a computer system associated with either party (e.g., the first party) or any other suitable client. After authentication, in step 756, the client computer 702 may send a join request or PDJ query (e.g., an SQL-style query) to the orchestrator 704.

[0173] Orchestrator 704 can parse PDJ queries and compile Apache Spark jobs for the first-party Spark cluster 726 and the second-party Spark cluster 728. These Spark jobs can correspond to the operations or steps to be performed by each cluster during the PDJ operation, including steps associated with the PPSI technology described above. The orchestrator can then send these Spark jobs, along with other relevant information such as dataset identifiers, join columns, network configurations, etc., to the first-party server 710 and the second-party server 712 via its Apache Livy interfaces 714 and 720.

[0174] Using Spark jobs and other relevant information, first-party server 710 and second-party server 712 can retrieve any relevant database tables from first-party database 722 and second-party database 724. From these database tables, first-party server 710 and second-party server 712 can extract any relevant datasets (e.g., first-party sets and second-party sets) on which PSI operations can be performed.

[0175] In step 758, the first-party server 710 and the second-party server 712 may perform a grouping technique (described in Section II above) on the first-party set and the second-party set. As described above, this may include first tokenizing these datasets to produce tokenized first and second-party datasets. Then, the first-party server 710 and the second-party server 712 may assign the tokenized elements to subsets (e.g., using a hash-based assignment function) to generate multiple first-party token subsets and multiple second-party token subsets. Subsequently, the first-party server 710 and the second-party server 712 may populate the token subsets with dummy values.

[0176] Subsequently, in step 760, the first-party server 710 and the second-party server 712 can initiate PSI execution and send a subset of tokens and any associated Spark code to the first-party Spark cluster and the second-party Spark cluster, respectively. The first-party server 710 and the second-party server 712 can use their respective Apache Livy

[45] interfaces 714 and 720 to internally manage Spark sessions and submit Spark code for determining the intersection of private sets. Spark drivers 730 and 732 can interpret this Spark code and assign Spark jobs or tasks related to PSI to worker nodes 738-744. The worker nodes 738-744 can then execute these tasks.

[0177] Additionally, first-party server 710 and second-party server 712 can act as "Kafka brokers" using their respective Apache Kafka frameworks 716 and 718, thereby establishing a secure data transmission channel between the first-party Spark cluster 726 and the second-party Spark cluster 728. These may include one or more "byte exchanges" 762 for performing specific steps of the underlying KKRT PSI protocol. While Apache Kafka has been chosen for implementing the communication pipeline in Spark-PSI, this architecture allows the parties to use any other suitable communication framework to read, write, and send data.

[0178] Advantages of the C.SPARK-PSI implementation architecture

[0179] Several advantages exist associated with the aforementioned Spark-PSI implementation. One advantage is that Spark-PSI requires no internal changes to Apache Spark, thus facilitating large-scale adoption and deployment. Other advantages relate to data security. While PDJ security is ensured through the use of the secure PSI protocol, the Spark-PSI architecture provides several additional security features. More specifically, in addition to Apache Spark's built-in security features, the Spark-PSI design ensures cluster isolation and session isolation, as described below.

[0180] Orchestrator 704 provides a protected virtual computing environment for each PSI or PDJ job, ensuring session isolation. While standard TLS can be used to protect communication between the first-party domain 706 and the second-party domain 708, orchestrator 704 provides additional communication protection, such as session-specific encryption and authentication keys, random and anonymous endpoints, managed allow and deny lists, and monitoring and / or prevention of DOS / DDOS attacks against the first-party server 710 and the second-party server 712. As mentioned above, the orchestrator also provides an additional layer of user authentication and authorization. All computing resources, including tasks, cached data, communication channels, and metadata, can be protected within the session. External users are prevented from viewing or altering the internal state of the session. The first-party Spark cluster 726 and the second-party Spark cluster 728 can be isolated from each other and can only report execution status to orchestrator 704 via the first-party server 710 and the second-party server 712.

[0181] Cluster isolation aims to protect the computing resources of each party from abuse during PSI or PDJ operations. To achieve this, orchestrator 704 may include a unique node in the Spark-PSI system that can access the end-to-end processing stream. Orchestrator 704 may also include a unique node in the Spark-PSI system with metadata corresponding to the first-party Spark cluster 726 and the second-party Spark cluster 728. Orchestrator 704 may reside outside the first-party domain 706 and the second-party domain 708 to allow orchestrator 704 to access the data flow pipeline between the first-party cluster 726 and the second-party cluster 728. However, even if orchestrator 704 is included in one party's domain, a separate secure communication channel is used between the first-party cluster 726 and the second-party cluster 728 via Apache Livy and Kafka, which prevents the parties from accessing the other Spark cluster, thus removing orchestrator 704 from the data flow pipeline. This secure communication channel also ensures that each Spark cluster is autonomous and requires little or no modification to participate in database connection protocols with other parties. The orchestrator 404 can also manage connection failures and uneven compute speeds to ensure out-of-the-box reusability of both first-party Spark clusters 726 and second-party Spark clusters 728.

[0182] Furthermore, the low-level APIs that call cryptographic libraries and exchange data between C++ instances and Spark DataFrames (e.g., Scala PSI libraries 734 and 736 and PSI worker libraries 746-752) reside in the first-party Spark cluster 726 and the second-party Spark cluster 728, thus introducing no information leakage. The high-level APIs can package secure Spark execution pipelines as services and can map individual jobs to each worker node 738-744 and collect results from the worker nodes.

[0183] In summary, the SPARK-PSI architecture provides theoretical security associated with underlying PSI protocols (e.g., KKRT). In other words, if one party is compromised by a hacker or other malicious user, the other party's data remains confidential, except for what is displayed in the output of PSI or PDJ operations.

[0184] D.Spark PSI Implementation Workflow

[0185] Figure 8 This illustrates the detailed data workflow within the SPARK-PSI framework instantiated using the KKRT protocol. Phases in each PSI instance (802 and 804) can be sequentially invoked by the orchestrator computer. The orchestrator can initiate KKRT execution by submitting metadata information about the first and second party sets to both parties.

[0186] In setup phase 806 (including steps 832-838), based on a request, the first-party computing system (including first-party server 820 and the Spark cluster including worker nodes 816) and the second-party computing system (including second-party server 822 and the Spark cluster including worker nodes 818) can begin executing their respective Spark code. This code can create new dataframes by loading the first-party and second-party collections using a supported Java Database Connectivity (JDBC) driver. As described above in Reference Section II, these dataframes can then be hashed to generate token dataframes. The token dataframes can then be mapped to m groups or subsets of token blocks. Using Apache Spark terminology, these groups of blocks may be referred to as “partitions”. Figure 8 The image shows four such block groups, the first-party token block group. i 824, First-party token block group m 828, Second-party token block group i 826 and second-party token block group m 830. These block groups can be populated with virtual tokens if needed. The token block groups can be distributed across worker nodes 816 and 818, enabling KKRT to be executed in parallel.

[0187] After the final setup is determined, PSI instances 802 and 804 can proceed to PSI phase 810. In this phase, the local KKRT protocol can be executed via a generic Java Native Interface (JNI) connected to Spark code. The JNI can operate based on round functions, thus operating regardless of the specific implementation of the KKRT PSI protocol. Note that the KKRT protocol has a one-time setup phase, which only needs to be performed once for a given pair of participants. This setup phase corresponds to steps 832-838. For more details on the setup phase, see

[41] . The online PSI phase (which determines the intersection between token block groups) corresponds to steps 840-848. Whenever a write operation is performed on either Kafka broker, both parties can mirror the data using first-party server 820 and second-party server 822.

[0188] Note that the main PSI phase involves sending the encrypted token dataset, which can be a performance bottleneck for Apache Kafka as it is optimized for small messages. To overcome this, worker nodes 816 and 818 can split the encrypted dataset into smaller chunks before sending it to the other party via first-party server 820 and second-party server 822. When receiving data from the other party via first-party server 820 and second-party server 822, worker nodes 816 and 818 can merge the chunks, replicate the encrypted token dataset, and allow them to execute the KKRT PSI protocol. Additionally, the intermediate data retention period of the Kafka broker can be shortened to address storage and security concerns.

[0189] Blocking data offers the added benefit of enabling streaming of the underlying PSI protocol messages. Note that the local KKRT implementation is designed to send and receive data immediately after its generation. Therefore, the SPARK-PSI implementation can continuously forward protocol messages to and from Kafka as they become available. This effectively results in additional parallelization since worker nodes 816 and 818 do not require blocking slow network I / O. Note that this implementation can cache token dataframes and instance address dataframes used across multiple stages to avoid any recomputation. In this way, the SPARK-PSI implementation can leverage Spark's lazy evaluation, which optimizes execution based on the persistence of Directed Acyclic Graphs (DAGs) and Resilient Distributed Datasets (RDDs).

[0190] E. Reusable components

[0191] The Spark-PSI implementation has several reusable components to parallelize PSI protocols other than KKRT. The code corresponding to the Spark-PSI implementation can be packaged into a Spark-Scala library, which includes an end-to-end example implementation of the native KKRT protocol. This library itself has several reusable components, such as a JDBC connector for use with multiple data sources, methods for tokenization and subset allocation, a generic C++ interface for linking other native PSI algorithms, and a generic JNI between Scala and C++. Each of these functionalities can be implemented in the library's base category, which can be reused for other native PSI implementations. Additionally, the library can decouple network methods from the actual PSI determination. This increases the framework's flexibility, allowing the use of other network channels when needed.

[0192] Most PSI protocols can be "plugged into" SPARK-PSI by exposing a C / C++ API that the framework can call. The API is built around the concepts of setup wheels and online wheels, therefore making no assumptions about the cryptographic protocols executed in these wheels. The API may include the following functions:

[0193] ●Get-setup-round-count()->count: Retrieves the total number of setup rounds required for this PSI implementation.

[0194] ●Setup(id,in-data)->out-data: Uses the data received from the other party in the previous setup round to call the appropriate party's round id and returns the data to be sent.

[0195] ●Get-online-round-count()->count: Retrieves the total number of online rounds required for this PSI implementation.

[0196] ●Psi-round(round id,in-data)->out-data: Uses the data received from the other party in the previous round of the PSI protocol to call the appropriate party's online round id and returns the data to be sent.

[0197] The data passed to the psi-round call can include data from a single tokenized subset, and SPARK-PSI can coordinate parallel calls to this API across all block groups. As an example, for a KKRT implementation, there are three setup rounds (labeled P1.setup1, P2.setup1, and P1.setup2) and three online rounds (labeled P1.psi1, P2.psi1, and P1.psi2). When running KKRT with 256 block groups, the setup rounds P1.setup1, P2.setup1, and P1.setup2 can each call setup once with the appropriate round ID, while the online rounds P1.psi1, P2.psi1, and P1.psi2 can each call psi-round256 times with the appropriate round ID.

[0198] VII. Experimental Evaluation

[0199] This section describes the results of PSI experiments performed using the SPARK-PSI system. Additionally, this section provides benchmarks for different steps in the PSI protocol (e.g., tokenization, setup rounds, etc.). This section also provides end-to-end performance results and details the impact of block size on runtime. Notably, when performing PSI on a set containing billions of elements, a runtime of 82.88 minutes was achieved. This result was obtained using 2048 block groups, corresponding to a value δ0 = 0.019 for a block size of approximately 500,000.

[0200] In similar Figure 7 and 8 These experiments were evaluated on the Spark-PSI setup described in [the document]. In the experiments, both the first and second parties used the Spark-PSI system to execute the KKRT-based PPSI protocol. Each party ran a separate, independent six-node Spark (v2.4.5) cluster with one driver server and five worker servers. Additionally, each party also ran a separate Kafka (v2.12-2.5.0) VM (acting as an edge server) for inter-cluster communication. The orchestrator server that triggered the PSI computation resided in the first party's domain. All servers had 8 vCPUs (2.6 GHz), 64 GB of RAM, and ran Ubuntu 18.04.4 LTS.

[0201] Table 1 summarizes the time required to perform various steps using the KKRT-based PPSI method with 2048 block groups (tokenized subsets) at different dataset sizes (i.e., 10 million, 50 million, and 100 million elements). P1.tokenize represents the time spent performing the block grouping technique, which involves tokenizing the first-party set, mapping these tokens to different tokenized subsets, and populating each tokenized subset. The tokenization step is performed in parallel by worker nodes.

[0202] P1.psi1 represents the amount of time spent sending the PSI byte set corresponding to the first party (i.e., in...). Figure 8 Step 540 in the table. In this step, the first-party computing system generates and transmits approximately 60n bytes of data (where n is the number of elements in the dataset) to the second-party computing system via a first-party (edge) server. Similarly, P2.psi1 represents the amount of time spent for each tokenized subset to receive the corresponding set of PSI bytes from the second party (e.g., in...). Figure 8(Step 546 in the original text). In this step, the second-party computing system generates data and transmits approximately 22n bytes of data back to the first-party computing system via a second-party (edge) server. P1.psi2 represents the amount of time required to receive the set of PSI bytes corresponding to the tokenized intersection subset, combine the tokenized intersection subset, and detoxify the tokenized intersection set to generate the intersection set.

[0203]

[0204] Table 1: Micro-benchmarking using KKRT PSI and SPARK-PSI with 2048 blocks.

[0205] Table 2 illustrates the impact of block size on the time spent performing inter-cluster communication, which includes reading and writing data via a data stream processor (e.g., Apache Kafka). The P1.psi1 step produces 9.1 GB of intermediate data, which is sent to the second-party computing system via a first-party (edge) server. The P2.psi1 step produces 3.03 GB of intermediate data, which is sent to the first-party computing system via a second-party (edge) server. The benchmarks in Table 2 clearly show that using more block groups improves network performance as message blocks become smaller. More specifically, when using 256 block groups, a single message of 35.55 MB is sent via the data stream processor during the P1.psi1 step. When using 2048 block groups, the corresponding single message size is only 4.44 MB.

[0206]

[0207] Table 2: Network latency for a dataset containing 100 million elements.

[0208] Table 3 compares the performance of Spark-PSI with that of an insecure join on a dataset containing 100 million elements. To evaluate and compare the performance of Spark-PSI and the insecure join, two variants of the insecure join were considered. In the first variant, called a “single-cluster Spark join,” a join is performed on two datasets, each containing 100 million elements, using a single compute cluster with six nodes (one driver node and five worker nodes). The join computation is performed by partitioning the data into multiple chunks and directly determining the intersection using a single Spark join call.

[0209] In the second variant, known as "Cross-Cluster Spark Join," two compute clusters are used, each consisting of six nodes (one driver node and five worker nodes), and each cluster contains a tokenized dataset of 100 million elements. To perform the join, each cluster partitions its dataset into multiple chunk groups. One cluster then sends the partitioned dataset to the other cluster, which then aggregates the received data into a single dataset and computes the final join using a single Spark Join call.

[0210] For insecure single-cluster connections, increasing the number of block groups increases the number of data shuffling operations (e.g., random read / write operations), which slows down execution. When insecure connections are split between two clusters, there is additional network communication overhead and additional scheduling operations on the target cluster, but parallelism increases because the computational resources of the two cluster systems are doubled.

[0211] When using SPARK-PSI, cross-cluster communication overhead is maintained, and PSI computation incurs additional overhead, but additional data shuffling is avoided (because the system uses broadcast joins). When the system uses a large number of block groups (e.g., 8,192 block groups), the impact of broadcast joins increases, making SPARK-PSI faster than insecure cross-cluster joins in some cases. Compared to insecure cross-cluster joins, the system introduces up to 77% overhead in the worst case.

[0212]

[0213] Table 3: Total execution time for different joins on a dataset with over 100 million elements. The fastest time in each column is shown in bold.

[0214] Table 4 details the relationship between runtime associated with SPARK-PSI and the number of block groups and dataset size. Runtime is also plotted on... Figure 9 In the middle. A noteworthy result is that for a dataset containing one billion elements and using 2048 block groups, the running time is 82.88 minutes, roughly 25 times faster than the previous study by Pinkas et al.

[60] . Also as shown in Table 4 and Figure 9 As indicated by the corresponding curve, Spark-PSI performance improves with increasing chunk size, then reaches an inflection point and declines. The initial improvement is a result of parallelization. A higher number of chunks results in smaller chunk sizes on Spark, which is preferable for larger datasets. However, as the number of chunks increases further, the task scheduling overhead in Spark (and the padding overhead of various grouping techniques) slows down execution. Using more worker nodes could potentially achieve even better performance, as this might allow for better parallelization.

[0215]

[0216]

[0217] Table 4: Total execution time of SPARK-PSI with different dataset sizes and block sizes. The fastest time in each column is shown in bold.

[0218] VIII. Relevant Research and Conclusions

[0219] Several protocols have been proposed to implement PSI, such as efficient but insecure naive hashing solutions, public-key cryptography-based protocols [4,12,18,25,26,29,37,46,64], protocols based on unintended transmissions [11,23,41,54,55,58], and other circuit-based solutions [7,36,56,57]. Another popular approach to PSI is to introduce a semi-trusted third party to help efficiently compute intersections [1,2,67]. For a more detailed overview of the various approaches to solving PSI, see

[59] . In addition, other variants of PSI have been extensively studied, such as multi-party PSI [35,42], PSI cardinality [13,39], PSI summation [38,39], threshold PSI [5,27], etc. Besides PSI, there is a range of work on performing other set operations, such as private unions [8,17,40,43].

[0220] Since the introduction of the MapReduce programming model

[20] , modern big data systems have demonstrated high scalability and high performance. This has brought both opportunities and challenges to large-scale datasets and secure distributed computing in cloud computing.

[0221] Dong et al.

[23] introduced a cryptographic Boolean filter to design an efficient PSI protocol for big data, implemented using the MapReduce framework. PSJoin

[22] utilizes differential privacy rights to build a MapReduce-based privacy-preserving similarity join. Hahn et al.

[30] designed a protocol for secure joins using searchable encryption and key policy attribute-based encryption, which leaks fine-grained access patterns and frequencies of elements selected for joins.

[0222] SMCQL[6] uses the backend ObliVM

[44] , based on cryptographic Boolean circuits, to compute query results from the union of several source databases without revealing sensitive information about individual tuples. Despite optimization, it introduces a daunting overhead. ConClave

[69] builds a secure query compiler based on ShareMind[9] and Obliv-C

[75] to improve scalability. ConClave works in a server-assisted model to reduce computational overhead. However, these systems still fall short in performing efficient and secure computation on big data. Furthermore, existing research is tailored to specific requirements and therefore cannot provide the same performance gains for arbitrary secure computation.

[0223] Another set of privacy-preserving frameworks utilizes hardware enclaves. Opaque

[76] is an inadvertent distributed data analytics platform that leverages Intel SGX hardware enclaves for robust security. OCQ

[16] further reduces Opaque’s communication and computation costs via an inadvertent planner. Unlike these approaches, SPARK-PSI is hardware-independent. Other recent research includes CryptDB

[61] and Seabed

[52] , which provide protocols for securely performing analytical queries on encrypted big data. Senate

[66] describes a framework for enabling privacy-preserving database queries in a multi-party setting.

[0224] In summary, this disclosure describes the analysis and application of methods that can be used to parallelize any PSI protocol, thereby significantly increasing the rate at which PSI can be determined. Using the methods according to embodiments, this disclosure demonstrates that the private set intersection of large (e.g., billion-element) datasets can be determined at a significantly greater speed. Furthermore, this disclosure describes a Spark framework and architecture for implementing these methods in PDJ applications. Experiments show that this framework is well-suited for real-world scenarios. Additionally, this framework provides reusable components that enable cryptographers to extend new PSI protocols to billion-element sets.

[0225] IX. Computer Systems

[0226] Any computer system mentioned in this article can use any suitable number of subsystems. Figure 10 Examples of such subsystems in computer system 1000 are shown. In some embodiments, the computer system includes a single computer device, wherein the subsystem may be a component of the computer device. In other embodiments, the computer system may include multiple computer devices, each of which is a subsystem having internal components. The computer system may include desktop and laptop computers, tablet computers, mobile phones, and other mobile devices.

[0227] Figure 10The subsystems shown are interconnected via system bus 1012. Additional subsystems are shown, such as printer 1008, keyboard 1018, storage device 1020, and a monitor 1024 coupled to display adapter 1014 (e.g., a display screen, such as an LED), etc. Peripheral devices and I / O devices coupled to input / output (I / O) controller 1002 can be connected to the computer system via various components known in the art, such as input / output (I / O) port 1016 (e.g., USB, etc.). For example, I / O port 1016 or external interface 1022 (e.g., Ethernet, Wi-Fi, etc.) can be used to connect computer system 1000 to a wide area network such as the Internet, a mouse input device, or a scanner. Interconnection via system bus 1012 allows central processing unit 1006 to communicate with each subsystem and control the execution of multiple instructions from system memory 1004 or storage device 1020 (e.g., a fixed disk, such as a hard disk drive or optical disk), as well as the exchange of information between subsystems. System memory 1004 and / or storage device 1020 may embody a computer-readable medium. Another subsystem is a data collection device 1010, such as a camera, microphone, accelerometer, etc. Any data mentioned herein can be output from one component to another and can be output to the user.

[0228] A computer system may include multiple identical components or subsystems connected together, for example, via an external interface 1022, an internal interface, or via a removable storage device that can be attached to and removed from one component to another. In some embodiments, the computer system, subsystem, or device may communicate via a network. In such cases, one computer may be considered a client, and another computer may be considered a server, where each computer may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.

[0229] Any computer system mentioned herein may use any suitable number of subsystems. In some embodiments, the computer system includes a single computer device, wherein the subsystem may be a component of the computer device. In other embodiments, the computer system may include multiple computer devices, each of which is a subsystem having internal components.

[0230] A computer system may include multiple components or subsystems connected together, for example, by external interfaces or internal interfaces. In some embodiments, the computer system, subsystem, or device may communicate via a network. In such cases, one computer may be considered a client, and another computer may be considered a server, where each computer may be part of the same computer system. The client and server may each include multiple systems, subsystems, or components.

[0231] It should be understood that any embodiment of the present invention can be implemented using hardware (e.g., application-specific integrated circuits or field-programmable gate arrays) and / or computer software in the form of control logic, wherein the general-purpose programmable processor is modular or integrated. As used herein, the processor includes a single-core processor, a multi-core processor on the same integrated chip, or multiple processing units on a single circuit board or networked thereon. Based on this disclosure and the teachings provided herein, those skilled in the art will know and understand other ways and / or methods of implementing embodiments of the present invention using hardware and combinations of hardware and software.

[0232] Any software component or function described in this application may be implemented as software code executed by a processor using any suitable computer language such as Java, C, C++, C#, Objective-C, Swift, or a scripting language such as Perl or Python, employing techniques such as conventional or object-oriented methods. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable media include random access memory (RAM), read-only memory (ROM), magnetic media (e.g., hard disk drives or floppy disks), or optical media (e.g., optical discs (CDs) or DVDs (Digital Universal Discs)), flash memory, etc. The computer-readable medium may be any combination of such storage or transmission means.

[0233] Such programs can also be encoded and transmitted using carrier signals suitable for transmission over wired, optical, and / or wireless networks conforming to various protocols, including the Internet. Therefore, computer-readable media according to embodiments of the invention can be created using data signals encoded with such programs. Computer-readable media encoded with program code can be packaged with a compatible device or provided separately from other devices (e.g., downloaded via the Internet). Any such computer-readable medium can reside on or within a single computer product (e.g., a hard disk drive, CD, or an entire computer system) and can exist on or within different computer products within a system or network. The computer system may include a monitor, printer, or other suitable display for providing any results mentioned herein to a user.

[0234] Any method described herein can be performed, wholly or partially, by a computer system including one or more processors that can be configured to perform these steps. Therefore, embodiments may relate to a computer system configured to perform the steps of any method described herein, possibly having different components that perform corresponding steps or groups of corresponding steps. Although presented as numbered steps, the steps of the methods herein may be performed simultaneously or in different orders. Furthermore, portions of these steps may be used in conjunction with portions of other steps from other methods. Similarly, all or part of a step may be optional. Additionally, any step of any method may be performed using modules, circuitry, or other means for performing these steps.

[0235] The specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of the embodiments of the invention. However, other embodiments of the invention may relate to specific embodiments associated with each individual aspect, or specific combinations of these individual aspects. The foregoing description of exemplary embodiments of the invention has been presented for purposes of illustration and description. It is not intended to be exhaustive, or to limit the invention to the precise forms described; many modifications and variations are possible in accordance with the teachings above. These embodiments were chosen and described in order to best explain the principles of the invention and its practical application, thereby enabling those skilled in the art to best utilize the invention in various embodiments and to make various modifications suitable for the particular intended use.

[0236] The above description is illustrative and not restrictive. Many variations of the invention will become apparent to those skilled in the art upon reading this disclosure. Therefore, the scope of the invention should not be determined by reference to the above description, but rather by reference to the pending claims and their full scope or equivalents.

[0237] Without departing from the scope of the invention, one or more features from any embodiment may be combined with one or more features from any other embodiment.

[0238] Unless explicitly indicated otherwise, the use of “one” or “the” is intended to mean “one or more”. The use of “or” is intended to mean “inclusive or” rather than “exclusive or” unless explicitly indicated otherwise.

[0239] All patents, patent applications, publications, and descriptions mentioned herein are incorporated herein by reference in their entirety for all purposes. This does not imply an admission that they are prior art.

[0240] X. References

[0241] [1] Aydin Abadi, Sotirios Terzis, and Changyu Dong, 2015, O-PSI: Delegated Private Set Intersection on Outsourced Datasets, in Proceedings of the 11th International Conference of the 30th International Conference on Information and Communication Technology, U.S. Securities and Exchange Commission, May 26-28, 2015, Hamburg, Germany, ICT Advances, Vol. 455, in Hannes Federath and Dieter Gollmann (eds.), Springer, pp. 3-17. https: / / doi.org / 10.1007 / 978-3-319-18467-8_1

[0242] [2] Aydin Abadi, Sotirios Terzis and Changyu Dong, 2016, VD-PSI: Verifiable Delegated Private Set Intersection on Outsourced Private Datasets. In the 20th International Conference on Financial Cryptography and Data Security, FC2016, Christ Church, Barbados, February 22-26, 2016, Jens Grossklages and Bart Preneel (eds.), Revised Proceedings (Lectures in Computer Science, Vol. 9603), Springer, pp. 149-168.

[0243] https: / / doi.org / 10.1007 / 978-3-662-54970-4_9

[0244] [3] Thomas Schneider N. Asokan Benny Pinkas Agnes Kiss, Jian Liu, 2017, Private Set Intersection for Uniqual Set Sizes with Mobile Applications. In the Proceedings of the Privacy Enhancement Technology Papers (4), pp. 177-197.

[0245] [4] Giuseppe Ateniese, Emiliano De Cristofaro, and Gene Tsudik, 2011, (If) Size Matters: Size-Hiding Private Set Intersection. In the proceedings of the 14th International Conference on Public-Key Cryptography Practice and Theory, PKC 2011, Taormina, Italy, March 6-9, 2011 (Lectures in Computer Science, Vol. 6571), Dario Catalano, Dario Catalano, Nelly Fazio, Rosario Gennaro, and Antonio Nicolosi (eds.), Springer, pp. 156-173. https: / / doi.org / 10.1007 / 978-3-642-19379-8_10

[0246] [5] Saikrishna Badrinarayanan, Peihan Miao, and Peter Rindal, 2020, Multi-Party Threshold Private Set Intersection with Sublinear Communication (IACR) Cryptography, ePrint Arch. 2020, p. 600. https: / / eprint.iacr.org / 2020 / 600

[0247] [6]Johes Bater, Gregory Elliott, Craig Eggen, Satyender Goel, Abel Kho and Jennie Rogers, SMCQL: secure querying for federated databases, VLDB Donated Papers 10, 6 (2017), pp. 673-684.

[0248] [7] Oleksandr Tkachenko Avishay Yanai Benny Pinkas, Thomas Schneider, 2019, Efficient Circuit-Based PSI with Linear Communication, European Conference on Cryptography 3, pp. 122-153.

[0249] [8] Marina Blanton and Everaldo Aguiar, 2016, Private and oblivious set and multiset operations, International Journal of Information Security, Chapter 15, 5 (2016), pp. 493-518. https: / / doi.org / 10.1007 / s10207-015-0301-1

[0250] [9] Dan Bogdanov, Sven Laur and Jan Willemson, 2008, Sharemind: A framework for fast privacy-preserving computations, Springer, pp. 192-206, at the European Symposium on Computer Security Research.

[0251]

[10] Justin Brickell, Donald E. Porter, Vitaly Shmatikov, and Emmett Witchel, 2007, Privacy-preserving remote diagnostics. In CCS,

[0252]

[11] Melissa Chase and Peihan Miao. Private Set Intersection in the Internet Setting from Lightweight Oblivious PRF, 2020, Proceedings of the 40th International Cryptocurrency Conference, CRYPTO 2020, Santa Barbara, CA, August 17-21, 2020, Part III (Lectures in Computer Science, Vol. 12172), Daniele Micciancio and Thomas Ristenpart (eds). Springer, pp. 34-63. https: / / doi.org / 10.1007 / 978-3-030-56877-1_2

[0253]

[12] Hao Chen, Kim Laine, and Peter Rindal, Fast Private Set Intersection from Homomorphic Encryption, 2017, in Proceedings of ACMSIGSAC Conference on Computer and Communications Security, CCS 2017, Dallas, TD, USA, October 30–November 3, 2017, Bhavani M. Thuraisingham, David Evans, Tal Malkin, and Dongyan Xu (eds). ACM, pp. 1243–1255. https: / / doi.org / 10.1145 / 3133956.3134061

[0254]

[13] Emilano De Cristofaro, Paolo Gasti and Gene Tsudik, Fast and Private Computation of Cardinality of Set Intersection and Union, 2012, in the proceedings of the 11th International Conference on Cryptography and Network Security, CANS 2012, Darmstadt, Germany, 12-14 December 2012, Josef Pieprzyk, Ahmad-Reza Sadeghi and Mark Manulis (eds.), Vol. 7712, Springer, pp. 218-231. https: / / doi.org / 10.1007 / 978-3-642-35404-5_17

[0255]

[14] Emilano De Cristofaro, Jihye Kim and Gene Tsudik, 2010, Linear-Complexity Private Set Intersection Protocols Secure in Malicious Model, in Proceedings of Cryptography - ASIACRYPT 2010 - The 16th International Conference on Cryptography and Information Security Theory and Applications, Singapore, December 5-9, 2010. Proceedings (Lectures in Computer Science, Vol. 6477), Masayuki Abe (ed.). Springer, pp. 213-231.

[0256] https: / / doi.org / 10.1007 / 978-3-642-17373-8_13

[0257]

[15] Thomas Schneider, Matthias Senker, Christian Weinert, Daniel Kales, and Christian Rechberger, 2019, Mobile Private Contact Discovery at Scale, USENIX Annual Technical Conference, pp. 1447-1464.

[0258]

[16] Ankur Dave, Chester Leung, Raluca Ada Popa, Joseph E Gonzalez and Ion Stoica, 2020, Oblivious coopetitive analytics using hardware enclaves, in the Proceedings of the 15th European Conference on Computer Systems, pp. 1-17.

[0259]

[17] Alex Davidson and Carlos Cid, 2017, An Efficient Toolkit for Computing Private Set Operations, in the proceedings of the 22nd Australasia Conference on Information Security and Privacy, ACISP 2017, Auckland, New Zealand, 3-5 July 2017, Part II (Computer Science Lectures, Vol. 10343), Josef Piepzyk and Suriadi (eds.). Springer, pp. 261-278. https: / / doi.org / 10.1007 / 978-3-319-59870-3_15

[0260]

[18] Emilano De Cristofaro and Gene Tsudik, 2010, Practical private set intersection protocols with linear complexity in FC.

[0261]

[19] Jeffrey Dean and Sanjay Ghemawat, 2004, MapReduce: Simplified Data Processing on Large Clusters, in OSDI'04: The Sixth Symposium on Operating System Design and Implementation, San Francisco, California, pp. 137-150.

[0262]

[20] Jeffrey Dean and Sanjay Ghemawat, MapReduce: simplified data processing on large clusters, ACM Communications 51, 1 (2008), pp. 107-113.

[0263]

[21] Daniel Demmler, Peter Rindal, Mike Rosulek, and Ni Trieu, 2018, PIR-PSI: Scaling Private Contact Discovery, Proceedings of the Privacy Enhancement Techniques 2018, 4 (2018), pp. 159–178. https: / / doi.org / 10.1515 / popets-2018-0037

[0264]

[22] Xiaofeng Ding, Wanlu Yang, Kim-Kwang Raymond Choo, Xiaoli Wang, and HaiJin, 2019, Privacy preserving similarity joins using MapReduce, Information Science 493 (2019), pp. 20–33. https: / / doi.org / 10.1016 / j.ins.2019.03.035

[0265]

[23] Changyu Dong, Liqun Chen and Zikai Wen, When private set intersection meets big data: an efficient and scalable protocol, in the proceedings of the 2013 ACMSIGSAC Conference on Computer and Communications Security, pp. 789-800.

[0266]

[24] Brett Hemenway Falk, Daniel Noble, and Rafail Ostrovsky, 2019, Private Set Inter-section with Linear Communication from General Assumptions, Proceedings of the 18th ACM Symposium on Electronic Social Privacy, WPES@CCS2019, London, UK, 11 November 2019, Lorenzo Cavallaro, Johannes Kinder, and Joseph Domingo Ferrer (eds.), ACM, pp. 14–25. https: / / doi.org / 10.1145 / 3338498.3358645

[0267]

[25] Michael J. Freedman, Carmit Hazay, Kobbi Nissim, and Benny Pinkas, 2016, Efficient Set Intersection with Simulation-Based Security, J. Cryptography 29, 1 (2016), pp. 115–155. https: / / doi.org / 10.1007 / s00145-014-9190-0

[0268]

[26] Michael J. Freedman, Kobbi Nissim and Benny Pinkas, 2004, Efficient private matching and set intersection, European Conference on Cryptography.

[0269]

[27] Satrajit Ghosh and Mark Simkin, 2019, The Communication Complexity of Threshold Private Set Intersection, in Proceedings of the 39th International Conference on Cryptography, Advances in Cryptography – Encryption 2019, Santa Barbara, CA, August 18–22, 2019, Proceedings, Part II (Lectures in Computer Science, Vol. 11693), Alexandra Boldyreva and Daniele Micciancio (eds). Springer, pp. 3–29. https: / / doi.org / 10.1007 / 978-3-030-26951-7_1

[0270]

[28] Thomas Schneider, Michael Zohner, Gilad Asharov, and Yehuda Lindell, 2013, More efficient oblivious transfer and extensions for faster secure computation, in CCS, pp. 535-548.

[0271]

[29] Gene Tsudik Giuseppe Ateniese, Emilano De Cristofaro, 2011, (if) size matters: Size-hiding private set intersection, in PKC. pp. 156-173.

[0272]

[30] Florian Hahn, Nicolas Loza, and Florian Kerschbaum, 2019, Joins Over Encrypted Data with Fine Granular Security, 35th IEEE International Conference on Data Engineering, ICDE 2019, Macau, April 8-11, 2019, IEEE, pp. 674-685. https: / / doi.org / 10.1109 / ICDE.2019.00066

[0273]

[31] Per A. Hallgren, Claudio Orlandi and Andrei Sabelfeld, 2017, PrivatePool: Privacy-Preserving Ridesharing, in CSF.

[0274]

[32] Kim Laine Hao Chen and Peter Rindal, 2017, Fast Private Set Intersection from Homomorphic Encryption, in CCS, pp. 1243-1255.

[0275]

[33] Kim Laine Hao Chen, Zhiccong Huang and Peter Rindal, 2018, Labeled PSI from Fully Homomorphic Encryption with Malicious Security, in CCS, pp. 1223-1237.

[0276]

[34] Carmit Hazay and Kobbi Nissim, 2010, Efficient Set Operations in the Presence of Malicious Adversaries, in PKC.

[0277]

[35] Carmit Hazay and Muthuramakrishnan Venkitasubramaniam. Scalable Multi-party Private Set-Intersection, in PKC, Serge Fehr (ed.). 2017.

[0278]

[36] Yan Huang, David Evans, Jonathan Katz, and Lior Malka, Faster Secure Two-Party Computation Using Garbled Circuits, 2011, Proceedings of the 20th USENIX Security Symposium, San Francisco, CA, August 8–12, 2011, USENIX Association. http: / / static.usenix.org / events / sec11 / tech / full_papers / Huang.pdf

[0279]

[37] Bernardo A. Huberman, Matthew K. Franklin, and Tad Hogg, 1999. Enhancing privacy and trust in electronic communities. In the proceedings of the First ACM Conference on Electronic Commerce (EC-99), Denver, DORC, November 3–5, 1999, Stuart I. Feldman and Michael P. Wellman (eds.), ACM, pp. 78–86. https: / / doi.org / 10.1145 / 336992.337012

[0280]

[38] Mihaela Ion, Ben Kreuter, Ahmet Erhan Nergiz, Sarvar Patel, Mariana Raykova, Shobhit Saxena, Karn Seth, David Shanahan, and Moti Yung, 2019. On Deploying Secure Computing Commercially: Private Intersection-Sum Protocols and their Business Applications. IACR Cryptol.ePrint Arch. 2019, p. 723. https: / / eprint.iacr.org / 2019 / 723

[0281]

[39] Mihaela Ion, Ben Kreuter, Erhan Nergiz, Sarvar Patel, Shobhit Saxena, Karn Seth, David Shanahan, and Moti Yung. Private Intersection-Sum Protocol with Applications to Attributing Aggregate Ad Conversions, 2017, ia.cr / 2017 / 735.

[0282]

[40] Lea Kissner and Dawn Song, 2005, Privacy-preservingset operations, in CRYPTO,

[0283]

[41] Vladimir Kolesnikov, Ranjit Kumaresan, Mike Rosulek and Ni Trieu, 2016, Efficient batched oblivious PRF with applications to private set intersection, in the proceedings of the 2016 ACM SIGSAC conference on computer and communications security, pp. 818-829.

[0284]

[42] Vladimir Kolesnikov, Naor Matania, Benny Pindas, Mike Rosulek and NiTrieu, 2017, Practical Multi-party Private Set Intersection from Symmetric-Key Techniques, in CCS.

[0285]

[43] Vladimir Kolesnikov, Mike Rosulek, Ni Trieu, and Xiao Wang, 2019, Scalable Private Set Union from Symmetric-Key Techniques, in Advances in Cryptography - ASIACRYPT 2019 - The 25th International Conference on Cryptography and Information Security Theory and Applications, Kobe, Japan, December 8-12, 2019, Proceedings, Part II (Lectures in Computer Science, Vol. 11922), Steven D. Galbraith and Shiho Moraii (eds.), Springer, pp. 636-666. https: / / doi.org / 10.1007 / 978-3-030-34621-8_23

[0286]

[44] Chang Liu, Xiao Shaun Wang, Kartik Nayak, Yan Huang and Elaine Shi, 2015, Oblivm: A programming framework for secure computation, at the IEEE Workshop on Security and Privacy 2015, IEE, pp. 359-376.

[0287]

[45] Apache Livy, 2017, Apache Livy, https: / / livy.apache.org /

[0288]

[46] Catherine A. Meadows, 1986, A More Efficient Cryptographic Matchmaking Protocol for Use in the Absence of a Continuously Available Third Party, in Proceedings of the 1986 IEEE Security and Privacy Workshop, Oakland, California, April 7-9, 1986, IEEE Computer Society, pp. 134-137. https: / / doi.org / 10.1109 / SP.1986.10022

[0289]

[47] Claudio Orlandi Michele Ciampi, 2018, Combining Private Set-Intersection with Secure Two-Party Computation, SCN, pp. 464-482.

[0290]

[48] ​​Benny Pinkas Moni Naor, 2001, Efficientoblivious transfer protocols, in SODA. pp. 448-457.

[0291]

[49] Shishir Nagaraja, Prateek Mittal, Chi-Yao Hong, Matthew Caesar and Nikita Borisov, 2010, BotGrep: Finding P2PBots with Structured Graph Analysis, USENIX Security Workshop.

[0292]

[50] Arvind Narayanan, Narendran Thiagarajan, Mugdha Lakhani, Michael Hamburg, and Dan Boneh, Location Privacy via Private Proximity Testing, in Proceedings of the Symposium on Security of Networks and Distributed Systems, NDSS 2011, San Diego, California, February 6–9, 2011, Internet Society. https: / / www.ndss-symposium.org / ndss2011 / privacy-private-proximity-testing-paper

[0293]

[51] Michele Orrua, Emmanuela Orsini and Peter Scholl, 2016, Actively Secure 1-out-of-N OT Extension with Application to Private Set Intersection, in CT-RSA.

[0294]

[52] Antonis Papadimitriou, Ranjita Bhagwan, Nishanth Chandran, Ramachandran Ramjee, Andreas Haeberlen, Harmeet Singh, Abhishek Modi and Saikrishna Badrinarayan, 2016, Big data analytics over encrypted datasets with seabed, 12th USENIX Operating System Design and Implementation Workshop (OSDI 16), pp. 587-602.

[0295]

[53] Phillipp Schoppmann Peter Rindal. [nd], VOLE-PSI: Fast OPRF and Circuit-PSI from Vector-OLE, at the European Conference on Cryptography.

[0296]

[54] Benny Pindas, Mike Rosulek, Ni Trieu and Avishay Yanai, 2019, Spot-light: Lightweight private setintersection from sparse OT extension, in CRYPTO.

[0297]

[55] Benny Pindas, Mike Rosulek, Ni Trieu and Avishay Yanai, PSI from PaXoS: Fast, Malicious Private Set Intersection, 2020, at the European Conference on Cryptography, Anne Canteaut and Yuval Ishai (eds).

[0298]

[56] Benny Pindas, Thomas Schneider, Gil Segev and Michael Zohner, 2015, Phasing: Private Set Intersection Using Permutation-based Hashing, in USENIX.

[0299]

[57] Benny Pindas, Thomas Schneider, Christian Weinert and Udi Wieder, Efficient Circuit-Based PSI via Cuckoo Hashing, 2018, European Conference on Cryptography.

[0300]

[58] Benny Pindas, Thomas Schneider and Michael Zohner, 2014, Faster Private Set Intersection Based on OT Extension, in USENIX.

[0301]

[59] Benny Pindas, Thomas Schneider, and Michael Zohner, 2016, Scalable Private Set Intersection Based on OT Extension, IACR Cryptol.ePrint Arch. 2016, p. 930. http: / / eprint.iacr.org / 2016 / 930

[0302]

[60] Benny Pindas, Thomas Schneider, and Michael Zohner, 2018, Scalable Private Set Intersection Based on OT Extension, ACM Trans. Priv. Secur. 21, 2 (2018), 7:1–7:35. https: / / doi.org / 10.1145 / 3154794

[0303]

[61] Raluca Ada Popa, Catherine MS Redfield, Nickolai Zeldovich and Hari Balakrishnan, 2011, CryptDB: protecting confidentiality with encrypted query processing, Proceedings of the 23rd ACM Symposium on Operating System Principles, pp. 85-100.

[0304]

[62] Amanda C. Davi Resende and Diego F. Aranha, 2017, Unbalanced Approximate Private Set Intersection, IACR Cryptol.ePrintArch. 2017, p. 677. http: / / eprint.iacr.org / 2017 / 677

[0305]

[63] Peter Rindal, [nd].libPSI: an efficient, portable, and easy-to-use Private Set IntersectionLibrary. https: / / github.com / osu-crypto / libPSI.

[0306]

[64] Peter Rindal and Mike Rosulek, 2017, Improved Private Set Intersection Against Malicious Adversaries, European Conference on Cryptography.

[0307]

[65] Peter Rindal and Mike Rosulek, 2017, Malicious-secure private set intersection via dual execution, in CCS.

[0308]

[66] Avishay Yanai Ryan Deng Raluca Ada Popa Joseph M. Hellerstein Rishabh Poddar, Sukrit Kalla, 2020, Seanate: A Maliciously-Secure MPC Platform for Collaborative Analytics, IACRCryptol.ePrint Arch.2020, p. 1350.

[0309]

[67] Mariana Raykova, Saeed Sadeghian, Seny Kamara, and Payman Mohassel, 2014, Scaling Private Set Intersection to Billion-Element Sets, in Financial Cryptography and Data Security. pp. 195-215.

[0310]

[68] Ranjit Kumaresan Vladimir Kolesnikov, 2013, Improved OT Extension for Transferring Short Secrets, in Crypto(2), pp. 54-70.

[0311]

[69] Nikolaj Volgushev, Malte Schwarzkopf, Ben Getchell, Mayank Varia, Andrei Lapets and Azer Bestavros, Conclave: secure multi-party computation on big data, Proceedings of the 14th European Systems Conference, 2019, pp. 1-18.

[0312]

[70] Wikipedia, 2020, Java Native Interface - Wikipedia. https: / / en.wikipedia.org / wiki / Java_Native_Interface

[0313]

[71] Song Jiang, Qiuyu Li, Shunde Cao, Pengfei Zuo, Yuan Yuan Sun, Yu Hua, 2017, SmartCuckoo: A Fast and Cost-Efficient Hashing Index Scheme for Cloud Storage Systems, USENIX Annual Technical Conference, pp. 553-565.

[0314]

[72] Kobbi Nissim Erez Petrank Yuval Ishai & Joe Kilian, 2003, Extending Oblivious Transfers Efficiently. Crypto, pp. 145-161.

[0315]

[73] Matei Zaharia, Mosharaf Chowdhury, Tathagata Das, Ankur Dave, Justin Ma, Murphy McCauly, Michael J Franklin, Scott Shenker and Ion Stoica, 2012, Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing, as part of the Ninth USENIX Workshop on Network Systems Design and Implementation (NSDI 12), pp. 15-28.

[0316]

[74] Matei Zaharia, Mosharaf Chowdhury, Michael J. Franklin, Scott Shenker and Ion Stoica, Spark: Cluster Computing with Working Sets, 2010. In Proceedings of the 2nd USENIX HotCloud'10 Conference (Boston, MA), USENIX Association, USA, p. 10.

[0317]

[75] Samee Zahur and David Evans, Obliv-C: A Language for Extensible Data-Oblivious Computation, IACRCryptol.ePrint Arch., 2015, p. 1153.

[0318]

[76] Wenting Zheng, Ankur Dave, Jethero G Beekman, Raluca Ada Popa, Joseph E Gonzalez and Ion Stoica, 2017, Opaque: Anoblivious and encrypted distributed analytics platform. In the 14th USENIX Network Systems Design and Implementation Workshop (NSDI 17), pp. 283-298.

Claims

1. A method comprising performing the following by a first-party computing system: The first-party set of data records is tokenized to generate a tokenized first-party set, wherein the tokenized first-party set includes multiple first-party tokens; The plurality of tokenized first-party tokens are generated by assigning each of the plurality of first-party tokens to a tokenized first-party subset of a plurality of tokenized first-party subsets using an assignment function; Each tokenized first-party subset is assigned to the corresponding first-party worker node among multiple first-party worker nodes; Generate multiple intersecting token subsets using the following steps: For each of the plurality of tokenized first-party subsets, the corresponding first-party worker node is used to: A private set intersection protocol is executed with the respective second-party worker nodes from a plurality of second-party worker nodes from a second-party computing system to privately determine the intersection between the tokenized first-party subset and the tokenized second-party subset assigned to the respective second-party worker node, the intersection providing an intersecting token subset of the plurality of intersecting token subsets; Combine the multiple intersecting token subsets to generate an intersecting token set; as well as Detoxify the intersecting token set to generate an intersecting set of data records. The first set corresponds to the set of join keys corresponding to the join query of the private database table, and the method further includes: Receive one or more filtered second-party database tables from the second-party computing system; The intersection set is used to filter one or more first-party database tables, thereby generating one or more filtered first-party database tables; and Join the one or more filtered first-party database tables with the one or more filtered second-party database tables to generate a join table.

2. The method of claim 1, wherein the allocation function comprises a dictionary sorting function, and wherein generating the plurality of tokenized first-party subsets by assigning each of the plurality of first-party tokens to one of the plurality of tokenized first-party subsets using the allocation function comprises: Each first-party token is assigned to a corresponding tokenized first-party subset based on the dictionary ordering of the multiple first-party tokens.

3. The method of claim 1, further comprising, for each of the plurality of tokenized first-party subsets: Determine the padding value, which includes the difference between the size of the first tokenized subset and the target value; Generate multiple random virtual tokens, wherein the multiple random virtual tokens include multiple random virtual tokens equal to the fill value; and The plurality of random virtual tokens are assigned to the first subset of tokenized tokens.

4. The method according to claim 1, wherein the first-party computing system includes a first-party server and a first-party computing cluster, and the first-party computing cluster includes a first-party driver node and the plurality of first-party worker nodes.

5. The method of claim 1, further comprising retrieving the first-party set from a first-party database.

6. The method of claim 1, wherein the first-party set comprises a plurality of first-party elements, wherein the plurality of first-party tokens comprises a plurality of hash values, and wherein tokenizing the first-party set comprises generating the plurality of hash values ​​by hashing each first-party element using a hash function.

7. The method according to claim 1, wherein, The intersection set of the data records includes the intersection of the first set and the second set corresponding to the second computing system, wherein the first set and the second set include an equal number of elements.

8. The method of claim 1, wherein the set of intersecting tokens comprises the union of the plurality of subsets of intersecting tokens, and wherein combining the plurality of subsets of intersecting tokens comprises determining the union of the plurality of subsets of intersecting tokens.

9. The method according to claim 1, further comprising: Before tokenizing the first set: Receive a request message from the orchestrator computer, the request message indicating the first party set, and Retrieve the first-party set from the first-party database; and The intersection set is sent to the orchestrator computer, which in turn sends the intersection set to the client computer.

10. The method of claim 1, wherein the plurality of tokenized first party subsets and the plurality of tokenized second party subsets comprise a predetermined number of subsets.

11. The method of claim 10, wherein each tokenized first-party subset is associated with a numeric identifier containing a predetermined number of end values ​​between the subset, wherein the allocation function is a hash function that produces a hash value containing a predetermined number of end values ​​between the subset, and wherein assigning each of the plurality of first-party tokens to one of the plurality of tokenized first-party subsets comprises: For each of the plurality of first-party tokens: The hash value is generated using the first-party token as input to the hash function; as well as The first-party token is assigned to a tokenized first-party subset, the tokenized first-party subset having the digital identifier equal to the hash value.

12. A method comprising performing the following by a first-party computing system: Receive a private database table join query, wherein the private database table join query identifies one or more first database tables and one or more attributes; Retrieve one or more first database tables from the first-party database; Multiple first-party connection keys are determined based on the one or more first database tables and the one or more attributes; The plurality of first-party connection keys are tokenized to generate a set of tokenized first-party connection keys, wherein the set of tokenized first-party connection keys includes a plurality of first-party tokens; The plurality of tokenized first-party tokens are generated by assigning each of the plurality of first-party tokens to a tokenized first-party subset of a plurality of tokenized first-party subsets using an assignment function; For each of the plurality of tokenized first-square subsets: Using the tokenized first-party subset and the tokenized second-party subset corresponding to the second-party computing system, a private set intersection protocol is executed with the second-party computing system, thereby executing multiple private set intersection protocols and generating multiple intersecting token subsets; Combine the multiple intersecting token subsets to generate an intersecting token set; The intersecting token set is detoxified to generate an intersecting connection key set; The one or more first database tables are filtered using the corresponding cross-connection key set to generate one or more filtered first database tables. Receive one or more filtered second database tables from the second-party computing system; and The one or more filtered first database tables are combined with the one or more filtered second database tables to generate a join database table.

13. The method of claim 12, wherein the private database table join query further includes an "wherein" clause, and wherein the method further includes pre-filtering the one or more first database tables based on the "wherein" clause before determining the plurality of first-party join keys.

14. The method of claim 12, wherein each of the plurality of first-party connection keys corresponds to an attribute among the one or more attributes, and wherein tokenizing the plurality of first-party connection keys comprises: For each of the one or more attributes, concatenate each first-party connection key corresponding to the attribute to generate a plurality of concatenated first-party connection keys; as well as Each of the plurality of concatenated first-party connection keys is hashed to generate a plurality of hash values, wherein the set of tokenized first-party connection keys includes the plurality of hash values.

15. The method of claim 12, further comprising sending the one or more filtered first database tables to a second third-party computing system, wherein the second third-party computing system combines the one or more filtered first database tables with the one or more filtered second database tables to generate the joined database table.

16. The method of claim 12, wherein the private database table join query is received from an orchestrator computer, and wherein the method further comprises sending the join database table to the orchestrator computer, wherein the orchestrator computer sends the join database table to a client computer.

17. The method of claim 12, further comprising: Generate a token column that includes the tokenized first-party connection keys in the set of tokenized first-party connection keys; as well as The token column is appended to the one or more first database tables.

18. The method of claim 17, wherein filtering the one or more first database tables using the associated cross-connection key set comprises: Based on the token column, remove one or more rows from the one or more first database tables, the one or more rows corresponding to one or more tokenized first-party connection keys that are not in the corresponding connection key set.

19. A first-party computing system, comprising: The first processor, and A first non-transient computer-readable medium coupled to the first processor, the first non-transient computer-readable medium comprising code executable by the first processor to implement the method according to any one of claims 1 to 18.

20. The first-party computing system of claim 19, further comprising a computing cluster including a driver node and a plurality of first-party worker nodes, wherein the first processor and the first non-transient computer-readable medium correspond to a first-party server, wherein the driver node includes a second processor and a second non-transient computer-readable medium coupled to the second processor, wherein each of the plurality of worker nodes includes a third processor among a plurality of third processors and a third non-transient computer-readable medium among a plurality of third non-transient computer-readable media, each third non-transient computer-readable medium being coupled to a corresponding third processor.

Citation Information

Patent Citations

  • Indexing databases for efficient relational querying

    US6507846B1