Methods for protecting membership and data in secure multi-party computing and communications, secure multi-party computing and communications systems, computer-readable media and computer programs
By generating padding datasets and performing set intersection operations with differential privacy, the method ensures secure multi-party computation protects membership and data privacy, thwarting attacker attempts to determine user membership.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- LEMON CO LTD
- Filing Date
- 2024-04-04
- Publication Date
- 2026-05-01
AI Technical Summary
Existing secure multi-party computation protocols fail to adequately protect membership and data privacy, particularly in private set intersection operations, allowing attackers to determine user membership and compromising privacy.
A method involving generating padding datasets based on privacy settings, upsampling, transforming, and performing set intersection operations to conceal the common dataset size, ensuring differential privacy and preventing user membership determination.
The solution effectively protects membership and data privacy by making it impossible for attackers to determine user membership, complying with privacy regulations and maintaining data confidentiality.
Smart Images

Figure 2026513966000001_ABST
Abstract
Description
Technical Field
[0001] [Cross - Reference to Related Applications] This application claims priority to U.S. Application No. 18 / 297,447, filed on April 7, 2023, with the invention title "Membership and Data Protection in Confidential Multi - Party Computation and / or Communication", and the disclosure of the application is incorporated herein by reference in its entirety.
[0002] The embodiments described in this specification generally relate to protecting membership privacy and data privacy. More specifically, the embodiments described herein relate to protecting membership privacy (such as elements, members, users, etc.) and data privacy (such as features or attributes associated with a user, etc.) in confidential multi - party computation and / or communication.
Background Art
[0003] Private Set Intersection (PSI) is one of the secure two - party or multi - party protocols or algorithms where common - part - related statistics are computed. With a PSI algorithm or protocol, two or more organizations are allowed to jointly compute a function (such as count, sum, etc.) on the common part of their respective data sets without explicitly exposing the common part to other parties. In applications, two parties may not want or may not be able to expose their underlying data to each other, but may still want to compute aggregate - level measurements. These two parties may want to do so while ensuring that the input data sets do not expose anything beyond these aggregate values related to individual users.
Summary of the Invention
Problems to be Solved by the Invention
[0004] The features of the embodiments disclosed herein can provide a secure multi-party computation (MPC) protocol or algorithm for facilitating effective and secure measurements. The features of the embodiments disclosed herein can also provide protection for membership and data, for example, through a differential privacy (DP) protocol or algorithm.
[0005] The features of the embodiments disclosed herein can further solve problems associated with privacy matching of two-party or multi-party data in two-party or multi-party secure computation and help generate a secure sharing of data relating to the common portion of a dataset from two or more parties.
[0006] Features of the embodiments disclosed herein include independently generating padding or filler elements for each party's dataset according to a pre-calibrated noise distribution, adding the padding elements to each dataset, and executing an MPC algorithm or protocol. Another feature of the embodiments disclosed herein is that the common size exposed in subsequent PSI operations is random and differentially concealed, making it virtually impossible for an attacker to determine a user's membership in a dataset or organization, thereby complying with privacy regulatory requirements. [Means for solving the problem]
[0007] In one exemplary embodiment, a method is provided for protecting membership and data in confidential multi-party computing and communications. The method includes generating a padding dataset, the size of which is determined based on data privacy settings. The method also includes upsampling a first dataset using the padding dataset, transforming the first dataset, dispatching the first dataset, generating a third dataset by performing a set intersection operation on the first and second datasets, generating a first share based on the third dataset, and constructing a result based on the first and second shares.
[0008] In another exemplary embodiment, a secure multiparty computing and communication system is provided. The system comprises memory for storing a first dataset and a processor for generating a padding dataset. The size of the padding dataset is determined based on data privacy settings. The processor further upsamples the first dataset using the padding dataset, transforms the first dataset, dispatches the first dataset, performs a set intersection operation on the first dataset and the second dataset to generate a third dataset, generates a first share based on the third dataset, and constructs a result based on the first share and the second share.
[0009] In yet another embodiment, a non-temporary computer-readable medium storing computer-executable instructions is provided. When the instructions are executed, one or more processors perform operations including generating a padding dataset. The size of the padding dataset is determined based on data privacy settings. The operations also include upsampling a first dataset using the padding dataset, transforming the first dataset, dispatching the first dataset, performing an intersection operation on the first dataset and the second dataset to generate a third dataset, generating a first share based on the third dataset, and constructing a result based on the first share and the second share. [Brief explanation of the drawing]
[0010] The accompanying drawings illustrate various embodiments of the systems and methods of the Disclosure, and embodiments of various other aspects of the Disclosure. Those skilled in the art will understand that the illustrated element boundaries in the drawings (e.g., boxes, groups of boxes, or other shapes) represent examples of boundaries. In some examples, one element may be designed as multiple elements, or multiple elements may be designed as one element. In some examples, an element shown as an internal component of one element may be realized as an external component of another element, and vice versa. The following description is non-limiting and non-exclusive, with reference to the drawings. The components in the drawings are not necessarily to scale, and the emphasis is on illustrating the principle. Since various changes and modifications may become apparent to those skilled in the art from the following detailed description, embodiments are described only as illustrations in the following detailed description. [Figure 1] This is a schematic diagram showing an exemplary secure computation and communication system arranged according to at least some embodiments described herein. [Figure 2]This flowchart shows an exemplary processing flow for a multi-identification matching algorithm according to at least some embodiments described herein. [Figure 3] This is a schematic diagram showing an example of the processing flow in Figure 2, according to at least some embodiments described herein. [Figure 4A] This flowchart shows an exemplary processing flow for protecting membership and data in confidential multi-party computing and communications, according to at least some embodiments described herein. [Figure 4B] This flowchart shows an exemplary processing flow for protecting membership and data in confidential multi-party computing and communications, according to at least some embodiments described herein. [Figure 5A] This figure shows a portion of a schematic diagram illustrating an example of the processing flow in Figures 4A and 4B, relating to at least some embodiments described herein. [Figure 5B] This figure shows a portion of a schematic diagram illustrating an example of the processing flow in Figures 4A and 4B, relating to at least some embodiments described herein. [Figure 5C] This figure shows a portion of a schematic diagram illustrating an example of the processing flow in Figures 4A and 4B, relating to at least some embodiments described herein. [Figure 5D] This figure shows a portion of a schematic diagram illustrating an example of the processing flow in Figures 4A and 4B, relating to at least some embodiments described herein. [Figure 5E] This figure shows a portion of a schematic diagram illustrating an example of the processing flow in Figures 4A and 4B, relating to at least some embodiments described herein. [Figure 5F] This figure shows a portion of a schematic diagram illustrating an example of the processing flow in Figures 4A and 4B, relating to at least some embodiments described herein. [Figure 6]This is a schematic diagram of an exemplary computer system applicable to realizing an electronic device, arranged according to at least some embodiments described herein. [Modes for carrying out the invention]
[0011] In the following detailed description, specific embodiments of the Disclosure will be described with reference to the accompanying drawings, which constitute part of the specification. In this specification and in the drawings, similar reference numerals represent elements capable of performing the same, similar, or equivalent functions, unless the context indicates otherwise. Furthermore, unless otherwise specified, the description of each successive drawing may refer to the features of one or more preceding drawings to provide a clearer context and a more substantial description of the current exemplary embodiment. However, the embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized and other modifications made without departing from the gist or scope of the subject matter presented herein. As is generally described herein and shown in the drawings, each aspect of the Disclosure may be designed in arrangement, substitution, combination, separation, and various different configurations, all of which are expressly assumed herein.
[0012] It should be understood that the disclosed embodiments are merely examples of the various forms in which the Disclosure can be embodied. To avoid obscuring the Disclosure with unnecessary detail, well-known functions or structures are not described in detail. Therefore, the specific structural and functional details described herein should be interpreted not as limitations, but merely as the basis for the claims and as representative grounds to instruct those skilled in the art to utilize the Disclosure in a variety of substantially appropriate detailed structures.
[0013] In addition, this disclosure may describe functional block components and various processing steps. It should be understood that such functional blocks may be implemented by any number of hardware and / or software components configured to perform a defined function.
[0014] The scope of the present disclosure should be determined by the appended claims and their legal equivalents, rather than by the examples given herein. For example, the steps recited in any method claim may be performed in any order and are not limited to the order recited in the claim. Also, there are no elements essential to the practice of the present disclosure unless specifically recited herein as "important" or "essential".
[0015] As described herein, "data set" or "dataset" is a term of art and may refer to an organized collection of data that is electronically stored and accessed. In an exemplary embodiment, a data set may refer to a database, a data table, a part of a database or a data table, etc. A data set may correspond to one or more data set tables where each column represents a particular variable or field and each row corresponds to a given record of the data set. The data set may list each of the variables and / or the values of each record of the data set. Additionally or alternatively, it should be understood that a data set may refer to a set of related data and the way in which the related data is organized. In an exemplary embodiment, each record of a data set may include fields or elements, such as one or more predefined or predetermined identification information (e.g., membership identification information, user identification information, such as a username, email address, phone number, unique ID of the user, etc.), and / or one or more attributes or characteristics or values associated with the one or more identification information. It should be understood that any user identification information and / or user data described herein is approved by the user, authorized, and / or otherwise approved for use in the embodiments described herein and their appropriate legally equivalent ones understood by those skilled in the art.
[0016] As used herein, "inner join" or "inner-join" is a technical term and may refer to an operation or function that involves combining records from data sets when there are values that match in fields common to the data sets. For example, by performing an inner join on a "department" data set and an "employee" data set, all employees within each department may be determined. It should be understood that in the resulting data set of the inner join operation (i.e., the "common part"), the inner join may include information from both related data sets. On the other hand, an outer join may also include information in the resulting data set that is not related to other data sets. A secret inner join may refer to an inner join operation on data sets of two or more parties that does not expose data within the common part of the data sets of the two or more parties.
[0017] As used herein, "hashing" may refer to an operation or function that converts or transforms an input (a key, such as a numerical value, a string, etc.) into an output (e.g., another numerical value, another string, etc.). It should be understood that hashing is a technical term and may be used in cyber security applications to access data in a slightly and approximately constant time for each search.
[0018] As used herein, "MPC" or "multi-party computation" is a technical term and may refer to the field of cryptography that aims to create a method for parties to jointly compute a function on their joint input while keeping their respective inputs secret. Different from conventional cryptographic tasks where cryptography can guarantee the security and integrity of communication or storage when adversaries are outside the participants' systems (e.g., eavesdroppers of senders and / or receivers), it should be understood that the cryptography in MPC can protect the privacy of participants from each other.
[0019] As used herein, “ECC” or “elliptic-curve cryptography” is a technical term and may refer to public-key cryptography based on the algebraic structure of elliptic curves over a finite field. It should be understood that ECC can tolerate smaller keys compared to non-EC cryptography in order to provide equal security. It should also be understood that “EC” or “elliptic curve” can be applied to key agreement, digital signatures, pseudorandom number generators, and / or other tasks. Elliptic curves can be used indirectly in encryption by combining key agreement between parties with symmetric encryption schemes. Elliptic curves can also be used in integer factorization algorithms based on elliptic curves that are applied to encryption.
[0020] As used herein, the “Deterministic Diffie-Hellman Assumption” or “DDH Assumption” is a technical term and may refer to a computational complexity assumption concerning a particular problem involving discrete logarithms in cyclic groups. It should be understood that the DDH assumption may be used as the basis for proving the security of many cryptographic protocols.
[0021] As used herein, “Elliptic Curve Diffie-Hellman” or “ECDH” is a technical term that may refer to a key agreement protocol or corresponding algorithm that enables two or more parties, each possessing an elliptic curve public-private key pair, to establish a shared secret over a non-secret channel. It should be understood that the shared secret can be used directly as a key or to derive another key. It should also be understood that the key, or a derived key, may be used to encrypt or encode subsequent communications using symmetric key ciphertext. Furthermore, it should be understood that ECDH may refer to a variation of the Diffie-Hellman protocol using elliptic curve cryptography.
[0022] As used herein, “homomorphic” encryption is a technical term and may refer to a form of encryption that allows a user to perform calculations on encrypted data without first decrypting the encrypted data. It should be understood that the resulting homomorphic encryption calculations remain encrypted and, when decrypted, yield the same output as if the calculations had been performed on unencrypted data. It should also be understood that homomorphic encryption can be used for privacy-preserving outsourced storage and computation, which may allow data to be encrypted and outsourced to a commercial cloud environment for processing in an encrypted state. Furthermore, it should be understood that additive homomorphic encryption or cryptographic system may refer to a form of encryption or cryptographic system that can compute the encryption of m1+m2 given only the public key and the encryption of messages m1 and m2.
[0023] As used herein, “secret sharing” or “secret partitioning” are technical terms and may refer to cryptographic operations or algorithms for generating a secret and dividing it into multiple shares for distribution among multiple parties, such that the secret can only be reconstructed when each party contributes its respective share. It should be understood that secret sharing may also refer to operations or algorithms for distributing a secret among groups such that no individual has any information that can be understood about the secret, but a sufficient number of individuals, by combining their “shares,” may be able to reconstruct the secret. It should also be understood that insecure secret sharing may allow an adversary to gain more information using each share, while secure secret sharing may be “all or nothing,” and “all” may mean any number of shares required.
[0024] As used herein, “secret set intersection operation” is a technical term and may refer to a secure multi-party computational cryptographic operation, algorithm, or function in which two or more parties, each holding a dataset, compare encrypted versions of these datasets in order to compute a common part. It should be understood that with a secret set intersection operation, neither party exposes to the other any data elements other than those in the common part.
[0025] As used herein, “shuffle,” “shuffling,” “rearrange,” or “sort” are technical terms and may refer to an operation or algorithm for rearranging and / or randomly rearranging the order of records (elements, rows, etc.) in, for example, arrays, datasets, databases, data tables, etc.
[0026] As used herein, “semi-honest” attacker is a technical term and may refer to a party that adheres to the established protocol but may attempt to corrupt the party. It should be understood that a “semi-honest” party may be a corrupted party that honestly executes the current protocol but attempts to learn messages received from another one or more parties for purposes beyond what the protocol intends.
[0027] As used herein, “Differential Privacy” or “DP” is a technical term and may refer to standards, protocols, systems, and / or algorithms for publicly sharing information about a dataset by describing patterns of groups of elements within the dataset while concealing information about individual users enumerated in the dataset. It should be understood that differential privacy may also refer to constraints on algorithms used to release aggregated information about a statistical dataset or database to users that limit the disclosure of personal information of individual records in the dataset or database.
[0028] The following are non-exclusive examples of the background, setting, or application of differential privacy. A trusted data owner (or data holder or curator, e.g., a social media platform, website, service provider, application, etc.) may store a dataset of sensitive information about a user or member (e.g., the dataset contains records / rows of a user or member). Each time the dataset is queried (or manipulated, e.g., analyzed, processed, used, stored, shared, accessed, etc.), there may be a possibility or probability of an individual's privacy being compromised (e.g., a probability of data privacy breach or privacy loss). Differential privacy can provide a rigorous framework and security definition for algorithms that manipulate sensitive data and expose aggregate statistics in order to prevent an individual's privacy from being compromised, for example, by resisting link attacks or auxiliary information and / or by providing limits on the quantitative measure of harm (privacy breach, privacy loss, etc.) caused by individual records in a dataset.
[0029] It should be understood that the above requirements for differential privacy protocols or algorithms may refer to a measure of "how much data privacy is granted when performing an operation or function (e.g., by a single query or operation on an input dataset)." The DP parameter "ε" may refer to a privacy budget (i.e., a limit on the amount of data privacy that is permissible in a leak), which represents, for example, the maximum difference between a query or operation on dataset A and the same query or operation on dataset A' (which differs from A by only one element or record). A smaller value of ε indicates stronger privacy protection for the multi-identity privacy protection mechanism. Another DP parameter "δ" may refer to a probability, such as the probability of information being leaked by chance. In an exemplary embodiment, the required or predetermined range of ε may be from about 1 to about 3. The required or predetermined range of δ may be about 10 -10 (or about 10 -8 ) from about 10 -6This may also be the case. Another DP parameter sensitivity may refer to a quantified amount indicating how much noise perturbation is required in the DP protocol or algorithm. It should be understood that determining sensitivity may require determining the maximum possible change in the results. That is, sensitivity may refer to the potential impact that changes in the underlying dataset may have on the results of queries to the dataset.
[0030] As used herein, “differential privacy synthesis” or “DP synthesis” is a technical term and may refer to the total or overall differential privacy when a particular dataset is queried (or operated on, e.g., analyzed, processed, used, stored, shared, accessed, etc.) two or more times. DP synthesis quantifies the overall differential privacy (which may be reduced by considering the DP of a single query or operation) when multiple separate queries or operations are performed on a single dataset. It should be understood that if a single query or operation on a dataset has a privacy loss L, the cumulative effect of N queries on data privacy (referred to as N-fold synthesis or N-fold DP synthesis) may be greater than L and less than L*N. In one exemplary embodiment, N-fold DP synthesis may be determined based on an N-fold convolution operation of the privacy loss distributions. For example, DP synthesis of two queries may be determined based on the convolution of the privacy loss distributions of the two queries. In one exemplary embodiment, the number N may be about 10, about 25, or any other suitable number. In exemplary embodiments, ε, δ, sensitivity, and / or number N may be predetermined to achieve a desired or predetermined data privacy protection objective or performance.
[0031] As referenced herein, the “binomial distribution” in probability theory and statistics is a technical term referring to the discrete probability distribution of successes in a sequence of n independent experiments, each asking a YES or NO question, and each having its own Boolean result of success (probability p) or failure (probability q = 1-p). It should be understood that Gaussian noise in the field of signal processing may refer to signal noise that has the same probability density function as a normal distribution (i.e., a Gaussian distribution). In other words, the values that Gaussian noise can take follow a normal distribution (i.e., a Gaussian distribution). Similarly, binomial noise may refer to signal noise that has the same probability density function as a binomial distribution.
[0032] It should be understood that differential privacy requirements can be achieved by intentionally adding or injecting noise into a dataset to form anonymized data. This allows data users to perform all possible or useful statistical analyses on the dataset without identifying any personal information. It should also be understood that adding controlled noise from a given distribution (such as a binomial, Lapeles, or normal / Gaussian distribution) can be a way to design differentially confidential algorithms. Furthermore, it should be understood that adding noise may be useful in designing confidentiality protection mechanisms for real-valued functions on sensitive data.
[0033] As referenced herein, in relation to cryptographic techniques, “lost communication” is a technical term that may refer to an algorithm, protocol, or operation in which the sender can transmit at least one of potentially many pieces of information to the receiver, but the sender does not know, is unaware of, or remains ignorant of which parts of the information (if any) have been transmitted. One form of lost communication is “1-2 lost communication” or “1-out-of-2 lost communication” for secret and confidential multi-party computation. For example, in a 1-2 lost communication protocol or algorithm, the sender has two messages m0 and m1 and wants to ensure that the receiver knows only one. The receiver has bit b (which may be 0 or 1) and receives m from the sender without the sender knowing b. b The sender wishes to receive the i-th message out of the n messages, but the sender wants to ensure that the receiver receives only one of the n messages. Lost communication may be generalized to "1-out-of-n lost communication," where the sender does not know which elements were queried, the receiver knows nothing about the other elements that were not retrieved, and the receiver receives exactly one dataset element. A 1-out-of-n lost communication protocol or algorithm may be defined as a natural generalization of a 1-out-of-2 lost communication protocol or algorithm. In one exemplary embodiment, the sender has n messages, the receiver has index i, and the receiver wishes to receive the i-th message out of the sender's messages without the sender knowing i, but the sender wants to ensure that the receiver receives only one of the n messages.
[0034] Figure 1 is a schematic diagram showing an exemplary secure computation and communication system 100 arranged according to at least some embodiments described herein.
[0035] System 100 may include terminal devices 110, 120, 130, and 140, a network 160, and a server 150. It should be understood that Figure 1 shows only an exemplary number of terminal devices, networks, and servers. The embodiments described herein are not limited to the number of terminal devices, networks, and / or servers described herein. That is, the number of terminal devices, networks, and / or servers described herein is provided for illustrative purposes only and is not intended to limit the number of terminal devices, networks, and / or servers described herein.
[0036] According to at least some embodiments, terminal devices 110, 120, 130, and 140 may be various electronic devices. These various electronic devices may include, but are not limited to, mobile devices such as smartphones, tablet computers, e-readers, laptop computers, desktop computers, and / or any other suitable electronic devices.
[0037] According to at least some exemplary embodiments, the network 160 may be a medium used to provide communication links between terminal devices 110, 120, 130, 140 and the server 150. The network 160 may be the internet, a local area network (LAN), a wide area network (WAN), a local interconnect network (LIN), a cloud, etc. The network 160 may be implemented by various types of connections such as wired communication links, wireless communication links, fiber optic cables, etc.
[0038] According to at least some embodiments, the server 150 may be a server that provides various services to users using one or more of the terminal devices 110, 120, 130, and 140. The server 150 may be implemented as a distributed server cluster including multiple instances of the server 150, or as a single server 150.
[0039] Users may interact with the server 150 via the network 160 using one or more of the terminal devices 110, 120, 130, and 140. Various applications or localized interfaces, such as social media applications and online shopping services, may be installed on the terminal devices 110, 120, 130, and 140.
[0040] Software applications or services relating to embodiments described herein, or services provided by a service provider, may be executed by server 150 and / or terminal devices 110, 120, 130, and 140 (which may be referred to herein as user devices). Accordingly, the equipment for the software applications and / or services may be located within server 150 and / or terminal devices 110, 120, 130, and 140.
[0041] Furthermore, it should be understood that if the service is not running remotely, system 100 may include only terminal devices 110, 120, 130, and 140 and / or server 150, without including network 160.
[0042] Furthermore, it should be understood that each of the terminal devices 110, 120, 130, and 140 and / or server 150 may also include one or more processors, memory, and a storage device for storing one or more programs. Each of the terminal devices 110, 120, 130, and 140 and / or server 150 may also include, respectively, an Ethernet connector, a wireless fidelity receptor, and the like. When the one or more programs are executed by the one or more processors, they may cause the one or more processors to perform the methods described in any embodiment described herein. It should also be understood that a computer-readable non-volatile medium may be provided according to the embodiments described herein. The computer-readable medium stores computer programs. When the computer programs are executed by the processors, they are used to perform the methods described in any embodiment described herein.
[0043] Figure 2 is a flowchart illustrating an exemplary processing flow 200 for a multi-identity matching algorithm according to at least some embodiments described herein.
[0044] Figure 3 is a schematic diagram showing an example of the processing flow 300 of Figure 2 according to at least some embodiments described herein. Therefore, the description of the processing flow 200 may refer to 310A, 310B, 320A, and 320B of schematic diagram 300.
[0045] It should be understood that the processing flow 200 disclosed herein may be performed by one or more processors (for example, the processors of one or more terminal devices among the terminal devices 110, 120, 130, and 140 in Figure 1, the processor of the server 150 in Figure 1, the central processor unit 605 in Figure 6, and / or any other suitable processor) unless otherwise specified.
[0046] It should also be understood that the processing flow 200 may include one or more operations, behaviors, or functions, as indicated by one or more of blocks 210, 220, 230, and 240. These various operations, behaviors, or actions may correspond, for example, to processor-executable software, program code, or program instructions that cause these functions to be performed. Although shown as discrete blocks, obvious modifications may be made depending on the desired implementation, for example, two or more of the blocks may be rearranged, more blocks may be added, and various blocks may be split into additional blocks, combined into fewer blocks, or removed. The processing flow 200 may begin in block 210.
[0047] In block 210 (initialization), the processor of each device may perform initialization functions or operations on, for example, system parameters and / or application parameters. The processor of each device may provide a dataset (e.g., 310A) for party 1 and / or provide a dataset (e.g., 310B) for party 2. It should be understood that datasets 310A and / or 310B may be upsampled datasets (e.g., 508A and / or 508B in Figure 5A) generated or acquired in block 420 of Figure 4A, as will be described in more detail below.
[0048] Furthermore, it should be understood that each dataset 310A or 310B may contain one or more identification (ID) fields or columns, and the number of identification fields or columns in dataset 310A may or may not be equal to the number of identification fields or columns in dataset 310B. As shown in Figure 3, each of datasets 310A and 310B contains two ID fields, id1 and id2.
[0049] In one exemplary embodiment, the processor of each device may shuffle dataset 310A for party 1 and / or shuffle dataset 310B for party 2. The processor may also transform the ID field of dataset 310A using a transformation scheme for party 1.
[0050] It should be understood that a function or operation for "transforming" or "transforming" one or more fields / columns (or records / rows) of a dataset or a portion thereof, such as one or more ID fields / columns (or records / rows), may also mean processing (e.g., encrypting, decrypting, coding, decrypting, manipulating, compressing, decompressing, converting, etc.) the dataset or a portion thereof. A "transformation method" may also mean an algorithm, protocol, or function that performs processing (e.g., encrypting, decrypting, coding, decrypting, manipulating, compressing, decompressing, converting, etc.) of the dataset or a portion thereof. In one embodiment, the processor may encrypt (or decrypt, code, decrypt, manipulate, compress, decompress, convert, etc.) the ID field of dataset 310A using, for example, the key of party 1, based on, for example, the ECDH algorithm or protocol.
[0051] The processor may also transform the ID field of dataset 310B using a transformation scheme for party 2. In one embodiment, the processor may encrypt (or decrypt, encode, decrypt, manipulate, compress, decompress, transform, etc.) the ID field of dataset 310B using, for example, the key of party 2, based on, for example, the ECDH algorithm or protocol.
[0052] For Party 1 and / or Party 2, the order in which the ID field of the dataset (310A or 310B) is transformed and the shuffling of the dataset (310A or 310B) may be switched or changed without affecting the purpose of the resulting dataset.
[0053] The processor of each device may further exchange dataset 310A and dataset 310B between party 1 and party 2. For party 1, the processor may dispatch or transmit dataset 310A to party 2 and receive or acquire dataset 310B from party 2. For party 2, the processor may dispatch or transmit dataset 310B to party 1 and receive or acquire dataset 310A from party 1. It should be understood that since datasets 310A and 310B have already been transformed (e.g., encoded), the corresponding receiving party may not know the actual data in the received dataset. It should now be understood that each party may have local copies of both dataset 310A and dataset 310B.
[0054] The processor of each device may further transform the ID field of the received transformed dataset 310B using a transformation scheme for Party 1. In one embodiment, the processor may encrypt (or decrypt, encode, decrypt, manipulate, compress, decompress, convert, etc.) the ID field of the received transformed dataset 310 using the key of Party 1, for example, based on the ECDH algorithm or protocol. The processor of each device may further transform the ID field of the received transformed dataset 310A using a transformation scheme for Party 2. In one embodiment, the processor may encrypt (or decrypt, encode, decrypt, manipulate, compress, decompress, convert, etc.) the ID field of the received transformed dataset 310A using the key of Party 2, for example, based on the ECDH algorithm or protocol.
[0055] The processor may also shuffle the received converted dataset 310A for party 2 and / or the received converted dataset 310B for party 1. The order of the conversion of the ID field of the received converted datasets (310A and / or 310B) and the shuffling of the received converted datasets (310A and / or 310B) for party 1 and / or party 2 may be switched or changed without affecting the purpose of the resulting dataset. The processor of each device may exchange the resulting shuffled dataset 310A (referred to as "310A" in blocks 220-240 for simplicity of explanation) and the resulting shuffled dataset 310B (referred to as "310B" in blocks 220-240 for simplicity of explanation) between party 2 and party 1. Processing may proceed from block 210 to block 220.
[0056] In block 220 (Sorting Datasets), the processor of each device may sort datasets 310A and / or dataset 310B for party 1 and / or party 2. For example, for party 1, the processor may sort the ID fields (id1, id2, etc.) of dataset 310A in an order (or sequence) corresponding to a predetermined importance or priority level of the ID fields. Dataset 310A may include ID fields such as username (e.g., having a priority level of 3), email address (e.g., having a priority level of 2), telephone number (e.g., having a priority level of 4), unique user ID (e.g., having a priority level of 1), etc. In one exemplary embodiment, the lower the priority level number, the more important the corresponding ID field. By sorting the ID field in dataset 310A, the user's unique ID (e.g., with a priority level of 1) is listed as the first field / column in dataset 310A, the email address (e.g., with a priority level of 2) is listed as the second field / column in dataset 310A, the username (e.g., with a priority level of 3) is listed as the third field / column in dataset 310A, and the phone number (e.g., with a priority level of 4) is listed as the fourth field / column in dataset 310A. In other words, in a non-restrictive example of dataset 310A, the ID field is sorted in ascending order of priority level numbers for the user's unique ID, email address, username, and user phone number.
[0057] For party 2, the processor may sort the ID fields (id1, id2, etc.) of dataset 310B in the same order (or sequence) as for dataset 310A for party 1, corresponding to a predetermined importance or priority level of the ID fields. It should be understood that the sorting of datasets 310A and 310B is for the purpose of preparing for the subsequent matching process. Processing may proceed from block 220 to block 230.
[0058] In block 230 (execution of matching logic), with datasets 310A and 310B sorted, the processor of each device may search for a match (or an inner join operation, etc.) between dataset 310A and dataset 310B for each ID field of dataset 310A (from the ID field with the lowest priority level number to the ID field with the highest priority level number), and obtain or generate a common part for party 1 (dataset 320A in Figure 3).
[0059] It should be understood that the search for a matching operation (or an inner join operation, etc.) involves, for each ID field in dataset 310A (from the ID field with the lowest priority level number to the ID field with the highest priority level number), and for each identifier in dataset 310A that matches an identifier in dataset 310B, deleting the record (or row) in dataset 310A that contains the matched identifier, and adding or appending the deleted record (or row) from dataset 310A to dataset 320A.
[0060] For example, as shown in Figure 3, for the ID field id1 in dataset 310A, records / rows containing "g", "c", and "e" each have corresponding matches in dataset 310B. Such records / rows may be deleted from dataset 310A, and the deleted records / rows may be added or appended to dataset 320A. For the ID2 in dataset 310A, records / rows containing "3" have corresponding matches in dataset 310B. Such records / rows may be deleted from dataset 310A, and the deleted records / rows may be added or appended to dataset 320A.
[0061] The processor of each device may search for a match (or perform an inner join operation, etc.) between dataset 310A and dataset 310B for each ID field of dataset 310B (from the ID field with the lowest priority level number to the ID field with the highest priority level number) and obtain or generate a common part for party 2 (dataset 320B in Figure 3).
[0062] It should be understood that the search for a matching operation (or an inner join operation, etc.) involves, for each ID field in dataset 310B (from the ID field with the lowest priority level number to the ID field with the highest priority level number), and for each identifier in dataset 310B that matches an identifier in dataset 310A, deleting the record (or row) in dataset 310B that contains the matched identifier, and adding or appending the deleted record (or row) from dataset 310B to dataset 255B.
[0063] For example, as shown in Figure 3, for the ID field id1 in dataset 310B, records / rows containing "g", "c", and "e" each have corresponding matches in dataset 310A. Such records / rows may be deleted from dataset 310B, and the deleted records / rows may be added or appended to dataset 320B. For the ID2 in dataset 310B, records / rows containing "3" have corresponding matches in dataset 310A. Such records / rows may be deleted from dataset 310B, and the deleted records / rows may be added or appended to dataset 320B.
[0064] It should be understood that the matching logic / algorithm calculation may be performed until all ID fields in dataset 310A have been processed for party 1 and / or until all ID fields in dataset 310B have been processed for party 2. Processing may proceed from block 230 to block 240.
[0065] In block 240 (common part generation), if all ID fields of dataset 310A have been processed, the processor of each device may generate a common part / dataset 320A for party 1. If all ID fields of dataset 310B have been processed, the processor of each device may generate a common part / dataset 320B for party 2.
[0066] It should be understood that the common parts 320A and / or 320B may be used for other MPC processing, such as generating secret shares based on common parts 320A and / or 320B, collecting secret shares, and / or generating results by combining collected secret shares. For details, see the explanations in Figures 4A to 5F.
[0067] Figures 4A and 4B are flowcharts showing the progress of exemplary processing flows 400A and 400B, respectively, for protecting membership and data in secure multi-party computing and communications, according to at least some embodiments described herein.
[0068] Figures 5A to 5F show the progressing portion (500A to 500F) of a schematic diagram illustrating an example of the processing flow of Figures 4A and 4B according to at least some embodiments described herein.
[0069] It should be understood that the processing flows (400A and 400B) disclosed herein may be carried out by one or more processors (e.g., the processors of one or more terminal devices among terminal devices 110, 120, 130, and 140 in Figure 1, the processor of server 150 in Figure 1, the central processor unit 605 in Figure 6, and / or any other suitable processor) unless otherwise specified.
[0070] Furthermore, the processing flows (400A and 400B) may include one or more operations, behaviors, or functions, as indicated by one or more of blocks 405, 410, 415, 420, 425, 430, 435, 440, 445, 450, 455, 460, 460, 465, 470, 475, and 480. These various operations, behaviors, or actions may correspond, for example, to processor-executable software, program code, or program instructions that cause these functions to be performed. Although shown as discrete blocks, obvious modifications may be made depending on the desired implementation, for example, two or more blocks may be rearranged, more blocks may be added, and various blocks may be split into additional blocks, combined into fewer blocks, or removed. It should be understood that operations, including initialization, may be performed before the processing flows (400A and 400B). For example, system parameters and / or application parameters may be initialized. Processing flows (400A and 400B) may be initiated in block 405.
[0071] In block 405 (determining size), the processor may determine a dataset size N (i.e., number) to be used to generate a padding / filled dataset in order to achieve a desired membership privacy protection objective or performance (described in further detail below). It should be understood that the size N is determined to ensure that membership privacy settings and / or privacy requirements are met or satisfied. In embodiments, such membership privacy settings and / or privacy requirements may include settings and / or requirements (described in further detail below) defined in a differential privacy protocol or algorithm. Processing may proceed from block 405 to block 410.
[0072] In block 410 (Padding Set Generation), the processor of each device may provide a dataset for party A (e.g., 502A in Figure 5A) and / or provide a dataset for party B (e.g., 502B). It should be understood that the operations or functions described in the processing flows (400A and 400B) may be symmetric for party A and party B. It should be understood that the format, content and / or arrangement of the datasets described herein are for illustrative purposes only and are not intended to be limiting.
[0073] For example, dataset 502A may have two or more ID fields (ID columns: idA1, idA2, etc.) and / or one or more features or attributes (columns, e.g., T1, etc.) associated with these ID fields. In one exemplary embodiment, ID field idA1 may represent a username, and ID field idA2 may represent an email address.
[0074] For example, dataset 502B may have two or more ID fields (ID columns: idB1, idB2, etc.) and / or one or more features or attributes (columns, e.g., T2, V, etc.) associated with these ID fields. In one exemplary embodiment, ID field idB1 may represent a username, and ID field idB2 may represent an email address.
[0075] For each ID field in dataset 502A (from the first ID field idA1 to the last ID field idA2) and / or 502B (from the first ID field idB1 to the last ID field idB2), the processor may generate the respective fields (e.g., idD1, idD2, etc.) in dataset (e.g., 504A and / or 504B in Figure 5A). The dataset (504A or 504B) may be a padding or fill dataset used or shared in common by both Party A and Party B (e.g., the processor may provide Party B with a local copy 504B of dataset 504A, and Party A with a local copy 504A of dataset 504B). In one exemplary embodiment, each of the datasets (504A, 504B) has a size of 2*N (see description in block 405). In other exemplary embodiments, each of the datasets (504A, 504B) may have a size of N or more.
[0076] It should be understood that the size of a dataset (e.g., 504A or 504B) may refer to the number of records (or rows, elements, etc.) in the dataset (e.g., 504A or 504B). It should also be understood that if each of the datasets (504A, 504B) has a size of 2*N, then subsequent operations, such as PSI or MPC operations on the upsampled datasets (e.g., 508A in Figure 5A for Party A and 508B in Figure 5A for Party B, as described in more detail below), can be guaranteed to be (ε,δ) differentially confidential (as described and / or defined below) for both Party A and / or Party B. In exemplary embodiments, ε and / or δ may be predetermined to achieve desired membership privacy protection objectives or performance.
[0077] Features in embodiments disclosed herein (e.g., a determined size N) may be "(ε,δ)-differentially confidential" with respect to predetermined ε and δ (i.e., "differentially confidential" based on ε and δ). That is, the size N may be determined based on predetermined ε and δ so that it is "(ε,δ)-differentially confidential" with respect to subsequent operations, such as PSI or MPC operations on an upsampled dataset (i.e., the subsequent operations are "differentially confidential" based on ε and δ).
[0078] It should be understood that the above settings or requirements for a differential privacy protocol or algorithm may refer to a measure of "how much data privacy is granted (e.g., by querying on an input dataset) in order to perform an operation or function." The measurable set E may refer to all potential outputs of the predictable M. The first parameter "ε" may refer to the privacy budget (i.e., the limit on how much privacy leakage is acceptable), which represents, for example, the maximum difference between a query on dataset A and the same query on dataset A'. A smaller value of ε indicates stronger privacy protection for the multi-identity privacy protection mechanism. The second parameter "δ" may refer to a probability, such as the probability of information being leaked by chance. In an exemplary embodiment, the required or predetermined range of ε may be about 1 to about 3. The requested or predetermined range of δ may be about 10 -10 (or about 10 -8 ) from about 10 -6 It may also be the case that, in order to achieve, satisfy, or guarantee the requirement of (ε,δ) differential confidentiality, the value of N may be several thousand or approximately several thousand.
[0079] In one exemplary embodiment, the relationship between ε, δ, and N may be determined by a pre-determined or predefined algorithm. That is, the size N may be determined, for example, based on a requested or pre-determined ε and δ according to a pre-calibrated or pre-determined noise distribution, so that "(ε,δ) difference-secretive" can be achieved for subsequent operations, such as PSI or MPC operations on an upsampled dataset.
[0080] Furthermore, it should be understood that the datasets (504A, 504B) are generated such that the intersection (e.g., the result of an inner join) between the ID field (idD1) in dataset (504A or 504B) and its corresponding ID field (idA1 or idB1) in dataset 502A for party A and dataset 502B for party B is empty (i.e., has a size of zero), and the intersection (i.e., has a size of zero) between the ID field (idD2) in dataset (504A or 504B) and its corresponding ID field (idA2 or idB2) in dataset 502A for party A and dataset 502B for party B is empty (i.e., has a size of zero). In other words, there are no common or shared elements between idD1 and idA1 (and / or idD1 and idB1), and there are no common or shared elements between idD2 and idA2 (and / or idD2 and idB2). Processing may proceed from block 410 to block 415.
[0081] In block 415 (Shuffling of Padding Sets), the processor of each device may independently shuffle (e.g., randomly rearrange) each ID field (idD1 and idD2) of the datasets (504A and 504B) for party A and party B to generate the corresponding shuffled dataset for party A (e.g., 506A in Figure 5A) and the corresponding shuffled dataset for party B (e.g., 506B in Figure 5A). Processing may proceed from block 415 to block 420.
[0082] In block 420 (data setup sampling), for each ID field (from the first ID field to the last ID field) in dataset 502A for party A and dataset 502B for party B, the processor of each device may upsample the corresponding ID fields in dataset 502A for party A and / or dataset 502B for party B. It should be understood that upsampling of the corresponding ID fields in dataset 502A may include (1) selecting or retrieving the first N elements (or records, rows, etc.) of each ID field (idD1, idD2) in dataset 506A, (2) generating the union of the corresponding ID fields in dataset 502A and the first N elements of each ID field (idD1, idD2) in dataset 506A (resulting in the corresponding ID fields in dataset 508A shown in Figure 5A), and (3) inserting N random numbers / elements into other fields in dataset 508A that are in the same records / rows as the added / inserted / appended first N elements of each ID field (idD1, idD2) in dataset 506A.
[0083] For example, as shown in Figure 5A, N is determined to be 2 in block 405. For idA1 in dataset 502A, the first N elements (or records, rows, etc.) of the ID field idD1 in dataset 506A are selected or retrieved. The union of the first N elements of the ID field idD1 in dataset 506A and the idA1 field in dataset 502A is generated and becomes the idA1 field in dataset 508A. N random numbers / elements are inserted into each of the other fields (e.g., idA2) in dataset 508A that are in the same record / row as the first N elements added / inserted / appended in the ID field idD1 of dataset 506A. It should be understood that one of these N random numbers / elements will have an empty intersection with any other element in the resulting dataset 508A for party A, and an empty intersection with any element in the upsampled dataset 508B for party B.
[0084] For idA2 in dataset 502A, the first N elements (or records, rows, etc.) of the ID field idD2 in dataset 506A are selected or retrieved. The union of the first N elements of the ID field idD2 in dataset 506A and the idA2 field in dataset 502A (extended by the inserted 1*N random numbers / elements) is generated and becomes the idA2 field in dataset 508A. The N random numbers / elements are inserted into each of the other fields in dataset 508A (e.g., idA1) that are in the same record / row as the first N elements added / inserted / appended in the ID field idD2 in dataset 506A. It should be understood that any one of these N random numbers / elements will have an empty intersection with any other element in dataset 508A resulting from party A, and an empty intersection with any element in the upsampled dataset 508B resulting from party B.
[0085] It should also be understood that the upsampled dataset 508A may be used as dataset 310A in Figure 3. Similarly, the ID fields (idB1, idB2) of dataset 502B for party B may be upsampled using independently shuffled ID fields (idD1, idD2) of dataset 506B to generate an upsampled dataset (e.g., 508B in Figure 5A or 310B in Figure 3).
[0086] Furthermore, it should be understood that for the V field of 508B, all added / inserted / appended elements can be zero (represented by "0"), and therefore for the T1 field of 508A and / or the T2 field of 508B, the added / inserted / appended elements can be any value (represented by "x") without affecting the final statistical result.
[0087] It should be understood that the processor of each device may process the upsampled dataset 508A for Party A and / or the upsampled dataset 508B for Party B to generate a common portion for further processing (without exposing the actual size of the common portion, as padding / filling elements and random numbers / elements are inserted into the upsampled datasets for Party A and / or Party B). Furthermore, by introducing the datasets (504A, 504B) and random numbers / elements for upsampling, the size of the common portion between the upsampled dataset 508A for Party A and the upsampled dataset 508B for Party B does not expose the actual common portion size of the original datasets (e.g., 502A for Party A and 502B for Party B). That is, the features of the embodiments disclosed herein make it possible to make the common portion size exposed in the subsequent PSI or MPC operation random and differentially confidential, making it virtually impossible for an attacker to determine a user's membership based on the size of the common portion.
[0088] As shown in Figure 5A, in one exemplary embodiment, dataset 508A includes multiple records (rows), each record including first member or user identification information (idA1), second member or user identification information (idA2), and a time (T1) indicating the time (e.g., start time or timestamp) when, for example, the member or user clicked a link on Party A's platform. Dataset 508B includes multiple records (rows), each record including first member or user identification information (idB1), second member or user identification information (idB2), a time (T2) indicating the time (e.g., start time or timestamp) when, for example, the member or user interacted with Party B's website, and a value (V) indicating the user's value for Party B. In one exemplary embodiment, the time (or timestamp) is expressed in units of "minutes". It should be understood that the format, content and / or arrangement of datasets 508A and / or 508B are for illustrative purposes only and are not intended to be limiting. For example, each dataset 508A or 508B may have one or more IDs (columns) and / or zero or one or more features or attributes (columns) associated with such one or more IDs.
[0089] In one exemplary embodiment, for various reasons relating to Party A and / or Party B, it may be reasonable to determine, for example, (1) the number of members or users who clicked, for example, a link on Party A's platform, thereby triggering an interaction with Party B's website and having a valuable interaction; (2) the number of members or users who clicked, for example, a link on Party A's platform, thereby triggering an interaction with Party B's website and having a valuable interaction within a certain period (e.g., within 7 minutes) after the user clicked the link on Party A's platform; and / or (3) the total number of all members or users who clicked, for example, a link on Party A's platform, thereby triggering an interaction with Party B's website and having a valuable interaction within a certain period (e.g., within 7 minutes) after the user clicked the link on Party A's platform.
[0090] It should be understood that, for various reasons, Party A and / or Party B may not wish to disclose to the other party at least some of the data in dataset 508A and / or dataset 508B, and / or the data in the common portion of dataset 508A and dataset 508B, respectively. Processing may proceed from block 420 to block 425.
[0091] In block 425 (Shuffling and Transformation), the processor may transform the ID fields (columns, idA1 and idA2) of dataset 508A (to obtain or generate dataset 505A in Figure 5B) using a transformation scheme for party A. It should be understood that the function or operation to "transform" or "transform" one or more columns (or rows) of a dataset, such as one or more identification fields / columns (or records / rows), may also refer to processing (e.g., encrypting, decrypting, coding, decrypting, manipulating, compressing, decompressing, transforming, etc.) the dataset or part thereof. It should be understood that "transformation scheme" may refer to an algorithm, protocol, or function that performs processing (e.g., encrypting, decrypting, coding, decrypting, manipulating, compressing, decompressing, transforming, etc.) of the dataset or part thereof. In one exemplary embodiment, the processor may encrypt (or decrypt, encode, decrypt, manipulate, compress, decompress, convert, etc.) the ID of dataset 508A (to obtain or generate dataset 505A) using, for example, the key of party A, based on an ECDH algorithm or protocol (represented by function D0(.)).
[0092] The processor may also transform the ID fields (idB2 and idB2) of dataset 508B (to obtain or generate dataset 505B in Figure 5B) using a transformation scheme for party B. In one exemplary embodiment, the processor may encrypt (or decrypt, encode, decode, manipulate, compress, decompress, transform, etc.) the IDs of dataset 508B (to obtain or generate dataset 505B) using, for example, party B's key, based on an ECDH algorithm or protocol (represented by function D1(.)).
[0093] The processor may further transform the data T1 (columns, features, attributes) of dataset 508A (to obtain or generate dataset 505A) using a transformation scheme for party A. In one exemplary embodiment, the processor may encrypt (or decrypt, encode, decode, manipulate, compress, decompress, transform, etc.) T1 of dataset 508A (to obtain or generate dataset 505A) using, for example, party A's key, based on an additive homomorphic encryption algorithm or protocol (represented by the function H0(.)).
[0094] The processor may also use a transformation scheme for party B to transform data T2 (columns, features, attributes) and data V (columns, features, attributes) of dataset 510B (in order to obtain or generate dataset 505B). In one exemplary embodiment, the processor may encrypt (or decrypt, encode, decode, manipulate, compress, decompress, transform, etc.) T2 and V of dataset 508B (in order to obtain or generate dataset 505B) using, for example, party B's key, based on an additive homomorphic encryption algorithm or protocol (represented by function H1(.)).
[0095] The processor in each device may shuffle dataset 505A for party A (e.g., by random rearrangement) and / or shuffle dataset 505B for party B.
[0096] In block 425, the order of dataset transformation and dataset shuffling may be switched or changed for party A and / or party B without affecting the purpose of the resulting dataset. For example, the processor may shuffle dataset 508A and then transform the shuffled dataset to obtain or generate dataset 505A for party A. The processor may also shuffle dataset 508B and then transform the shuffled dataset to obtain or generate dataset 505B for party B. Processing may proceed from block 425 to block 430.
[0097] In block 430 (exchange, shuffling, and transformation), the processor of each device may exchange (shuffled) dataset 505A and (shuffled) dataset 505B between party A and party B. For party A, the processor may dispatch or transmit (shuffled) dataset 505A to party B and receive or acquire (shuffled) dataset 505B from party B as dataset 510A (see Figure 5B). For party B, the processor may dispatch or transmit (shuffled) dataset 505B to party A and receive or acquire (shuffled) dataset 505A from party A as dataset 510B (see Figure 5B). It should be understood that because datasets 505A and 505B have already been transformed (e.g., encoded), the corresponding receiving party may not know the actual data in the received dataset.
[0098] The processor may further transform the ID field (idB1) of dataset 510A using a transformation scheme for party A. In one exemplary embodiment, the processor may encrypt (or decrypt, encode, decrypt, manipulate, compress, decompress, convert, etc.) the ID (idB1) of dataset 510A using party A's key based on an ECDH algorithm or protocol (represented by function D0(.)). The processor may further transform the ID field (idA1) of dataset 510B using a transformation scheme for party B. In one exemplary embodiment, the processor may encrypt (or decrypt, encode, decrypt, manipulate, compress, decompress, convert, etc.) the ID field (idA1) of dataset 510B using party B's key based on an ECDH algorithm or protocol (represented by function D1(.)). It should be understood that the results of functions D1(D0(p)) and D0(D1(p)) may be the same for the same parameter "p".
[0099] The processor may also shuffle dataset 510A for party A and / or shuffle dataset 510B for party B. In block 430, the order of transforming the ID field of the dataset and shuffling the dataset may be switched or changed for party A and / or party B without affecting the purpose of the resulting dataset. For example, the processor may shuffle dataset 510A and then transform the shuffled dataset 510A for party A. The processor may also shuffle dataset 510B and then transform dataset 510B for party B. Processing may proceed from block 430 to block 435.
[0100] In block 435 (transformation and matching), the processor of each device may extract the ID field (idA1) of the (shuffled) dataset 510B to obtain or generate dataset 515A for party A, and / or extract the ID field (idB1) of the (shuffled) dataset 510A to obtain or generate dataset 515B for party B. The processor of each device may also exchange the extracted dataset 510A (shuffled idB1 field) with the extracted dataset 510B (shuffled idA1 field) between party A and party B. For party A, the processor may dispatch or send the extracted dataset 510A (shuffled idB1 field) to party B, and receive or obtain the extracted dataset 510B (shuffled idA1 field) from party B as dataset 515A. With respect to Party B, the processor may dispatch or send the extracted dataset 510B (the idA1 field after shuffling) to Party A, and receive or acquire the extracted dataset 510A (the idB1 field after shuffling) from Party A as dataset 515B.
[0101] The processor may also perform a search for matching (or an inner join operation, etc.) between dataset 510A and dataset 515A to obtain or generate a common part (dataset 520A in Figure 5C) for party A. It should be understood that the above operations include, for each identifier in dataset 515A that matches an identifier in dataset 510A, adding or appending the record (or row) from dataset 510A containing the matched identifier to dataset 520A, and removing the record (or row) containing the matched identifier from dataset 510A to obtain or generate the resulting dataset 525A.
[0102] The processor may also perform a search for matching (or an inner join operation, etc.) between dataset 510B and dataset 515B to obtain or generate a common part (dataset 520B in Figure 5C) for party B. It should be understood that the above operations include, for each identifier in dataset 515B that matches an identifier in dataset 510B, adding or appending the record (or row) from dataset 510B containing the matched identifier to dataset 520B, and removing the record (or row) containing the matched identifier from dataset 510B to obtain or generate the resulting dataset 525B.
[0103] In one exemplary embodiment, it should be understood that the idB2 field / common part 520A in the dataset may be optional because the matching is based on idB1 (which has a higher priority than idB2). The idA2 field / common part 520B in the dataset may also be optional because the matching is based on idA1 (which has a higher priority than idA2). Additionally, dataset 525A contains all unmatched records (rows) from dataset 510A. Dataset 525B contains all unmatched records (rows) from dataset 510B.
[0104] It should be understood that for Party A, the data in the common section 520A is also transformed (e.g., encoded) by Party B (via D1(.) and H1(.)), so Party A may not know the actual data in the common section 520A. For Party B, the data in the common section 520B is also transformed (e.g., encoded) by Party A (via D0(.) and H0(.)), so Party B may not know the actual data in the common section 520B. In other words, the matching or inner join operation performed as described above is a "confidential" matching or inner join operation. The processor performs confidential identity matching without revealing the common section of the datasets of these two parties. Processing may proceed from block 435 to block 440.
[0105] In block 440 (transformation, shuffling, and exchange), the processor of each device may transform the ID field (column, idB2) of dataset 525A (to obtain or generate dataset 530A in Figure 5C) using a transformation scheme for party A. In one exemplary embodiment, the processor may encrypt (or decrypt, encode, decode, manipulate, compress, decompress, transform, etc.) the ID field idB2 of dataset 525A (to obtain or generate dataset 530A) using, for example, another key of party A, based on an ECDH algorithm or protocol (represented by function D3(.)).
[0106] The processor may also transform the ID field (idA2) of dataset 525B (to obtain or generate dataset 530B in Figure 5C) using a transformation scheme for party B. In one exemplary embodiment, the processor may encrypt (or decrypt, encode, decode, manipulate, compress, decompress, transform, etc.) the ID field idA2 of dataset 525B (to obtain or generate dataset 530B) using, for example, another key of party B, based on an ECDH algorithm or protocol (represented by function D4(.)).
[0107] The processor in each device may shuffle dataset 530A for party A (e.g., by random rearrangement) and / or shuffle dataset 530B for party B. The processor in each device may also record, save, retain, or otherwise maintain the shuffled rearrangement of dataset 530A and / or the shuffled rearrangement of dataset 530B (in preparation for the unshuffling process in block 445).
[0108] In block 440, the order of dataset transformation and dataset shuffling may be switched or changed for party A and / or party B, without affecting the purpose of the resulting dataset. For example, the processor may shuffle dataset 530A and then transform the shuffled dataset 530A for party A. The processor may also shuffle dataset 530B and then transform the shuffled dataset 530B for party B.
[0109] The processor in each device may exchange (shuffled) dataset 530A and (shuffled) dataset 530B between party A and party B. For party A, the processor may dispatch or transmit (shuffled) dataset 530A to party B and receive or acquire (shuffled) dataset 530B from party B as dataset 535A (see Figure 5C). For party B, the processor may dispatch or transmit (shuffled) dataset 530B to party A and receive or acquire (shuffled) dataset 530A from party A as dataset 535B (see Figure 5C). It should be understood that because datasets 530A and 530B have already been transformed (e.g., encoded), the corresponding receiving party may not know the actual data in the received dataset. Processing may proceed from block 440 to block 445.
[0110] In block 445 (transformation, exchange, and deshuffling), the processor of each device may transform dataset 535A (to obtain or generate dataset 540A in Figure 5D) using a transformation scheme for party A. In one exemplary embodiment, the processor may decrypt (or encrypt, encode, decrypt, manipulate, compress, decompress, transform, etc.) dataset 535A using, for example, party A's key, based on an ECDH algorithm or protocol (represented by function D0(.)), and then encrypt (or decrypt, encode, decrypt, manipulate, compress, decompress, transform, etc.) dataset 535A using, for example, another key of party A, based on an ECDH algorithm or protocol (represented by function D3(.)). That is, dataset 535A is detransformed (e.g., key D0(.) is removed), and then transformed again (key D3(.) is added) to obtain or generate dataset 540A.
[0111] The processor may also transform dataset 535B (to obtain or generate dataset 540B in Figure 5D) using a transformation scheme for party B. In one exemplary embodiment, the processor may decrypt (or encrypt, encode, decrypt, manipulate, compress, decompress, transform, etc.) dataset 535B using, for example, party B's key, based on an ECDH algorithm or protocol (represented by function D1(.)), and then encrypt (or decrypt, encode, decrypt, manipulate, compress, decompress, transform, etc.) dataset 535B using, for example, another key of party B, based on an ECDH algorithm or protocol (represented by function D4(.)). That is, the transformation of dataset 535B is undone (e.g., key D1(.) is removed), and then transformed again (key D4(.) is added) to obtain or generate dataset 540B.
[0112] The processor of each device may exchange dataset 540A and dataset 540B between party A and party B. For party A, the processor may dispatch or transmit dataset 540A to party B and receive or acquire dataset 540B from party B as dataset 545A (see Figure 5D). For party B, the processor may dispatch or transmit dataset 540B to party A and receive or acquire dataset 540A from party A as dataset 545B (see Figure 5D).
[0113] The processor of each device may unshuffle dataset 545A for party A so that the records (rows) in dataset 545A and dataset 530A have the same order or sequence (other than the conversion method used for records / rows), based on the sorting (of the shuffling of dataset 530A) maintained in block 440. The processor of each device may also unshuffle dataset 545B for party B so that the records (rows) in dataset 545B and dataset 530B have the same order or sequence (other than the conversion method used for records / rows), based on the sorting (of the shuffling of dataset 530B) maintained in block 440. Processing may proceed from block 445 to block 450.
[0114] In block 450 (matching and joining), the processor of each device may perform a search for matching (or an inner join operation, etc.) between dataset 540A and dataset 545A to obtain or generate a common part (dataset 550A in Figure 5D) for party A. It should be understood that the above operations include adding or appending the record (or row) from dataset 545A containing the matched identifier for each identifier in dataset 545A that matches an identifier in dataset 540A, and adding or appending the features, attributes, or data (e.g., T2 and V) of the corresponding record (or row) in dataset 525A to dataset 550A. It should be understood that the features, attributes, or data (e.g., T2 and V) of the corresponding record (or row) in dataset 525A are associated with the corresponding ID in dataset 545A, since the records in dataset 545A have the same order or sequence as the records in dataset 530A (extracted from dataset 525A). It should be understood that the results of functions D3(D4(p)) and D4(D3(p)) may be the same for the same parameter "p".
[0115] The processor may also perform a search for matching (or an inner join operation, etc.) between dataset 540B and dataset 545B to obtain or generate a common part (dataset 550B in Figure 5D) for party B. It should be understood that the above operations include adding or appending the record (or row) from dataset 545B containing the matched identifier for each identifier in dataset 545B that matches an identifier in dataset 540B, and adding or appending the feature, attribute, or data (e.g., T1) of the corresponding record (or row) in dataset 525B to dataset 550B. It should be understood that the feature, attribute, or data (e.g., T1) of the corresponding record (or row) in dataset 525B is associated with the corresponding ID in dataset 545B, since the records in dataset 545B have the same order or sequence as the records in dataset 530B (extracted from dataset 525B).
[0116] In one exemplary embodiment, the idB1 field / common portion 550A in the dataset may be optional because the matching is based on idB2 (after idB1 is unmatched in dataset 525A). The idA1 field / common portion 550B in the dataset may be optional because the matching is based on idA2 (after idA1 is unmatched in dataset 525B).
[0117] It should be understood that for Party A, the data in the common portion 550A is also transformed (e.g., encoded) by Party B (via D4(.) and H1(.)), so Party A may not know the actual data in the common portion 550A. For Party B, the data in the common portion 550B is also transformed (e.g., encoded) by Party A (via D3(.) and H0(.)), so Party B may not know the actual data in the common portion 550B. In other words, the matching or inner join operation performed as described above is a "confidential" matching or inner join operation. The processor performs confidential identity matching without revealing the common portion of the datasets of these two parties.
[0118] The processor of each device may, for party A, combine the records / rows of dataset 520A and the records / rows of dataset 550A to obtain or generate dataset 555A. It should be understood that in dataset 555A, a blank value for idB2 indicates that such a value is not important (because the higher-priority ID field idB1 was matched), and a blank value for idB1 indicates that such a value is not important (because the ID field idB1 was not matched, but the ID field idB2 was matched, and the matched records are for the same user / member).
[0119] The processor of each device may also combine the records / rows of dataset 520B and the records / rows of dataset 550B for party B to obtain or generate dataset 555B. It should be understood that in dataset 555B, a blank value for idA2 indicates that such a value is not important (because the higher-priority ID field idA1 was matched), and a blank value for idA1 indicates that such a value is not important (because the ID field idA1 was not matched, but the ID field idA2 was matched, and the matched records are for the same user / member). Processing may proceed from block 450 to block 455.
[0120] In block 455 (share generation), the processor of each device may generate a secret share for party A for each attribute, feature, or data in dataset 555A (e.g., an element that is not an identifier in an ID field / column) to obtain or generate dataset 560A. The processor may generate a secret share for party B for each attribute, feature, or data in dataset 555B in order to obtain or generate dataset 560B.
[0121] In one exemplary embodiment, the processor of each device may generate a corresponding mask and obtain or generate dataset 560A by masking each attribute or feature in dataset 555A with its corresponding mask using a masking scheme. In one exemplary embodiment, each mask is a random number or random plaintext (e.g., having a length of 64 bits). In one exemplary embodiment, the masking scheme is a homomorphic operation or computation (e.g., addition, subtraction, etc.) in an additive homomorphic encryption algorithm or protocol. For example, as shown in Figure 5E, the processor may homomorphically compute T2 data S2(4) in dataset 560A by masking T2 data H1(4) in dataset 555A with a mask (represented as S1(4) in dataset 560A) and subtracting the mask S1(4) from H1(4), where the mask S1(4) is generated for and corresponds to T2 time "4". It should be understood that the combination of the mask S1(4) and the homomorphic result S2(4) is the T2 data H1(4) transformed or encrypted by party B.
[0122] Similarly, for each attribute, feature, or data in dataset 555B for party B, the processor may generate a corresponding mask and obtain or generate dataset 560B by masking each attribute or feature in dataset 555B with its corresponding mask using a masking scheme. In one exemplary embodiment, each mask is a random number or random plaintext (e.g., having a length of 64 bits). In one exemplary embodiment, the masking scheme is a homomorphic operation or computation (e.g., addition, subtraction, etc.) in an additive homomorphic encryption algorithm or protocol. For example, as shown in Figure 5E, the processor may homomorphically compute T1 data S3(3) in dataset 560B by masking T1 data H0(3) in dataset 555B with a mask (represented as S4(3) in dataset 560B) and subtracting the mask S4(3) from H0(3), where the mask S4(3) is generated for and corresponds to T1 time "3". It should be understood that the combination of mask S4(3) and homomorphic result S3(3) is the T1 data H0(3) transformed or encrypted by party A. Processing may proceed from block 455 to block 460.
[0123] In block 460 (share exchange), the processor of each device may exchange dataset 565A and dataset 565B between party A and party B. For party A, the processor may dispatch or send dataset 565B (homomorphic result of features or attributes) to party B. For party B, the processor may dispatch or send dataset 565A (homomorphic result of features or attributes) to party A. For party A, the processor may also detransform or decrypt dataset 565A (homomorphic result of features or attributes) using party A's key based on an additive homomorphic encryption algorithm or protocol, for example, to remove key H0(.). For party B, the processor may also detransform or decrypt dataset 565B (homomorphic result of features or attributes) using party B's key based on an additive homomorphic encryption algorithm or protocol, for example, to remove key H1(.). Processing may proceed from block 460 to block 465.
[0124] In block 465 (share construction), the processor may construct a secret share (dataset 570A) about party A by combining the masks of dataset 565A and dataset 560A (for example, by performing a union operation). The processor may construct a secret share (dataset 570B) about party B by combining the masks of dataset 565B and dataset 560B (for example, by performing a union operation). It should be understood that the datasets (570A, 570B) contain records or elements represented by random numbers (for example, a mask that is a random number, or the result that remains a random number after subtracting or adding the random number mask to an actual value) without transformation (for example, encryption). Furthermore, it should be understood that combinations of datasets or secret shares (570A and 570B, 575A and 575B, 580A and 580B, 585A and 585B, 590A and 590B, or 595A and 595B) may result in actual data (with or without noise). See the explanation in block 40 for details. Processing may proceed from block 465 to block 470.
[0125] In block 470 (Execution of Secret MPC), the processor of each device may perform a secret multiparty computation (see description below) on Party A's secret share and / or perform a secret multiparty computation (see description below) on Party B's secret share.
[0126] In one exemplary embodiment, the processor may obtain or generate dataset 575A for party A by subtracting T1 from T2 for dataset 570A, and / or obtain or generate dataset 575B for party B by subtracting T1 from T2 for dataset 570B.
[0127] In one exemplary embodiment, the processor may determine whether T2 is greater than 0 and less than a predetermined value (e.g., less than 7) for the dataset 575A, and then obtain or generate the dataset 580A for party A. If T2 is greater than 0 and less than the predetermined value, the processor may set the Flag value in the dataset 580A to a secret share of 1, which is a random number for party A (S1(1), representing "true"). If T2 is less than or equal to 0, or greater than or equal to the predetermined value, the processor may set the Flag value in the dataset 580A to a secret share of 0, which is a random number for party A (S1(0), representing "false").
[0128] In one exemplary embodiment, the processor may determine whether T2 is greater than 0 and less than a predetermined value (e.g., less than 7) for dataset 575B, and then obtain or generate dataset 580B for party B. If T2 is greater than 0 and less than the predetermined value, the processor may set the Flag value in dataset 580B to a secret share of 1, which is a random number for party B (S2(1), representing "true"). If T2 is less than or equal to 0, or greater than or equal to the predetermined value, the processor may set the Flag value in dataset 580B to a secret share of 0, which is a random number for party B (S2(0), representing "false").
[0129] In one exemplary embodiment, the determination of whether T2 is greater than 0 and less than a predetermined value for Party A and / or Party B may be made, for example, via a lost-communication algorithm or protocol (or secret-comparison algorithm or protocol) based on a lost-communication algorithm or protocol. For example, the lost-comparison algorithm or protocol may receive the T2 field of dataset 575A as input from Party A and the T2 field of dataset 575B as another input from Party B, and generate an output (the result of the T2 determination) to Party A and / or Party B.
[0130] In one exemplary embodiment, the processor may further sum the "V" fields of dataset 580A for records / rows that do not have a secret share of 0 in either the flag field or the V field (i.e., S1(0)), and store or save the result in the "Sum" field of dataset 585A to generate dataset 585A for party A. The processor may also sum the "V" fields of dataset 580B for records / rows that do not have a secret share of 0 in either the flag field or the V field (i.e., S2(0)), and store or save the result in the "Sum" field of dataset 585B to generate dataset 585B for party B.
[0131] At this point, it should be understood that Party A possesses a dataset 585A (Secret Share) that represents the total number of users who clicked, for example, a link on Party A's platform, navigated to Party B's website within a certain period (e.g., within 7 minutes) after clicking the link on Party A's platform, and engaged in a valuable interaction. It should also be understood that, because the Secret Share is a random value, Party A does not know the actual data within dataset 585A.
[0132] At this point, it should be understood that Party B has a dataset 585B (secret share) that represents the total number of users who clicked, for example, a link on Party A's platform, and then, within a certain period (e.g., within 7 minutes) after clicking the link on Party A's platform, navigated to Party B's website and performed a valuable interaction. It should also be understood that, because the secret share is a random value, Party B does not know the actual data in dataset 585B. Processing may proceed from block 470 to block 475. Furthermore, it should be understood that combinations of the secret shares (585A and 585B) can result in the actual data ("19").
[0133] In block 475 (noise generation), the processor of each device may execute or perform a lost communication algorithm or protocol to jointly generate a noise share with Party A and Party B. In one exemplary embodiment, Party A's processor may generate two messages (M0, M1) based, for example, randomly generated noise data and / or parameters of a differential privacy algorithm or protocol, and Party B's processor may generate random bits indicating which message may be received from Party A. The processor of each device may also perform 1-2 lost communication with the messages (M0, M1) and the random bits as input to generate shares of noise R (R1 and R2) for Party A and Party B, respectively. For example, the processor of each device may generate a share of noise R R1 for Party A and / or a share of noise R R2 for Party B. In one exemplary embodiment, noise R may be constructed, for example, by adding the noise shares (R1 and R2, etc.).
[0134] The processor in each device may add a noise share (e.g., R1) to a data share (e.g., a secret share in dataset 585A in Figure 5F) to generate a data share with the added noise for party A (e.g., a secret share in 590A). The processor in each device may add a corresponding noise share (e.g., R2) to a data share (e.g., a secret share in dataset 585B) to generate a data share with the added noise for party B (e.g., a secret share in 590B). Due to the lost communication occurring in block 475, party A may not know party B's noise share R1, and party B may not know party A's noise share R1. In one exemplary embodiment, the process in block 475 may be optional. Processing may proceed from block 475 to block 480.
[0135] In block 480 (Result Construction), the processor of each device may exchange dataset 585A and dataset 585B between Party A and Party B. For Party A, the processor may dispatch or send dataset 585A to Party B and / or receive or acquire dataset 585B from Party B. The processor may also construct the result ("19") by, for example, adding the data in dataset 585A and the data in the received dataset 585B to generate dataset 595A. That is, the total value of all users who clicked, for example, a link on Party A's platform, and then navigated to Party B's website within a certain period (e.g., within 7 minutes) after clicking the link on Party A's platform and performed a valuable interaction is "19". In one exemplary embodiment, if the process in block 475 is not optional, datasets 590A and 590B are used (instead of 585A and 585B) to construct a result in which the actual result (e.g., value 19) is supplemented with noise R jointly generated by party A (via R1) and by party B (via R2).
[0136] With respect to Party B, the processor may dispatch or send dataset 585B to Party A's processor and / or receive or acquire dataset 585A from Party A. The processor may also construct the result ("19") by, for example, adding the data in dataset 585B to the data in the received dataset 585A. That is, the sum of all users who clicked, for example, a link on Party A's platform, navigated to Party B's website within a certain period (e.g., within 7 minutes) after clicking the link on Party A's platform, and performed a valuable interaction is "19", which is the same result as determined by Party A. In one exemplary embodiment, if the process in block 475 is not optional, datasets 590A and 590B (instead of 585A and 585B) are used to construct a result which is the actual result (e.g., the value 19) plus noise R. It should be understood that after constructing the results (through the process in block 475), either party A or party B, or both parties A and B, can obtain the same result (for example, the actual value 19 plus noise R).
[0137] Furthermore, it should be understood that other results can be constructed or determined by combining secret shares (570A and 570B), secret shares (575A and 575B), secret shares (580A and 580B), secret shares (585A and 585B), secret shares (590A and 590B), etc. It should also be understood that other results can be further constructed or determined by performing other MPC calculations on secret shares to obtain the secret shares of Party A and Party B, and by combining the secret shares of both Party A and Party B.
[0138] Figure 6 is a schematic diagram of an exemplary computer system 600 applicable to realizing an electronic device (e.g., one of the servers or terminal devices shown in Figure 1), arranged according to at least some embodiments described herein. It should be understood that the computer system shown in Figure 6 is provided for illustrative purposes only and does not limit the functions and applications of the embodiments described herein.
[0139] As shown in the figure, the computer system 600 may include a central processing unit (CPU) 605. The CPU 605 may perform various operations and processes based on programs stored in read-only memory (ROM) 610 or programs loaded from storage device 640 into random access memory (RAM) 615. The RAM 615 may also store various data and programs required for the operation of the system 600. The CPU 605, ROM 610, and RAM 615 may be connected to each other via a bus 620. An input / output (I / O) interface 625 may also be connected to the bus 620.
[0140] The components connected to the I / O interface 625 may further include an input device 630, such as a keyboard, mouse, digital pen, or drawing pad; an output device 635, such as a display like a liquid crystal display (LCD) or a speaker; a storage device 640, such as a hard disk; and a communication device 645, such as a network interface card like a LAN card or a modem. The communication device 645 may perform communication processing via a network, such as the Internet, WAN, LAN, LINE, or cloud. In one embodiment, a driver 650 may also be connected to the I / O interface 625. A removable medium 655, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, may be mounted to the driver 650 as needed so that a computer program read from the removable medium 655 may be installed in the storage device 640.
[0141] It should be understood that the processes described with reference to the flowcharts in Figures 2, 4A and 4B and / or the processes described in other figures may be implemented as a computer software program or in hardware. The computer program product may include a computer program stored on a computer-readable non-volatile medium. The computer program includes program code for performing the methods shown in the flowcharts and / or GUI. In this embodiment, the computer program may be downloaded and installed from a network via the communication device 645, or it may be installed from a removable medium 655. When the computer program is executed by the central processing unit (CPU) 605, it can perform the functions defined in the methods in the embodiments disclosed herein.
[0142] It should be understood that the disclosed and other solutions, examples, embodiments, modules, and functional operations described herein may be implemented within digital electronic circuits, or within computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or within one or more combinations thereof. The disclosed embodiments and other embodiments may be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium, to be executed by or to control the operation of a data processing device. The computer-readable medium may be a machine-readable storage device, a machine-readable storage board, a memory device, a composition of a material that affects machine-readable propagating signals, or one or more combinations thereof. "Data processing device" includes all equipment, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the device may include code that generates the execution environment of the computer program being discussed, such as processor firmware, a protocol stack, a database management system, an operating system, or code that constitutes one or more combinations thereof.
[0143] Computer programs (also referred to as programs, software, software applications, scripts, or code) may be written in any form of programming language, including compiled or interpreted languages, and may be deployed in any form, as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. Computer programs do not necessarily correspond to files in a file system. A program may be stored in part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program, or in a group of collaborative files (e.g., a file containing one or more modules, subprograms, or parts of code). Computer programs may be deployed to run on one computer, located in one site, or distributed across multiple sites and interconnected by a communication network.
[0144] The processing and logic flows described herein can perform their functions by manipulating input data and generating outputs, which are executed by one or more programmable processors running one or more computer programs. The processing and logic flows may also be executed by dedicated logic circuits, such as field-programmable gate arrays and application-specific integrated circuits, and devices may also be implemented as such.
[0145] Processors suitable for executing computer programs include, for example, both general-purpose microprocessors and dedicated microprocessors, and any one or more processors of any type of digital computer. Generally, a processor receives instructions and data from read-only memory or random-access memory or both. Essential elements of a computer are a processor for executing instructions and one or more storage devices for storing instructions and data. Generally, a computer also includes or is operablely coupled to one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks or optical disks, to receive or transfer data or both. However, a computer is not required to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including, for example, erasable programmable read-only memory, electrically erasable programmable read-only memory and semiconductor memory devices such as flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks and compact disk read-only memory and digital video disk read-only memory disks. The processor and memory may be complemented by dedicated logic circuits, or they may be incorporated into dedicated logic circuits.
[0146] It should be understood that different features, variations, and multiple different embodiments are illustrated and described in various details. In this application, what is described with respect to a particular embodiment is done for illustrative purposes only and is not intended to limit or suggest that what has been devised is only one specific embodiment or particular embodiment. It should be understood that this disclosure is not limited to any single specific embodiment or enumerated variation. A person skilled in the art will conceive of many modifications, variations, and other embodiments which are intended and actually covered by this disclosure. The scope of this disclosure is actually intended to be determined by the appropriate legal interpretation and structure of the disclosure, including equivalents, as a person skilled in the art will understand by relying on the complete disclosure available at the time of filing.
[0147] Pattern:
[0148] It will be understood that any one of the embodiments may be combined.
[0149] Appearance 1, A method for protecting membership and data in confidential multi-party computation and communication, the method comprising: generating a padding dataset whose size is determined based on data privacy settings; upsampling a first dataset using the padding dataset; transforming the first dataset; dispatching the first dataset; generating a third dataset by performing an intersection operation on the first dataset and a second dataset; generating a first share based on the third dataset; and constructing a result based on the first share and the second share.
[0150] Appearance 2, A method according to Embodiment 1, further comprising shuffling the padding dataset before upsampling the first dataset using the padding dataset.
[0151] Appearance 3, A method according to Embodiment 1 or Embodiment 2, further comprising shuffling a first identification information field of the second dataset and recording the order of the shuffling of the first identification information field.
[0152] Appearance 4, A method according to embodiment 3, further comprising receiving a second identification information field and unshuffling the second identification information field based on a recorded sorting.
[0153] Appearance 5, A method according to any one of embodiments 1 to 4, wherein the intersection of the padding dataset and the first dataset is empty.
[0154] Appearance 6, The method according to any one of embodiments 1 to 5, wherein the first dataset includes a first identification information field and a second identification information field, the first identification information field having a higher priority than the second identification information field, and the second dataset includes a third identification information field and a fourth identification information field, the third identification information field having a higher priority than the fourth identification information field.
[0155] Appearance 7, A method according to embodiment 6, further comprising, for each piece of identification information in the first identification information field that matches the identification information in the third identification information field, deleting the row having the matched identification information from the first dataset and adding the deleted row to the first intersection.
[0156] Appearance 8, A method according to embodiment 7, further comprising, for each piece of identification information in the second identification information field that matches the identification information in the fourth identification information field, deleting the row having the matched identification information from the first dataset and adding the deleted row to the first intersection.
[0157] Appearance 9, A method according to any one of embodiments 1 to 8, wherein the data privacy setting includes a first parameter and a second parameter, and the size of the padding dataset is determined based on the first parameter and the second parameter such that the intersection operation is differentially confidential.
[0158] Appearance 10, The method according to embodiment 9, wherein the size of the padding dataset is determined based on the number of identification information fields of the first dataset.
[0159] Embodiment 11, The method according to embodiment 10, wherein the size of the padding dataset is further determined based on the number of intersection operations.
[0160] Appearance 12, A method according to any one of embodiments 1 to 11, further comprising: performing lost communication to generate first noise data; and applying the first noise data to the first share.
[0161] Embodiment 13, A method according to embodiment 12, wherein the execution of the lost communication includes generating second noise data, and the method further includes applying the second noise data to the second share.
[0162] Appearance 14, A secure multi-party computing and communication system comprising: a memory for storing a first dataset; and a processor for generating a padding dataset whose size is determined based on data privacy settings; upsampling the first dataset using the padding dataset; transforming the first dataset; dispatching the first dataset; performing an intersection operation on the first dataset and the second dataset to generate a third dataset; generating a first share based on the third dataset; and constructing a result based on the first share and the second share.
[0163] Appearance 15, A system according to embodiment 14, wherein the processor further shuffles the padding dataset before upsampling the first dataset using the padding dataset.
[0164] Appearance 16, A system according to Embodiment 14 or Embodiment 15, wherein the processor further shuffles the first identification information field of the second dataset, records the shuffled order of the first identification information field, receives the second identification information field, and unshuffles the second identification information field based on the recorded order.
[0165] Appearance 17, A non-temporary computer-readable medium storing computer-executable instructions, wherein when an instruction is executed, one or more processors perform operations including: generating a padding dataset whose size is determined based on data privacy settings; upsampling a first dataset using the padding dataset; transforming the first dataset; dispatching the first dataset; generating a third dataset by performing an intersection operation on the first dataset and a second dataset; generating a first share based on the third dataset; and constructing a result based on the first share and the second share.
[0166] Appearance 18, A computer-readable medium according to embodiment 17, wherein the operation further includes shuffling the padding dataset before upsampling the first dataset using the padding dataset.
[0167] Appearance 19, A computer-readable medium according to embodiment 17 or embodiment 18, wherein the operation further includes shuffling a first identification information field of the second dataset and recording the order of the shuffling of the first identification information field.
[0168] Appearance 20, A computer-readable medium according to embodiment 19, wherein the operation further includes receiving a second identification information field and unshuffling the second identification information field based on a recorded sorting.
[0169] The terms used herein are intended to describe, and not limit, specific embodiments. The terms “one,” “one,” and “the” include the plural form unless expressly indicated. The terms “includes” and / or “equipment,” as used herein, presuppose the presence of the described features, integers, steps, operations, elements, and / or components, but do not presuppose the presence or addition of one or more other features, integers, steps, operations, elements, and / or components.
[0170] It should be understood that, with respect to the above description, modifications may be made to details, particularly the materials used, shapes, sizes, and arrangements of components, without departing from the scope of this disclosure. The embodiments described herein and described herein are illustrative only, and the true scope and essence of this disclosure are given by the following claims.
Claims
1. A method for protecting membership and data in secure multi-party computing and communications, This involves generating a padding dataset whose size is determined based on data privacy settings, Upsampling the first dataset using the aforementioned padding dataset, Converting the first dataset described above, Dispatching the first dataset, The process involves generating a third dataset by performing an intersection operation based on the first dataset and the second dataset, To generate a first share based on the third dataset, Constructing results based on the first share and the second share, A method that includes this.
2. Before upsampling the first dataset using the padding dataset, shuffling the padding dataset, The method according to claim 1, further comprising:
3. Shuffling the first identification information field of the second dataset, Record the shuffling order of the first identification information field, The method according to claim 1, further comprising:
4. Receiving a second identification information field, Based on the recorded sorting, the shuffling of the second identification information field is unshrunk, The method according to claim 3, further comprising:
5. The intersection of the padding dataset and the first dataset is empty. The method according to claim 1.
6. The first dataset includes a first identification information field and a second identification information field, wherein the first identification information field has a higher priority than the first and second identification information fields. The second dataset includes a third identification information field and a fourth identification information field, wherein the third identification information field has a higher priority than the fourth identification information field. The method according to claim 1.
7. For each piece of identification information in the first identification information field that matches the identification information in the third identification information field, the row containing the matched identification information is deleted from the first dataset, and the deleted row is added to the first common part. The method according to claim 6, further comprising:
8. For each piece of identification information in the second identification information field that matches the identification information in the fourth identification information field, the row containing the matched identification information is deleted from the first dataset, and the deleted row is added to the first common part. The method according to claim 7, further comprising:
9. The aforementioned data privacy setting includes a first parameter and a second parameter, The size of the padding dataset is determined based on the first and second parameters such that the intersection operation is differentially secret. The method according to claim 1.
10. The size of the padding dataset is determined based on the number of identification information fields in the first dataset. The method according to claim 9.
11. The size of the padding dataset is further determined based on the number of intersection operations. The method according to claim 10.
12. Performing lost communication to generate the first noise data, Applying the first noise data to the first share, The method according to claim 1, further comprising:
13. The execution of the aforementioned lost communication includes generating second noise data, The method further includes applying the second noise data to the second share. The method according to claim 12.
14. A secure multi-party computing and communication system, Memory for storing the first dataset, It is a processor, Generate a padding dataset whose size is determined based on data privacy settings. The first dataset is upsampled using the padding dataset. The first dataset described above is transformed, The first dataset is dispatched, A third dataset is generated by performing an intersection operation on the first dataset and the second dataset. Based on the third dataset, a first share is generated. The results are constructed based on the first and second shares mentioned above. Processor and A system equipped with these features.
15. The aforementioned processor further, Before upsampling the first dataset using the padding dataset, the padding dataset is shuffled. The system according to claim 14.
16. The aforementioned processor further, The first identification information field of the second dataset is shuffled, The shuffling order of the first identification information field is recorded, Upon receiving the second identification information field, Based on the recorded sorting, the shuffling of the second identification information field is unshrunk. The system according to claim 14.
17. A non-temporary computer-readable medium in which computer-executable instructions are stored, wherein when the instructions are executed, one or more processors are configured to: This involves generating a padding dataset whose size is determined based on data privacy settings, Upsampling the first dataset using the aforementioned padding dataset, Converting the first dataset described above, Dispatching the first dataset, The process involves generating a third dataset by performing an intersection operation based on the first dataset and the second dataset, To generate a first share based on the third dataset, Constructing results based on the first share and the second share, A computer-readable medium that allows the execution of operations including [specific actions].
18. The aforementioned operation is, The further step includes shuffling the padding dataset before upsampling the first dataset using the padding dataset. The computer-readable medium according to claim 17.
19. The aforementioned operation is, Shuffling the first identification information field of the second dataset, The further includes recording the shuffling order of the first identification information field. The computer-readable medium according to claim 17.
20. The aforementioned operation is, Receiving a second identification information field, The further includes unshuffling the second identification information field based on the recorded sorting. The computer-readable medium according to claim 19.