Data processing method and device and electronic equipment
Through the packet streaming sharding algorithm and multi-phase modulus index algorithm, the dynamic privacy set interception of RSA blind signature and pre-computed dynamic privacy set interception is solved, and the problem of inefficient computing efficiency in the existing technology is realized, efficient and fast privacy set interception is protected, and data privacy is protected.
Patent Information
- Application Number
- CN202410058076.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-22
AI Technical Summary
The existing privacy set interception method is not efficient enough, especially in large-scale data interaction scenarios, resulting in high computing overhead.
Modal index calculation is performed using packet streaming sharding algorithm and multi-phase modular index algorithm. The data security and privacy interaction between the server and the client is realized through RSA blind signature technology, and information search and interception is used by Bloom filter.
It greatly improves the computing efficiency, reduces the calculation time of modulus index, optimizes the overall running time of privacy set interception, protects the non-intersection privacy data of both parties from leaking, and realizes efficient and fast privacy set interception.
Smart Images

Figure CN120354440A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a data processing method, apparatus, and electronic device. Background Art
[0002] With the continuous development of computer technologies, different users can interact data through various devices. For example, individual users interact data through terminal devices such as smart phones and personal computers, and enterprise users interact data through devices such as servers. In the process of data interaction, issues of data privacy and data security are often involved. In scenarios such as social software for finding friends, antivirus software, and finding websites with leaked passwords, both users have a need for data interaction, but there may be a situation where one party does not want to expose its own data to the other party. In the prior art, the privacy set intersection method is usually adopted to protect the privacy of users while satisfying the interaction of data. However, the existing privacy set intersection methods have the problem of low computational efficiency. Summary of the Invention
[0003] In view of this, the present disclosure provides a data processing method, apparatus, electronic device, storage medium, and computer program product.
[0004] According to an aspect of the present disclosure, there is provided a data processing method applied to a server, including:
[0005] Calculating the modular exponentiation of each data in a first data set under an RSA private key based on a preset algorithm; wherein, the first data set includes user data for which the server is to perform a privacy set intersection; the preset algorithm is a grouped streaming sharding algorithm or a multi - base modular exponentiation algorithm, the grouped streaming sharding algorithm is used to evenly group multiple triple data and perform parallel processing on different groups obtained by the even grouping; the triple data includes target data, target exponent, and target modulus; the multi - base modular exponentiation algorithm is used to evaluate modular exponentiation based on a multi - base target exponent;
[0006] Inserting the modular exponentiation of each data under the RSA private key into a Bloom filter and sending the Bloom filter;
[0007] Receive a third data set, where the third data set includes the modular exponentiation product of each data in the second data set calculated based on the grouped streaming sharding algorithm and the corresponding random number in the random number set under the RSA public key, or an exponent value retrieval table of the modular exponentiation of each random number in the random number set constructed based on the multi - base modular exponentiation algorithm; wherein, the second data set includes the user data for which the client is to perform private set intersection, the random number set is generated by the client, and each random number corresponds one - to - one with each data in the second data set;
[0008] Calculate the modular exponentiation of each data in the third data set under the RSA private key based on the preset algorithm to obtain a fourth data set;
[0009] Send the fourth data set.
[0010] In a possible implementation manner, the averaging and grouping of multiple triple data includes:
[0011] Calculate the indexes of multiple groups corresponding to the target data in the hash table, and determine the target group with the least current stored data volume among the multiple groups;
[0012] Store the triple data corresponding to the target data in the target group;
[0013] The parallel processing of different groups obtained by the averaging and grouping includes:
[0014] Divide the target group into multiple shards;
[0015] Perform parallel processing on the multiple shards.
[0016] In a possible implementation manner, the modular exponentiation evaluation based on the multi - base target exponent includes:
[0017] Divide the target exponent according to a preset multi - base;
[0018] Construct an exponent value retrieval table of the target data set;
[0019] Use the divided target exponent to perform modular exponentiation evaluation based on the exponent value retrieval table of the target.
[0020] In a possible implementation manner, when the preset algorithm is the grouped streaming sharding algorithm,
[0021] The method further includes:
[0022] Obtain the data that has changed in the first data set, and calculate the modular exponentiation of the changed data under the RSA private key based on the preset algorithm;
[0023] Determine the target position of the modular exponent of the changed data under the RSA private key in the Bloom filter;
[0024] Send the target position and the modular exponent of the changed data under the RSA private key.
[0025] According to another aspect of the present disclosure, a data processing method is provided, which is applied to a client and includes: generating a set of random numbers, and calculating the modular exponent of each random number in the set of random numbers under the RSA public key and the modular inverse of each random number based on a preset algorithm; wherein, each random number corresponds one-to-one with each data in the second data set, and the second data set includes user data for which the client is to perform private set intersection; the preset algorithm is a grouped streaming sharding algorithm or a multi - base modular exponent algorithm, the grouped streaming sharding algorithm is used to evenly group multiple triple data and perform parallel processing on different groups obtained by the even grouping; the triple data includes target data, target exponent, and target modulus; the multi - base modular exponent algorithm is used to evaluate the modular exponent based on the multi - base target exponent;
[0026] Receive a Bloom filter; the Bloom filter stores the modular exponent of each data in the first data set under the RSA private key; the first data set includes user data for which the server is to perform private set intersection;
[0027] Send a third data set to the server; the third data set includes the modular exponent of the product of the modular exponents of each data in the second data set and the corresponding random number in the set of random numbers under the RSA public key calculated based on the grouped streaming sharding algorithm, or an index value table of the modular exponents of each random number in the set of random numbers under the RSA public key constructed based on the multi - base modular exponent algorithm;
[0028] Receive a fourth data set, the fourth data set includes the modular exponent of each data in the third data set calculated based on the preset algorithm under the RSA private key;
[0029] Judge whether the data in the fourth data set is in the Bloom filter to complete the private set intersection processing.
[0030] In a possible implementation manner, the evenly grouping of the multiple triple data includes:
[0031] Calculate the indexes of multiple groups corresponding to the target data in the hash table, and determine the target group with the least amount of currently stored data among the multiple groups;
[0032] Store the triple data corresponding to the target data in the target group;
[0033] Performing parallel processing on the different groups obtained by the average grouping includes:
[0034] Dividing the target group into multiple shards;
[0035] Performing parallel processing on the multiple shards.
[0036] In a possible implementation manner, the modulo exponentiation evaluation based on the multi - base target exponent includes:
[0037] Dividing the target exponent according to a preset multi - base;
[0038] Constructing a target exponent value retrieval table for the target data set;
[0039] Using the divided target exponent, performing modulo exponentiation evaluation based on the target exponent value retrieval table.
[0040] In a possible implementation manner, when the preset algorithm is the grouped streaming sharding algorithm,
[0041] Judging whether the data in the fourth data set is in the Bloom filter includes:
[0042] Calculating the modulo exponent of the product of the modulo inverses of each data in the fourth data set and the corresponding random numbers of each data in the random number set based on the grouped streaming sharding algorithm;
[0043] Judging whether the modulo exponent of the product is in the Bloom filter;
[0044] Or,
[0045] The method further includes:
[0046] Obtaining the modulo exponent of the changed data in the first data set under the RSA private key and the target position in the Bloom filter;
[0047] Modifying the value at the target position in the Bloom filter.
[0048] According to another aspect of the present disclosure, there is provided a data processing apparatus, the apparatus comprising: a first computing module, configured to: calculate the modular exponentiation of each data in the first data set under the RSA private key based on a preset algorithm; wherein, the first data set includes user data for which the server is to perform private set intersection; the preset algorithm is a grouped streaming sharding algorithm or a multi - radix modular exponentiation algorithm, the grouped streaming sharding algorithm is used to evenly group multiple triple data and perform parallel processing on different groups obtained by the even grouping; the triple data includes a target data, a target exponent, and a target modulus; the multi - radix modular exponentiation algorithm is used to evaluate modular exponentiation based on a multi - radix target exponent; a first communication module, configured to insert the modular exponentiation of each data under the RSA private key into a Bloom filter and send the Bloom filter; the first communication module is further configured to receive a third data set, the third data set includes the modular exponentiation of the product of each data in the second data set calculated based on the grouped streaming sharding algorithm and the corresponding random number in the random number set under the RSA public key, or an exponential value retrieval table of the modular exponentiation of each random number in the random number set under the RSA public key constructed based on the multi - radix modular exponentiation algorithm; wherein, the second data set includes user data for which the client is to perform private set intersection, the random number set is generated by the client, and each random number corresponds one - to - one with each data in the second data set; the first computing module is further configured to calculate the modular exponentiation of each data in the third data set under the RSA private key based on the preset algorithm to obtain a fourth data set; the first communication module is further configured to send the fourth data set;
[0049] Alternatively, the apparatus includes: a second computing module, configured to generate the random number set and calculate the modular exponentiation of each random number in the random number set under the RSA public key and the modular inverse of each random number based on the preset algorithm; a second communication module, configured to receive the Bloom filter; the Bloom filter stores the modular exponentiation of each data in the first data set under the RSA private key; the second communication module is further configured to send the third data set to the server; the second communication module is further configured to receive the fourth data set, the fourth data set includes the modular exponentiation of each data in the third data set calculated based on the preset algorithm; the second computing module is further configured to determine whether the data in the fourth data set is in the Bloom filter to complete the private set intersection process.
[0050] According to another aspect of the present disclosure, there is provided an electronic device, comprising: a processor; a memory for storing processor - executable instructions; wherein, the processor is configured to implement the above - mentioned data processing method when executing the instructions stored in the memory.
[0051] According to another aspect of the present disclosure, there is provided a computer-readable storage medium having computer program instructions stored thereon, wherein when the computer program instructions are executed by a processor, the above-mentioned data processing method is implemented.
[0052] According to another aspect of the present disclosure, there is provided a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above-mentioned data processing method.
[0053] Through various aspects in the embodiments of the present disclosure, based on the dynamic private set intersection algorithm of RSA blind signature and precomputation, a preset algorithm is used for modular exponentiation operations; among them, based on the RSA blind signature technology, secure privacy interaction between the server and the client (such as Bloom filters, third data sets, fourth data sets, etc.) is realized, and the Bloom filter is used to implement information search and intersection on the client side; during the process of RSA blind signature and precomputation, a preset algorithm is used for modular exponentiation operations (such as calculating the modular exponent of each data in the first data set under the RSA private key, calculating the modular exponent of each data in the third data set under the RSA private key, calculating the modular exponent of each random number in the random number set under the RSA public key, etc.), which greatly improves the operation efficiency; for example, the grouped streaming sharding algorithm is used to calculate the modular exponentiation of a large amount of data involved by both parties (such as the client and the server), effectively reducing the modular exponentiation calculation time of both parties in each stage, optimizing the overall running time of the private set intersection, and thus realizing efficient and fast private set intersection; for another example, the multi-radix modular exponentiation algorithm is used for fast calculation of the modular exponentiation of a large amount of data involved by both parties, greatly saving the time for both parties to perform calculations and interactions online, and more effectively protecting the non-disclosure of private data in the non-intersection sets of both parties, thereby realizing efficient and fast private set intersection.
[0054] According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. Brief Description of the Drawings
[0055] The accompanying drawings, which are included in and constitute a part of this specification, illustrate exemplary embodiments, features, and aspects of the present disclosure together with the specification, and are used to explain the principles of the present disclosure.
[0056] Figure 1 The structural schematic diagram of a private set intersection system according to an embodiment of the present disclosure is shown.
[0057] Figure 2 The flowchart of a data processing method according to an embodiment of the present disclosure is shown.
[0058] Figure 3 The flowchart of a data processing method according to an embodiment of the present disclosure is shown.
[0059] Figure 4 The flowchart of a data processing method according to an embodiment of the present disclosure is shown.
[0060] Figure 5 The flowchart of a method for optimizing the intersection of dynamic private sets based on RSA blind signature and pre - calculation through a grouped streaming sharding algorithm according to an embodiment of the present disclosure is shown.
[0061] Figure 6 The flowchart of a data processing method according to an embodiment of the present disclosure is shown.
[0062] Figure 7 The flowchart of a method for optimizing the intersection of dynamic private sets based on RSA blind signature and pre - calculation through a multi - base modular exponentiation algorithm according to an embodiment of the present disclosure is shown.
[0063] Figure 8 All show the flowchart of a method for optimizing the intersection of dynamic private sets based on RSA blind signature and pre - calculation through a multi - base modular exponentiation algorithm according to an embodiment of the present disclosure.
[0064] Figure 9 The structural diagram of a data processing device according to an embodiment of the present disclosure is shown.
[0065] Figure 10 The structural diagram of a data processing device according to an embodiment of the present disclosure is shown.
[0066] Figure 11 The schematic structural diagram of an electronic device according to an embodiment of the present disclosure is shown. Detailed implementation manners
[0067] The following will describe various exemplary embodiments, features, and aspects of the present disclosure in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0068] References to "one embodiment" or "some embodiments" etc. described in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present disclosure. Thus, statements such as "exemplary", "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.
[0069] In the present disclosure, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may mean: including the case where A exists alone, where A and B exist simultaneously, and where B exists alone, where A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (item)" or a similar expression thereof refers to any combination of these items, including any combination of a single item or plural items. For example, at least one (item) of a, b, or c may mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or plural.
[0070] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following specific implementation manners. Those skilled in the art should understand that the present disclosure can also be implemented without certain specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0071] During the process of data interaction, issues related to data privacy and data security often arise. In some scenarios, such as finding friends on social software, antivirus software, and finding websites with leaked passwords, both users have a need for data interaction, but there may also be a situation where they do not want to expose their own data to the other party. As an example, taking the platform's recommendation of possible friends to users as an example, after user A logs in to a certain social, shopping, or multimedia platform, the platform can recommend user B who has a friendship with user A to user A, so that user A can quickly add friends on the platform and increase user stickiness. Usually, with the authorization and permission of user A, based on the contact information in the client's address book, the possible friends of user A on the platform can be determined. For example, users usually use phone numbers or email addresses as the login accounts for the platform. Therefore, the phone numbers of all users who have registered on the platform can be intersected with the phone numbers of the contacts in the client's address book, so as to determine the phone numbers of the contacts in user A's address book who have registered on the platform. The user corresponding to this phone number is the possible friend of user A on the platform. To ensure that the information in user A's address book and the information of the registered users on the platform are not leaked, privacy set intersection (PSI) needs to be performed during the intersection operation.
[0072] Privacy set intersection is a specific secure multi-party computing problem, aiming to obtain the intersection of the data held by both parties without leaking any additional information. Here, the additional information refers to any information other than the intersection of the data of both parties; privacy set intersection allows two parties holding their own data sets to jointly calculate the intersection operation of the two data sets; at the end of the two-party interaction, one party or both parties should obtain the correct intersection and will not obtain any information in the other party's set outside the intersection. Among them, the current dynamic privacy set intersection algorithm (KLS-RSA) based on RSA blind signature and pre-computation is as follows: Before the interaction between the server and the client, certain pre-operations are performed by both parties, and the information interaction between the server and the client is realized based on the RSA blind signature technology, and the Bloom filter is used to realize the intersection operation of the information search on the client side. This algorithm is mainly applicable to the privacy set intersection calculation where the data volume scales of the two parties differ greatly. For example, the privacy set intersection calculation of the million level versus the thousand level. Moreover, the algorithm structure is simple and has strong practicability. However, this algorithm requires a large number of large modular exponentiation calculations, which will bring high computational overhead and make the computational efficiency of this algorithm not high enough.
[0073] To solve the above technical problems, an embodiment of the present disclosure proposes a data processing method (for detailed description, see below). Based on the RSA blind signature and pre-computation-based dynamic private set intersection algorithm, a preset algorithm is used for modular exponentiation operations. For example, the grouped streaming sharding algorithm is used to perform modular exponentiation evaluation calculations on a large amount of data involved in both parties (such as a client and a server), effectively reducing the modular exponentiation calculation time in each stage of the two parties involved, optimizing the overall running time of the private set intersection, and thus achieving an efficient and fast private set intersection; for another example, the multi-radix modular exponentiation algorithm is used to quickly calculate the modular exponentiation evaluation of a large amount of data involved in both parties, greatly saving the time for both parties to perform calculations and interactions online, and more effectively protecting the non-disclosure of private data in the non-intersection sets of both parties, thus achieving an efficient and fast private set intersection.
[0074] Among them, RSA blind signature is a digital signature scheme that allows a signer to sign a message without knowing the specific content of the message. In the RSA blind signature scheme, the message owner (such as a client) first blinds the message, converts the message into a blinded message, and then sends the blinded message to the signer (such as a server). The signer signs the blinded message and sends the signature result to the message owner. The message owner obtains the signature on the original message by unblinding the signature result. Blind signature has the properties of unforgeability and untraceability. Unforgeability means that no one other than the signer himself can generate a valid blind signature in his name. Untraceability means that even if the signer retains the signature data, it is impossible to find the internal connection between the signature and the original message, and thus it is impossible to trace the message owner.
[0075] Among them, the grouped streaming sharding algorithm is a multi-threaded calculation algorithm. The grouped streaming sharding algorithm is used to evenly group multiple triple data and perform parallel processing on different groups obtained by the even grouping, so as to effectively accelerate batch modular exponentiation calculations. Especially for scenarios with batch modular exponentiation calculations of no less than 2^20, it has high calculation efficiency.
[0076] The grouped streaming sharding algorithm assumes the following: there are n large integers X = (x1,…,x n ), large exponents D = (d1,…,d n ), large moduli P = (p1,…,p n ); then for x i ∈X, i∈[n], calculate its corresponding modular exponent and form the result set R = (r1,…,r n ); the triple data (x i , d i , p i), i ∈ [n], including the target data x i , the target exponent d i , and the target modulus p i . Exemplarily, the process of calculating modular exponentiation by the grouped streaming sharding algorithm may include: grouping and hash table construction, parallel streaming sharding calculation:
[0077] First step, grouping and hash table construction. A two-dimensional k-Choice hash table structure T of W * L can be generated. The number of bins in the hash table T is W, and the maximum filling quantity of each bin is Each bin is a group. Let there be random hash functions H1, …, H h ∈ {0, 1} * → [W]. First, evenly group multiple triple data, that is, evenly distribute the triple data (x i , d i , p i ), i ∈ [n] into each bin of the two-dimensional hash table T; Exemplarily, the indexes of multiple groups corresponding to the target data in the hash table can be calculated, and the target group with the least current stored data volume among the multiple groups can be determined, and the triple data corresponding to the target data is stored in the target group; For example, for each target data x i ∈ X, use H1, …, H h respectively for mapping to calculate the indexes of h bins, and select the bin with the least current stored data quantity among these h bins to store the triple data (x i , d i , p i ), i ∈ [n], that is, each triple data (x i , d i , p i ), i ∈ [n] is stored only once in the hash table T; Such a data mapping method will construct a two-dimensional hash table T with almost the same data volume in each bin.
[0078] Second step, parallel streaming sharding calculation. Parallel process different groups obtained by average grouping, that is, parallelly calculate the modular exponentiation of the data stored in each bin to improve the calculation efficiency. Exemplarily, the target group can be divided into multiple shards; Parallelly process multiple shards. For example, when the data filled in a certain bin reaches the maximum filling quantity, then shard the bin. When processing the j-th bin, divide this one-dimensional bin into m shards (m ≤ L + 1), and further parallelly process the modular exponentiation calculation of the m shards to obtain the modular exponentiation calculation result R. It can be understood that at the same time, multiple bins with the filled data reaching the maximum filling quantity can be parallelly processed.
[0079] Among them, the multi - base modular exponentiation algorithm, also known as the m - base exponentiation algorithm, the value of m can be set according to requirements. For example, it can be 64. This multi - base modular exponentiation algorithm is used to evaluate modular exponentiation based on the target exponent in multi - base; the multi - base modular exponentiation algorithm can be divided into three stages: exponent division, construction of the exponent value lookup table, and modular exponentiation evaluation based on the exponent value lookup table; exemplarily, the target exponent can be divided according to a preset multi - base (i.e., m - base); construct the target exponent value lookup table of the target data set; use the divided target exponent to perform modular exponentiation evaluation based on the target exponent value lookup table. This multi - base modular exponentiation algorithm can be used for private set intersection based on multi - base RSA blind signature modular exponentiation calculation under the semi - honest security model, where the semi - honest security model means that each participating party strictly executes according to the protocol, but is still somewhat curious about other private data in the process; while improving the efficiency of private set intersection, it can better protect the privacy of the data in the non - intersection sets of the server and the client from leakage, especially applicable to the private set intersection in the scenario where the server data set has no less than 2^20 data and the client data set has no more than 2^10 data.
[0080] First, an exemplary description of the possible application scenarios of the data processing method in the embodiments of the present disclosure will be given below.
[0081] Figure 1 Show a schematic structural diagram of a private set intersection system according to an embodiment of the present disclosure. As Figure 1 shown, the private set intersection system may include: a first terminal 10 and a second terminal 20.
[0082] Among them, the first terminal 10 is the device used by the user and can also be called the client. The number of the first terminals 10 can be one or more (only one is shown in the figure). The first terminal 10 may include, but is not limited to: devices such as smart phones, tablet computers, portable personal computers, mobile Internet devices, etc.; the first terminal 10 is often configured with a display device, and the display device can also be a monitor, a display screen, a touch screen, etc. The touch screen can also be a touch - control screen, a touch - control panel, etc. The embodiments of the present application do not make limitations.
[0083] The second terminal 20 is a server that provides services to users and can also be called the server - side. The second terminal 20 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0084] Exemplarily, the first terminal 10 and the second terminal 20 can be directly or indirectly connected through wired or wireless communication means, and the embodiments of the present disclosure do not limit this.
[0085] As an example, the first terminal 10 is a smart phone of user A, and the second terminal 20 is a server of a certain social platform B. After user A logs in to social platform B, social platform B can send the common friends of user A on this social platform B to user A's smart phone. By executing the data processing method provided by the embodiments of the present disclosure, the first terminal 10 and the second terminal 20 can efficiently determine the common friends of user A on this social platform B while ensuring the data privacy and security in user A and social platform B, that is, user A cannot learn the registered user data in social platform B based on the intersection data, and social platform B cannot learn the contact information of user A based on the summation data.
[0086] It should be noted that the above-described application scenarios described in the embodiments of the present disclosure are for more clearly explaining the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Those of ordinary skill in the art can know that for the emergence of other similar or new scenarios, the technical solutions provided by the embodiments of the present disclosure are equally applicable to similar technical problems.
[0087] The implementation process of the data processing method provided by the embodiments of the present disclosure will be described in detail below.
[0088] Figure 2 The flowchart of a data processing method according to an embodiment of the present disclosure is shown. This method can be applied to a server and a client, where the server can be the first terminal 10 in the above Figure 1 system, and the client can be the second terminal 20 in the above Figure 1 system; as Figure 2 shown, this method can include the following steps:
[0089] Step 201, the server calculates the modular exponentiation of each data in the first data set under the RSA private key based on a preset algorithm.
[0090] Wherein, the preset algorithm is a grouped streaming sharding algorithm or a multi - base modular exponentiation algorithm.
[0091] The grouped streaming sharding algorithm is used to evenly group multiple triple data and perform parallel processing on different groups obtained from the even grouping; the triple data includes target data, a target exponent, and a target modulus; in a possible implementation, the step of evenly grouping multiple triple data includes: calculating indexes of multiple groups corresponding to the target data in a hash table, and determining a target group with the least amount of currently stored data among the multiple groups; storing the triple data corresponding to the target data in the target group; the step of performing parallel processing on different groups obtained from the even grouping includes: dividing the target group into multiple shards; performing parallel processing on the multiple shards. The specific description of calculating the modular exponent by this grouped streaming sharding algorithm can refer to the relevant expressions in the previous text.
[0092] In some scenarios, taking the server as a single server as an example, an exemplary description of calculating the modular exponent based on the grouped streaming sharding algorithm is given. First, the number of bins W of the k-Choice Hash hash table T can be set to the number of physical threads of this single server. For each target data x i ∈X, use H1,…,H h respectively for mapping to calculate the indexes of h bins, and select the bin with the least amount of currently stored data among these h bins to store the triple data (x i ,d i ,p i ), i ∈ [n], so as to store the triple data corresponding to the target data in the hash table T. Then, allocate a main thread for the hash table T, and allocate a bucket thread for each bin for processing. Each bucket thread is responsible for sharding its corresponding bin, and then allocate each shard to a sub-thread for calculation to achieve parallel streaming sharding calculation; at the same time, the bucket thread aggregates the calculation results of each sub-thread and submits them to the main thread, so as to implement multi-threaded modular exponent operation.
[0093] In other scenarios, taking the server as a server cluster as an example, an exemplary description of calculating the modular exponent based on the grouped streaming sharding algorithm is given. First, the main server in the server cluster sets the number of bins W of the k-Choice Hash table T to the total number of servers used. For each target data x i ∈X, use H1,…,H h respectively for mapping to calculate the indexes of h bins, and select the bin with the least amount of currently stored data among these h bins to store the triple data (x i ,d i ,p i), i ∈ [n], so as to store the triple data corresponding to the target data in the hash table T. Then the master server distributes the calculation of each bin to each computing server. The computing server is responsible for sharding the obtained corresponding bin and allocating each shard to a sub-thread to achieve parallel streaming sharding calculation. At the same time, the master server can also aggregate the calculation results of each computing server, thereby realizing multi-threaded modular exponentiation operation.
[0094] As an example, the server can calculate the modular exponentiation of each data in the first data set under the RSA private key based on the grouped streaming sharding algorithm.
[0095] The multi-radix modular exponentiation algorithm is used to evaluate the modular exponentiation based on the multi-radix target exponent. In a possible implementation, the evaluation of the modular exponentiation based on the multi-radix target exponent includes: dividing the target exponent according to a preset multi-radix; constructing a target exponent value retrieval table for the target data set; and using the divided target exponent to evaluate the modular exponentiation based on the target exponent value retrieval table. The specific description of calculating the modular exponentiation by this multi-radix modular exponentiation algorithm can refer to the relevant expressions above.
[0096] As another example, the server can calculate the modular exponentiation of each data in the first data set under the RSA private key based on the multi-radix modular exponentiation algorithm.
[0097] Among them, the first data set includes the user data for which the server needs to perform private set intersection. The first data set is the data set of the server, and this data set usually contains a large number of data. Exemplarily, the data can also be called elements and can be any type such as phone numbers, emails, images, documents, etc., and no limitation is made here. For example, taking the server of a social platform as an example, the first data set can include all user account information registered on this social platform. Among them, the user account information can be phone numbers, email addresses, etc. As an example, the first data set can include D1, D2, D3... Dr, a total of r phone numbers, where r is a positive integer.
[0098] Exemplarily, before executing this step, the server can establish a communication connection with the client. The client and the server can establish a communication connection based on an existing communication protocol. The client can successfully log in to the server based on the user login protocol so that the server and the client can interact data subsequently. For example, the server can send a Bloom filter to the client subsequently.
[0099] Exemplarily, before performing this step, the server and the client reach an agreement on information such as the RSA modulus, the RSA public key (common exponent), and the error rate of the Bloom filter. For example, the same RSA modulus N, RSA public key e, and error rate ε of the Bloom filter can be pre-configured on the server and the client, so as to make both parties reach an agreement. Alternatively, the server and the client can negotiate by interacting with each other to determine the RSA modulus N, the RSA public key e, and the error rate ε of the Bloom filter; wherein, the type of the Bloom filter can be configured according to actual needs, and the embodiments of the present disclosure do not limit this.
[0100] Exemplarily, before performing this step, the server can also generate an RSA private key (secret key) using the RSA algorithm.
[0101] Step 202: The server inserts the modular exponent of each data under the RSA private key into the Bloom filter, and sends the Bloom filter.
[0102] Exemplarily, the server can initialize the Bloom filter, and then insert the modular exponent of each data in the first data set under the RSA private key into the Bloom filter.
[0103] Step 203: The client generates a set of random numbers, and calculates the modular exponent of each random number in the set of random numbers under the RSA public key and the modular inverse of each random number based on a preset algorithm.
[0104] Wherein, each random number in the set of random numbers corresponds one-to-one with each data in the second data set. The second data set includes the user data for which the client needs to perform a private set intersection, that is, the second data set is the data set of the client, and this data set usually contains a small amount of data. As an example, the second data set may include E1, E2, E3…Et, a total of t phone numbers. Where t is a positive integer. Exemplarily, r is selected to be greater than t. For example, t can be 100 and r can be 1000000; the set of random numbers can include F1, F2, F3…Ft, a total of t random numbers, wherein each random number corresponds one-to-one with each data in the second data set, that is, F1 corresponds to E1, F2 corresponds to E2…Ft corresponds to Et.
[0105] As an example, the client can calculate the modular exponent of each random number in the set of random numbers under the RSA public key and the modular inverse of each random number based on the grouped streaming sharding algorithm.
[0106] As another example, the client can calculate the modular exponent of each random number in the set of random numbers under the RSA public key and the modular inverse of each random number based on the multi - base modular exponent algorithm.
[0107] Step 204, the client receives the Bloom filter; each data in the first data set stored in the Bloom filter is the modular exponent under the RSA private key.
[0108] Step 205, the client sends the third data set to the server.
[0109] As an example, the third data set includes the modular exponent of the product of each data in the second data set and the corresponding random number in the random number set under the RSA public key calculated based on the grouped streaming sharding algorithm. The client can calculate the modular exponent of the product of each data in the second data set and the corresponding random number under the RSA public key based on the grouped streaming sharding algorithm, so as to generate the third data set, and then the client sends the third data set to the server.
[0110] As another example, the third data set includes an exponential value retrieval table of the modular exponents of each random number in the random number set under the RSA public key constructed based on the multi - base modular exponent algorithm.
[0111] Step 206, the server receives the third data set.
[0112] Step 207, the server calculates the modular exponent of each data in the third data set under the RSA private key based on the preset algorithm to obtain a fourth data set.
[0113] As an example, the server can calculate the modular exponent of each data in the third data set under the RSA private key based on the grouped streaming sharding algorithm, so as to obtain the fourth data set.
[0114] As another example, the server can calculate the modular exponent of each data in the third data set under the RSA private key based on the multi - base modular exponent algorithm, so as to obtain the fourth data set.
[0115] Step 208, the server sends the fourth data set.
[0116] Step 209, the client receives the fourth data set.
[0117] Wherein, the fourth data set includes the modular exponent of each data in the third data set calculated based on the preset algorithm under the RSA private key.
[0118] Step 210, the client determines whether the data in the fourth data set is in the Bloom filter to complete the private set intersection processing.
[0119] In a possible implementation, when the preset algorithm is the grouped streaming sharding algorithm, this step may include: The client calculates the modular exponent of the product of each data in the fourth data set and the modular inverse of the random number corresponding to each data in the random number set based on the grouped streaming sharding algorithm; Determine whether the modular exponent of the product is in the Bloom filter.
[0120] Exemplarily, since the fourth data set includes the modular exponent of each data in the third data set under the RSA private key, and at the same time the third data set includes the modular exponent of the product of each data in the second data set and the random number under the RSA public key, then if the modular exponent B of the product of a certain data A in the fourth data set and the modular inverse of the random number corresponding to this data is in the Bloom filter, it indicates that the user data corresponding to this modular exponent B in the original data set corresponding to the Bloom filter (i.e., the first data set) is the same as the user data corresponding to this data A in the original data set corresponding to the fourth data set (i.e., the second data set); that is, this user data is in the intersection of the first data set and the second data set; if the modular exponent B of the product of a certain data A in the fourth data set and the modular inverse of the random number corresponding to this data is not in the Bloom filter, it indicates that the user data corresponding to this data A in the original data set corresponding to the fourth data set (i.e., the second data set) is different from all user data in the original data set corresponding to the Bloom filter (i.e., the first data set), that is, this user data is not in the intersection of the first data set and the second data set; In this way, by traversing all the modular exponents of the products, the private set intersection of the first data set and the second data set can be obtained, and the private set intersection processing is completed.
[0121] In the embodiments of the present disclosure, based on the RSA blind signature and pre-computation-based dynamic private set intersection algorithm, a preset algorithm is used for modular exponentiation operations. Among them, based on the RSA blind signature technology, data security and privacy interactions between the server and the client (such as Bloom filters, third data sets, fourth data sets, etc.) are realized, and the Bloom filter is used to implement information search and intersection on the client side. During the RSA blind signature and pre-computation process, a preset algorithm is used for fast operations of a large number of modular exponents (such as calculating the modular exponent of each data in the first data set under the RSA private key, calculating the modular exponent of each data in the third data set under the RSA private key, calculating the modular exponent of each random number in the random number set under the RSA public key, etc.), which greatly improves the operation efficiency. For example, the grouped streaming sharding algorithm is used to calculate the modular exponent value of a large amount of data involved by both parties (such as the client and the server), effectively reducing the modular exponent calculation time in each stage of both parties and optimizing the overall running time of the private set intersection, thereby realizing efficient and fast private set intersection. For another example, the multi-base modular exponentiation algorithm is used for fast calculation of the modular exponent value of a large amount of data involved by both parties, greatly saving the time for both parties to perform calculations and interactions online, and more effectively protecting the non-disclosure of private data in the non-intersection sets of both parties, thereby realizing efficient and fast private set intersection.
[0122] The speed of the dynamic private set intersection based on RSA blind signature and pre-computation is optimized by the grouped streaming sharding algorithm or the multi-base modular exponentiation algorithm. For example, when the server has 2 20 private data, the client has 2 10 private data, the data of both parties is 128 bits, the random number is set to 2 10 pieces, and the error rate of the Bloom filter is set to 0.001, and the RSA private key is 2048 bits long, when using the unoptimized private set intersection method, it takes 109 minutes to calculate the intersection of the server's private data and the client's private data. After using the dynamic private set intersection based on RSA blind signature and pre-computation optimized by the grouped streaming sharding algorithm or the multi-base modular exponentiation algorithm, it only takes 7 minutes to calculate the intersection of the same set of data, and the optimization effect is 93%.
[0123] Figure 3 FIG. shows a flowchart of a data processing method according to an embodiment of the present disclosure. This method can be applied to the server and the client, where the server can be the first terminal 10 in the above Figure 1 system, and the client can be the second terminal 20 in the above Figure 1 system; as Figure 3 shown, this method may include the following steps:
[0124] Step 301, the server calculates the modular exponent of each data in the first data set under the RSA private key based on a preset algorithm.
[0125] Step 302: The server inserts the modular exponent of each piece of data under the RSA private key into the Bloom filter and sends the Bloom filter.
[0126] Step 303: The client generates a set of random numbers and calculates the modular exponent of each random number in the set of random numbers under the RSA public key and the modular inverse of each random number based on a preset algorithm.
[0127] Step 304: The client receives the Bloom filter.
[0128] Step 305: The client sends a third data set to the server.
[0129] Step 306: The server receives the third data set.
[0130] Step 307: The server calculates the modular exponent of each piece of data in the third data set under the RSA private key based on the preset algorithm to obtain a fourth data set.
[0131] Step 308: The server sends the fourth data set.
[0132] Step 309: The client receives the fourth data set.
[0133] Step 310: The client determines whether the data in the fourth data set is in the Bloom filter to complete the private set intersection processing.
[0134] The above steps 301 - 310 are the same as Figure 2 steps 201 - 210 in
[0135] Step 311: The server obtains the data that has changed in the first data set and calculates the modular exponent of the changed data under the RSA private key based on the preset algorithm.
[0136] Exemplarily, the data that has changed can be data with numerical changes or newly added data, etc.
[0137] Step 312: The server determines the target position of the modular exponent of the changed data under the RSA private key in the Bloom filter.
[0138] Step 313: The server sends the target position and the modular exponent of the changed data under the RSA private key.
[0139] Step 314: The client obtains the modular exponent of the data that has changed in the first data set under the RSA private key and the target position in the Bloom filter.
[0140] Step 315: The client modifies the value at the target position in the Bloom filter.
[0141] In the embodiments of the present disclosure, there is a feature of dynamic update of private data. If the server needs to update data, it can quickly calculate the modular exponent of the updated data under its RSA private key through the grouped streaming sharding algorithm or the multi - base modular exponentiation algorithm, and then calculate the positions where these modular exponents are inserted into the Bloom filter, and send these positions to the client. After receiving these positions, the client only needs to modify the values at these positions in its Bloom filter, which is convenient and fast. For example, when new registered user information appears in the registered user information stored in the server, the server can calculate the modular exponent of the new user information under the above - mentioned RSA private key through a preset algorithm. Thus, a private set intersection with the characteristics of dynamic update of private data based on RSA blind signature and pre - calculation is realized through the grouped streaming sharding algorithm or the multi - base modular exponentiation algorithm.
[0142] Further, taking the preset algorithm as the grouped streaming sharding algorithm as an example, the possible implementation manners of the data processing method in the above - mentioned embodiments are described, so as to apply the grouped streaming sharding algorithm to the dynamic private set intersection based on RSA blind signature and pre - calculation to improve the calculation efficiency of the private set intersection.
[0143] Figure 4 The flowchart of a data processing method according to an embodiment of the present disclosure is shown. This method can be applied to the server and the client, where the server can be the first terminal 10 in the above - mentioned Figure 1 system, and the client can be the second terminal 20 in the above - mentioned Figure 1 system; as Figure 4 shown, this method may include the following steps:
[0144] Step 401: The server calculates the modular exponent of each data in the first data set under the RSA private key based on the grouped streaming sharding algorithm.
[0145] Step 402: The server inserts the modular exponent of each data under the RSA private key into the Bloom filter and sends the Bloom filter.
[0146] Step 403: The client generates a set of random numbers and calculates the modular exponent of each random number in the set of random numbers under the RSA public key and the modular inverse of each random number based on the grouped streaming sharding algorithm.
[0147] Step 404: The client receives the Bloom filter; the Bloom filter stores the modular exponent of each data in the first data set under the RSA private key.
[0148] Step 405: The client sends the third data set to the server; the third data set includes the modular exponentiation of the product of each data in the second data set calculated based on the grouped streaming sharding algorithm and the corresponding random number in the random number set under the RSA public key.
[0149] Step 406: The server receives the third data set.
[0150] Step 407: The server calculates the modular exponentiation of each data in the third data set under the RSA private key based on the grouped streaming sharding algorithm to obtain the fourth data set;
[0151] Step 408: The server sends the fourth data set.
[0152] Step 409: The client receives the fourth data set.
[0153] Step 410: The client determines whether the data in the fourth data set is in the Bloom filter to complete the private set intersection processing.
[0154] Exemplarily, the client calculates the modular exponentiation of the product of each data in the fourth data set and the modular inverse of each data and the corresponding random number in the random number set; determines whether the modular exponentiation of the product is in the Bloom filter. It can be understood that the Bloom filter is the Bloom filter received by the client in the above step 404.
[0155] Step 411: The server obtains the data that has changed in the first data set and calculates the modular exponentiation of the changed data under the RSA private key based on the grouped streaming sharding algorithm;
[0156] Step 412: The server determines the target position of the modular exponentiation of the changed data under the RSA private key in the Bloom filter.
[0157] Step 413: The server sends the target position and the modular exponentiation of the changed data under the RSA private key.
[0158] Step 414: The client obtains the modular exponentiation of the data that has changed in the first data set under the RSA private key and the target position in the Bloom filter;
[0159] Step 415: The client modifies the value at the target position in the Bloom filter.
[0160] In the embodiments of the present disclosure, the grouped streaming sharding algorithm utilizes the multi-threading capabilities of the server and / or the client to be able to execute multiple tasks simultaneously; by the grouped streaming sharding algorithm, the speed of the dynamic private set intersection based on RSA blind signature and pre-computation is optimized. Based on the algorithm for dynamic private set intersection with RSA blind signature and pre-computation, the grouped streaming sharding algorithm is used to perform modular exponentiation evaluation calculations on a large amount of data involved in both parties (such as the client and the server), effectively reducing the modular exponentiation calculation time in each stage of the participating parties and optimizing the overall running time of the private set intersection, thereby achieving an efficient and fast private set intersection.
[0161] For example, in combination with Figure 5 An exemplary description is given of the process of optimizing the dynamic private set intersection based on RSA blind signature and pre-computation through the batch modular exponentiation grouped streaming sharding algorithm. The grouped streaming sharding algorithm can be applied to modular exponentiation calculations in the process of private set intersection. Figure 5 The flowchart shows a method for optimizing the dynamic private set intersection based on RSA blind signature and pre-computation through the grouped streaming sharding algorithm according to an embodiment of the present disclosure. As Figure 5 shown, the server and the client first reach an agreement on the RSA modulus N, the RSA public key e, and the Bloom filter error rate ε. The first data set is X = (x1,..., x Ns ), and the second data set is Y = (y1,..., y Nc ). The data processing stage can be divided into a basic stage, a preparation stage, an online stage, and an update stage. Among them, the basic stage and the preparation stage are the stages before the server and the client perform online private summation. Exemplarily, in these two stages, the server can execute steps 401 and 402 in Figure 4 above, and the client can execute steps 403 and 404 in Figure 4 above. In the online stage, the server can execute steps 406, 407, and 408 in Figure 4 above, and the client can execute steps 405, 409, and 410 in Figure 4 above; in the update stage, the server can execute steps 411, 412, and 413 in Figure 4 above, and the client can execute steps 414 and 415 in Figure 4 above.
[0162] I. Basic stage:
[0163] The server generates the RSA private key d.
[0164] The client can generate a set of random numbers and calculate the modular exponentiation of each random number r in the set of random numbers R and each random number r i under the RSA public key e based on the grouped streaming sharding algorithm iThe modular inverse; Exemplarily, for Calculate the modular exponentiation through the grouped streaming sharding algorithm and the modular inverse
[0165] II. Preparation Phase
[0166] The server calculates the modular exponentiation of each data x in the first data set X under the RSA private key d based on the grouped streaming sharding algorithm, and inserts the modular exponentiation of each data x i under the RSA private key d into the Bloom filter and sends it to the client. Exemplarily, for i ∈ [Ns], calculate the modular exponentiation through the grouped streaming sharding algorithm i Furthermore, for i ∈ [Ns], insert it into the Bloom filter BF.Insert(A[i]), and send BF to the client. For i ∈ [Ns], insert it into the Bloom filter BF.Insert(A[i]), and send BF to the client.
[0167] The client receives the Bloom filter BF.
[0168] III. Online Phase
[0169] The client calculates the modular exponentiation of the product of each data y in the second data set Y and the corresponding random number r in the random number set R i under the RSA public key e. Exemplarily, for i ∈ [Nc], calculate the modular exponentiation B[i] = y i * r' i under the RSA public key e. B[i] is the third data set, and send the third data set B[i] to the server. i * r' i mod N, and B[i] is the third data set, and send the third data set B[i] to the server.
[0170] The server receives the third data set B[i], and calculates the modular exponentiation of each data in the third data set B[i] under the RSA private key d based on the grouped streaming sharding algorithm. Exemplarily, for i ∈ [Nc], calculate the modular exponentiation C[i] = B[i] d mod N; C[i] is the fourth data set, and send the fourth data set C[i] to the client.
[0171] The client receives the fourth data set C[i], and calculates the modular exponentiation of the product of each data in the fourth data set C[i] and the corresponding random number of each data in the random number set. Furthermore, the client determines whether the modular exponentiation of the product is in the Bloom filter; Exemplarily, for i ∈ [Nc], calculate the modular exponentiation D[i] = C[i] * r i inv of the product.i inv mod N, and then obtain the intersection BF.Check(e i ·D[i]) to output the data y in the second data set Y corresponding to the intersection i .
[0172] IV. Update Phase
[0173] The server obtains the changed data, calculates the modular exponentiation of the changed data under the RSA private key based on the grouped streaming sharding algorithm, determines the target position of the modular exponentiation of the changed data in the Bloom filter under the RSA private key, and sends the target position and the modular exponentiation of the changed data under the RSA private key to the client; Exemplarily, the changed data is (u1,…,u Nu ), for i ∈ [Nu], calculate the modular exponentiation through the grouped streaming sharding algorithm Furthermore, for i ∈ [Nu], calculate the target position BF.Pos(U[i]) in the Bloom filter, and the position set corresponding to (u1,…,u Nu ) is P, and send P to the client.
[0174] The client obtains the modular exponentiation of the changed data in the first data set under the RSA private key and the target position in the Bloom filter, and modifies the value at the target position in the Bloom filter; Exemplarily, adjust the value at the corresponding position in the Bloom filter BF in P according to the modular exponentiation U[i].
[0175] In this way, pre-computation is performed in the above basic phase and preparation phase to transfer most of the computational burden before the online phase, thereby minimizing the running time of the online phase. At the same time, the grouped streaming sharding algorithm is used to improve the computational efficiency of the massive modular exponentiation calculations performed in the basic phase, preparation phase, online phase, and update phase, optimizing the overall running time of the private set intersection, thereby achieving an efficient and fast private set intersection.
[0176] Furthermore, taking the preset algorithm as the multi - radix modular exponentiation algorithm as an example, the possible implementation manners of the data processing method in the above embodiments are described, so as to apply the multi - radix modular exponentiation algorithm to the dynamic private set intersection based on RSA blind signature and pre - calculation to improve the computational efficiency of the private set intersection.
[0177] Figure 6 Shows a flowchart of a data processing method according to an embodiment of the present disclosure. This method can be applied to the server and the client, where the server can be the first terminal 10 in the above Figure 1 system, and the client can be the second terminal 20 in the above Figure 1 system; As Figure 6As shown, the method may include the following steps:
[0178] Step 601: The server calculates the modular exponentiation of each data in the first data set under the RSA private key based on the multi - base modular exponentiation algorithm.
[0179] Exemplarily, the server divides the RSA private key according to a preset multi - base; constructs an index value retrieval table for the first data set; and uses the divided RSA private key to perform modular exponentiation evaluation based on this index value retrieval table, so as to obtain the modular exponentiation of each data in the first data set under the RSA private key.
[0180] Step 602: The server inserts the modular exponentiation of each data under the RSA private key into a Bloom filter and sends the Bloom filter.
[0181] Step 603: The client generates a set of random numbers and calculates the modular exponentiation of each random number in the set of random numbers under the RSA public key and the modular inverse of each random number based on the multi - base modular exponentiation algorithm.
[0182] Exemplarily, the server divides the RSA public key according to a preset multi - base; constructs an index value retrieval table for the set of random numbers; and uses the divided RSA public key to perform modular exponentiation evaluation based on this index value retrieval table, so as to obtain the modular exponentiation of each random number in the set of random numbers under the RSA public key.
[0183] Step 604: The client receives the Bloom filter; each data in the first data set has its modular exponentiation under the RSA private key stored in the Bloom filter.
[0184] Step 605: The client sends a third data set to the server; the third data set includes an index value retrieval table of the modular exponentiation of each random number in the set of random numbers under the RSA public key constructed based on the multi - base modular exponentiation algorithm.
[0185] Step 606: The server receives the third data set.
[0186] Step 607: The server calculates the modular exponentiation of each data in the third data set under the RSA private key based on the multi - base modular exponentiation algorithm to obtain a fourth data set;
[0187] Exemplarily, an index value retrieval table for the third data set is constructed; and using the RSA private key divided according to a preset multi - base, modular exponentiation evaluation is performed based on this index value retrieval table, so as to obtain the modular exponentiation of each data in the third data set under the RSA private key.
[0188] Step 608: The server sends the fourth data set.
[0189] Step 609, the client receives the fourth data set.
[0190] Step 610, the client determines whether the data in the fourth data set is in the Bloom filter to complete the private set intersection processing.
[0191] Step 611, the server obtains the data that has changed in the first data set, and calculates the modular exponentiation of the changed data under the RSA private key based on the multi - base modular exponentiation algorithm.
[0192] Exemplarily, construct an exponent value retrieval table for the changed data; use the RSA private key divided by a preset multi - base, and perform modular exponentiation evaluation based on this exponent value retrieval table to obtain the modular exponentiation of the changed data under the RSA private key.
[0193] Step 612, the server determines the target position of the modular exponentiation of the changed data under the RSA private key in the Bloom filter.
[0194] Step 613, the server sends the target position and the modular exponentiation of the changed data under the RSA private key.
[0195] Step 614, the client obtains the modular exponentiation of the data that has changed in the first data set under the RSA private key and the target position in the Bloom filter.
[0196] Step 615, the client modifies the value at the target position in the Bloom filter.
[0197] In the embodiments of the present disclosure, the multi - base modular exponentiation algorithm optimizes the speed of dynamic private set intersection based on RSA blind signature and pre - calculation. The multi - base modular exponentiation algorithm is used for the fast calculation of modular exponentiation evaluation of a large amount of data involved by both parties, greatly saving the time for both parties to perform calculations and interactions online, and more effectively protecting the non - disclosure of private data in the non - intersection sets of both parties, thereby realizing efficient and fast private set intersection.
[0198] For example, in combination with Figure 7 and Figure 8 exemplarily illustrate the process of optimizing the dynamic private set intersection based on RSA blind signature and pre - calculation through the multi - base modular exponentiation algorithm. The multi - base modular exponentiation algorithm can be applied to the modular exponentiation calculation in the process of private set intersection. Figure 7 and Figure 8 both show the flowchart of the method for optimizing the dynamic private set intersection based on RSA blind signature and pre - calculation through the multi - base modular exponentiation algorithm according to an embodiment of the present disclosure, as Figure 7 and Figure 8As shown, the server and the client first reach an agreement on the RSA modulus N, the RSA public key e, the m - base, and the Bloom filter error rate ε. The first data set is X = (x1, …, x Ns ), and the second data set is Y = (y1, …, y Nc ). The data processing stage can be divided into a basic stage, a preparation stage, an online stage, and an update stage. Among them, the basic stage and the preparation stage are the stages before the server and the client perform online private summation. Exemplarily, in these two stages, the server can execute steps 601, 602, and 606 in the above Figure 6 ; the client can execute steps 603, 604, and 605 in the above Figure 6 . In the online stage, the server can execute steps 607 and 608 in the above Figure 6 ; the client can execute steps 609 and 610 in the above Figure 6 . In the update stage, the server can execute steps 611, 612, and 613 in Figure 6 , and the client can execute steps 614 and 615 in the above Figure 6 .
[0199] I. Basic Stage
[0200] The server generates the RSA private key d and divides it according to the base m. At the same time, it constructs the exponential value retrieval table A of the first data set X. Exemplarily, the server generates the RSA private key d and divides the private key d = (d t d t-1 …d1d0) m , where m = 2 k , k ≥ 1; then processes x i ∈X one by one, and let Calculate successively from j = 1 to j = m - 1 to form the exponential value retrieval table A of the first data set X. The i - th row stored in the exponential value retrieval table A is
[0201] The client generates a set of random numbers and calculates the set of modular inverses R inv ; and divides the RSA public key e according to the base m, and constructs the exponential value retrieval table B of the set of random numbers R. Exemplarily, the client generates a set of random numbers For calculate the modular inverse r i inv = r i -1 mod N; divide the RSA public key e = (e t e t-1 …e1e0) m , where m = 2k , where \(k\geq1\); process \(r\) one by one i \(\in R\), let Calculate successively from \(j = 1\) to \(j = m - 1\) Form an exponential value retrieval table \(B\) of the random number set \(R\). The \(i\)-th row stored in the exponential value retrieval table \(B\) is
[0202] II. Preparation stage
[0203] The server uses the partitioned private key \(d\) to perform modular exponentiation evaluation based on the exponential value retrieval table \(A\), and inserts the evaluated modular exponentiation value into the Bloom filter; sends the Bloom filter to the client. Exemplarily, the server processes \(x\) one by one i \(\in X\), let Calculate successively from \(j = t\) to \(j = 0\) Insert the finally obtained Into the Bloom filter Then send the Bloom filter \(BF\) to the client.
[0204] The client uses the partitioned RSA public key \(e\) to perform modular exponentiation evaluation based on the exponential value retrieval table \(B\), thereby forming a result set \(C\). The result set \(C\) includes the evaluated modular exponentiation value; constructs an exponential value retrieval table \(D\) of \(C\) and sends it to the server. The numerical retrieval table \(D\) is the third data set. Exemplarily, the client processes \(r\) one by one i \(\in R\), let Calculate successively from \(j = t\) to \(j = 0\) Form a result set \(C\), and the \(i\)-th row stored in it is \(r\) i e mod \(N\); and further, process \(c\) one by one i \(\in C\), let Calculate successively from \(j = 1\) to \(j = m - 1\) Form an exponential retrieval table \(D\), and the \(i\)-th row stored in it is After that, send the exponential retrieval table \(D\) to the server.
[0205] III. Online stage
[0206] The server uses the partitioned private key \(d\) to perform modular exponentiation evaluation based on the exponential retrieval table \(D\), and obtains a result set \(E\). The result set \(E\) is the fourth data set, and sends the result set \(E\) to the client. Exemplarily, for each row of the exponential retrieval table \(D\) on the server It calculates the corresponding Insert the finally obtained \(f\) i Construct the result set \(E\), that is, \(f\) i \(\in E\). Send the result set \(E\) to the client.
[0207] The client checks whether the data in the result set E exists in the Bloom filter BF to obtain the private set intersection X ∩ Y. Exemplarily, the client uses e i ∈ E to check whether BF.Check(e i ·r i inv mod N) holds. If it holds, the corresponding y i ∈ X ∩ Y.
[0208] IV. Update Phase
[0209] The server obtains the changed data, calculates the modular exponentiation of the changed data under the RSA private key based on the multi - radix modular exponentiation algorithm, determines the target position of the modular exponentiation of the changed data in the Bloom filter under the RSA private key, and sends the target position and the modular exponentiation of the changed data under the RSA private key to the client. Exemplarily, the changed data is (u1,…,u Nu ). For i ∈ [Nu], construct the exponent value retrieval table U[i] of (u1,…,u Nu ), and use the partitioned private key d to perform modular exponentiation evaluation based on the exponent retrieval table U[i], so as to determine the modular exponentiation of (u1,…,u Nu ) under the RSA private key. Furthermore, for i ∈ [Nu], calculate the target position BF.Pos(U[i]) in the Bloom filter. The position set corresponding to (u1,…,u Nu ) is P, and send P to the client.
[0210] The client obtains the modular exponentiation of the changed data in the first data set under the RSA private key and the target position in the Bloom filter, and modifies the value at the target position in the Bloom filter. Exemplarily, adjust the value at the corresponding position in the Bloom filter BF in P according to the modular exponentiation U[i].
[0211] In this way, through pre - calculation in the above - mentioned basic phase and preparation phase, most of the computational burden is transferred before the online phase, thus effectively saving the computational time in the online phase. At the same time, the multi - radix modular exponentiation algorithm is used to improve the computational efficiency of the massive modular exponentiation operations in the basic phase, preparation phase, online phase, and update phase, optimize the overall running time of the private set intersection, and effectively protect the non - disclosure of the private data in the non - intersection sets of both parties, thereby achieving an efficient and fast private set intersection.
[0212] Based on the same inventive concept as the above - mentioned method embodiments, the embodiments of the present disclosure also provide a data processing device, which can be used to execute the technical solutions described in the above - mentioned data processing method embodiments.
[0213] Figure 9The structural diagram of a data processing device according to an embodiment of the present disclosure is shown, which is applied to a server side, such as Figure 9 As shown, the device may include:
[0214] A first calculation module 901, configured to: calculate the modular exponentiation of each data in the first data set under the RSA private key based on a preset algorithm; wherein, the first data set includes user data to be subjected to private set intersection by the server side; the preset algorithm is a grouped streaming sharding algorithm or a multi - base modular exponentiation algorithm, the grouped streaming sharding algorithm is used to evenly group multiple triple data and perform parallel processing on different groups obtained by the even grouping; the triple data includes target data, target exponent, and target modulus; the multi - base modular exponentiation algorithm is used to evaluate modular exponentiation based on a multi - base target exponent;
[0215] A first communication module 902, configured to insert the modular exponentiation of each data under the RSA private key into a Bloom filter and send the Bloom filter;
[0216] The first communication module 901 is further configured to receive a third data set, where the third data set includes the modular exponentiation of the product of the modular exponentiations of each data in the second data set and the corresponding random numbers in the random number set under the RSA public key calculated based on the grouped streaming sharding algorithm, or an exponential value retrieval table of the modular exponentiations of each random number in the random number set under the RSA public key constructed based on the multi - base modular exponentiation algorithm; wherein, the second data set includes user data to be subjected to private set intersection by the client side, the random number set is generated by the client side, and each random number corresponds one - to - one with each data in the second data set;
[0217] The first calculation module 902 is further configured to calculate the modular exponentiation of each data in the third data set under the RSA private key based on the preset algorithm to obtain a fourth data set;
[0218] The first communication module is further configured to send the fourth data set;
[0219] In a possible implementation manner, the first calculation module 902 is further configured to: calculate indexes of multiple groups corresponding to the target data in the hash table and determine a target group with the least current stored data volume among the multiple groups; store the triple data corresponding to the target data in the target group; and is further configured to: divide the target group into multiple shards; and perform parallel processing on the multiple shards.
[0220] In a possible implementation manner, the first calculation module 902 is further configured to: divide the target exponent according to a preset number system; construct a target exponent value retrieval table for the target data set; and perform modular exponentiation evaluation based on the divided target exponent and the target exponent value retrieval table.
[0221] In a possible implementation manner, the first calculation module 902 is further configured to: when the preset algorithm is the grouped streaming sharding algorithm, obtain the data that has changed in the first data set, and calculate the modular exponent of the changed data under the RSA private key based on the grouped streaming sharding algorithm; determine the target position of the modular exponent of the changed data under the RSA private key in the Bloom filter; and send the target position and the modular exponent of the changed data under the RSA private key.
[0222] Figure 10 The structure diagram of a data processing device according to an embodiment of the present disclosure is shown, which is applied to a client, as Figure 10 shown, the device may include:
[0223] A second calculation module 1001, configured to generate the random number set, and calculate the modular exponent of each random number in the random number set under the RSA public key and the modular inverse of each random number based on the preset algorithm;
[0224] A second communication module 1002, configured to receive the Bloom filter; each data in the first data set has a modular exponent under the RSA private key stored in the Bloom filter;
[0225] The second communication module 1002 is further configured to send the third data set to the server;
[0226] The second communication module 1002 is further configured to receive the fourth data set, where the fourth data set includes the modular exponent of each data in the third data set calculated based on the preset algorithm under the RSA private key;
[0227] The second calculation module 1001 is further configured to determine whether the data in the fourth data set is in the Bloom filter to complete the private set intersection processing.
[0228] In a possible implementation manner, the second calculation module 1001 is further configured to: calculate the indexes of multiple groups corresponding to the target data in the hash table, and determine the target group with the least current stored data volume among the multiple groups; store the triple data corresponding to the target data in the target group; and is further configured to: divide the target group into multiple shards; and perform parallel processing on the multiple shards.
[0229] In a possible implementation, the second calculation module 1001 is further configured to: divide the target exponent according to a preset multiple number system; construct a target exponent value retrieval table for the target data set; and perform modular exponentiation evaluation based on the divided target exponent and the target exponent value retrieval table.
[0230] In a possible implementation, the second calculation module 1001 is further configured to: when the preset algorithm is the grouped streaming sharding algorithm, calculate the modular exponent of the product of the modular inverses of each data in the fourth data set and the corresponding random numbers in the random number set based on the grouped streaming sharding algorithm; determine whether the modular exponent of the product is in the Bloom filter; or, is further configured to: obtain the modular exponent of the changed data in the first data set under the RSA private key and the target position in the Bloom filter; and modify the value at the target position in the Bloom filter.
[0231] The above Figure 9 and Figure 10 The device shown above performs modular exponentiation operations using a preset algorithm based on the RSA blind signature and pre-computation-based dynamic private set intersection algorithm; wherein, based on the RSA blind signature technology, data security and privacy interaction between the server and the client (such as Bloom filter, third data set, fourth data set, etc.) is realized, and information search and intersection on the client side is realized using the Bloom filter; during the RSA blind signature and pre-computation process, modular exponentiation operations are performed using a preset algorithm (such as calculating the modular exponent of each data in the first data set under the RSA private key, calculating the modular exponent of each data in the third data set under the RSA private key, calculating the modular exponent of each random number in the random number set under the RSA public key, etc.), which greatly improves the operation efficiency; for example, using the grouped streaming sharding algorithm to perform modular exponentiation evaluation calculations on a large amount of data involved by both parties (such as the client and the server), effectively reducing the modular exponent calculation time for each stage of both parties and optimizing the overall running time of the private set intersection, thereby realizing an efficient and fast private set intersection; for another example, using the multiple number system modular exponentiation algorithm to perform fast calculations of the modular exponent of a large amount of data involved by both parties, greatly saving the time for both parties to perform calculations and interactions online, and more effectively protecting the non-disclosure of private data in the non-intersection sets of both parties, thereby realizing an efficient and fast private set intersection.
[0232] The above Figure 9 and Figure 10 For the technical effects and specific descriptions of various possible implementation manners in the device shown above, reference may be made to the above method embodiments, which will not be elaborated here.
[0233] It should be understood that the division of each module in the above device is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. In addition, the modules in the device can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or the functions of each module of the device. The processor is, for example, a general-purpose processor, such as a central processing unit (CPU) or a microprocessor, and the memory is a memory inside or outside the device. Alternatively, the modules in the device can be implemented in the form of a hardware circuit, and the functions of some or all of the modules can be implemented by designing the hardware circuit. The hardware circuit can be understood as one or more processors. For example, in one implementation, the hardware circuit is an application-specific integrated circuit (ASIC), and the functions of some or all of the above modules are implemented by designing the logical relationship between the components in the circuit. Again, for example, in another implementation, the hardware circuit can be implemented by a programmable logic device (PLD). Taking a field programmable gate array (FPGA) as an example, it can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured through a configuration file, so as to implement the functions of some or all of the above modules. All modules of the above device can be fully implemented in the form of a processor calling software, or fully implemented in the form of a hardware circuit, or partially implemented in the form of a processor calling software, and the remaining part is implemented in the form of a hardware circuit.
[0234] In the embodiments of the present disclosure, a processor is a circuit with the ability to process signals. In one implementation, the processor can be a circuit with the ability to read and execute instructions, such as a CPU, a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a neural-network processing unit (NPU), a tensor processing unit (TPU), etc.; in another implementation, the processor can implement certain functions through the logical relationship of hardware circuits, and the logical relationship of the hardware circuits is fixed or can be reconfigured. For example, the processor is a hardware circuit implemented by an ASIC or a PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the configuration of the hardware circuit can be understood as the process of the processor loading instructions to implement the functions of some or all of the above modules.
[0235] It can be seen that each module in the above device can be one or more processors (or processing circuits) configured to implement the methods of the above embodiments. For example: CPU, GPU, NPU, TPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms. In addition, each module in the above device can be fully or partially integrated together, or can be independently implemented, and this is not limited.
[0236] The embodiments of the present disclosure also provide an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor is configured to implement the methods of the above embodiments when executing the instructions. Exemplarily, it can execute the steps of the methods shown above Figures 2 - 8 in the above.
[0237] Figure 11 The structural schematic diagram of an electronic device according to an embodiment of the present disclosure is shown. Exemplarily, the electronic device can be the first terminal 10 or the second terminal 20 shown above Figure 1 in the above; as Figure 11 shown, the electronic device can include: at least one processor 801, a communication line 802, a memory 803, and at least one communication interface 804.
[0238] The processor 801 may be a general-purpose central processing unit, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits for controlling the execution of the programs of the present disclosure; the processor 801 may also include a heterogeneous computing architecture of multiple general-purpose processors. For example, it may be a combination of at least two of CPU, GPU, microprocessor, DSP, ASIC, and FPGA; as an example, the processor 801 may be CPU+GPU or CPU+ASIC or CPU+FPGA.
[0239] The communication line 802 may include a path for transmitting information between the above components.
[0240] The communication interface 804 uses any device such as a transceiver for communicating with other devices or communication networks, such as Ethernet, RAN, wireless local area networks (WLAN), etc.
[0241] The memory 803 may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or it may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may exist independently and be connected to the processor through the communication line 802. The memory may also be integrated with the processor. The memory provided in the embodiments of the present disclosure generally has non-volatility. Among them, the memory 803 is used to store the computer execution instructions for executing the present disclosure solution, and is controlled by the processor 801 for execution. The processor 801 is used to execute the computer execution instructions stored in the memory 803, so as to implement the method provided in the above embodiments of the present disclosure; exemplarily, it may execute the steps of the method shown above Figures 2 - 8 in the
[0242] Optionally, the computer execution instructions in the embodiments of the present disclosure may also be referred to as application code, and the embodiments of the present disclosure do not make specific limitations thereon.
[0243] Exemplarily, the processor 801 may include one or more CPUs. For example, Figure 11 the CPU0 in Figure 11 ; the processor 801 may also include a CPU and any one of a GPU, an ASIC, and an FPGA. For example,
[0244] the CPU0+GPU0 or CPU 0+ASIC0 or CPU0+FPGA0 in Figure 8 . Exemplarily, the electronic device may include multiple processors. For example,
[0245] the processor 801 and the processor 807 in
[0246] . Each of these processors may be a single-core (single-CPU) processor, a multi-core (multi-CPU) processor, or a heterogeneous computing architecture including multiple general-purpose processors. The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). Figures 2 - 8
[0247] Figures 2 - 8 Embodiments of the present disclosure provide a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the methods in the above embodiments are implemented. Exemplarily, the steps of the methods shown in the above Figures 2 - 8 can be executed.
[0248] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement aspects of the present disclosure.
[0249] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not to be construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0250] The computer-readable program instructions described herein may be downloaded to respective computing / processing devices from a computer-readable storage medium or may be downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0251] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0252] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0253] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture comprising instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0254] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0255] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0256] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the marketplace, or to enable other ordinary skilled artisans in the art to understand the embodiments disclosed herein.
Claims
1. A data processing method, characterized in that, Applied to the server side, including: Calculating the modular exponentiation of each data in the first data set under the RSA private key based on a preset algorithm; wherein, the first data set includes user data for which the server is to perform private set intersection; the preset algorithm is a grouped streaming sharding algorithm or a multi - radix modular exponentiation algorithm, the grouped streaming sharding algorithm is used to evenly group multiple triple data and perform parallel processing on different groups obtained by the even grouping; the triple data includes target data, target exponent, and target modulus; the multi - radix modular exponentiation algorithm is used to perform modular exponentiation evaluation based on a multi - radix target exponent; Inserting the modular exponentiation of each data under the RSA private key into a Bloom filter and sending the Bloom filter; Receiving a third data set, the third data set including the modular exponentiation of the product of the modular exponentiations of each data in the second data set calculated based on the grouped streaming sharding algorithm and the corresponding random numbers in the random number set under the RSA public key, or an index table of the exponent values of the modular exponentiations of each random number in the random number set constructed based on the multi - radix modular exponentiation algorithm under the RSA public key; wherein, the second data set includes user data for which the client is to perform private set intersection, the random number set is generated by the client, and each random number corresponds one - to - one with each data in the second data set; Calculating the modular exponentiation of each data in the third data set under the RSA private key based on the preset algorithm to obtain a fourth data set; Sending the fourth data set.
2. The method according to claim 1, characterized in that, The evenly grouping of multiple triple data includes: Calculating the indexes of multiple groups corresponding to the target data in the hash table and determining the target group with the least amount of currently stored data among the multiple groups; Storing the triple data corresponding to the target data in the target group; The parallel processing of different groups obtained by the even grouping includes: Dividing the target group into multiple shards; Performing parallel processing on the multiple shards.
3. The method according to claim 1, characterized in that, The performing of modular exponentiation evaluation based on a multi - radix target exponent includes: Dividing the target exponent according to a preset multi - radix; Constructing an index table of the target exponent values of the target data set; Performing modular exponentiation evaluation based on the divided target exponent and the index table of the target exponent values.
4. The method according to any one of claims 1-3, characterized in that, The method further includes: Obtaining the changed data in the first data set and calculating the modular exponentiation of the changed data under the RSA private key based on the preset algorithm; Determining the target position of the modular exponentiation of the changed data under the RSA private key in the Bloom filter; Sending the target position and the modular exponentiation of the changed data under the RSA private key.
5. A data processing method, characterized in that, Applied to the client side, including: Generate a set of random numbers, and calculate the modular exponentiation of each random number in the set of random numbers under the RSA public key and the modular inverse of each random number based on a preset algorithm; wherein, each random number corresponds one-to-one with each data in the second data set, and the second data set includes user data for which the client is to perform private set intersection; the preset algorithm is a grouped streaming sharding algorithm or a multi - base modular exponentiation algorithm, the grouped streaming sharding algorithm is used to evenly group multiple triple data and perform parallel processing on different groups obtained by the even grouping; the triple data includes target data, target exponent, and target modulus; the multi - base modular exponentiation algorithm is used to evaluate modular exponentiation based on a multi - base target exponent; Receive a Bloom filter; the Bloom filter stores the modular exponentiation of each data in the first data set under the RSA private key; the first data set includes user data for which the server is to perform private set intersection; Send a third data set to the server; the third data set includes the modular exponentiation of the product of the modular exponentiations of each data in the second data set and the corresponding random numbers in the set of random numbers under the RSA public key calculated based on the grouped streaming sharding algorithm, or an index table of the exponent values of the modular exponentiations of each random number in the set of random numbers under the RSA public key constructed based on the multi - base modular exponentiation algorithm; Receive a fourth data set, the fourth data set includes the modular exponentiation of each data in the third data set calculated based on the preset algorithm under the RSA private key; Judge whether the data in the fourth data set is in the Bloom filter to complete the private set intersection process.
6. The method according to claim 5, characterized in that, The step of evenly grouping multiple triple data includes: Calculate the indexes of multiple groups corresponding to the target data in the hash table, and determine the target group with the least amount of currently stored data among the multiple groups; Store the triple data corresponding to the target data in the target group; The step of performing parallel processing on different groups obtained by the even grouping includes: Divide the target group into multiple shards; Perform parallel processing on multiple shards.
7. The method according to claim 5, wherein The step of evaluating modular exponentiation based on a multi - base target exponent includes: Divide the target exponent according to a preset multi - base; Construct an index table of the exponent values of the target data set; Use the divided target exponent to evaluate modular exponentiation based on the index table of exponent values.
8. The method according to any one of claims 5-7, characterized in that, When the preset algorithm is the grouped streaming sharding algorithm, The step of judging whether the data in the fourth data set is in the Bloom filter includes: Calculate the modular exponentiation of the product of each data in the fourth data set and the modular inverse of the corresponding random number of each data in the set of random numbers based on the grouped streaming sharding algorithm; Judge whether the modular exponentiation of the product is in the Bloom filter; Or, The method further includes: Obtain the modular exponentiation of the changed data in the first data set under the RSA private key and the target position in the Bloom filter; Modify the value at the target position in the Bloom filter.
9. A data processing device, characterized in that, The device includes: A first computing module, configured to: calculate the modular exponentiation of each data in the first data set under the RSA private key based on a preset algorithm; wherein, the first data set includes user data to be used for private set intersection by the server; the preset algorithm is a grouped streaming sharding algorithm or a multi - radix modular exponentiation algorithm, the grouped streaming sharding algorithm is used to evenly group multiple triple data and perform parallel processing on different groups obtained by the even grouping; the triple data includes a target data, a target exponent, and a target modulus; the multi - radix modular exponentiation algorithm is used to evaluate the modular exponentiation based on a multi - radix target exponent; A first communication module, configured to insert the modular exponentiation of each data under the RSA private key into a Bloom filter and send the Bloom filter; The first communication module is further configured to receive a third data set, the third data set includes the modular exponentiation of the product of each data in the second data set calculated based on the grouped streaming sharding algorithm and the corresponding random number in the random number set under the RSA public key, or an exponential value retrieval table of the modular exponentiation of each random number in the random number set constructed based on the multi - radix modular exponentiation algorithm under the RSA public key; wherein, the second data set includes user data to be used for private set intersection by the client, the random number set is generated by the client, and each random number corresponds one - to - one with each data in the second data set; The first computing module is further configured to calculate the modular exponentiation of each data in the third data set under the RSA private key based on the preset algorithm to obtain a fourth data set; The first communication module is further configured to send the fourth data set; Or, The device includes: A second computing module, configured to generate the random number set and calculate the modular exponentiation of each random number in the random number set and the modular inverse of each random number under the RSA public key based on the preset algorithm; A second communication module, configured to receive the Bloom filter; the Bloom filter stores the modular exponentiation of each data in the first data set under the RSA private key; The second communication module is further configured to send the third data set to the server; The second communication module is further configured to receive the fourth data set, the fourth data set includes the modular exponentiation of each data in the third data set calculated based on the preset algorithm under the RSA private key; The second computing module is further configured to determine whether the data in the fourth data set is in the Bloom filter to complete the private set intersection process.
10. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to implement the method according to any one of claims 1 to 4 or 5 to 8 when executing the instructions stored in the memory.