Method and device for safely searching top k data by two parties
By dividing data into multiple buckets and using secret sharing and obfuscating circuit protocols, the problem of large communication overhead when two parties securely search for top k data in the prior art is solved, and faster end-to-end running time and more efficient privacy data processing are achieved.
Patent Information
- Application Number
- CN202510242303.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-06-20
AI Technical Summary
When the prior art searches for the top k data on both sides securely, the communication overhead is large and the end-to-end running time is long, which is difficult to effectively reduce.
By dividing n secret shared data into b buckets, using a comparison protocol based on inadvertent transmission under secret sharing, the minimum value in each bucket is obtained, and k data ranking the top k minimum value is determined from the b data through an obfuscated circuit protocol.
Reduces communication overhead and achieves faster end-to-end runtime, which is suitable for secure search of private data in multi-party data sharing scenarios.
Smart Images

Figure CN120180443A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of computers, and in particular, to a method and apparatus for securely finding the top k data between two parties. Background Art
[0002] Currently, the data held by different data holders may contain users' privacy information, and data sharing between data holders may violate users' privacy. In order to enable data circulation among multiple parties, secure multi-party computing is used to support joint computing among multiple parties, extract the value of data, and ensure that the plaintext information of each party's private data is not leaked during multi-party interaction.
[0003] Secure multi-party computing enables multiple mutually untrusted participants to securely compute a given function without revealing inputs and intermediate computation results other than the result. Secret sharing is a method of dispersing a secret among different participants, where each party obtains a part of the secret, called a shard. Only when enough shards are held can the secret be restored; a single shard cannot restore the secret.
[0004] Due to its high efficiency for arithmetic calculations and linear algebra operations, secret sharing is widely used in secure computing for various scenarios. Computations based on secret sharing often involve securely finding the top k data between two parties. Due to the sensitivity of the data, neither party wants to disclose their respective data. In the prior art, when implementing the above secure search, a high communication overhead is required and the end-to-end running time is long. Therefore, an improved solution is needed to reduce the communication overhead and achieve a faster end-to-end running time when implementing the above secure search. Summary of the Invention
[0005] One or more embodiments of this specification describe a method and apparatus for securely finding the top k data between two parties, which can reduce the communication overhead and achieve a faster end-to-end running time.
[0006] In a first aspect, a method for securely finding the top k data between two parties is provided. Any data is distributed between a first party and a second party in the form of secret sharing. The method is executed by the first party and includes:
[0007] For n secret-shared data, divide its own shards into b buckets according to the manner agreed with the second party; where b is greater than k;
[0008] For each of the b buckets, jointly execute a comparison protocol based on oblivious transfer under secret sharing with the second party to obtain the own shards of the b data corresponding to the minimum value in each bucket among the b buckets;
[0009] Execute a garbled circuit protocol jointly with a second party to determine k data with the top k minimum values from the b data.
[0010] In a possible implementation, n is divisible by b, and each bucket has n / b data.
[0011] In a possible implementation, for each of the b buckets, jointly execute a comparison protocol based on oblivious transfer under secret sharing with the second party, including:
[0012] Execute each round of the comparison protocol for the b buckets in parallel; wherein, the communication content of the same round in different buckets is merged and transmitted.
[0013] In a possible implementation, obtaining the local shards of the b data corresponding to the minimum value in each of the b buckets, including:
[0014] Divide the data in a single bucket into several groups in pairs;
[0015] For each pair of data in a group, jointly execute the comparison protocol with the second party to obtain the local shard of the smaller data in each group;
[0016] Iteratively execute the comparison results in pairs and jointly execute the comparison protocol with the second party until the local shard of the minimum value data in the bucket is obtained.
[0017] Furthermore, for each pair of data in a group, jointly execute the comparison protocol with the second party in a parallel processing manner.
[0018] Furthermore, for each pair of data in a group, jointly execute the comparison protocol with the second party, including:
[0019] According to the local shard of the first data and the local shard of the second data held by the local party, obtain the local shard of the difference between the first data and the second data;
[0020] Based on the local shard of the difference, jointly with the second party, judge whether the difference is greater than 0 based on the comparison protocol to obtain a first comparison result;
[0021] According to the first comparison result, jointly with the second party, select the first data or the second data as the smaller data of the two based on the oblivious transfer protocol.
[0022] In a possible implementation, the determining k data with the top k minimum values from the b data includes:
[0023] From the b data, use a garbled circuit for bitonic network sorting to determine k data with the top k minimum values.
[0024] Further, the use of garbled circuits for double-tone network sorting includes:
[0025] Partition each k of the local shards of b data into a group in a manner agreed upon with the second party, and jointly execute a full sorting of the data within each group according to their values with the second party; the full sorting includes a comparison and swap process for two data using garbled circuits;
[0026] Form an ordered sequence from the local shards of the k data in each group after full sorting, and obtain a double-tone sequence of length k by jointly iterating with the second party to pairwise merge the ordered sequences; the pairwise merging of ordered sequences includes performing the comparison and swap process on the data in two sequences;
[0027] Adjust the double-tone sequence of length k to a fully ordered sequence by jointly executing a double-tone sequence merging algorithm with the second party; the double-tone sequence merging algorithm includes the comparison and swap process.
[0028] Further, the pairwise merging of ordered sequences includes:
[0029] Input two monotonically increasing ordered sequences v1 and v2 of length k, and jointly execute a sequential traversal of v1 and a reverse traversal of v2 with the second party, and after merging, obtain a double-tone sequence of length 2k;
[0030] The comparison and swap process includes:
[0031] Jointly execute with the second party that if the value of the i-th element of the ordered sequence v1 is greater than the value of the i-th element of the ordered sequence v2, then swap these two elements, otherwise, keep them unchanged. After pairwise comparison of the elements, retain the k elements of the ordered sequence v1.
[0032] On the second aspect, a device for securely finding the top k data for two parties is provided. Any data is distributed between the first party and the second party in a secret sharing form. This device is set in the first party and includes:
[0033] A bucketing unit for partitioning the local shards of n secret sharing data into b buckets in a manner agreed upon with the second party; where b is greater than k;
[0034] A minimum / maximum finding unit for jointly executing a comparison protocol based on oblivious transfer under secret sharing for each of the b buckets obtained by the bucketing unit with the second party to obtain the local shards of the b data corresponding to the minimum values in each bucket;
[0035] A result determination unit for jointly executing a garbled circuit protocol with the second party to determine k data with the top k minimum values from the b data obtained by the minimum / maximum finding unit.
[0036] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method of the first aspect.
[0037] In a fourth aspect, a computing device is provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method of the first aspect is implemented.
[0038] Through the method and device provided in the embodiments of this specification, in the first aspect, for n secret-shared data, the first party first divides its own shards into b buckets in a manner agreed with the second party; where b is greater than k; then for each of the b buckets, jointly execute the oblivious transfer-based comparison protocol under secret sharing with the second party to obtain the first-party shards of the b data corresponding to the minimum value in each bucket; finally, jointly execute the garbled circuit protocol with the second party to determine k data among the b data that are the top k minimum values. As can be seen from the above, in the embodiments of this specification, when solving for the minimum value in each bucket, the oblivious transfer-based comparison protocol under secret sharing is adopted, which has lower communication volume but higher communication rounds compared to garbled circuits, and the communication rounds can be reduced through a parallel computing process; when determining k data among the b data that are the top k minimum values, garbled circuits are adopted, taking advantage of the constant communication rounds of garbled circuits to achieve a highly serialized comparison process. This solution can reduce communication overhead and achieve a faster end-to-end running time. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0040] Figure 1 It is a schematic diagram of an implementation scenario of an embodiment disclosed in this specification;
[0041] Figure 2 It shows a flowchart of a method for two-party secure search for the top k data according to an embodiment;
[0042] Figure 3 It shows a schematic diagram of the processing process of a bitonic network sorting according to an embodiment;
[0043] Figure 4 It shows a schematic block diagram of a device for two-party secure search for the top k data according to an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] The solution provided in this specification will be described below with reference to the accompanying drawings.
[0045] Figure 1 It is a schematic diagram of the implementation scenario of an embodiment disclosed in this specification. This implementation scenario involves two-party secure search for the top k data. It can be understood that the above two-party secure search is specifically used to find the top k data from n data. Each data includes a keyword key and its corresponding unique identifier id. The basis for sorting each data is the keyword key. The above n data are all private data. Any data is distributed to two parties in a secret sharing manner. That is to say, the keyword key and the unique identifier id of this data are respectively distributed to two parties in a secret sharing manner. After the two-party secure search, the result obtained by one party can include the local shard of the unique identifier id corresponding to each of the top k data, and the local shard of the keyword key. Among them, the two-party secure search needs to ensure that private data will not be leaked, that is, neither the keyword key of the k data included in the search result nor its corresponding unique identifier id can be known by any party.
[0046] As Figure 1 shown, the scenario of two-party secure search for the top k data involves Party A and Party B, also known as the first party and the second party, or Party A and Party B, or P0 and P1. Each participating party can be implemented as any device, platform, server, or cluster of devices with computing and processing capabilities. All parties need to jointly implement the two-party secure search for the top k data while protecting data privacy.
[0047] In the embodiments of this specification, the aforementioned n data can form an input array with a length of n. Each of the two parties has a shard of the input array, and its shard of the array is represented as <(key1, id1),..., (key n , id n )>. Each data in the input array is composed of a keyword and its unique identifier. Among them, key represents the keyword, and id represents the unique identifier corresponding to this keyword. After performing the two-party secure search for the top k data, each of the two parties returns a shard of the output array with a length of k, and its shard of the array is represented as <(o1, ID1),..., (o k , I D k)>. The output array contains k data composed of the smallest k keywords o and their corresponding identifiers ID.
[0048] In addition, it should be noted that for the two-party secure search for the top k data, specifically, it can be to find the data with the largest k keywords key from n data, or to find the data with the smallest k keywords key from n data. The implementation methods of the two are similar and can both be called the Top-k algorithm. In the embodiments of this specification, only the example of finding the data with the smallest k keywords key from n data is used for illustration.
[0049] The Top-k algorithm is an algorithm for selecting the top k elements from a large amount of data and is widely used in business scenarios such as information retrieval, recommendation systems, and data stream processing. These scenarios usually do not require precise Top-k calculation results and can accept approximate solutions. In many business scenarios, the calculation and data interaction involve two parties. Due to the sensitivity of the data, neither party wants to disclose their respective data. In this case, how to efficiently calculate the Top-k problem through secure multi-party computing technology has become the key. The embodiments of this specification aim to reduce the communication cost required for the Top-k problem in secure multi-party computing, thereby improving the performance of the overall solution.
[0050] The basic concepts involved in the embodiments of this specification are briefly described below.
[0051] Secret sharing (SS): That is, a confidential data x is divided into two random numbers [x]1 and [x]2, and each random number is stored and managed by a different entity.
[0052] Oblivious transfer (OT): A cryptographic protocol. The sender has multiple messages m1, m2,..., m n , and the receiver has a selection variable i. After the protocol is executed, the receiver can only obtain the message m i , and has no knowledge of other messages. In addition, the sender cannot know which message the receiver has selected.
[0053] Garbled circuit (GC): A cryptographic protocol for two-party secure computing. Its basic principle is that one party (referred to as the garbler) generates an encrypted truth table for the circuit to be calculated for the other party (referred to as the evaluator) to evaluate.
[0054] Figure 2 The flowchart of the method for two-party secure search for the top k data according to an embodiment is shown. This method can be based on Figure 1 the implementation scenario shown. Any data is distributed between the first party and the second party in the form of secret sharing, and this method is executed by the first party. As Figure 2As shown in the figure, the method for securely finding the top k data between two parties in this embodiment includes the following steps:
[0055] Step 21: For n secretly shared data, divide its own shards into b buckets according to the method agreed with the second party; where b is greater than k. Step 22: Jointly execute the comparison protocol based on oblivious transfer under secret sharing for each of the b buckets with the second party to obtain the own shards of the b data corresponding to the minimum value in each bucket. Step 23: Jointly execute the garbled circuit protocol with the second party to determine k data with the top k minimum values from the b data. The specific implementation methods of the above steps are described below.
[0056] First, in step 21, for n secretly shared data, divide its own shards into b buckets according to the method agreed with the second party; where b is greater than k. It can be understood that the first party and the second party jointly implement dividing the n secretly shared data into b buckets.
[0057] In the embodiments of this specification, the two parties can pre-agree on a random bucketing scheme, including the number of buckets in the scheme. In specific implementations, a common random number seed can be used to ensure the same randomness of the two parties.
[0058] In one example, n is divisible by b, and each bucket has n / b data.
[0059] For example, assume that the input n sequentially random secretly shared data are evenly divided into b buckets, and each bucket has n / b elements. The value of b is set according to the specific usage scenario and requirements. The larger the value of b, the higher the accuracy of the approximate result.
[0060] Then, in step 22, jointly execute the comparison protocol based on oblivious transfer under secret sharing for each of the b buckets with the second party to obtain the own shards of the b data corresponding to the minimum value in each bucket. It can be understood that for the data in each bucket, find the data with the minimum value, and b data are obtained for the b buckets.
[0061] In the embodiments of this specification, the two parties hold the input data under secret sharing after bucketing and expect to find the minimum value in each bucket through the comparison protocol. In the calculation process, the input data can be compared pairwise, and the smaller input data can be selected. By continuously executing this step, the minimum value can be finally obtained. The core is to input two secretly shared data and return the smaller data in the form of secret sharing. If the two data are equal, any one of the data is returned. The implementation of this process uses the comparison protocol based on oblivious transfer under secret sharing.
[0062] The specific process of the above comparison protocol is as follows. The first party holds the input data x0, and the second party holds the input data x1. The two parties transfer information based on the input data through the oblivious transfer protocol to compare the magnitudes of the two data x0 and x1, and return the smaller value of the two in the form of secret sharing. Transferring information through the oblivious transfer protocol can ensure that neither party can obtain any information about the other party's input. Moreover, the oblivious transfer protocol requires less additional communication volume than the garbled circuit protocol itself in the process of comparing two data, but requires a higher number of communication rounds.
[0063] In one example, for each of the b buckets, jointly execute the oblivious transfer-based comparison protocol under secret sharing with the second party, including:
[0064] Execute each round of the comparison protocol in parallel for the b buckets; wherein, the communication content of the same round in different buckets is merged and transmitted.
[0065] In this example, find the minimum value for each of the n / b elements in each bucket, and use the oblivious transfer protocol under secret sharing to complete the operation of the comparison process. Since the calculation processes in the b buckets are completely independent, their calculation processes can be executed in parallel. By merging the communication content of the same round in different buckets, the number of communication rounds only depends on the number of communication rounds required for the calculation in a single bucket. There will not be too many communication rounds during the process of finding the minimum value, and at the same time, it can effectively utilize the characteristic of small communication volume when using the oblivious transfer protocol under secret sharing for comparison.
[0066] In one example, obtain the local shards of the b data corresponding to the minimum values in each of the b buckets, including:
[0067] Divide the data within a single bucket into several groups in pairs;
[0068] For each pair of data in each group, jointly execute the comparison protocol with the second party to obtain the local shard of the smaller data in each group;
[0069] Iteratively execute the operation of taking the comparison results in pairs and jointly execute the comparison protocol with the second party until obtaining the local shard of the minimum value data in the bucket.
[0070] In this example, for the process of finding the minimum value for a single bucket, adopt a tree-shaped comparison protocol. For example, in the process of finding the minimum value for n values, divide the n values into n / 2 groups in pairs, and execute the above comparison protocol for each pair of data in each group to obtain n / 2 smaller results. Then execute the above comparison protocol for the results obtained for the first time in pairs, and so on, until finally obtaining the minimum value within a single bucket. This process is similar to the state of a binary tree gradually comparing in pairs from the leaves and merging to the root node, which is called tree-shaped comparison and can effectively reduce the number of times of pairwise comparison of data.
[0071] Further, for the two data in each group, the comparison protocol is jointly executed with the second party in a parallel processing manner.
[0072] In this example, the aforementioned tree comparison is more conducive to parallelization compared to sequential traversal.
[0073] Further, jointly executing the comparison protocol with the second party for the two data in each group includes:
[0074] Based on the local shards of the first data and the local shards of the second data held by this party, obtain the local shard of the difference between the first data and the second data;
[0075] Based on the local shard of the difference, jointly with the second party, determine whether the difference is greater than 0 based on the comparison protocol to obtain a first comparison result;
[0076] According to the first comparison result, jointly with the second party, select the first data or the second data as the smaller data between the two based on the oblivious transfer protocol.
[0077] For example, when implementing the comparison of two input data x and y for secret sharing, the problem is transformed into whether z = x - y is greater than 0. In the implementation of determining whether the secretly shared z is greater than 0, the first party holds a shard z0 of the data z, and the second party holds another shard z1 of the data z. Each party locally splits its shard into two parts: z0 = MSB(z0) || z0′, z1 = MSB(z1) || z1′, where MSB represents the most significant bit. The problem of determining whether z is greater than 0 can be transformed into determining whether the most significant bit of z is 1 in binary representation. It can be judged by whether the sum of the exclusive OR of the most significant bit of z0, the exclusive OR of the most significant bit of z1, and z0′ and z1′ generates a carry, that is, by judging whether z0′ + z1′ is greater than or equal to 2 l-1 , where 2 l is the size of the ring where the secret sharing is located. Just set the inputs of the comparison protocol to z0′ and 2 l-1 -z1′. Obtain the protocol result through the comparison protocol implemented based on oblivious transfer, and it can be determined whether z is greater than 0. Then select x and y through it. The selection process is implemented through the oblivious transfer protocol, and finally obtain the smaller value between x and y.
[0078] Finally, in step 23, jointly execute the garbled circuit protocol with the second party to determine k data among the b data that are the top k minimum values. It can be understood that determining k data among the b data that are the top k minimum values is an approximate result, which may not be exactly the same as the top k data among the n data. This approximate result meets the requirements of certain business scenarios.
[0079] In the embodiments of this specification, the above b data are in the form of secret sharing, and the obtained k data are also in the form of secret sharing, that is, the data held by both parties are indistinguishable from random numbers, and the result is secret without information leakage.
[0080] In one example, the determining of the k data with the top k minimum values from the b data includes:
[0081] From the b data, use a garbled circuit to perform a bitonic network sort to determine the k data with the top k minimum values.
[0082] In this example, in terms of the specific number of comparisons, using a bitonic sort network has more advantages than simply finding the minimum value k times. A lower number of comparisons can significantly reduce the communication and computing overhead when using a garbled circuit for calculation, and at the same time give play to the advantage of the constant number of comparisons of the garbled circuit in complex circuits.
[0083] Further, the performing of the bitonic network sort using a garbled circuit includes:
[0084] Divide every k of the local shards of the b data into a group according to the method agreed with the second party, and jointly execute a full sort of the data within each group according to their values with the second party; the full sort includes the comparison and swap processing for two data using a garbled circuit;
[0085] Form an ordered sequence from the local shards of the k data in each group after the full sort, and jointly iterate with the second party to pairwise merge the ordered sequences to obtain a bitonic sequence of length k; the pairwise merging of the ordered sequences includes performing the comparison and swap processing on the data in the two sequences;
[0086] Jointly execute a bitonic sequence merging algorithm with the second party to adjust the bitonic sequence of length k to a fully ordered sequence; the bitonic sequence merging algorithm includes the comparison and swap processing.
[0087] For example, the bitonic network sort can be expressed as the following code:
[0088]
[0089] Among them, the bitonic network sort inputs an array of length n, and k represents the desired k minimum data to be obtained. After the bitonic network sort, the algorithm returns an array of length k. Since in the embodiments of this specification, the k data with the top k minimum values are determined from the b data, the specific value of the above n is b. The NetworkTopK function is used to implement the function of the bitonic network sort, and the remaining PartialSort function, KMerge function, and BitonicMerge function are sub-functions to be called in this process.
[0090] In the embodiments of this specification, an array of length n is input into the PartialSort function. The purpose is to evenly divide the array of length n into groups of k data, and perform a full permutation on the data within each group. After completing the partial sorting, each group of k data forms an ordered sequence, with a total of n / k groups. It is necessary to pairwise merge the ordered sequences. The core step is to select two groups of data and use the KMerge function for merging. The KMerge function outputs a bitonic sequence of length k. Through the BitonicMerge function, any bitonic sequence can be adjusted to a completely ordered sequence. Among them, after merging every two ordered sequences of length k, the number of locally ordered sequences is reduced to half of the original. By repeatedly calling the merging process, a set of locally ordered sequences is finally obtained, which is the smallest k data.
[0091] Further, the pairwise merging of the ordered sequences includes:
[0092] Input two monotonically increasing ordered sequences v1 and v2 of length k, and jointly execute with the second party to traverse v1 in order while traversing v2 in reverse order. After merging, a bitonic sequence of length 2k is obtained;
[0093] The comparison and swap process includes:
[0094] Jointly execute with the second party. If the value of the i-th element of the ordered sequence v1 is greater than the value of the i-th element of the ordered sequence v2, then swap these two elements; otherwise, keep them unchanged. After pairwise comparison of the elements, keep the k elements of the ordered sequence v1.
[0095] For example, the PartialSort function can be expressed as the following code:
[0096]
[0097] Among them, the function of the PartialSort function is to divide the original array into n / k blocks, and each block is sorted by keywords internally. In the embodiments of this specification, the sorting algorithm for each part adopts the Odd Even Merge Sort in the specific implementation process, with fewer comparison times.
[0098] The KMerge function can be expressed as the following code:
[0099]
[0100] Among them, the KMerge function takes as input two monotonically increasing ordered sequences v1 and v2 of length k. When sequentially traversing v1 and reversely traversing v2, it satisfies the definition of a bitonic sequence, first monotonically increasing and then monotonically decreasing. By comparing and swapping the elements in the two sequences, if v1[i].key > v2[i].key, then swap v1[i] and v2[i]; otherwise, keep them unchanged. According to the properties of the bitonic sequence, the output v1 is still a bitonic sequence, and the keywords key of all elements in v1 are less than the keywords key in v2. Since the goal of this function is to find the smallest k values, the data in the v2 part can be directly excluded without any further calculation.
[0101] The BitonicMerge function can be expressed as the following code:
[0102]
[0103] Among them, the BitonicMerge function can adjust any bitonic sequence into a completely ordered sequence.
[0104] The CompSwap function can be expressed as the following code:
[0105]
[0106] Among them, the CompSwap function can adjust the order of two data, and the result after adjustment is that the smaller one is in the front and the larger one is in the back.
[0107] In the embodiments of this specification, through the optimization of the BitonicMerge function, it can support adjusting any bitonic sequence of length k back to a monotonic sequence, reducing the extra comparison times brought by filling the sequence to a predetermined value, and being more suitable for the scenario of encrypted state calculation. Currently, the algorithm requires n to be divisible by k. And the process of dividing the data into blocks and performing local sorting is completed based on the odd-even merge sorting network. The odd-even merge sorting network has fewer comparison times compared to the bitonic sorting network. Combining with the encrypted state calculation under the garbled circuit protocol can obtain lower communication and calculation overhead.
[0108] Figure 3 Shows a schematic diagram of the bitonic network sorting process according to an embodiment. Refer to Figure 3, its input is b data, where b = 15, and it is necessary to determine k data with the top k minimum values, where k = 3. The array of length 15 is input into the PartialSort function, and its purpose is to evenly divide the array of length 15 into groups of k data and perform a full permutation on the data within each group. After completing the partial sorting, each group of k data forms an ordered sequence, with a total of 5 groups. It is necessary to pairwise merge the ordered sequences. The core step is to select two groups of data and use the KMerge function (i.e., Top-k Merge) to merge them. The KMerge function outputs a bitonic sequence of length 3. Through the BitonicMerge function, any bitonic sequence can be adjusted to a fully ordered sequence. Among them, after merging every two ordered sequences of length k, the number of locally ordered sequences is reduced to half of the original. By repeatedly calling the merging process, a group of locally ordered sequences is finally obtained, which is the k data with the smallest values. Among them, each vertical line with an arrow represents a call to the CompSwap function.
[0109] In the embodiment of this specification, since the entire bitonic network sorting process, except for the comparison and swap process, the rest of the operations are independent of the input data. For secret-shared data, the comparison and swap process can be implemented through garbled circuits, and the rest of the operations can be completed locally. All the overheads are determined by the number of comparison and swap processes. The number of comparison and swap processes used in the entire bitonic network sorting process is lower than that of the usual algorithms, and a lower communication volume can be achieved during the implementation through garbled circuits.
[0110] Through the method provided by the embodiment of this specification, the first party first divides its own shards of n secret-shared data into b buckets in a manner agreed with the second party; where b is greater than k; then jointly executes a comparison protocol based on oblivious transfer under secret sharing for each of the b buckets with the second party to obtain the first-party shards of the b data corresponding to the minimum value in each bucket; finally, jointly executes a garbled circuit protocol with the second party to determine k data with the top k minimum values from the b data. As can be seen from the above, in the embodiment of this specification, when solving for the minimum value in each bucket, a comparison protocol based on oblivious transfer under secret sharing is adopted, which has a lower communication volume but a higher number of communication rounds compared to garbled circuits, and the number of communication rounds can be reduced through a parallel computing process; when determining k data with the top k minimum values from b data, garbled circuits are adopted, taking advantage of the constant number of communication rounds of garbled circuits to implement a highly serialized comparison process. This solution can reduce communication overhead and achieve a faster end-to-end running time.
[0111] It should be noted that the method for obtaining k minimum values in each embodiment can be similarly and equivalently applied to obtaining k maximum values.
[0112] According to an embodiment of another aspect, there is also provided an apparatus for two-party secure search for the top k data. This apparatus is used to execute the method provided by the embodiments described in this specification. Figure 2 Any data is distributed between the first party and the second party in the form of secret sharing. This apparatus is set in the first party. Figure 4 The schematic block diagram of an apparatus for two-party secure search for the top k data according to an embodiment is shown. As Figure 4 shown, the apparatus 400 includes:
[0113] A bucketing unit 41, configured to divide the shards of its own party of n secret-sharing data into b buckets in a manner agreed with the second party; where b is greater than k.
[0114] A minimum value finding unit 42, configured to jointly execute a comparison protocol based on oblivious transfer under secret sharing with the second party for each of the b buckets obtained by the bucketing unit 41, to obtain the shards of the data corresponding to the minimum value in each of the b buckets.
[0115] A result determination unit 43, configured to jointly execute a garbled circuit protocol with the second party to determine k data with the top k minimum values from the b data obtained by the minimum value finding unit 42.
[0116] Optionally, as an embodiment, n is divisible by b, and each bucket has n / b data.
[0117] Optionally, as an embodiment, the minimum value finding unit 42 is specifically configured to execute each round of the comparison protocol for the b buckets in parallel; where the communication content in the same round for different buckets is merged and transmitted.
[0118] Optionally, as an embodiment, the minimum value finding unit 42 includes:
[0119] A grouping subunit, configured to group the data within a single bucket into several groups in pairs.
[0120] A comparison and exchange subunit, configured to jointly execute the comparison protocol with the second party for the two data in each group obtained by the grouping subunit to obtain the shards of the smaller data in each group.
[0121] An iterative processing subunit, configured to iteratively execute grouping the comparison results obtained by the comparison and exchange subunit in pairs and jointly execute the comparison protocol with the second party until the shards of the minimum value data in the bucket are obtained.
[0122] Further, the comparison and exchange subunit adopts a parallel processing method.
[0123] Further, the comparison and exchange subunit is specifically configured to:
[0124] Obtain the local shard of the difference between the first data and the second data based on the local shards of the first data and the local shards of the second data held by this party;
[0125] Based on the local shard of the difference, jointly with the second party, determine whether the difference is greater than 0 based on the comparison protocol to obtain the first comparison result;
[0126] According to the first comparison result, jointly with the second party, select the first data or the second data as the smaller data of the two based on the oblivious transfer protocol.
[0127] Optionally, as an embodiment, the result determination unit 43 is specifically configured to perform a bitonic network sorting on the b data using a garbled circuit, and determine k data with the top k minimum values.
[0128] Further, the result determination unit 43 includes:
[0129] A grouping and sorting subunit, configured to divide every k of the local shards of the b data into a group in a manner agreed with the second party, and jointly with the second party, perform a full sorting on the data within each group according to their values; the full sorting includes a comparison and exchange process for two data using a garbled circuit;
[0130] A merging processing subunit, configured to form an ordered sequence by combining the local shards of the k data in each group after the full sorting obtained by the grouping and sorting subunit, and jointly with the second party, iteratively perform pairwise merging of the ordered sequences to obtain a bitonic sequence with a length of k; the pairwise merging of the ordered sequences includes performing the comparison and exchange process on the data in the two sequences;
[0131] A bitonic merging subunit, configured to jointly with the second party, perform a bitonic sequence merging algorithm to adjust the bitonic sequence with a length of k obtained by the merging processing subunit into a completely ordered sequence; the bitonic sequence merging algorithm includes the comparison and exchange process.
[0132] Further, the merging processing subunit is specifically configured to input two monotonically increasing ordered sequences v1 and v2 with a length of k, jointly with the second party, perform a sequential traversal of v1 while performing a reverse traversal of v2, and after merging, obtain a bitonic sequence with a length of 2k;
[0133] The comparison and exchange process includes:
[0134] Jointly with the second party, perform if the value of the i-th element of the ordered sequence v1 is greater than the value of the i-th element of the ordered sequence v2, then exchange these two elements, otherwise, keep them unchanged. After pairwise comparison of the elements, retain the k elements of the ordered sequence v1.
[0135] Through the device provided by the embodiments of this specification, the first party first divides its own shards of the n secret sharing data into b buckets by the bucketing unit 41 in a manner agreed with the second party; where b is greater than k; then the minimum value finding unit 42 jointly executes the oblivious transfer-based comparison protocol under secret sharing with the second party for each of the b buckets to obtain the shards of the b data corresponding to the minimum value in each bucket; finally, the result determination unit 43 jointly executes the garbled circuit protocol with the second party to determine k data among the top k minimum values from the b data. As can be seen from the above, in the embodiments of this specification, when finding the minimum value in each bucket, the oblivious transfer-based comparison protocol under secret sharing is adopted, which has lower communication volume but higher communication rounds compared with the garbled circuit, and the communication rounds can be reduced through a parallel computing process; when determining k data among the top k minimum values from the b data, the garbled circuit is adopted, taking advantage of the constant communication rounds of the garbled circuit to achieve a highly serialized comparison process. This solution can reduce the communication overhead and achieve a faster end-to-end running time.
[0136] According to an embodiment of another aspect, there is also provided a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed in a computer, the computer is caused to execute the method described in conjunction with Figure 2 what is described.
[0137] According to an embodiment of still another aspect, there is also provided a computing device including a memory and a processor, where an executable code is stored in the memory, and when the processor executes the executable code, the method described in conjunction with Figure 2 what is described is implemented.
[0138] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0139] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for two parties to securely search for top k ranked data, wherein any data is distributed between a first party and a second party in a secret sharing form, and the method is performed by the first party, comprising: For n secret shared data, divide its own shards into b buckets according to the method agreed with the second party; where b is greater than k; For each of the b buckets, jointly execute the comparison protocol based on oblivious transfer under secret sharing with the second party to obtain the b data fragments corresponding to the minimum value in each of the b buckets; The garbled circuit protocol is jointly executed with the second party to determine the k data with the top k minimum values from the b data.
2. The method of claim 1, wherein: The n is divisible by b, and each bucket contains n / b data.
3. The method of claim 1, wherein: For each of the b buckets, jointly execute a comparison protocol based on oblivious transfer under secret sharing with the second party, including: Each round of the comparison protocol is executed in parallel for b buckets; wherein the communication contents of the same round in different buckets are combined for transmission.
4. The method of claim 1, wherein: Get the local slices of b data corresponding to the minimum value in each bucket of b buckets, including: Divide the data in a single bucket into several groups in pairs; For each group of two data, jointly executing the comparison protocol with the second party to obtain the smaller slice of each group of data; The iterative execution groups the comparison results in pairs, and jointly executes the comparison protocol with the second party until the local fragment with the minimum value data in the bucket is obtained.
5. The method of claim 4, wherein: For each group of two data, the comparison protocol is jointly executed with the second party in a parallel processing manner.
6. The method of claim 4, wherein: The step of executing the comparison protocol jointly with the second party for each group of two data includes: Obtaining a local slice of the difference between the first data and the second data according to the local slice of the first data and the local slice of the second data held by the local party; Based on the difference, the own party's slice and the second party jointly determine whether the difference is greater than 0 based on the comparison protocol to obtain a first comparison result; According to the first comparison result, the first data or the second data is selected as the smaller one of the two based on the oblivious transfer protocol in conjunction with the second party.
7. The method of claim 1, wherein: The step of determining the k data with the top k minimum values from the b data includes: From the b data, a bitonic network sorting is performed using a confusion circuit to determine the top k smallest values of k data.
8. The method of claim 7, wherein: The method of using a garbled circuit to perform a dual-tuned network sorting comprises: Divide each k of the b data slices into a group in accordance with the method agreed with the second party, and jointly perform with the second party a full sorting of the data in each group according to their values; the full sorting includes a comparison and exchange process of two data using a garbled circuit; The k pieces of data in each group after the full sorting are combined into an ordered sequence, and the ordered sequences are merged in pairs by joint iteration with the second party to obtain a bitonic sequence with a length of k; the pairwise merging of the ordered sequences includes performing the comparison and exchange processing on the data in the two sequences; By jointly executing a bitonic sequence merging algorithm with the second party, the bitonic sequence with a length of k is adjusted to a completely ordered sequence; the bitonic sequence merging algorithm includes the comparison and exchange process.
9. The method of claim 8, wherein: The two-by-two merged ordered sequences include: Input two monotonically increasing ordered sequences v1 and v2 of length k, and jointly perform the sequential traversal of v1 and the reverse traversal of v2 with the second party, and after merging, a bitonic sequence of length 2k is obtained; The comparison and exchange process comprises: Execute jointly with the second party: If the value of the i-th element of the ordered sequence v1 is greater than the value of the i-th element of the ordered sequence v2, then swap the two elements; otherwise, keep them unchanged. After the pairwise comparison of the elements is completed, retain the k elements of the ordered sequence v1.
10. A method for two parties to securely search for top k ranked data, wherein any data is distributed between a first party and a second party in a secret sharing form, and the method is performed by the first party, comprising: For n secret shared data, divide its own shards into b buckets according to the method agreed with the second party; where b is greater than k; For each of the b buckets, jointly execute the comparison protocol based on oblivious transfer under secret sharing with the second party to obtain the b data fragments corresponding to the maximum value in each of the b buckets; The garbled circuit protocol is jointly executed with the second party to determine the k data with the top k maximum values from the b data.
11. A device for two parties to securely search for top k ranked data, wherein any data is distributed between a first party and a second party in a secret sharing form, and the device is arranged on the first party, comprising: The bucketing unit is used to divide the n secret shared data into b buckets according to the method agreed with the second party; where b is greater than k; A maximum value finding unit is used to jointly execute a comparison protocol based on oblivious transmission under secret sharing with the second party for each of the b buckets obtained by the bucketing unit, so as to obtain the local fragments of b data corresponding to the minimum value in each of the b buckets; The result determination unit is used to jointly execute the confusion circuit protocol with the second party to determine the k data with the top k minimum values from the b data obtained by the maximum value finding unit.
12. A device for two parties to securely search for top k ranked data, wherein any data is distributed between a first party and a second party in a secret sharing form, and the device is disposed on the first party, comprising: The bucketing unit is used to divide the n secret shared data into b buckets according to the method agreed with the second party; where b is greater than k; A maximum value finding unit is used to jointly execute a comparison protocol based on oblivious transmission under secret sharing with the second party for each of the b buckets obtained by the bucketing unit, so as to obtain the local fragments of b data corresponding to the maximum value in each of the b buckets; The result determination unit is used to jointly execute the confusion circuit protocol with the second party to determine the k data with the top k maximum values from the b data obtained by the maximum value determination unit.
13. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 10.
14. A computing device, comprising a memory and a processor, wherein the memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 10 is implemented.