Privacy set intersection method and system based on optimized inadvertent pseudo-random function

CN122839433APending Publication Date: 2026-09-29LINGSHU TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611316086.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-28
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

其基本问题是:两方或多方各自持有一个私有数据集,如何在保护各方数据隐私的前提下,计算出各方数据集的交集

Benefits of technology

[0017]上述技术方案具有如下有益效果:通过将不经意伪随机函数值分割派生出定位编码值和验证编码值,使双方针对相同元素能够在本地同步获得相同的定位编码值,且仅有相同的定位编码值才能进行正确解码;基于此,发起方无需将本地的每个不经意伪随机函数值与发送方发来的全部不经意伪随机函数值进行全局比对,而是以本地的定位编码值为键对查找编码向量解码后,针对解码结果进行一对一比对。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122839433A_ABST
    Figure CN122839433A_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a privacy set intersection method and system based on optimization of an inadvertent pseudo-random function, wherein the method comprises the following steps: an initiator and a participant generate first and second inadvertent pseudo-random function value sequences according to first and second original data sets respectively; the initiator derives a respective first positioning encoding value and a first verification encoding value for each value in the first inadvertent pseudo-random function value sequence according to a preset rule; the participant derives a respective second positioning encoding value and a second verification encoding value for each value in the second inadvertent pseudo-random function value sequence according to the preset rule; the participant constructs a lookup encoding vector according to the second positioning encoding value and the second verification encoding value and sends the vector to the initiator; and the initiator decodes the lookup encoding vector according to each first positioning encoding value and performs consistency determination, and adds elements determined as consistent to an intersection set. The method can reduce communication volume while ensuring data security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of privacy computing, and in particular to a method and system for finding intersections of privacy sets based on optimizing an unintentional pseudo-random function. Background Technology

[0002] Private Set Intersection (PSI) is an important branch of secure multi-party computation. Its fundamental problem is: given two or more parties each holding a private dataset, how can we compute the intersection of their datasets while protecting the privacy of each party's data? Specifically, assume the initiating party holds the first dataset... The participants hold the second dataset Both parties hope that after the agreement is implemented, the initiating party will only be informed of the intersection. Without revealing any information outside the intersection.

[0003] Privacy-preserving set intersection techniques have wide applications in real-world scenarios. With the rapid development of the internet, big data, and other new technologies, an increasing amount of distributed data requires multi-party collaborative processing, making data privacy protection a growing concern. Privacy-preserving set intersection techniques can unlock data value while protecting data privacy, showing promising application prospects in areas such as privacy-protected real-name authentication, joint risk control, data exploration, contact tracing, private address book lookup, and online advertising effectiveness calculation.

[0004] Currently, various technical solutions have been proposed for privacy set intersection. Among existing technologies, high-performance privacy set intersection protocols are typically based on Vector Inadvertent Linear Estimation (VOLE) and Inadvertent Key-Value Storage (OKVS). Participating parties need to send all their OPRF values ​​to the initiator, who then compares each of their local OPRF values ​​with the received OPRF values ​​to determine the intersection. To ensure that the probability of false positives does not exceed a preset security target, the length of the OPRF values ​​needs to be determined based on statistical security parameters and the size of both parties' datasets. For example, with a statistical security parameter of 40 and both sets having a size of [missing value], [missing value]... In this scenario, each OPRF value needs to reach In bits, the amount of communication during the transmission of OPRF values ​​is proportional to the length of the OPRF value, resulting in significant communication overhead.

[0005] Therefore, how to reduce the communication volume in the intersection phase of the privacy set intersection protocol has become an urgent technical problem to be solved. Summary of the Invention

[0006] To address the aforementioned issues, this invention provides a privacy set intersection method and system based on an optimized unintentional pseudo-random function. By splitting the unintentional pseudo-random function value into a location-coded value and a verification-coded value, the intersection matching method is changed from global comparison to a location-based one-to-one comparison, thereby reducing the communication volume of a single data entry while maintaining the same security level.

[0007] To achieve the above objectives, this invention provides a privacy set intersection method based on an optimized unintentional pseudo-random function, comprising: an initiator generating a first unintentional pseudo-random function value sequence based on a first original dataset held by itself; a participant generating a second unintentional pseudo-random function value sequence based on a second original dataset held by itself; the initiator deriving a corresponding first location code value and a first verification code value for each value in the first unintentional pseudo-random function value sequence according to a preset rule; the participant deriving a corresponding second location code value and a second verification code value for each value in the second unintentional pseudo-random function value sequence according to the preset rule; the participant constructing a lookup code vector based on all the second location code values ​​and the second verification code values ​​and sending it to the initiator; the initiator decoding the lookup code vector based on each first location code value to obtain a decoding return value, and performing a consistency determination between the decoding return value and the corresponding first verification code value, adding the elements in the first original dataset corresponding to the decoded return values ​​that are determined to be consistent to the intersection.

[0008] Further optionally, the initiator generates a first unintentional pseudo-random function value sequence based on its own first original dataset, including: calculating a first hash value for each element in the first original dataset according to a first preset hash function, and calculating a second hash value according to a second preset hash function; performing unintentional key-value storage encoding based on all the first hash values ​​and the second hash values ​​to obtain an initial encoding vector; performing XOR masking on the initial encoding vector based on the first random vector to obtain a mask vector and sending it to the participants; using the first hash value corresponding to each element in the first original dataset as the key, performing unintentional key-value storage decoding on the second random vector to obtain the first unintentional pseudo-random function value sequence; wherein, the initiator and the participants pre-call a vector unintentional linear estimation protocol, so that the initiator obtains the first random vector and the second random vector, and the participants obtain a random scalar value and a third random vector.

[0009] Optionally, the participating party generates a second unintentional pseudo-random function value sequence based on the second original dataset it holds, including: calculating a third hash value and a fourth hash value for each element in the second original dataset according to the first preset hash function; and performing unintentional key-value storage decoding and XOR operations based on the mask vector, the third random vector, the random scalar value, and the third and fourth hash values ​​corresponding to each element in the second original dataset to obtain the second unintentional pseudo-random function value sequence.

[0010] Optionally, the initiator derives a corresponding first location code value and a first verification code value for each value in the first unintentional pseudo-random function value sequence according to a preset rule, including: for each unintentional pseudo-random function value in the first unintentional pseudo-random function value sequence, truncating the first preset length as a first truncation value, concatenating the first truncation value with the corresponding unintentional pseudo-random function value and performing a hash calculation, and using the hash calculation result as the first location code value; truncating a second preset length after the first preset length from the unintentional pseudo-random function value as a second truncation value, and using the second truncation value as the first verification code value; wherein, the sum of the first preset length and the second preset length is equal to the total length of the unintentional pseudo-random function value.

[0011] Optionally, the initiator decodes the lookup encoding vector based on each of the first location encoding values ​​to obtain a decoding return value, and performs a consistency determination between the decoding return value and the corresponding first verification encoding value, adding the elements in the first original dataset corresponding to the decoded return value that is determined to be consistent to the intersection. This includes: performing unintentional key-value storage decoding on the lookup encoding vector using each of the first location encoding values ​​as keys to obtain the corresponding decoding return value; wherein the lookup encoding vector is obtained by the participants performing unintentional key-value storage encoding based on the second location encoding value and the second verification encoding value; determining whether the decoding return value is equal to the corresponding first verification encoding value; if they are equal, adding the elements in the first original dataset corresponding to the decoding return value to the intersection; if they are not equal, determining that the elements corresponding to the decoding return value do not belong to the intersection.

[0012] On the other hand, the present invention also provides a privacy set intersection system based on an optimized unintentional pseudo-random function, comprising: a first sequence generation module, used by an initiator to generate a first unintentional pseudo-random function value sequence based on a first original dataset held by itself; a second sequence generation module, used by a participant to generate a second unintentional pseudo-random function value sequence based on a second original dataset held by itself; a first encoding value derivation module, used by the initiator to derive a corresponding first location encoding value and a first verification encoding value for each value in the first unintentional pseudo-random function value sequence according to a preset rule; a second encoding value derivation module, used by a participant to derive a corresponding second location encoding value and a second verification encoding value for each value in the second unintentional pseudo-random function value sequence according to the preset rule; a vector encoding module, used by a participant to construct a lookup encoding vector based on all the second location encoding values ​​and the second verification encoding values ​​and send it to the initiator; and an intersection module, used by the initiator to decode the lookup encoding vector based on each first location encoding value to obtain a decoding return value, and to determine the consistency between the decoding return value and the corresponding first verification encoding value, and to add the elements in the first original dataset corresponding to the decoded return value that is determined to be consistent to the intersection.

[0013] Further optionally, the first sequence generation module includes: a first hash submodule, used to calculate a first hash value for each element in the first original dataset according to a first preset hash function, and calculate a second hash value according to a second preset hash function; an encoding submodule, used to perform unintentional key-value storage encoding based on all the first hash values ​​and the second hash values ​​to obtain an initial encoding vector; a masking submodule, used to perform XOR masking processing on the initial encoding vector based on a first random vector to obtain a mask vector and send it to the participants; and a first sequence generation submodule, used to perform unintentional key-value storage decoding on the second random vector using the first hash value corresponding to each element in the first original dataset as the key to obtain a first unintentional pseudo-random function value sequence; wherein, the initiator and the participants pre-call a vector unintentional linear estimation protocol, so that the initiator obtains a first random vector and a second random vector, and the participants obtain a random scalar value and a third random vector.

[0014] Further optionally, the second sequence generation module includes: a second hash calculation submodule, used to calculate a third hash value and a fourth hash value for each element in the second original dataset according to the first preset hash function; and a second sequence generation submodule, used to perform unintentional key-value storage decoding operation and XOR operation according to the mask vector, the third random vector, the random scalar value, and the third and fourth hash values ​​corresponding to each element in the second original dataset, to obtain the second unintentional pseudo-random function value sequence.

[0015] Optionally, the first encoding value derivation module includes: a positioning encoding value calculation submodule, used to, for each unintentional pseudo-random function value in the first unintentional pseudo-random function value sequence, extract a first preset length from the beginning of it as a first truncation value, concatenate the first truncation value with the corresponding unintentional pseudo-random function value, perform a hash calculation, and use the hash calculation result as the first positioning encoding value; and a verification encoding value calculation submodule, used to, extract a second preset length from the unintentional pseudo-random function value after the first preset length as a second truncation value, and use the second truncation value as a first verification encoding value; wherein, the sum of the first preset length and the second preset length is equal to the total length of the unintentional pseudo-random function value.

[0016] Further optionally, the intersection module includes: a decoding return value calculation submodule, used to perform unintentional key-value storage decoding on the lookup encoding vector with each of the first positioning encoding values ​​as keys to obtain the corresponding decoding return value; wherein the lookup encoding vector is obtained by the participants performing unintentional key-value storage encoding based on the second positioning encoding value and the second verification encoding value; and an intersection calculation submodule, used to determine whether the decoding return value is equal to the corresponding first verification encoding value. If they are equal, the elements in the first original dataset corresponding to the decoding return value are added to the intersection; if they are not equal, the elements corresponding to the decoding return value are determined not to belong to the intersection.

[0017] The above technical solution has the following beneficial effects: by deriving location code values ​​and verification code values ​​from the unintentional pseudo-random function values, both parties can synchronously obtain the same location code value locally for the same element, and only the same location code value can be correctly decoded; based on this, the initiator does not need to perform a global comparison between each unintentional pseudo-random function value locally and all unintentional pseudo-random function values ​​sent by the sender, but instead uses the local location code value as the key to look up the code vector and decode it, and then performs a one-to-one comparison on the decoding result.

[0018] Because the matching method has changed from global comparison to one-to-one comparison, each initiating element triggers only one comparison, and the total number of comparisons in the protocol has increased. The second reduction to Therefore, under the same statistical safety parameters Under these conditions, the length of the verification code value only needs to meet the following requirements. Compared to a complete, unintentional pseudo-random function, the length of the function value is reduced. This reduces the communication overhead of the intersection phase while maintaining the same level of security. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a flowchart of the privacy set intersection method based on optimizing an unintentional pseudo-random function provided in an embodiment of the present invention; Figure 2 This is a flowchart of the first unintentional pseudo-random function value sequence generation method provided in the embodiments of the present invention; Figure 3 This is a flowchart of the second unintentional pseudo-random function value sequence generation method provided in the embodiments of the present invention; Figure 4 This is a flowchart of the encoded value generation method provided in the embodiments of the present invention; Figure 5 This is a flowchart of the intersection generation method provided in the embodiments of the present invention; Figure 6 This is a schematic diagram of the privacy set intersection system based on optimized unintentional pseudo-random functions provided in an embodiment of the present invention; Figure 7 This is a schematic diagram of the structure of the first sequence generation module provided in an embodiment of the present invention; Figure 8 This is a schematic diagram of the structure of the second sequence generation module provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of the structure of the first encoded value derivation module provided in an embodiment of the present invention; Figure 10 This is a schematic diagram of the intersection module provided in an embodiment of the present invention.

[0021] Reference numerals: 100-First sequence generation module; 1001-First hash submodule; 1002-Encoding submodule; 1003-Mask processing submodule; 1004-First sequence generation submodule; 200-Second sequence generation module; 2001-Second hash calculation submodule; 2002-Second sequence generation submodule; 300-First encoded value derivation module; 3001-Location encoded value calculation submodule; 3002-Verification encoded value calculation submodule; 400-Second encoded value derivation module; 500-Vector encoding module; 600-Intersection module; 6001-Decoding return value calculation submodule; 6002-Intersection calculation submodule. Detailed Implementation

[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0023] To address the technical problem of effectively reducing the communication volume of privacy set intersection protocols while ensuring the accuracy of intersection determination and data privacy security in existing technologies, this invention provides a privacy set intersection method based on optimizing an unintentional pseudo-random function. Figure 1 This is a flowchart of the privacy set intersection method based on optimizing an unintentional pseudo-random function provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes: S1. The initiator generates a first unintentional pseudo-random function value sequence based on the first original dataset it holds.

[0024] The initiator, R, owns the first original dataset. The initiator R holds the first original dataset As input, interactive computation is performed with participant S to obtain the first original dataset. Each original data element generates a corresponding OPRF (unintentional pseudo-random function) value, which is the first unintentional pseudo-random function value.

[0025] The first unintentional pseudo-random function value sequence That is: the initiator R on the first original dataset The sequence is formed by all the first unintentional pseudo-random function values ​​obtained after each original data element in the dataset is calculated using an unintentional pseudo-random function.

[0026] Without knowing the secret parameters held by participant S, initiator R calculates deterministic pseudo-random output values ​​for S's original data elements by executing an interaction protocol. These output values ​​are unique; the function values ​​obtained for the same original data element are consistent between the two parties, while the function values ​​for different original data elements are distinct and computationally indistinguishable.

[0027] Therefore, each value in the first unintentional pseudo-random function value sequence can serve as a unique and unforgeable digital fingerprint of the corresponding original data element, replacing the original data element itself in subsequent privacy set intersection operations, thereby ensuring the accuracy of intersection determination without exposing the original data.

[0028] S2. The participants generate a second unintentional pseudo-random function value sequence based on the second original dataset they hold.

[0029] Participant S itself holds the second original dataset. Participant S used this second original dataset. As input, an interactive computation is performed with the initiator R to generate a second original dataset. Each original data element in the dataset generates a corresponding second unintentional pseudo-random function value.

[0030] The second unintentional pseudo-random function value sequence That is: Participant S on the second original dataset The sequence is formed by all the second unintentional pseudo-random function values ​​obtained after each original data element in the dataset is calculated using an unintentional pseudo-random function.

[0031] Without knowing the input data held by the initiator R, participant S calculates a deterministic pseudo-random output value for its original data elements by executing an interaction protocol with the initiator R. This output value has the same properties as the first unintentional pseudo-random function value: for the same original data element, both parties obtain the same function value; for different original data elements, their corresponding function values ​​are different and computationally indistinguishable.

[0032] In this sequence, if a participant's original data is the same as a participant's original data, the unintentional pseudo-random function values ​​generated by both parties for that same data will necessarily be the same; if the participants' original data are different, the corresponding unintentional pseudo-random function values ​​will be different, thus ensuring the accuracy of subsequent intersection determination.

[0033] S3. The initiator derives the corresponding first positioning code value and first verification code value for each value in the first unintentional pseudo-random function value sequence according to the preset rules.

[0034] Initiator R obtains the first unintentional pseudo-random function value sequence Then, for each unintentional pseudo-random function value in the sequence, a first positioning code value is derived for each value according to the preset derivation rules agreed upon by both parties during the initialization phase. and a first verification code value Thus, the first positioning code value sequence and the first verification code value sequence are obtained.

[0035] The first location encoding value is used to retrieve the key of the encoding vector in a subsequent query as the initiator, and the first verification encoding value is used to determine consistency with the decoded return value. Through the above derivation rules, each unintentional pseudo-random function value is transformed into a pair of encoding values ​​with different uses.

[0036] It should be noted that the above derivation rules are deterministic. That is, the same input value will necessarily derive the same location code value and verification code value; different input values ​​will not probabilistically derive the same location code value.

[0037] S4. The participants derive their respective second positioning code value and second verification code value for each value in the second unintentional pseudo-random function value sequence according to preset rules.

[0038] Participant S obtains the second unintentional pseudo-random function value sequence Then, for each unintentional pseudo-random function value in the sequence, a second location code value is derived for each value according to the same preset derivation rule as the initiator. and a second verification code value Thus, the second positioning code value sequence and the second verification code value sequence are obtained.

[0039] Since both parties use the same derivation rules, for the same original data element, the derived positioning code value and verification code value are the same; for different original data elements, the derived code values ​​are different.

[0040] S5. The participating parties construct a lookup encoding vector based on all the second location encoding values ​​and the second verification encoding values ​​and send it to the initiator.

[0041] After deriving the second location code value sequence and the second verification code value sequence, participant S uses all the second location code values ​​as keys and the second verification code values ​​corresponding to each key as values ​​to perform the unintentional key-value storage encoding OKVS, generating a lookup encoding vector. .Right now .

[0042] The lookup encoding vector is essentially a compressed lookup structure. For any given lookup key, the value corresponding to the key can be recovered from the lookup encoding vector through decoding. For keys that were not input as keys during encoding, the decoding result is unpredictable.

[0043] Participant S will use the lookup encoding vector The lookup encoding vector is sent to the initiator R. In this step, the participant only sends the lookup encoding vector, not the second location encoding value or the second verification encoding value itself. Due to the one-way nature of the unintentional key-value storage encoding, even if the initiator obtains the lookup encoding vector, it cannot deduce the participant's second location encoding value, second verification encoding value, or any element from the second original dataset, thus ensuring the privacy and security of the participant's data.

[0044] S6. The initiator decodes the lookup encoding vector according to each first positioning encoding value to obtain the decoding return value, and performs a consistency judgment between the decoding return value and the corresponding first verification encoding value. The elements in the first original dataset corresponding to the decoding return value that is judged to be consistent are added to the intersection.

[0045] Initiator R receives the lookup encoding vector sent by participant S Then, each of its own derived first positioning code values Using the key, perform an unintentional key-value storage decoding operation on the lookup encoding vector to obtain the decoding return value corresponding to each first location encoding value.

[0046] The initiator performs a judgment operation on each element, that is, compares the decoding return value of the element with the first verification code value corresponding to the element. If the two are consistent, the original data element is determined to be an intersection element and the element is added to the intersection result; if the two are inconsistent, the original data element is determined not to belong to the intersection.

[0047] The correctness of the above determination process is based on the following property: For an element that exists in both parties' datasets, when the participating parties construct the lookup encoding vector, they encode it with the second location encoding value corresponding to the element as the key and the second verification encoding value as the value. The initiating party decodes it with the same first location encoding value (which is the same as the second location encoding value), and will inevitably obtain a decoding return value that is the same as the second verification encoding value. Since this value is consistent with the initiating party's first verification encoding value, the determination is successful. For an element that exists only in the initiating party's dataset, its first location encoding value is not encoded as a key when the participating parties construct the lookup encoding vector, and the decoding return value cannot match the first verification encoding value, so the determination fails.

[0048] The above methods for finding intersections of privacy sets can be applied to, but are not limited to, the following scenarios: Privacy-preserving contact matching: In social applications, users can discover other registered contacts without revealing their own contact list, thus protecting the privacy of non-contacts from being exposed. Joint risk control and anti-fraud: Financial institutions and data service providers can identify common risk users without exchanging their respective customer lists, which can be used for credit assessment or fraud detection; Privacy-protected medical data sharing: Medical institutions can collaborate with pharmaceutical companies or disease control centers without disclosing patient identity information for purposes such as disease tracking or drug efficacy analysis; Online advertising performance attribution: Advertisers and media outlets can calculate the intersection of converting users without disclosing their respective user profiles, thereby evaluating the effectiveness of advertising campaigns and avoiding the leakage of user privacy.

[0049] Taking medical data sharing as an example, a medical institution is the task initiator (R), and a disease control center (CDC) is the task participant (S). The medical institution holds a patient dataset (X), and the CDC holds a close contact dataset (Y). Both parties want to find the intersection elements that exist in both datasets without disclosing their original data, namely, the list of confirmed patients who are also close contacts.

[0050] In practice, R (the medical institution) and S (the CDC) first generate unintentional pseudo-random function value sequences for their respective datasets through an interactive protocol. This ensures that the same person generates the same function value, different people generate different function values, and neither party can access the other's original data. Subsequently, both parties derive location and verification codes from each function value according to the same rules. S (the CDC) constructs a lookup code vector using all its location and verification codes and sends it to R (the medical institution). R (the medical institution) queries this lookup code vector using its own location code as the key. If the decoded return value matches the corresponding verification code value, the patient is considered to belong to the intersection, meaning the person exists in both the patient dataset and the close contact dataset; otherwise, they are not considered to belong to the intersection.

[0051] As an optional implementation method, Figure 2 This is a flowchart of the first unintentional pseudo-random function value sequence generation method provided in the embodiments of the present invention, such as... Figure 2 As shown, the initiator generates a first unintentional pseudo-random function value sequence based on the first original dataset it holds, including: S101. For each element in the first original dataset, calculate the first hash value according to the first preset hash function and calculate the second hash value according to the second preset hash function.

[0052] In the process of generating the first unintentional pseudo-random function value sequence by the initiator R, the first original dataset is first processed. Each raw data element in Each uses a pre-agreed first hash function. Second preset hash function Perform calculations to obtain the first hash value corresponding to each element. Second hash value .

[0053] The first and second preset hash functions are two different cryptographic hash functions, both of which output hash values ​​of fixed length, such as... .

[0054] S102. Perform unintentional key-value storage encoding based on all first hash values ​​and second hash values ​​to obtain the initial encoding vector.

[0055] The initiator R takes all the calculated first hash values ​​and their corresponding second hash values ​​as input to construct... There are 10 key-value pairs. In each key-value pair, for each element in the first original dataset... Using the first hash value corresponding to the element As the key, use the second hash value corresponding to the element. As values, forming key-value pairs The initiator will... The key-value pair input is inadvertently stored in the encoder, which performs encoding operations to generate an encoding vector, denoted as the initial encoding vector. .

[0056] Unintentional key-value storage encoding is a data structure encoding method that encodes a set of key-value pairs into a fixed-length vector. For each key-value pair input during encoding, the corresponding value can be correctly restored when the encoded vector is decoded using that key. For keys not input during encoding, the decoding result has pseudo-randomness and will not reveal any information about the stored key-value pairs.

[0057] For parameters used in unintentional key-value stores, the definition is completed during the initialization phase. These parameters include expansion factors. , encoding vector length ,bandwidth Where N is the number of elements, and in this embodiment, an expansion factor is selected. .

[0058] S103. Perform XOR masking on the initial encoding vector according to the first random vector to obtain the mask vector and send it to the participants.

[0059] Initiator R will use the initial encoding vector With the first random vector Perform an XOR operation, that is, calculate the result of bitwise XORing the initial encoded vector with the first random vector, and use it as the mask vector. The first random vector is obtained in advance by the initiator R and the participant S through negotiation using the VOLE protocol (Vector Unintentional Linear Valuation Protocol).

[0060] The initiator R will calculate the mask vector It is sent to participant S via a communication channel for subsequent generation of a second unintentional pseudo-random function value sequence.

[0061] S104. Using the first hash value corresponding to each element in the first original dataset as the key, perform unintentional key-value storage decoding on the second random vector to obtain the first unintentional pseudo-random function value sequence; wherein, the initiator and the participants pre-call the vector unintentional linear estimation protocol so that the initiator obtains the first random vector and the second random vector, and the participants obtain random scalar values ​​and the third random vector.

[0062] The initiator uses each first hash value As the key, with the second random vector (Obtained beforehand via an unintentional linear estimation protocol) as the target decoding vector, an unintentional key-value storage decoding operation is performed on each element in the first original dataset. For each key... The decoder starts from the second random vector. Extract the value at the corresponding position and output a fixed-length decoding result, which is the first unintentional pseudo-random function value corresponding to that element. After performing the above decoding operation on all elements in the first original dataset, the initiator obtains a sequence of first unintentional pseudo-random function values. .

[0063] In the above steps, before generating the OPRF sequence, the initiator R and the participant S invoke the Vector Unintentional Linear Estimation (VOLE) protocol. This protocol is a two-party random correlation generation protocol. After execution, the initiator R obtains the first random vector. Second random vector Participant S receives a random scalar value. and the third random vector Furthermore, the outputs of both parties satisfy a preset linear correlation, such as: .

[0064] As an optional implementation method, Figure 3 This is a flowchart of the second unintentional pseudo-random function value sequence generation method provided in the embodiments of the present invention, such as... Figure 3 As shown, the participants generate a second unintentional pseudo-random function value sequence based on the second original dataset they possess, including: S201. For each element in the second original dataset, calculate the third hash value according to the first preset hash function and the fourth hash value according to the second preset hash function.

[0065] In the process of participant S generating the second unintentional pseudo-random function value sequence, the second original dataset is first processed. Each raw data element in The third hash value for each element is calculated using the same first and second preset hash functions as the initiator R. and the fourth hash value Both parties use the same set of hash functions for hash calculation to ensure that the same original data will produce the same hash value in both calculations.

[0066] S202. Based on the mask vector, the third random vector, the random scalar value, and the third and fourth hash values ​​corresponding to each element in the second original dataset, perform unintentional key-value storage decoding and XOR operations to obtain the second unintentional pseudo-random function value sequence.

[0067] Participant S receives the mask vector sent by initiator R. Then, based on the third hash value and the fourth hash value And a third random vector obtained through the vector inadvertent linear estimation protocol. and random scalar values The second unintentional pseudo-random function value sequence is calculated for the second original dataset in the following manner. : .

[0068] As an optional implementation method, Figure 4 This is a flowchart of the encoded value generation method provided in the embodiments of the present invention, such as... Figure 4 As shown, the initiator derives a corresponding first positioning code value and a first verification code value for each value in the first unintentional pseudo-random function value sequence according to a preset rule, including: S301. For each unintentional pseudo-random function value in the first unintentional pseudo-random function value sequence, the first preset length of its beginning is taken as the first truncation value. The first truncation value is concatenated with the corresponding unintentional pseudo-random function value and then a hash calculation is performed. The hash calculation result is taken as the first positioning code value.

[0069] Initiator R obtains the first unintentional pseudo-random function value sequence Then, for each unintentional pseudo-random function value in the sequence Perform the following derivation operations: First, from the value of this unintentional pseudo-random function Starting from the initial position, the first preset length of its front portion is extracted. The bit, as the first truncation value, is denoted as... .

[0070] Then, the first truncation value With complete unintentional pseudo-random function values By concatenating the bits, we obtain a concatenated bit string with a length of [length missing]. ,in The total length of the unintentional pseudo-random function values.

[0071] Finally, perform a hash calculation on the concatenated bit string, and use the hash result as the first location code value corresponding to the unintentional pseudo-random function value. .

[0072] S302. From the unintentional pseudo-random function value, extract a second preset length after the first preset length as the second truncation value, and use the second truncation value as the first verification code value; wherein, the sum of the first preset length and the second preset length is equal to the total length of the unintentional pseudo-random function value.

[0073] Initiator R for the first unintentional pseudo-random function value sequence Each unintentional pseudo-random function value in After extracting the first truncation value, subsequent bits are extracted from the same unintentional pseudo-random function value to form the second truncation value. The second truncation value is the first bit extracted from the same unintentional pseudo-random function value. Starting from the first preset length (i.e., immediately following the first preset length), continuously extract the second preset length. Bit, denoted as This involves truncating the last t bits of the unintentional pseudo-random function value. The initiator directly uses this second truncated value as the first verification code value corresponding to the unintentional pseudo-random function value. .

[0074] First preset length With the second preset length The sum equals the total length of the unintentional pseudo-random function values. ,Right now Therefore, for each unintentional pseudo-random function value, after the above two-step truncation operation, all bits of the value are completely divided into two parts: the front part... The first bit is used as the truncated value to generate the location code value, and the subsequent bits... The bit is used as the second truncation value and directly as the verification encoding value.

[0075] Participant S on the second unintentional pseudo-random function value sequence For each value in the sequence, a derivation operation is performed according to the same rules as the initiator R to obtain the corresponding second location code value. Second verification code value Since both parties use the same truncation position and truncation length for the same unintentional pseudo-random function value, the derived positioning code value and verification code value are the same for the same original data element.

[0076] In a preferred embodiment, statistical security parameters are determined by both parties during the initialization phase. The first preset length is 40. Second preset length In this case, since the intersection matching method changes from global matching to one-to-one matching, the amount of communication required to transmit a single piece of data in the intersection stage increases from the length of the original complete unintentional pseudo-random function value. Reduce to .

[0077] Based on the original dataset size of both parties In this case, the total communication volume of the original protocol is In this embodiment, OKVS encoding is performed based on a 64-bit verification code value, reducing the communication volume to... bit, here .

[0078] As an optional implementation method, Figure 5 This is a flowchart of the intersection generation method provided in the embodiments of the present invention, such as... Figure 5 As shown, the initiator decodes the lookup encoding vector based on each first positioning encoding value to obtain a decoding return value. Then, it performs a consistency check between the decoding return value and the corresponding first verification encoding value, adding the elements from the first original dataset corresponding to the decoded return values ​​that are determined to be consistent to the intersection, including: S601. Using each first positioning code value as a key, perform unintentional key-value storage decoding on the lookup code vector to obtain the corresponding decoding return value; wherein, the lookup code vector is obtained by the participants performing unintentional key-value storage encoding based on the second positioning code value and the second verification code value.

[0079] Initiator R receives the lookup encoding vector sent by participant S Then, for each first location code value derived from itself , with that As a query key, for the lookup encoding vector Perform an unintentional key-value storage decoding operation to obtain the corresponding decoding return value.

[0080] The lookup encoding vector is pre-constructed by participant S. Participant S uses all its derived second-position encoding values. For each key, a second verification code value corresponding to that key is used. The lookup encoding vector is generated after performing unintentional key-value storage encoding on the value. .

[0081] S602. Determine whether the decoding return value is equal to the corresponding first verification code value. If they are equal, add the element in the first original dataset corresponding to the decoding return value to the intersection. If they are not equal, determine that the element corresponding to the decoding return value does not belong to the intersection.

[0082] The decoded return value obtained from decoding is compared with the first verification code value corresponding to that element. Perform a consistency comparison to determine if the two are equal. If the decoded value matches the first verification code value... If they match, the element is determined to be an intersection element and is added to the intersection result set; otherwise, the element is determined not to belong to the intersection and is not added to the intersection result set.

[0083] The final intersection .

[0084] This invention also provides a privacy set intersection system based on optimizing an unintentional pseudo-random function. Figure 6 This is a schematic diagram of the privacy set intersection system based on optimized unintentional pseudo-random functions provided in an embodiment of the present invention, as shown below. Figure 6 As shown, the system includes: The first sequence generation module 100 is used by the initiator to generate a first unintentional pseudo-random function value sequence based on the first original dataset it holds.

[0085] The initiator, R, owns the first original dataset. The initiator R holds the first original dataset As input, interactive computation is performed with participant S to obtain the first original dataset. Each original data element generates a corresponding OPRF (unintentional pseudo-random function) value, which is the first unintentional pseudo-random function value.

[0086] The first unintentional pseudo-random function value sequence That is: the initiator R on the first original dataset The sequence is formed by all the first unintentional pseudo-random function values ​​obtained after each original data element in the dataset is calculated using an unintentional pseudo-random function.

[0087] Without knowing the secret parameters held by participant S, initiator R calculates deterministic pseudo-random output values ​​for S's original data elements by executing an interaction protocol. These output values ​​are unique; the function values ​​obtained for the same original data element are consistent between the two parties, while the function values ​​for different original data elements are distinct and computationally indistinguishable.

[0088] Therefore, each value in the first unintentional pseudo-random function value sequence can serve as a unique and unforgeable digital fingerprint of the corresponding original data element, replacing the original data element itself in subsequent privacy set intersection operations, thereby ensuring the accuracy of intersection determination without exposing the original data.

[0089] The second sequence generation module 200 is used by the participants to generate a second unintentional pseudo-random function value sequence based on the second original dataset they hold.

[0090] Participant S itself holds the second original dataset. Participant S used this second original dataset. As input, an interactive computation is performed with the initiator R to generate a second original dataset. Each original data element in the dataset generates a corresponding second unintentional pseudo-random function value.

[0091] The second unintentional pseudo-random function value sequence That is: Participant S on the second original dataset The sequence is formed by all the second unintentional pseudo-random function values ​​obtained after each original data element in the dataset is calculated using an unintentional pseudo-random function.

[0092] Without knowing the input data held by the initiator R, participant S calculates a deterministic pseudo-random output value for its original data elements by executing an interaction protocol with the initiator R. This output value has the same properties as the first unintentional pseudo-random function value: for the same original data element, both parties obtain the same function value; for different original data elements, their corresponding function values ​​are different and computationally indistinguishable.

[0093] In this sequence, if a participant's original data is the same as a participant's original data, the unintentional pseudo-random function values ​​generated by both parties for that same data will necessarily be the same; if the participants' original data are different, the corresponding unintentional pseudo-random function values ​​will be different, thus ensuring the accuracy of subsequent intersection determination.

[0094] The first encoding value derivation module 300 is used by the initiator to derive the corresponding first positioning encoding value and first verification encoding value for each value in the first unintentional pseudo-random function value sequence according to preset rules.

[0095] Initiator R obtains the first unintentional pseudo-random function value sequence Then, for each unintentional pseudo-random function value in the sequence, a first positioning code value is derived for each value according to the preset derivation rules agreed upon by both parties during the initialization phase. and a first verification code value Thus, the first positioning code value sequence and the first verification code value sequence are obtained.

[0096] The first location encoding value is used to retrieve the key of the encoding vector in a subsequent query as the initiator, and the first verification encoding value is used to determine consistency with the decoded return value. Through the above derivation rules, each unintentional pseudo-random function value is transformed into a pair of encoding values ​​with different uses.

[0097] It should be noted that the above derivation rules are deterministic. That is, the same input value will necessarily derive the same location code value and verification code value; different input values ​​will not probabilistically derive the same location code value.

[0098] The second code value derivation module 400 is used by the participants to derive their respective second positioning code value and second verification code value for each value in the second unintentional pseudo-random function value sequence according to preset rules.

[0099] Participant S obtains the second unintentional pseudo-random function value sequence Then, for each unintentional pseudo-random function value in the sequence, a second location code value is derived for each value according to the same preset derivation rule as the initiator. and a second verification code value Thus, the second positioning code value sequence and the second verification code value sequence are obtained.

[0100] Since both parties use the same derivation rules, for the same original data element, the derived positioning code value and verification code value are the same; for different original data elements, the derived code values ​​are different.

[0101] The vector encoding module 500 is used by the participants to construct a lookup encoding vector based on all the second positioning encoding values ​​and the second verification encoding values, and then send it to the initiator.

[0102] After deriving the second location code value sequence and the second verification code value sequence, participant S uses all the second location code values ​​as keys and the second verification code values ​​corresponding to each key as values ​​to perform the unintentional key-value storage encoding OKVS, generating a lookup encoding vector. .Right now .

[0103] The lookup encoding vector is essentially a compressed lookup structure. For any given lookup key, the value corresponding to the key can be recovered from the lookup encoding vector through decoding. For keys that were not input as keys during encoding, the decoding result is unpredictable.

[0104] Participant S will use the lookup encoding vector The lookup encoding vector is sent to the initiator R. In this step, the participant only sends the lookup encoding vector, not the second location encoding value or the second verification encoding value itself. Due to the one-way nature of the unintentional key-value storage encoding, even if the initiator obtains the lookup encoding vector, it cannot deduce the participant's second location encoding value, second verification encoding value, or any element from the second original dataset, thus ensuring the privacy and security of the participant's data.

[0105] The intersection module 600 is used by the initiator to decode the lookup encoding vector according to each first positioning encoding value, obtain the decoding return value, and make a consistency judgment between the decoding return value and the corresponding first verification encoding value, and add the elements in the first original dataset corresponding to the decoding return value that is determined to be consistent to the intersection.

[0106] Initiator R receives the lookup encoding vector sent by participant S Then, each of its own derived first positioning code values Using the key, perform an unintentional key-value storage decoding operation on the lookup encoding vector to obtain the decoding return value corresponding to each first location encoding value.

[0107] The initiator performs a judgment operation on each element, that is, compares the decoding return value of the element with the first verification code value corresponding to the element. If the two are consistent, the original data element is determined to be an intersection element and the element is added to the intersection result; if the two are inconsistent, the original data element is determined not to belong to the intersection.

[0108] The correctness of the above determination process is based on the following property: For an element that exists in both parties' datasets, when the participating parties construct the lookup encoding vector, they encode it with the second location encoding value corresponding to the element as the key and the second verification encoding value as the value. The initiating party decodes it with the same first location encoding value (which is the same as the second location encoding value), and will inevitably obtain a decoding return value that is the same as the second verification encoding value. Since this value is consistent with the initiating party's first verification encoding value, the determination is successful. For an element that exists only in the initiating party's dataset, its first location encoding value is not encoded as a key when the participating parties construct the lookup encoding vector, and the decoding return value cannot match the first verification encoding value, so the determination fails.

[0109] The above methods for finding intersections of privacy sets can be applied to, but are not limited to, the following scenarios: Privacy-preserving contact matching: In social applications, users can discover other registered contacts without revealing their own contact list, thus protecting the privacy of non-contacts from being exposed. Joint risk control and anti-fraud: Financial institutions and data service providers can identify common risk users without exchanging their respective customer lists, which can be used for credit assessment or fraud detection; Privacy-protected medical data sharing: Medical institutions can collaborate with pharmaceutical companies or disease control centers without disclosing patient identity information for purposes such as disease tracking or drug efficacy analysis; Online advertising performance attribution: Advertisers and media outlets can calculate the intersection of converting users without disclosing their respective user profiles, thereby evaluating the effectiveness of advertising campaigns and avoiding the leakage of user privacy.

[0110] Taking medical data sharing as an example, a medical institution is the task initiator (R), and a disease control center (CDC) is the task participant (S). The medical institution holds a patient dataset (X), and the CDC holds a close contact dataset (Y). Both parties want to find the intersection elements that exist in both datasets without disclosing their original data, namely, the list of confirmed patients who are also close contacts.

[0111] In practice, R (the medical institution) and S (the CDC) first generate unintentional pseudo-random function value sequences for their respective datasets through an interactive protocol. This ensures that the same person generates the same function value, different people generate different function values, and neither party can access the other's original data. Subsequently, both parties derive location and verification codes from each function value according to the same rules. S (the CDC) constructs a lookup code vector using all its location and verification codes and sends it to R (the medical institution). R (the medical institution) queries this lookup code vector using its own location code as the key. If the decoded return value matches the corresponding verification code value, the patient is considered to belong to the intersection, meaning the person exists in both the patient dataset and the close contact dataset; otherwise, they are not considered to belong to the intersection.

[0112] As an optional implementation method, Figure 7 This is a schematic diagram of the structure of the first sequence generation module provided in an embodiment of the present invention, as shown below. Figure 7 As shown, the first sequence generation module 100 includes: The first hash submodule 1001 is used to calculate a first hash value for each element in the first original dataset according to a first preset hash function and a second hash value according to a second preset hash function.

[0113] In the process of generating the first unintentional pseudo-random function value sequence by the initiator R, the first original dataset is first processed. Each raw data element in Each uses a pre-agreed first hash function. Second preset hash function Perform calculations to obtain the first hash value corresponding to each element. Second hash value .

[0114] The first and second preset hash functions are two different cryptographic hash functions, both of which output hash values ​​of fixed length, such as... .

[0115] Encoding submodule 1002 is used to perform unintentional key-value storage encoding based on all first hash values ​​and second hash values ​​to obtain an initial encoding vector.

[0116] The initiator R takes all the calculated first hash values ​​and their corresponding second hash values ​​as input to construct... There are 10 key-value pairs. In each key-value pair, for each element in the first original dataset... Using the first hash value corresponding to the element As the key, use the second hash value corresponding to the element. As values, forming key-value pairs The initiator will... The key-value pair input is inadvertently stored in the encoder, which performs encoding operations to generate an encoding vector, denoted as the initial encoding vector. .

[0117] Unintentional key-value storage encoding is a data structure encoding method that encodes a set of key-value pairs into a fixed-length vector. For each key-value pair input during encoding, the corresponding value can be correctly restored when the encoded vector is decoded using that key. For keys not input during encoding, the decoding result has pseudo-randomness and will not reveal any information about the stored key-value pairs.

[0118] For parameters used in unintentional key-value stores, the definition is completed during the initialization phase. These parameters include expansion factors. , encoding vector length ,bandwidth Where N is the number of elements, and in this embodiment, an expansion factor is selected. .

[0119] The mask processing submodule 1003 is used to perform XOR masking on the initial encoding vector according to the first random vector to obtain the mask vector and send it to the participants.

[0120] Initiator R will use the initial encoding vector With the first random vector Perform an XOR operation, that is, calculate the result of bitwise XORing the initial encoded vector with the first random vector, and use it as the mask vector. The first random vector is obtained in advance by the initiator R and the participant S through negotiation using the VOLE protocol (Vector Unintentional Linear Valuation Protocol).

[0121] The initiator R will calculate the mask vector It is sent to participant S via a communication channel for subsequent generation of a second unintentional pseudo-random function value sequence.

[0122] The first sequence generation submodule 1004 is used to perform unintentional key-value storage decoding on the second random vector using the first hash value corresponding to each element in the first original dataset as the key, to obtain the first unintentional pseudo-random function value sequence; wherein, the initiator and the participants pre-call the vector unintentional linear valuation protocol, so that the initiator obtains the first random vector and the second random vector, and the participants obtain random scalar values ​​and the third random vector.

[0123] The initiator uses each first hash value As the key, with the second random vector (Obtained beforehand via an unintentional linear estimation protocol) as the target decoding vector, an unintentional key-value storage decoding operation is performed on each element in the first original dataset. For each key... The decoder starts from the second random vector. Extract the value at the corresponding position and output a fixed-length decoding result, which is the first unintentional pseudo-random function value corresponding to that element. After performing the above decoding operation on all elements in the first original dataset, the initiator obtains a sequence of first unintentional pseudo-random function values. .

[0124] In the above steps, before generating the OPRF sequence, the initiator R and the participant S invoke the Vector Unintentional Linear Estimation (VOLE) protocol. This protocol is a two-party random correlation generation protocol. After execution, the initiator R obtains the first random vector. Second random vector Participant S receives a random scalar value. and the third random vector Furthermore, the outputs of both parties satisfy a preset linear correlation, such as: .

[0125] As an optional implementation method, Figure 8 This is a schematic diagram of the structure of the second sequence generation module provided in an embodiment of the present invention, as shown below. Figure 8 As shown, the second sequence generation module 200 includes: The second hash calculation submodule 2001 is used to calculate a third hash value for each element in the second original dataset according to a first preset hash function and a fourth hash value according to a second preset hash function.

[0126] In the process of participant S generating the second unintentional pseudo-random function value sequence, the second original dataset is first processed. Each raw data element in The third hash value for each element is calculated using the same first and second preset hash functions as the initiator R. and the fourth hash value Both parties use the same set of hash functions for hash calculation to ensure that the same original data will produce the same hash value in both calculations.

[0127] The second sequence generation submodule 2002 is used to perform unintentional key-value storage decoding operation and XOR operation based on the mask vector, the third random vector, the random scalar value, and the third hash value and the fourth hash value corresponding to each element in the second original dataset, to obtain the second unintentional pseudo-random function value sequence.

[0128] Participant S receives the mask vector sent by initiator R. Then, based on the third hash value and the fourth hash value And a third random vector obtained through the vector inadvertent linear estimation protocol. and random scalar values The second unintentional pseudo-random function value sequence is calculated for the second original dataset in the following manner. : .

[0129] As an optional implementation method, Figure 9 This is a schematic diagram of the structure of the first encoded value derivation module provided in an embodiment of the present invention, as shown below. Figure 9 As shown, the first encoded value derivation module 300 includes: The positioning code value calculation submodule 3001 is used to extract a first preset length as a first truncation value for each unintentional pseudo-random function value in the first unintentional pseudo-random function value sequence, concatenate the first truncation value with the corresponding unintentional pseudo-random function value, perform hash calculation, and use the hash calculation result as the first positioning code value.

[0130] Initiator R obtains the first unintentional pseudo-random function value sequence Then, for each unintentional pseudo-random function value in the sequence Perform the following derivation operations: First, from the value of this unintentional pseudo-random function Starting from the initial position, the first preset length of its front portion is extracted. The bit, as the first truncation value, is denoted as... .

[0131] Then, the first truncation value With complete unintentional pseudo-random function values By concatenating the bits, we obtain a concatenated bit string with a length of [length missing]. ,in The total length of the unintentional pseudo-random function values.

[0132] Finally, perform a hash calculation on the concatenated bit string, and use the hash result as the first location code value corresponding to the unintentional pseudo-random function value. .

[0133] The verification code value calculation submodule 3002 is used to extract a second preset length after a first preset length from the unintentional pseudo-random function value as a second truncated value, and use the second truncated value as the first verification code value; wherein, the sum of the first preset length and the second preset length is equal to the total length of the unintentional pseudo-random function value.

[0134] Initiator R for the first unintentional pseudo-random function value sequence Each unintentional pseudo-random function value in After extracting the first truncation value, subsequent bits are extracted from the same unintentional pseudo-random function value to form the second truncation value. The second truncation value is the first bit extracted from the same unintentional pseudo-random function value. Starting from the first preset length (i.e., immediately following the first preset length), continuously extract the second preset length. Bit, denoted as This involves truncating the last t bits of the unintentional pseudo-random function value. The initiator directly uses this second truncated value as the first verification code value corresponding to the unintentional pseudo-random function value. .

[0135] First preset length With the second preset length The sum equals the total length of the unintentional pseudo-random function values. ,Right now Therefore, for each unintentional pseudo-random function value, after the above two-step truncation operation, all bits of the value are completely divided into two parts: the front part... The first bit is used as the truncated value to generate the location code value, and the subsequent bits... The bit is used as the second truncation value and directly as the verification encoding value.

[0136] Participant S on the second unintentional pseudo-random function value sequence For each value in the sequence, a derivation operation is performed according to the same rules as the initiator R to obtain the corresponding second location code value. Second verification code value Since both parties use the same truncation position and truncation length for the same unintentional pseudo-random function value, the derived positioning code value and verification code value are the same for the same original data element.

[0137] In a preferred embodiment, statistical security parameters are determined by both parties during the initialization phase. The first preset length is 40. Second preset length In this case, since the intersection matching method changes from global matching to one-to-one matching, the amount of communication required to transmit a single piece of data in the intersection stage increases from the length of the original complete unintentional pseudo-random function value. Reduce to .

[0138] Based on the original dataset size of both parties In this case, the total communication volume of the original protocol is In this embodiment, OKVS encoding is performed based on a 64-bit verification code value, reducing the communication volume to... bit, here .

[0139] As an optional implementation, the intersection module 600 includes: The decoding return value calculation submodule 6001 is used to perform unintentional key-value storage decoding on the lookup encoding vector with each first positioning encoding value as the key to obtain the corresponding decoding return value; wherein, the lookup encoding vector is obtained by the participants performing unintentional key-value storage encoding based on the second positioning encoding value and the second verification encoding value.

[0140] Initiator R receives the lookup encoding vector sent by participant S Then, for each first location code value derived from itself , with that As a query key, for the lookup encoding vector Perform an unintentional key-value storage decoding operation to obtain the corresponding decoding return value.

[0141] The lookup encoding vector is pre-constructed by participant S. Participant S uses all its derived second-position encoding values. For each key, a second verification code value corresponding to that key is used. The lookup encoding vector is generated after performing unintentional key-value storage encoding on the value. .

[0142] The intersection calculation submodule 6002 is used to determine whether the decoding return value is equal to the corresponding first verification code value. If they are equal, the element in the first original dataset corresponding to the decoding return value is added to the intersection. If they are not equal, the element corresponding to the decoding return value is determined not to belong to the intersection.

[0143] The decoded return value obtained from decoding is compared with the first verification code value corresponding to that element. Perform a consistency comparison to determine if the two are equal. If the decoded value matches the first verification code value... If they match, the element is determined to be an intersection element and is added to the intersection result set; otherwise, the element is determined not to belong to the intersection and is not added to the intersection result set.

[0144] The final intersection .

[0145] The above technical solution has the following beneficial effects: by deriving location code values ​​and verification code values ​​from the unintentional pseudo-random function values, both parties can synchronously obtain the same location code value locally for the same element, and only the same location code value can be correctly decoded; based on this, the initiator does not need to perform a global comparison between each unintentional pseudo-random function value locally and all unintentional pseudo-random function values ​​sent by the sender, but instead uses the local location code value as the key to look up the code vector and decode it, and then performs a one-to-one comparison on the decoding result.

[0146] Because the matching method has changed from global comparison to one-to-one comparison, each initiating element triggers only one comparison, and the total number of comparisons in the protocol has increased. The second reduction to Therefore, under the same statistical safety parameters Under these conditions, the length of the verification code value only needs to meet the following requirements. Compared to a complete, unintentional pseudo-random function, the length of the function value is reduced. This reduces the communication overhead of the intersection phase while maintaining the same level of security.

[0147] The above-described specific embodiments of the invention further illustrate the purpose, technical solution, and beneficial effects of the invention. It should be understood that the above content is only for specific embodiments of the invention and is not intended to limit the scope of protection of the invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the invention should be included within the scope of protection of the invention.

Claims

1. A privacy set intersection method based on optimizing an unintentional pseudo-random function, characterized in that, include: The initiator generates a first unintentional pseudo-random function value sequence based on the first original dataset it holds; The participants generate a second unintentional pseudo-random function value sequence based on the second original dataset they possess; The initiator derives a first location code value and a first verification code value for each value in the first unintentional pseudo-random function value sequence according to a preset rule. The participants derive their respective second positioning code value and second verification code value for each value in the second unintentional pseudo-random function value sequence according to the preset rules. The participating parties construct a lookup encoding vector based on all the second location encoding values ​​and the second verification encoding values ​​and send it to the initiator; The initiator decodes the lookup encoding vector according to each of the first positioning encoding values ​​to obtain a decoding return value, and performs a consistency determination between the decoding return value and the corresponding first verification encoding value, adding the elements in the first original dataset corresponding to the decoding return value that is determined to be consistent to the intersection.

2. The privacy set intersection method based on optimized unintentional pseudo-random functions according to claim 1, characterized in that, The initiator generates a first unintentional pseudo-random function value sequence based on the first original dataset it holds, including: For each element in the first original dataset, a first hash value is calculated according to a first preset hash function, and a second hash value is calculated according to a second preset hash function. Based on all the first hash value and the second hash value, perform unintentional key-value storage encoding to obtain the initial encoding vector; The initial encoded vector is XORed and masked according to the first random vector to obtain a mask vector, which is then sent to the participants. Using the first hash value corresponding to each element in the first original dataset as the key, perform unintentional key-value storage decoding on the second random vector to obtain the first unintentional pseudo-random function value sequence; wherein, the initiator and the participants pre-call the vector unintentional linear valuation protocol, so that the initiator obtains the first random vector and the second random vector, and the participants obtain random scalar values ​​and the third random vector.

3. The privacy set intersection method based on optimized unintentional pseudo-random functions according to claim 2, characterized in that, The participating parties generate a second unintentional pseudo-random function value sequence based on the second original dataset they possess, including: For each element in the second original dataset, calculate the third hash value according to the first preset hash function and the fourth hash value according to the second preset hash function; Based on the mask vector, the third random vector, the random scalar value, and the third and fourth hash values ​​corresponding to each element in the second original dataset, perform unintentional key-value storage decoding and XOR operations to obtain the second unintentional pseudo-random function value sequence.

4. The privacy set intersection method based on optimized unintentional pseudo-random functions according to claim 1, characterized in that, The initiator derives a corresponding first positioning code value and a first verification code value for each value in the first unintentional pseudo-random function value sequence according to a preset rule, including: For each unintentional pseudo-random function value in the first unintentional pseudo-random function value sequence, the first preset length of its beginning is taken as the first truncation value. The first truncation value is concatenated with the corresponding unintentional pseudo-random function value and then a hash calculation is performed. The hash calculation result is taken as the first positioning code value. From the unintentional pseudo-random function value, a second preset length is extracted after the first preset length as a second truncation value, and the second truncation value is used as a first verification code value; wherein, the sum of the first preset length and the second preset length is equal to the total length of the unintentional pseudo-random function value.

5. The privacy set intersection method based on optimized unintentional pseudo-random functions according to claim 1, characterized in that, The initiator decodes the lookup encoding vector according to each of the first location encoding values ​​to obtain a decoding return value, and performs a consistency determination between the decoding return value and the corresponding first verification encoding value. Elements in the first original dataset corresponding to the decoded return values ​​that are determined to be consistent are added to the intersection, including: Using each of the first location encoding values ​​as keys, perform unintentional key-value storage decoding on the lookup encoding vector to obtain the corresponding decoding return value; wherein, the lookup encoding vector is obtained by the participants performing unintentional key-value storage encoding based on the second location encoding value and the second verification encoding value; Determine whether the decoded return value is equal to the corresponding first verification code value. If they are equal, add the element in the first original dataset corresponding to the decoded return value to the intersection. If they are not equal, determine that the element corresponding to the decoded return value does not belong to the intersection.

6. A privacy set intersection system based on optimizing an unintentional pseudo-random function, characterized in that, include: The first sequence generation module is used by the initiator to generate a first unintentional pseudo-random function value sequence based on the first original dataset it holds. The second sequence generation module is used by the participants to generate a second unintentional pseudo-random function value sequence based on the second original dataset they hold. The first encoding value derivation module is used by the initiator to derive the corresponding first positioning encoding value and first verification encoding value for each value in the first unintentional pseudo-random function value sequence according to preset rules. The second encoding value derivation module is used by the participants to derive their respective second positioning encoding value and second verification encoding value for each value in the second unintentional pseudo-random function value sequence according to the preset rules. The vector encoding module is used by the participants to construct a lookup encoding vector based on all the second positioning encoding values ​​and the second verification encoding values, and then send it to the initiator. The intersection module is used by the initiator to decode the search encoding vector according to each of the first positioning encoding values ​​to obtain the decoding return value, and to determine the consistency between the decoding return value and the corresponding first verification encoding value, and to add the elements in the first original dataset corresponding to the decoding return value that is determined to be consistent to the intersection.

7. The privacy set intersection system based on optimized unintentional pseudo-random functions according to claim 6, characterized in that, The first sequence generation module includes: The first hash submodule is used to calculate a first hash value for each element in the first original dataset according to a first preset hash function and a second hash value according to a second preset hash function. The encoding submodule is used to perform unintentional key-value storage encoding based on all the first hash values ​​and the second hash values ​​to obtain an initial encoding vector; The masking submodule is used to perform XOR masking on the initial encoding vector according to the first random vector to obtain a mask vector and send it to the participants. The first sequence generation submodule is used to perform unintentional key-value storage decoding on the second random vector using the first hash value corresponding to each element in the first original dataset as the key, to obtain the first unintentional pseudo-random function value sequence; wherein, the initiator and the participants pre-call the vector unintentional linear valuation protocol, so that the initiator obtains the first random vector and the second random vector, and the participants obtain random scalar values ​​and the third random vector.

8. The privacy set intersection system based on optimized unintentional pseudo-random functions according to claim 7, characterized in that, The second sequence generation module includes: The second hash calculation submodule is used to calculate a third hash value for each element in the second original dataset according to the first preset hash function and a fourth hash value according to the second preset hash function. The second sequence generation submodule is used to perform unintentional key-value storage decoding operation and XOR operation based on the mask vector, the third random vector, the random scalar value, and the third hash value and the fourth hash value corresponding to each element in the second original dataset, to obtain the second unintentional pseudo-random function value sequence.

9. The privacy set intersection system based on optimized unintentional pseudo-random functions according to claim 6, characterized in that, The first encoded value derivation module includes: The positioning code value calculation submodule is used to extract a first preset length as a first truncation value for each unintentional pseudo-random function value in the first unintentional pseudo-random function value sequence, concatenate the first truncation value with the corresponding unintentional pseudo-random function value, perform hash calculation, and use the hash calculation result as the first positioning code value. The verification code value calculation submodule is used to extract a second preset length after the first preset length from the unintentional pseudo-random function value as a second truncated value, and use the second truncated value as a first verification code value; wherein, the sum of the first preset length and the second preset length is equal to the total length of the unintentional pseudo-random function value.

10. The privacy set intersection system based on optimized unintentional pseudo-random functions according to claim 6, characterized in that, The intersection module includes: The decoding return value calculation submodule is used to perform unintentional key-value storage decoding on the lookup encoding vector with each of the first positioning encoding values ​​as keys to obtain the corresponding decoding return value; wherein, the lookup encoding vector is obtained by the participants performing unintentional key-value storage encoding based on the second positioning encoding value and the second verification encoding value; The intersection calculation submodule is used to determine whether the decoded return value is equal to the corresponding first verification code value. If they are equal, the element in the first original dataset corresponding to the decoded return value is added to the intersection. If they are not equal, the element corresponding to the decoded return value is determined not to belong to the intersection.