A multi-party secure subgraph matching method based on differential privacy acceleration
Patent Information
- Application Number
- CN202610873542.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2046-06-17
AI Technical Summary
[0008]为应对现有多方安全计算子图匹配中因采用理论最坏情况界限进行填充而导致的中间结果爆炸与计算效率低下问题,本发明提供了一种基于差分隐私加速的多方安全子图匹配方法,通过构建无损差分隐私索引与差分隐私最大频率统计量,结合混合界限估算机制,实现在密文环境下对中间结果规模的紧致估算与精确截断,从而显著削减无效计算开销,大幅提升多方安全子图匹配的执行效率
1.本发明通过构建包含双向边信息的分布式边匹配表,利用秘密共享协议确保了原始数据在多方计算环境下的严格隐私性。
Smart Images

Figure CN122413480B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of information security and privacy computing technology, specifically relating to a multi-party secure subgraph matching method based on differential privacy acceleration. Background Technology
[0002] With the development of big data technology, graph data has become an important carrier for expressing complex relational structures and is widely used in key areas such as social network analysis and financial fraud prevention. Subgraph matching, as an important query in graph analysis, aims to find all substructures in large-scale graph data that are isomorphic to a given query graph structure; for example, in financial regulation, regulatory agencies may need to query specific fund flow patterns across transaction data from multiple banks.
[0003] However, in practical applications, graph data often contains sensitive personal privacy or trade secrets. Due to legal and regulatory considerations and commercial interests, data owners cannot directly share plaintext data. Therefore, how to achieve secure subgraph matching without disclosing the privacy of the original data has become a research hotspot.
[0004] Secure Multi-Party Computation (MPC), especially secret-sharing-based MPC protocols, is considered an ideal solution to the above problems. It allows participants to collaboratively complete computations while keeping data encrypted (or fragmented), ensuring that no input information is leaked except for the final result.
[0005] While MPC offers robust security guarantees, its computational efficiency faces significant challenges when handling complex subgraph matching tasks. This is primarily because the MPC protocol must meet the inadvertence requirement, meaning the program's execution path and memory access patterns cannot depend on the input data to prevent side-channel attacks. In multi-table join operations involved in subgraph matching, this means the size of intermediate results cannot be exposed. To hide the true number of intermediate results, existing MPC schemes typically employ a worst-case padding strategy, which fills all intermediate results to the theoretically maximum possible size based on the size of the input table and the theoretical upper bound of the graph topology.
[0006] However, in graph data, the theoretical worst-case scenario is often much larger than the actual number of intermediate results. This results in a large number of dummy cells marked as invalid during the computation process. The system must process, transmit, and store this massive amount of invalid data in encrypted form. This redundant computation leads to huge communication overhead and computational latency, making MPC-based subgraph matching extremely impractical when processing large-scale graph data, and even unable to complete it within an acceptable time.
[0007] In summary, existing technologies lack solutions that can both strictly guarantee data privacy and effectively avoid the significant performance overhead caused by worst-case padding. How to design a solution that can accurately and securely estimate the size of intermediate results in a encrypted environment, thereby significantly reducing redundant computation, is a key technical problem that urgently needs to be solved to achieve efficient multi-party secure subgraph matching. Summary of the Invention
[0008] To address the problems of intermediate result explosion and low computational efficiency caused by using theoretical worst-case limits for filling in existing multi-party secure computational subgraph matching, this invention provides a multi-party secure subgraph matching method accelerated by differential privacy. By constructing a lossless differential privacy index and a differential privacy maximum frequency statistic, combined with a hybrid limit estimation mechanism, it achieves compact estimation and accurate truncation of intermediate result size in a ciphertext environment, thereby significantly reducing invalid computational overhead and greatly improving the execution efficiency of multi-party secure subgraph matching.
[0009] A multi-party secure subgraph matching method based on differential privacy acceleration includes the following steps: (1) Obtain the original graph data to construct an edge matching table containing bidirectional edge information, and distribute the edge matching table to a multi-party computing server using a secret sharing protocol; (2) Perturb the hierarchical histogram based on the attribute value range using the one-sided Laplace mechanism to generate a lossless differential privacy index; (3) Calculate the true maximum frequency of the connection key attribute and inject differential privacy noise to generate differential privacy maximum frequency statistics; (4) The lossless differential privacy index and the differential privacy maximum frequency statistic are used to perform the connection operation in parallel, compress the connection result, and accelerate the execution of multi-round secure connections.
[0010] Furthermore, the specific implementation of step (1) is as follows: S11: Obtain the original graph data ,in V For a set of nodes, E Let it be the set of edges; S12: For the edge set E any undirected edge in Construct the corresponding reverse join tuple , u and v These are two nodes connected by an undirected edge; S13: Concatenate and merge the original list of undirected edges with the constructed list of reverse-connected tuples to build an edge matching table containing complete bidirectional connections. S14: Based on the replication secret sharing protocol, split any attribute value in the edge matching table into three random pieces that satisfy the XOR reconstruction relationship; S15: Distribute the random shards to three computing servers according to the ring replication strategy, ensuring that each computing server holds two copies of the shards, thereby constructing a distributed edge matching table of ciphertext state in a multi-party computing environment.
[0011] Furthermore, the specific implementation of step (2) is as follows: A1. Construct a noise-free, true hierarchical histogram tree based on the edge matching table of the original graph data to provide an accurate baseline structure for subsequent perturbations; A2. Using differential privacy budgeting, a sign-constant one-sided Laplace noise is generated and injected into the real hierarchical histogram tree to obtain upper and lower bound histograms for estimating cumulative frequency, laying a conservative estimation foundation for lossless indexing. A3. Utilizing the conservative offset characteristics of upper and lower bound histograms, through interval partitioning, range querying, and truncation correction, an index range is generated for each logical interval to ensure no record is missed, and finally, a lossless index set with differential privacy protection is output.
[0012] Furthermore, the specific implementation process of step A1 is as follows: S21: Determine the target attribute used to construct the index in the edge matching table, and then determine the discrete value range space of the attribute; S22: Check the size of the discrete value range space. If the size is not a power of 2, then fill and expand the value range by adding dummy buckets with a count of zero until the total length meets the construction requirements of a complete binary tree. S23: Statistically analyze the true frequency distribution of the target attribute in the edge matching table, and construct a true hierarchical histogram tree as the baseline, where each leaf node stores the true count of a single attribute value, and each internal node stores the sum of the counts of its child nodes.
[0013] Furthermore, the specific implementation process of step A2 is as follows: S24: Obtain the privacy budget for index building, distribute the privacy budget evenly to each level according to the height of the real hierarchical histogram tree, and calculate the privacy budget for a single level. S25: Based on a single-layer privacy budget, a probability distribution model is constructed using a one-sided Laplace mechanism. Two independent noise vectors are generated by sampling, where all components of the first noise vector are constrained to be non-negative and all components of the second noise vector are constrained to be non-positive. S26: Inject the second set of noise vectors into the node counts of the true hierarchical histogram tree to construct a lower bound histogram specifically for estimating the lower bound of the cumulative frequency. S27: Inject the first set of noise vectors into the node counts of the true hierarchical histogram tree to construct an upper bound histogram specifically for estimating the upper bound of the cumulative frequency.
[0014] Furthermore, the specific implementation process of step A3 is as follows: S28: Divide the value space of the target attribute into multiple non-overlapping continuous logical intervals, which serve as the basic scheduling unit for subsequent parallel connection operations; S29: For any logical interval, perform a range query on the lower bound histogram, accumulate the noise count of all buckets from the start of the value range to the start point of the interval, and obtain the estimated starting offset of the interval. S210: Perform a non-negative truncation operation on the estimated starting offset, correcting values less than zero to zero, and obtain the effective index starting position of the interval in the ciphertext table; S211: For any logical interval, perform a range query on the upper bound histogram, accumulate the noise count of all buckets from the start of the value range to the end of the interval, and obtain the estimated end offset of the interval. S212: Perform boundary truncation operation on the estimated end offset, correct the value exceeding the total number of rows in the edge matching table to the maximum number of rows in the table, and obtain the effective index end position of the interval in the ciphertext table; based on the effective index start position and effective index end position of each logical interval, output the final lossless index set.
[0015] Furthermore, the specific implementation of step (3) is as follows: S31: Obtain the edge matching statement before secret sharing, and determine the connection key attribute column that needs to be statistically analyzed; S32: Utilizing the structural characteristics of the edge matching table, which is composed of the original edge and its transpose edge, we can confirm that the source node column and the target node column in the table have symmetrical consistency in statistical distribution, and thus select only one of the columns as the statistical object. S33: Perform aggregation statistics on the node identifiers in the selected column to calculate the frequency of each node in the edge matching table in the original graph. This frequency is semantically equivalent to the degree of the corresponding node. S34: Compare the frequency values of all nodes, filter out the maximum value, and define it as the true maximum frequency of the connection key attribute; S35: Obtain a privacy budget dedicated to the maximum frequency statistics, and generate non-negative Laplace noise based on the privacy budget; S36: Add Laplace noise to the true maximum frequency to generate the maximum frequency after differential privacy perturbation; S37: Encapsulate the maximum frequency after differential privacy perturbation into a statistical digest, and distribute the digest as a public parameter to the multi-party computation server.
[0016] Furthermore, the specific implementation of step (4) is as follows: B1. Logical bucketing and parallel join of data are achieved through differential privacy indexing, and the amount of result filling is precisely controlled by hybrid upper bound, which accelerates the join operation while ensuring security. B2. Based on the results of this round of connections, dynamically adjust the differential privacy index and the maximum frequency statistic to provide an accurate and tight statistical upper bound for the next round of secure connections.
[0017] Furthermore, the specific implementation process of step B1 is as follows: S41: In the secure connection input table and At that time, lossless differential privacy index is used to divide the edge matching table ciphertext logic into multiple independent logical buckets according to the index range. The sparsity of the index is used to identify and skip empty buckets, thereby accelerating I / O at the data access level. S42: For any non-empty logical bucket, execute the bucket input table in parallel across multiple computing servers. and The connection operations are optimized to reduce the computational load at a single point. S43: Calculate the upper bound of the statistical estimate based on the actual data distribution using the differential privacy maximum frequency statistic; S44: Analyze the current graph query topology, using the AGM theoretical limit formula, based solely on the input table. and The size and connection mode are used to calculate the theoretical worst-case upper bound; S45: Implement the hybrid bounds optimization strategy, compare the statistically estimated upper bound with the theoretical worst-case upper bound, and select the minimum of the two as the safe hybrid upper bound for this connection operation. ; S46: Perform controlled invalid tuple filling based on a safe hybrid upper bound, if the actual number of join results... Less than Then fill A dummy tuple marked as invalid is forcibly removed, and redundant padding exceeding the limit is eliminated, thus accelerating computation.
[0018] Furthermore, the specific implementation process of step B2 is as follows: S47: After generating the join result table, update the lossless differential privacy index and the differential privacy maximum frequency. First, for the join key attribute reused in the join result, count the total number of tuples actually generated in each bucket, including dummy elements. S48: Dynamically recalculate the differential privacy index range of the join key attribute in the join result table based on the total number of tuples, and update the lossless differential privacy index structure to adapt to the changes in the result table; S49: Calculate the join key attributes for the input table respectively. and The maximum frequency of differential privacy is obtained by multiplying the two sets of frequencies to get a new maximum frequency of differential privacy. S410: The updated lossless differential privacy index and differential privacy maximum frequency are encapsulated as input parameters for the next round of connection operations, ensuring that efficient mixing limit estimation can always be maintained by utilizing effective statistical information during multi-round continuous subgraph matching.
[0019] A computer device includes a memory and a processor, wherein the memory stores a computer program and the processor executes the computer program to implement the above-described multi-party secure subgraph matching method based on differential privacy acceleration.
[0020] A computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described multi-party secure subgraph matching method based on differential privacy acceleration.
[0021] Based on the above technical solution, the present invention has the following beneficial technical effects: 1. This invention constructs a distributed edge matching table containing bidirectional edge information and utilizes a secret sharing protocol to ensure the strict privacy of the original data in a multi-party computing environment.
[0022] 2. This invention utilizes a lossless differential privacy index generated by a one-sided Laplace mechanism. While ensuring that the index range completely covers the real data, it achieves bucket-based parallel data access and skipping of invalid I / O, significantly improving efficiency.
[0023] 3. This invention provides a tight statistical estimate of the size of intermediate results in a ciphertext environment by injecting differential privacy noise to generate differential privacy maximum frequency statistics, breaking the limitation of traditional methods that must rely on theoretical worst-case blind filling.
[0024] 4. This invention comprehensively utilizes the index and statistics to perform secure connection operations in parallel. By compressing the connection results, it significantly reduces redundant dummy cells and corresponding invalid computation and communication overhead, thus significantly accelerating the execution process of multi-round secure connections and providing an efficient solution for large-scale privacy-preserving graph computation. Attached Figure Description
[0025] Figure 1 This is a schematic diagram of the multi-party secure subgraph matching method based on differential privacy acceleration according to the present invention.
[0026] Figure 2 This is a schematic diagram of the lossless differential privacy index constructed in this invention.
[0027] Figure 3 This is a schematic diagram illustrating the generation of the maximum frequency statistic for differential privacy in this invention. Detailed Implementation
[0028] To describe the present invention in more detail, the technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0029] This embodiment specifically addresses a social network analysis application scenario to achieve sensitive relationship subgraph matching. For example, a company possesses a complete social network graph, where each user corresponds to a node in the graph, and the friend relationships between users correspond to edges. The company wants to identify user groups that conform to specific structural patterns, such as a triangular subgraph formed by three interconnected friends, while ensuring user privacy. The original social network graph can be represented as a graph... , Represents a set of users. This represents a set of friend relationships. To achieve efficient and secure subgraph matching, the system will... Convert to an edge matching table, where each row records the two user identifiers and their attribute information for each edge; query the graph. This represents a target pattern subgraph, where nodes and edges correspond to users and relationships to be matched. The system's goal is to identify all matching relationships from the edge matching table. The matching subgraph of structural constraints, without revealing any user's sensitive information. For example... Figure 1 As shown, the specific implementation process of this embodiment is as follows: (1) Obtain the original graph data to construct an edge matching table containing bidirectional edge information, and distribute the edge matching table to multiple servers using a secret sharing protocol.
[0030] This step establishes the data foundation for secure multi-party computation. Specifically, to perform subgraph matching using relational algebra join operators in an encrypted environment, unstructured graph data needs to be transformed into structured relational tables. A unified view containing bidirectional join information is constructed through preprocessing, transforming complex graph traversal into standardized equi-join operations. Finally, a replication-based secret sharing protocol is used to securely outsource the computational task to a multi-party server cluster, ensuring that data holders do not disclose the original data. The specific implementation of this step may include the following sub-steps: S11: Obtain the original graph data held by the data owner. .
[0031] First, obtain the original map data from the data owner's local machine. ,in For the set of vertices, For the edge set, the data is still in plaintext at this stage and is only parsed within the controlled domain of the data owner.
[0032] S12: For sets Every undirected edge Construct the corresponding reverse join tuple This is to capture reverse topology information in the graph structure.
[0033] For each edge in the original edge set Explicitly generate its symmetric tuple Logically, undirected edges are decoupled into bidirectional directed edges, eliminating the dependence on the source or target node position during ciphertext queries. This eliminates the need for subsequent algorithms to design complex conditional branches for the query direction, thereby reducing the design complexity of the unintentional algorithm.
[0034] S13: Concatenate and merge the original list of undirected edges with the constructed list of reverse-connected tuples to build an edge matching table containing complete bidirectional connections. .
[0035] Edge matching table It contains two columns: source node and target node, covering all possible connection paths in the graph. By constructing... The graph's topology is mapped to a single relational table, which allows complex subgraph isomorphism problems to be transformed into standardized SQL (Structured Query Language) join sequences for that table, providing a unified data interface for subsequent processing of graph data using general MPC primitives.
[0036] S14: Based on the replication secret sharing protocol, the edge matching table any attribute value in Split into XOR reconstruction relationships Three random partitions.
[0037] This implementation uses a (2, 3)-replication secret sharing scheme. Encryption is performed for any attribute value in the table. Generate three random partitions , , ,satisfy XOR-based secret sharing not only possesses information-theoretic security, but also involves only simple bit operations, reducing the computational overhead of ciphertext computation on massive graph data.
[0038] S15: Distribute the random shards to three computing servers according to a ring replication strategy, ensuring that each server holds two copies of the shards, thereby constructing a distributed edge matching table of ciphertext state in a multi-party computing environment. .
[0039] Establish a distributed encrypted environment and combine the fragments according to a ring redundancy strategy. , and The data is sent to three computing servers. This allocation method ensures that each server holds only a partial view of the data and cannot independently reconstruct the original graph structure, while guaranteeing that any two servers can collude to recover the data. At this point, data ownership and computational rights are separated, forming a ciphertext edge matching table. .
[0040] (2) The hierarchical histogram constructed based on the attribute value range is perturbed by the one-sided Laplace mechanism to generate a lossless differential privacy index.
[0041] To achieve fast location functionality similar to plaintext database indexes, an index can be built to improve efficiency. Building a differential privacy index involves range queries on a differentially private hierarchical histogram. The start position of each interval is the cumulative count of the previous bucket, and the end position is the cumulative count of the last bucket in that interval. Due to the post-processing characteristics of differential privacy, these calculations do not incur additional privacy costs. However, directly adding Laplace noise to the histogram count may cause noise interval boundaries to exclude actual data points; the noise start position may exceed the true start position, or the noise end position may be lower than the true end position. This is unacceptable for subgraph matching that requires precise results. Therefore, this step builds an index as follows: Figure 2 The lossless differential privacy index shown utilizes hierarchical histograms to support efficient range queries and employs a one-sided Laplace mechanism to construct a lower bound tree (limiting noise to strictly non-positive values) and an upper bound tree (limiting noise to strictly non-negative values). This asymmetric noise injection ensures that the generated index range never excludes real data points, thus protecting privacy while guaranteeing the integrity of data access. The specific implementation of this step may include the following sub-steps: S21: Determine the target attribute in the edge matching table used to build the index. And determine the discrete range space of the attribute. .
[0042] In this implementation, the index object is clearly defined, and the system selects the connection attribute that needs to be accelerated as the target attribute based on the query. At the same time, determine all possible values of the attribute, i.e., the value range. In graph data scenarios, target attributes This can be a node identifier attribute in the edge matching table, and its value can be represented as... , , , Node identifiers, each node identifier The number of times it appears in the edge matching table is denoted as . This number corresponds to a node in the graph structure semantics. The degree. Figure 2 middle , , , Representing nodes respectively , , , In target attributes The actual frequency on.
[0043] S22: Check the size of the value range space. If the size is not a power of 2, then expand the value range by adding dummy buckets with a count of zero until the total length meets the requirements for constructing a complete binary tree.
[0044] A hierarchical histogram is essentially a complete binary tree structure designed to support efficient range summation queries. To ensure the balance of the tree structure and the recursive applicability of the index building algorithm, the number of leaf nodes at the bottom level must be equal to the number of nodes in the tree. .
[0045] S23: Statistically analyze the true frequency distribution of the target attribute in the edge matching table, and construct a true hierarchical histogram tree as the baseline. Each leaf node stores the actual count of a single attribute value, and each internal node stores the sum of the counts of its child nodes.
[0046] This implementation method performs a full table scan to count the actual number of times each attribute value appears in the table, and fills these counts into the leaf nodes of the tree. Then, a bottom-up aggregation method is used to calculate the value of each parent node as the sum of the values of its left and right child nodes. In this structure, each node in the tree represents the total number of tuples within a specific binary interval, and the root node stores the total number of rows in the entire table. This un-noiseed baseline tree reflects the true distribution of the data. Figure 2 As shown, the numbers 2, 3, 3, and 2 in the leaf nodes represent... , , , The actual count of nodes, and the count of internal nodes, represents the cumulative count of the set of nodes they cover.
[0047] S24: Obtain the privacy budget for index building And based on the tree height of the hierarchical histogram The privacy budget is evenly distributed across each layer, and the privacy budget for a single layer is calculated. .
[0048] This implementation method, based on the sequence combination property of differential privacy, allocates a privacy budget for index construction. Distributed equally to each level of the tree This allocation strategy ensures that any query path from root to leaf satisfies differential privacy constraints, balancing the noise impact at each level while controlling overall privacy risks.
[0049] S25: Based on a single-layer privacy budget, a probability distribution model is constructed using a one-sided Laplace mechanism. Two independent noise vectors are generated by sampling. All components of the first noise vector are constrained to be non-negative (i.e., positive one-sided noise), and all components of the second noise vector are constrained to be non-positive (i.e., negative one-sided noise).
[0050] Achieving "lossless" performance requires breaking the symmetry of the standard Laplace noise, which can be either positive or negative, potentially causing the estimated interval to be smaller than the true interval. To address this, we employ a one-sided Laplace mechanism: starting from the distribution... Mid-sampling yields non-negative noise used for overestimation. and the non-positive noise used for underestimation These two sets of uncorrelated noise vectors will be used to construct two independent views of the "starting point" and "ending point" of the control interval, respectively. Figure 2 The Laplace noise in the figure represents random perturbations injected into the node count of the hierarchical histogram, where Used to generate an upper bound histogram , Used to generate a lower bound histogram .
[0051] S26: Inject the second set of noise vectors into the true hierarchical histogram tree. In the node counts, a lower bound histogram is constructed specifically for estimating the lower bound of the cumulative frequency. .
[0052] In this embodiment, negative noise is... Construct a lower bound histogram by superimposing the actual counts. , Figure 2 The Middle Layer The noise count of each node is denoted as , This indicates that the node is located at the bottom-up position. layer, This indicates the node's index within that layer. For example... , , , They represent Corresponding in the leaf node layer , , , Counting after perturbation, express Medium coverage The cumulative count after the interval disturbance, express Medium coverage The cumulative count after the interval disturbance, express The count is performed after the root node is disturbed. Because... Constructed from non-positive noise, the cumulative prefix sum query result performed on this structure will always be less than or equal to the actual value, thereby shifting the estimated index starting position toward the table head and eliminating the risk of missing data due to an excessively late starting point.
[0053] S27: Inject the first set of noise vectors into the true hierarchical histogram tree. In the node counts, an upper bound histogram is constructed specifically for estimating the upper bound of the cumulative frequency. .
[0054] In this embodiment, positive noise is... Superimposed on the true count to construct an upper bound histogram , Figure 2 The Middle Layer The noise count of each node is denoted as ,For example , , , They represent Corresponding in the leaf node layer , , , Counting after perturbation, express Medium coverage The cumulative count after the interval disturbance, express Medium coverage The cumulative count after the interval disturbance, express The count is performed after the root node is disturbed. Because... Constructed from non-negative noise, the cumulative prefix sum query result performed on this structure will always be greater than or equal to the actual value, thereby shifting the estimated end position of the index towards the end of the table and eliminating the risk of missing data due to the end point being too early.
[0055] S28: Divide the value space of the target attribute into... Non-overlapping continuous logical intervals It serves as the basic scheduling unit for subsequent parallel connection operations.
[0056] This implementation divides the continuous attribute value range into multiple discrete logical segments. These logical segments are task scheduling units for multi-party computation, allowing subsequent connection operations to be distributed to different servers or threads for parallel execution, thereby improving throughput. Figure 2 middle Indicates the first logical interval coverage node and , Indicates the second logical interval coverage node and The "ID" column identifies the number of each logical interval, and the "Range Query" column indicates the range to be calculated when determining the index boundaries for that interval. and For the prefix range query executed above, the "DP" column is used to represent the index range obtained from the differential privacy histogram, and the "actual" column is used to represent the actual range of the logical interval in the real sort table without noise.
[0057] S29: For each logical interval In the lower bound histogram Perform a range query, accumulate the noise counts of all buckets from the start of the value range to the start point of the interval, and obtain the estimated starting offset of the interval.
[0058] S210: Perform a non-negative truncation operation on the estimated starting offset, correcting values less than zero to zero, to obtain the effective starting position of the index for this interval in the ciphertext table. .
[0059] Negative noise may cause the calculated physical offset to be negative, requiring execution of... The truncation operation corrects invalid negative values to 0 in the table header position, ensuring that the generated starting index is valid and usable in physical storage.
[0060] S211: For each logical interval In the upper bound histogram Perform a range query, accumulate the noise counts of all buckets from the start of the value range to the end of the interval, and obtain the estimated end offset of the interval.
[0061] S212: Perform a boundary truncation operation on the estimated end offset, correcting the value exceeding the total number of rows in the edge matching table to the maximum number of rows in the table, and obtain the valid index end position of the interval in the ciphertext table. And output the final lossless index set {( ,[ ))}.
[0062] Positive noise may cause the calculated offset to exceed the total number of rows in the table, requiring execution. The truncation operation results in the final output index range [ It physically forms a security envelope slightly larger than the real area, ensuring lossless data retrieval under any privacy budget.
[0063] (3) Calculate the true maximum frequency of the connection key attribute and inject differential privacy noise to generate differential privacy maximum frequency statistics.
[0064] This step aims to provide a secure statistical basis for estimating intermediate results in subsequent multi-party secure computation. When performing a join operation, the magnitude of the intermediate result directly depends on the frequency distribution of the join key. If the maximum frequency of the join key (i.e., the degree of the node with the highest degree in the graph) can be predicted, then... A tighter upper bound than the theoretical worst-case scenario is estimated. However, the true maximum frequency is sensitive information. By calculating the true maximum node degree and injecting positive Laplace noise, a statistical summary is generated that satisfies differential privacy constraints while probabilistically covering the true value. This summary will be published as a public parameter, enabling multiple servers to collaboratively calculate the connection truncation threshold without prying into the specific data distribution. The specific implementation of this step may include the following sub-steps: S31: Obtain the plaintext edge matching table before secret sharing Determine the join key attribute column that needs to be statistically analyzed.
[0065] Statistics are typically calculated before the data owner distributes the data in encrypted form. Based on subsequent subgraph matching queries, the key attribute columns that will serve as join conditions are identified. For example... Figure 3 As shown, This indicates the edges in the original graph. The edge matching table is constructed, where each row represents the node pair corresponding to an edge. , , , Represents the node identifiers in the original graph, and the edge matching table. The node identifier column can be used as the connection key attribute column for subsequent secure connection operations, and is used to count the frequency of each node appearing in this attribute column.
[0066] S32: Utilizing the structural characteristic that the edge matching table is composed of the original edge and its transpose edge, we can confirm that the source node column and the target node column in the table have symmetrical consistency in statistical distribution, thereby selecting only one column as the statistical object and reducing redundant calculations.
[0067] Figure 3The first table on the left represents the original edge table, where each row records node pairs according to the original direction; the second table on the left represents the transposed edge table, where each row swaps the positions of the two nodes corresponding to the original edge to record the reverse node pairs; the symbol ∪ indicates that the original edge table and the transposed edge table are concatenated and merged to form an edge matching table containing bidirectional edge information. .
[0068] S33: Perform aggregation statistics on the node identifiers in the selected column to calculate the frequency of each node in the edge matching table, which is semantically equivalent to the degree of the corresponding node.
[0069] This implementation performs a grouping and counting operation, scanning the selected attribute columns and counting the occurrences of each unique node identifier. In the context of the edge matching table, the number of rows in which a node ID appears in the table strictly corresponds to the degree of that node in the graph topology. This step compresses the massive edge table data into a frequency mapping table of "node-degree", extracting the macroscopic density characteristics of the graph data.
[0070] S34: Traverse the frequency set, compare all frequency values, filter out the maximum value, and define it as the true maximum frequency of the connection key attribute. .
[0071] like Figure 3 As shown, this embodiment sorts the aggregated degree set and identifies the maximum degree value in the graph. This value represents the upper limit of the influence of the most densely connected hub node in the graph, and is a key factor in determining the number of tuples generated by a single connection operation in the worst case.
[0072] S35: Obtain a privacy budget dedicated to maximum frequency statistics Its budget is independent of the index building budget.
[0073] In order to precisely control the risk of privacy breaches, a separate privacy budget needs to be allocated for the generation of statistics; Figure 3 middle Indicates the protection of the true maximum frequency. The differential privacy budget is used to control subsequent payments to... We added the intensity of Laplace noise. We chose to increase the index privacy budget. and the maximum frequency of privacy budget Separate, instead of using directly The maximum number of leaves in the index is used as the maximum frequency of noise addition. This is because only the index budget is used. This results in very small budgets at each level (because it must be divided into) This significantly increases the noise added at the maximum frequency, thus reducing its effectiveness.
[0074] S36: Based on privacy budget Generate non-negative Laplace noise .
[0075] To ensure that the upper bound estimated subsequently based on this statistic is safe—that is, to prevent underestimation from causing the actual result to overflow the buffer—we base our approach on a privacy budget. Generate non-negative Laplace noise .
[0076] S37: The non-negative random noise value Add to the true maximum frequency In the middle, the maximum frequency after the disturbance is obtained. And satisfy .
[0077] S38: Encapsulate the generated differential privacy maximum frequency into a statistical digest, and publish this digest as a public parameter to each computing node in the multi-party computing environment.
[0078] Maximum frequency of disturbance As metadata distributed along with encrypted fragments, this statistic satisfies differential privacy constraints, and its public disclosure will not reveal the true degree information of any specific node. When the multi-party servers subsequently execute the parallel connection protocol, they will directly read this public parameter and substitute it into the hybrid boundary formula to collaboratively complete a rapid estimation of the size of the intermediate result without the need for additional encrypted communication interaction.
[0079] (4) The differential privacy index and differential privacy statistics are used to perform the connection operation in parallel, compress the connection result, and accelerate the execution of multiple rounds of secure connection.
[0080] This step is the execution phase that transforms differential privacy technology into performance gains for multi-party secure computation. Traditional MPC join operations are limited by data invisibility and must perform full filling according to the worst-case scenario of AGM (Atserias-Grohe-Marx) theory, causing the computational load to grow exponentially with the data size. This step innovatively introduces a "hybrid bound" mechanism: dynamically comparing the estimated upper bound based on statistics with the theoretical worst-case upper bound at runtime, and selecting the tighter bound as the truncation threshold; this not only significantly reduces invalid dummy fills but also achieves bucket-level parallelism and I / O filtering through indexing. In addition, for the multi-round join characteristics of subgraph matching, this step executes "summary update" logic to maintain the index and statistical distribution of intermediate results in real time, ensuring that the acceleration effect is maintained throughout the entire query cycle, rather than just limited to the first round of joins. The specific implementation of this step may include the following sub-steps: S41: In the security connection table and When matching the encrypted edge table, the differential privacy index is used based on the index range. Divide the ciphertext logic of the edge matching table into... A separate bucket By leveraging the sparsity of indexes to identify and skip empty buckets, I / O acceleration is achieved at the data access level.
[0081] This implementation method calls the differential privacy index to obtain the range. Logically divide the ciphertext table into By checking the interval length, empty buckets that do not contain data can be directly identified and skipped, avoiding invalid reads of sparse data areas. This optimizes full table scans into index-based on-demand data block access, significantly reducing disk I / O and network transmission overhead.
[0082] S42: For each non-empty logic bucket Perform bucket join operations in parallel across multiple servers. This reduces the computational load at a single point.
[0083] This implementation leverages the independent data characteristics between buckets to handle connection tasks from non-empty buckets. The computation time is evenly distributed to multiple cores of the multi-party computing cluster for parallel execution, so that the computation time depends on the amount of data in the largest bucket rather than the total size of the entire table.
[0084] S43: Calculate the upper bound of the statistical estimate based on the actual data distribution using the differential privacy maximum frequency statistic. .
[0085] This implementation method will maximize the frequency of differential privacy. Substitute into the formula ,because Including positive one-sided noise, the calculation result can statistically cover the maximum possible size of the actual connection results, thus providing an estimation benchmark that is compact and secure based on the actual data distribution.
[0086] S44: Analyze the current graph query topology, using the AGM bounds formula, based solely on the size of the input table. and And the upper bound of the worst-case scenario that might occur in connection mode computation theory. .
[0087] Applying the AGM theoretical bounds, based solely on the total number of rows in the input table and the number of join keys, we derive the theoretically maximum number of tuples that can be generated. This is a deterministic worst-case bound that does not depend on the data distribution. Although the value is usually large, it exists as a safety threshold in the hybrid bound strategy.
[0088] S45: Implement the hybrid bounds optimization strategy, compare the statistically estimated upper bound with the theoretical worst-case upper bound, and select the minimum of the two as the safe hybrid upper bound for this connection operation. .
[0089] Implement pruning decisions Given the skewed distribution of the actual graph data, Typically much smaller Selecting the minimum value means actively eliminating redundant spaces that theoretically exist but are actually extremely unlikely to occur, thus significantly compressing the size of intermediate results.
[0090] S46: Based on a secure hybrid upper bound Perform controlled invalid tuple filling if the actual number of join results is low. Then fill A dummy tuple marked as invalid is forcibly removed, and redundant padding exceeding the limit is eliminated, thus accelerating computation.
[0091] S47: After generating the join result table, update the differential privacy index and differential privacy statistics, first targeting the join key attributes reused in the join result. Count the total number of tuples containing dummy elements actually generated in each bucket. .
[0092] Initiate the summary update process, scan the encrypted result table, and count the number of rows actually occupied by each logical bucket in the output table (including valid data and dummy cells). This step captures the actual distribution of intermediate results and provides basic data for rebuilding the index structure in the next round.
[0093] S48: Based on the total number of tuples Dynamically recalculate the differential privacy index range of the join key attribute in the result table. Update the index structure to adapt to changes in the result table.
[0094] Based on the actual number of rows in each bucket, a prefix sum calculation is performed to deduce the new index range of this attribute in the result table. The updated index accurately reflects the data layout after hybrid bounds compression, ensuring that the next round of join operations can continue to use the index for accurate bucket-level positioning and skipping.
[0095] S49: For the maximum frequency update, calculate the new differential privacy maximum frequency in the result table. .
[0096] This implementation provides a connection. ,application The updated statistical summary, a multiplicative update strategy, ensures that the maximum frequency of propagation remains an effective upper bound of the true maximum frequency in the connection results, providing an immediate statistical factor for the next round of estimation.
[0097] S410: The updated differential privacy index and the new differential privacy maximum frequency are encapsulated as input parameters for the next round of connection operations, ensuring that efficient mixing limit estimation can always be maintained using effective statistical information during multi-round continuous subgraph matching.
[0098] By constructing a closed-loop feedback, the updated lossless differential privacy index and differential privacy maximum frequency are passed to the next stage. This ensures that no matter how long the query path is, each round of connections can perform hybrid boundary optimization based on the latest statistical information, preventing the accumulation of estimation errors and achieving continuous acceleration throughout the entire process.
[0099] The above description of the embodiments is provided to enable those skilled in the art to understand and apply the present invention. Those skilled in the art can readily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative effort. Therefore, the present invention is not limited to the above embodiments, and any improvements and modifications made to the present invention by those skilled in the art based on the disclosure thereof should be within the scope of protection of the present invention.
Claims
1. A multi-party secure subgraph matching method based on differential privacy acceleration, characterized in that, Includes the following steps: (1) Obtain the original graph data to construct an edge matching table containing bidirectional edge information, and distribute the edge matching table to a multi-party computing server using a secret sharing protocol; the original graph data is a social network consisting of multiple users and their friend relationships. (2) The hierarchical histogram constructed based on the attribute value range is perturbed using the one-sided Laplace mechanism to generate a lossless differential privacy index. The specific implementation method is as follows: A1. Construct a noise-free, true hierarchical histogram tree based on the edge matching table of the original graph data to provide an accurate baseline structure for subsequent perturbations; A2. Using differential privacy budgeting, a sign-constant one-sided Laplace noise is generated and injected into the real hierarchical histogram tree to obtain upper and lower bound histograms for estimating cumulative frequency, laying a conservative estimation foundation for lossless indexing. A3. Utilizing the conservative offset characteristics of upper and lower bound histograms, through interval partitioning, range querying, and truncation correction, an index range is generated for each logical interval to ensure no record is missed, and finally, a lossless index set with differential privacy protection is output. The specific implementation process of step A2 is as follows: S24: Obtain the privacy budget for index building, distribute the privacy budget evenly to each level according to the height of the real hierarchical histogram tree, and calculate the privacy budget for a single level. S25: Based on a single-layer privacy budget, a probability distribution model is constructed using a one-sided Laplace mechanism. Two independent noise vectors are generated by sampling, where all components of the first noise vector are constrained to be non-negative and all components of the second noise vector are constrained to be non-positive. S26: Inject the second set of noise vectors into the node counts of the true hierarchical histogram tree to construct a lower bound histogram specifically for estimating the lower bound of the cumulative frequency. S27: Inject the first set of noise vectors into the node counts of the real hierarchical histogram tree to construct an upper bound histogram specifically for estimating the upper bound of the cumulative frequency. (3) Calculate the true maximum frequency of the connection key attribute and inject differential privacy noise to generate differential privacy maximum frequency statistics; (4) The lossless differential privacy index and the differential privacy maximum frequency statistic are used to perform the join operation in parallel, compress the join result, and accelerate the execution of multi-round secure joins. The specific implementation method is as follows: B1. Logical bucketing and parallel join of data are achieved through differential privacy indexing, and the amount of result filling is precisely controlled by hybrid upper bound, which accelerates the join operation while ensuring security. B2. Based on the results of this round of connections, dynamically adjust the differential privacy index and the maximum frequency statistic to provide an accurate and tight statistical upper bound for the next round of secure connections.
2. The multi-party secure subgraph matching method based on differential privacy acceleration according to claim 1, characterized in that, The specific implementation method of step (1) is as follows: S11: Obtain the original graph data G =( V , E ),in V For a set of nodes, E Let be a set of edges, where nodes correspond to users and edges correspond to the friend relationships between users; S12: For the edge set E Any undirected edge in ( u , v Construct the corresponding reverse join tuple ( ). v , u ), u and v These are two nodes connected by an undirected edge; S13: Concatenate and merge the original list of undirected edges with the constructed list of reverse-connected tuples to build an edge matching table containing complete bidirectional connections. S14: Based on the replication secret sharing protocol, split any attribute value in the edge matching table into three random pieces that satisfy the XOR reconstruction relationship; S15: Distribute the random shards to three computing servers according to the ring replication strategy, ensuring that each computing server holds two copies of the shards, thereby constructing a distributed edge matching table of ciphertext state in a multi-party computing environment.
3. The multi-party secure subgraph matching method based on differential privacy acceleration according to claim 1, characterized in that, The specific implementation process of step A1 is as follows: S21: Determine the target attribute used to construct the index in the edge matching table, and then determine the discrete value range space of the attribute; S22: Check the size of the discrete value range space. If the size is not a power of 2, then fill and expand the value range by adding dummy buckets with a count of zero until the total length meets the construction requirements of a complete binary tree. S23: Statistically analyze the true frequency distribution of the target attribute in the edge matching table, and construct a true hierarchical histogram tree as the baseline, where each leaf node stores the true count of a single attribute value, and each internal node stores the sum of the counts of its child nodes.
4. The multi-party secure subgraph matching method based on differential privacy acceleration according to claim 1, characterized in that, The specific implementation process of step A3 is as follows: S28: Divide the value space of the target attribute into multiple non-overlapping continuous logical intervals, which serve as the basic scheduling unit for subsequent parallel connection operations; S29: For any logical interval, perform a range query on the lower bound histogram, accumulate the noise count of all buckets from the start of the value range to the start point of the interval, and obtain the estimated starting offset of the interval. S210: Perform a non-negative truncation operation on the estimated starting offset, correcting values less than zero to zero, and obtain the effective index starting position of the interval in the ciphertext table; S211: For any logical interval, perform a range query on the upper bound histogram, accumulate the noise count of all buckets from the start of the value range to the end of the interval, and obtain the estimated end offset of the interval. S212: Perform boundary truncation operation on the estimated end offset, correct the value exceeding the total number of rows in the edge matching table to the maximum number of rows in the table, and obtain the effective index end position of the interval in the ciphertext table; based on the effective index start position and effective index end position of each logical interval, output the final lossless index set.
5. The multi-party secure subgraph matching method based on differential privacy acceleration according to claim 1, characterized in that, The specific implementation method of step (3) is as follows: S31: Obtain the edge matching statement before secret sharing, and determine the connection key attribute column that needs to be statistically analyzed; S32: Utilizing the structural characteristics of the edge matching table, which is composed of the original edge and its transpose edge, we can confirm that the source node column and the target node column in the table have symmetrical consistency in statistical distribution, and thus select only one of the columns as the statistical object. S33: Perform aggregation statistics on the node identifiers in the selected column to calculate the frequency of each node in the edge matching table in the original graph. This frequency is semantically equivalent to the degree of the corresponding node. S34: Compare the frequency values of all nodes, filter out the maximum value, and define it as the true maximum frequency of the connection key attribute; S35: Obtain a privacy budget dedicated to the maximum frequency statistics, and generate non-negative Laplace noise based on the privacy budget; S36: Add Laplace noise to the true maximum frequency to generate the maximum frequency after differential privacy perturbation; S37: Encapsulate the maximum frequency after differential privacy perturbation into a statistical digest, and distribute the digest as a public parameter to the multi-party computation server.
6. The multi-party secure subgraph matching method based on differential privacy acceleration according to claim 1, characterized in that, The specific implementation process of step B1 is as follows: S41: In the secure connection input table R 1 and R At time 2, lossless differential privacy index is used to divide the edge matching table ciphertext logic into multiple independent logical buckets according to the index range. The sparsity of the index is used to identify and skip empty buckets, thereby accelerating I / O at the data access level. S42: For any non-empty logical bucket, execute the bucket input table in parallel across multiple computing servers. R 1 and R 2. Connectivity operations to reduce the computational load at a single point; S43: Calculate the upper bound of the statistical estimate based on the actual data distribution using the differential privacy maximum frequency statistic; S44: Analyze the current graph query topology, using the AGM theoretical limit formula, based solely on the input table. R 1 and R The size of 2 and the upper bound of the worst-case calculation of the connection mode are used. S45: Implement the hybrid bounds optimization strategy, compare the statistically estimated upper bound with the theoretical worst-case upper bound, and select the minimum of the two as the safe hybrid upper bound for this connection operation. U h ; S46: Perform controlled invalid tuple filling based on a safe hybrid upper bound, if the actual number of join results... N Less than U h Then fill U h - N A dummy tuple marked as invalid is forcibly removed, and redundant padding exceeding the limit is eliminated, thus accelerating computation.
7. The multi-party secure subgraph matching method based on differential privacy acceleration according to claim 1, characterized in that, The specific implementation process of step B2 is as follows: S47: After generating the join result table, update the lossless differential privacy index and the differential privacy maximum frequency. First, for the join key attribute reused in the join result, count the total number of tuples actually generated in each bucket, including dummy elements. S48: Dynamically recalculate the differential privacy index range of the join key attribute in the join result table based on the total number of tuples, and update the lossless differential privacy index structure to adapt to the changes in the result table; S49: Calculate the join key attributes for the input table respectively. R 1 and R The maximum frequency of differential privacy is obtained by multiplying the two sets of frequencies to get a new maximum frequency of differential privacy. S410: The updated lossless differential privacy index and differential privacy maximum frequency are encapsulated as input parameters for the next round of connection operations, ensuring that efficient mixing limit estimation can always be maintained by utilizing effective statistical information during multi-round continuous subgraph matching.
Citation Information
Patent Citations
Accurate histogram publishing method based on differential privacy
CN109492047A
Non-equidistant histogram publishing method based on differential privacy
CN110795758A