Efficient and safe database connection and grouping aggregation method based on server auxiliary architecture
By introducing a three-party server auxiliary architecture in database connection and packet aggregation, and using the auxiliary server to perform calculation-intensive tasks, the problems of high communication overhead and number of rounds are solved, and efficient database query performance is achieved.
Patent Information
- Application Number
- CN202510332297.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art solves the problem of high communication overhead or high communication round count when solving database connection and packet aggregation queries, resulting in inefficiency in large-scale and high-latency scenarios.
An efficient and secure database connection and packet aggregation method based on a tripartite server auxiliary architecture is proposed. By introducing auxiliary server P0, it performs calculation-intensive tasks, such as inadvertent sorting and merging calculations, and generates randomness to reduce redundant interaction overhead.
It significantly reduces cross-party communication load, reduces traffic volume, and provides low-round efficient aggregation computing solutions to adapt to database query requirements in large-scale and high-latency scenarios.
Smart Images

Figure CN120179684A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of secure multi-party computation (MPC) and secure database joint analysis, and mainly relates to a group of efficient and secure Join and Group-by-Aggregation protocols under a semi-honest secure third-party assistance architecture. More specifically, it is an efficient and secure database connection and group aggregation method based on a server assistance architecture. Background Art
[0002] Currently, the demand for cross-institutional data joint analysis in fields such as financial risk control, medical research, and government governance has increased exponentially. However, due to the increasing attention to privacy issues and the introduction of relevant laws and regulations, the data silo problem has become the core obstacle to such collaborations. Secure multi-party computation (MPC), as a secure execution method, allows participating parties to jointly compute functions without revealing the original data, providing a solution for collaborative analysis of sensitive databases for mutually distrustful institutions. However, the direct use of MPC technology has a high cost, and there are serious performance bottlenecks in the face of actual application scenarios. There have been many works that have implemented a secure collaborative analysis framework for databases through secure multi-party computation (MPC), but the core bottleneck of its query performance still focuses on the secure design and development of key operators such as Join and Group-by-Aggregation. Such operations involve cross-computation of multi-party data, often resulting in high communication overhead and multi-round interaction latency.
[0003] More information related to the above technical solutions can also be found in the following literature. Bater et al. proposed the first secure database joint analysis system SMCQL, which is based on the semi-honest secure garbled circuit technology. SMCQL allows the data holder to preprocess the plaintext data locally by decomposing the query plan, thus reducing the use of complex MPC operators and the computing scale. For the implementation of the join operation, the data holder first sorts the local data, then aligns the data through the oblivious merge circuit, and finally performs comparison on adjacent elements to complete the join. For grouped aggregation, the system sequentially traverses all elements to complete the grouped statistics. The architecture of SMCQL that reduces the usage frequency of high-overhead MPC operators provides an important paradigm reference for the design of secure database systems. This "local preprocessing + encrypted state collaboration" hybrid architecture effectively reduces the call frequency of MPC and compresses the core computing scale to the sub-linear level, providing an important paradigm reference for the design of early privacy computing systems. Poddar et al. proposed a maliciously secure joint database analysis system Senate that supports n parties. By disassembling the multi-party protocol into two sub-circuits and then gradually merging them, the efficiency is improved through parallelism. Its join and grouped aggregation operations are similar to the execution strategies of SMCQL. Wang et al. proposed a secure variant of the classic Yannakakis algorithm for query scenarios. Their scheme realizes the secure join operation through circuit-based private set intersection (PSI), while the grouped aggregation operation requires the data holder to first sort the grouped attribute columns locally. Then, the oblivious extension permutation (OEP) is used to align the payload columns of the secret sharing to complete the grouping, and finally the aggregation calculation is performed through the garbled circuit. However, there are two major limitations in such schemes: First, the data holder needs to participate in the online calculation throughout the process and cannot adapt to the offline delegation scenario; Second, the core operators based on garbled circuits (such as PSI and OEP) result in a large communication overhead, seriously restricting the feasibility of actual deployment.
[0004] To meet the requirements of actual scenarios, Secrecy, Scape, and AHK+ etc. use the outsourcing computing mode to share data among computing nodes, allowing the data holder not to participate in the calculation, improving the stability of the service and reducing the communication overhead. Conclave designed a hybrid MPC-plaintext protocol that exposes the plaintext values of the join and grouped attribute columns to the selectively trusted party during the calculation, significantly reducing the communication overhead. The three-party semi-honest system Secrecy constructed by Liagouris et al. uses a double-loop full-scale matching strategy to implement the join operation, resulting in O(n 2) high communication overhead at the [level]. Its packet aggregation relies on an oblivious sorting circuit based on Bitonic sorting and parity aggregation. For the maliciously secure three-party scenario, the Scape system proposed by Feng Han et al. encodes and sorts the secret-sharing vectors through a shared oblivious pseudorandom function, uses the oblivious permutation and merging protocol to align the same encoded values, and finally completes the join operation by comparing adjacent elements. Its packet operation is also based on oblivious Bitonic sorting, and the aggregation uses an oblivious traversal protocol with a parallel network structure. The AHK+ scheme proposed by Asharov et al. locates the matching items for the unique key join through oblivious sorting and uses the 0 / 1 flag column to identify and propagate the payload. The Sum aggregation is achieved based on oblivious sorting and the prefix sum property, and the Max / Min aggregation requires secondary sorting of the grouping and aggregation attributes.
[0005] However, the prior art still has the following problems: There are defects of high communication overhead or high communication rounds when solving join and grouped aggregation queries, which makes them inefficient for large-scale and high-latency services. On the one hand, multi-table join faces a trade-off between communication overhead and communication rounds. Secrecy uses a double loop to compare all pairs, which can be executed in parallel, but this method results in O(n 2 ) communication overhead. Scape and AHK+ use expensive MPC operators, such as oblivious merging or oblivious sorting, to group the rows with the same elements in the join columns, thus reducing the comparison, but this method leads to more communication rounds. On the other hand, grouped aggregation heavily relies on sorting, which is an operation that requires a large number of communication rounds. Secrecy and Scape use a Bitonic sorting network, where oblivious sorting involves many comparison and swap operations, as well as hierarchical dependencies, resulting in high communication costs and communication rounds. AHK+ uses radix sorting, where oblivious sorting sorts the data bit by bit, resulting in communication rounds that are linear with the size of the input elements but with a large multiplicative factor.
[0006] Considering these problems, to address the performance bottleneck in secure database queries in large-scale and high-latency scenarios, the challenge for multi-table join is to reduce the communication overhead of O(n 2 ) comparisons while maintaining high parallelism similar to Secrecy. The challenge for grouped aggregation is to achieve efficient-round grouping with a low communication overhead similar to AHK+ and avoid high-round sorting. Summary of the Invention
[0007] To this end, the present invention proposes a set of communication-efficient and round-efficient secure join and grouped aggregation protocols. While ensuring data privacy and security, it significantly improves the efficiency in actual high-latency large-scale scenarios, enabling it to adapt to the requirements of efficient database queries.
[0008] To achieve the above goals, the present invention is based on a three-party server-assisted computing architecture. Besides the traditional two-party computing parties (P1 and P2), an auxiliary server P0 is introduced to be responsible for computationally intensive tasks such as oblivious sorting and merging computations, and to generate randomness, which is converted into local computations to eliminate redundant interaction overheads.
[0009] To achieve this goal, the present invention is based on a three-party server-assisted architecture. Outside the traditional two-party computing model, a dedicated auxiliary server is introduced to reduce the computational tasks of high-cost operations, thereby converting the high-overhead MPC operations of multiple rounds of interactions commonly used in existing solutions into local computations. By delegating these tasks to the auxiliary server, the present invention effectively avoids redundant overheads and significantly reduces the cross-party communication load. Under this architecture, the present invention designs an efficient communication protocol for connection and grouping. While ensuring data security and efficient rounds, this protocol reduces the communication volume and provides a low-round efficient solution for aggregation computations, which is specifically optimized for common aggregation functions such as Sum, Avg, and Max / Min.
[0010] An efficient and secure database connection and grouping aggregation method based on a server-assisted architecture, characterized in that the construction of the three-party server-assisted architecture includes:
[0011] Computer nodes P1 and P2, which are responsible for storing the secret sharing shares provided by the data owners and performing secure computing tasks;
[0012] Auxiliary node P0: It does not hold any data secret shares and is responsible for performing locally computationally intensive operations;
[0013] Data owner: After generating secret shares from the original data, it distributes them to computer nodes P1 and P2 and can then go offline;
[0014] Data analyst: Submits a query request and obtains the final analysis result without contacting the original data and the intermediate computing process throughout the process;
[0015] The efficient and secure database connection and grouping aggregation method based on the server-assisted architecture specifically includes the following steps:
[0016] S1. Data distribution: The data owner shares the data with computer nodes P1 and P2 and then goes offline. Computer nodes P1 and P2 respectively hold data shares;
[0017] S2. Query trigger: The data analyst submits a query request to the system, and computer nodes P1 and P2 and the auxiliary node P0 start computing tasks;
[0018] S3. Collaborative computing: The computer nodes P1, P2 and the auxiliary node P0 cooperate to complete secure computing through random coding and oblivious shuffling operations;
[0019] S4. Result return: The computer nodes P1, P2 send the result shares to the data analyst, and the data analyst reconstructs the result.
[0020] For further optimization of this technical solution, the collaborative computing in step S3 includes the following two protocols
[0021] A. Secure connection protocol: Used for efficient alignment of cross-table private data, which solves the O(n 2 ) communication overhead problem in the traditional solution while maintaining a constant number of rounds;
[0022] B. Group aggregation protocol: Supports encrypted aggregation calculation of grouped key values, reducing the round complexity from O(n) to O(logn).
[0023] For further optimization of this technical solution, it is characterized in that the secure connection protocol is specifically as follows:
[0024] A1. Data preparation: In the data distribution stage of the system, the data owner splits the key value columns of the two tables X and Y to be connected into secret sharing shares, and transmits them to the computing nodes P1 and P2 respectively. The computing nodes P1 and P2 generate a shared key k for subsequent encryption operations;
[0025] A2. Random coding mapping: The computing nodes P1 and P2 perform secure coding mapping on the key values based on the shared key k using the LowMC block cipher algorithm. All encoded values are transmitted to the auxiliary node P0, and the auxiliary node P0 cannot reverse-derive the original key values through the encoding, ensuring privacy security;
[0026] A3. Local matching and permutation rule generation: The auxiliary node P0 locally compares the encoded values e X and e Y of table X and table Y pair by pair, identifies all attribute pairs with equal encodings, and generates permutation rules π1 and π2. The design goal of the permutation rules is to align the matching rows to the same index position, and the unmatched rows are uniformly randomly arranged at the end of the table;
[0027] A4. Composite permutation and secure shuffling: To hide the data alignment path, the computing nodes P1 and P2 perform an oblivious shuffle on tables X and Y respectively based on the matching permutations π1 and π2, each time performing two random permutations and three re-randomization operations;
[0028] A5. Result generation and output: The computing nodes P1 and P2 send the shuffled share tables <x>and <y>Return to the data analyst through the secure channel. The data analyst reconstructs the table shares to obtain the joined result table.
[0029] For a further optimization of this technical solution, in step A4, compound permutation and secure shuffle: To hide the data alignment path, computing nodes P1 and P2 perform an oblivious shuffle on tables X and Y respectively based on the matching permutations π1 and π2. Each execution involves two random permutations and three re-randomization operations. Taking table X as an example:
[0030] A41. First random permutation: Participating computing nodes P1 and P2 re-randomize the share <x K > to generate <h>, and submit the share of computing node P1 to the auxiliary node P0. The auxiliary node P0 generates a random permutation σ1 and synchronizes it to computing node P2. The two parties jointly perform the σ1 permutation on the share and re-randomize it again to generate <c>;
[0031] A42. Composite permutation generation: The auxiliary node P0 calculates the composite permutation based on the target permutation π1 Guide the auxiliary node P0 and the calculation node P1 for the share <c>Perform replacement and re-randomization again to obtain <g>, the auxiliary node P0 sends the share to the computing node P1, and the computing nodes P1 and P2 pair with <g>Finally, perform one round of re-randomization to obtain <w>As an accidental shuffling result, the whole process equivalently realizes the original permutation π through the superposition of two permutations;
[0032] A43. Re-randomization injection: After each permutation, the computing nodes P1 and P2 add random masks to the shares to ensure that no participant can infer data correlation through the intermediate state.
[0033] For further optimization of this technical solution, the grouped aggregation protocol includes a grouping stage and an aggregation stage.
[0034] The steps of the grouping stage are as follows:
[0035] B1. Data preparation: The computing nodes P1 and P2 hold the secret shares of the grouped key value column t of the data table T obtained in the system distribution stage and generate a shared key k. K and generate a shared key k;
[0036] B2. Random coding mapping: Similar to the connection scheme, the three parties execute a random coding function based on the LowMC algorithm to assist the node P0 to obtain the encoded value e K of t K , and the auxiliary node P0 dynamically generates a grouping permutation rule π according to the equality of ex, so that the rows with the same encoding are arranged adjacent to each other, and the ungrouped rows are uniformly arranged at the end of the table;
[0037] B3. Secure shuffle alignment grouping: To hide the grouping logic, the computing nodes P1 and P2 perform two random permutations, and through the same oblivious shuffle operation as the connection protocol, the auxiliary node P0 and the computing nodes P1 and P2 jointly complete the rearrangement of the table T based on π to ensure that the secretly shared data is aligned in the same grouping order;
[0038] B5. Grouping result verification: The auxiliary node P0 generates a grouping indication array d to mark the starting position of each group, and d is secretly shared to the computing nodes P1 and P2 for subsequent aggregation operations;
[0039] In the aggregation stage, for the grouped secretly shared table R, based on the Ladner-Fischer parallel prefix network, a multi-level aggregation unit is designed to efficiently calculate the aggregation value, and the steps are as follows:
[0040] B6. Data preprocessing:
[0041] ο The grouped secretly shared table R is aligned according to the grouping attribute <r K >>, and the computing nodes P1 and P2 hold the secret shares of the aggregation attribute <r S >>;
[0042] ο If the length of the secretly shared table R is not a power of 2, dummy data is supplemented;
[0043] B7. Parallel prefix network construction:
[0044] ο Network structure: The grouped and aligned table is divided into multiple sub - blocks. Each sub - block independently calculates the local aggregation value, and then recursively merges the local results;
[0045] ο Recursive merging: Starting from the bottom - layer sub - blocks, the results are merged layer by layer upwards;
[0046] ο Basic unit definition: Given a grouping indicator array d ∈ {0, 1} n and an aggregation attribute r S ∈ R n , define AggNetUnit as Δ g:h =(a, b), which represents the aggregation information of the last group within the index range [g, h], where g, h ∈ {1, …, n}. In AggNetUnit Δ g:h =(a, b), the grouping information a ∈ {0, 1} indicates whether all items within the range [g - 1, h] belong to the same group, and b ∈ R records the aggregation result of the last group within the range [g, h];
[0047] B8. Aggregation calculation logic: Build a parallel comparison tree Max / Min or a summation tree Sum / Avg within each group, and recursively select the extreme values; Compare the encrypted numerical values through a secure comparison protocol to avoid plaintext leakage;
[0048] B9. Optimize Sum / Avg: For the Sum operation, calculate the prefix - sum array u. The auxiliary node P0 inputs the permutation π generated based on the grouping result identification array d. The three parties perform an oblivious shuffle to shuffle the secretly - shared prefix - sum array u according to the permutation π, so that after shuffling, the last row of each group in the prefix - sum array is arranged in order at the beginning of the array, and the sum within the group is extracted by taking the difference between adjacent prefix - sums; For the Avg operation, this optimization also applies;
[0049] B10. Result generation and output: The computing nodes P1 and P2 return the aggregation result share table to the data analyst through a secure channel, and the data analyst reconstructs the table shares.
[0050] For further optimization of this technical solution, noise is injected into the original key - value column before data sharing to ensure that the data for subsequent processing is the noise - added result, meeting the (∈, δ) - differential privacy requirement, specifically as follows:
[0051] · Noise generation: The data owner uses the truncated Laplace mechanism to generate noise;
[0052] · Expand the coding domain: To avoid key - value conflicts caused by virtual data, expand the original key - value coding domain from 64 bits to 128 bits to ensure that the random coding of virtual key - values does not overlap with the real key - values.
[0053] A further optimization of this technical solution also includes generating an identification array of hidden statistical information based on the noised data to avoid directly exposing the number of matching rows or the grouping scale, specifically as follows:
[0054] Unique key connection: The auxiliary node P0 generates an array of matching row identifiers composed of 0 / 1: 1 indicates a match, and 0 indicates no match, instead of directly returning the number of matching rows; this array is distributed to the computing nodes P1 and P2 through secret sharing, and only the valid matching row data is extracted after they jointly decrypt it;
[0055] Group aggregation: The auxiliary node P0 generates an array of group boundary identifiers, marking the continuity within the group and the separation points between groups, and secretly shares the array to the computing nodes P1 and P2 through oblivious shuffling.
[0056] Different from the prior art, the present invention has made key optimizations in the connection and group aggregation calculations:
[0057] · For the multi-table join problem, the present invention introduces a random coding mapping outsourcing mechanism, entrusting the random coding of the join key to the auxiliary server P0 for processing, enabling it to perform pairwise comparisons locally. This method converts the high communication complexity comparison of the original O(n 2 ) level into the local calculation of P0, while maintaining a low communication round and significantly reducing the data transmission volume between the computing parties.
[0058] · For the group aggregation calculation problem, the present invention breaks through the dependence on sorting in the traditional scheme, allowing the auxiliary server P0 to group the random coding of elements locally without performing expensive oblivious sorting. Since the random coding hides the relative order of the original data, this method only requires a constant number of rounds (O(1)) to complete the grouping calculation, while still ensuring that the grouping results are correctly arranged. This method supports grouping but not sorting, avoiding the additional communication overhead in the existing methods and making our scheme more efficient.
[0059] · In addition, the present invention further optimizes the aggregation calculation by extending the parallel prefix network to the aggregation function,
[0060] significantly reducing the number of calculation rounds, enabling aggregation calculations such as Sum, Avg, Max / Min, etc. to be completed at a lower communication cost.
[0061] · For the problem that the auxiliary node may infer sensitive data through statistical information, the present invention adopts differential privacy noise injection and data obfuscation technologies, injecting noise into the original key value column before data sharing, and using a dynamic identification array to hide the number of matching rows and the grouping scale, ensuring that the leakage risk is controlled at the (∈,δ)-differential privacy level (within the acceptable range of the scenario). Description of the Drawings
[0062] Figure 1 It is a server-assisted three-party architecture diagram;
[0063] Figure 2 It is a connection operation flow chart;
[0064] Figure 3 It is a parallel prefix network diagram extended to grouped summation;
[0065] Figure 4 It is a comparison chart of connection protocol performance under local area network (LAN) and wide area network (WAN);
[0066] Figure 5 It is a comparison chart of grouped Max / Min protocol performance under local area network (LAN) and wide area network (WAN);
[0067] Figure 6 It is a comparison chart of grouped summation protocol performance under local area network (LAN) and wide area network (WAN). Detailed implementation
[0068] 1. Overall architecture design
[0069] The present invention designs a new server-assisted architecture, which breaks the symmetric structure between traditional computing nodes, redistributes computing tasks, effectively converts high-cost multi-party oblivious operations and random correlation generation into local computing, eliminates redundant sorting steps in connection and grouping operations, avoids a large amount of security computing overhead, and makes the overall computing process more efficient, improving the actual application performance. For this reason, the present invention further constructs a server-assisted three-party computing model. Refer to Figure 1 As shown, it is a server-assisted three-party architecture diagram, showing the task division and cooperation mode between different nodes.
[0070] This embodiment is implemented based on a server-assisted architecture of three-party cooperation, and includes the following core roles:
[0071] · Computing nodes P1, P2: responsible for storing the secret sharing shares provided by the data owner and performing security computing tasks (such as encryption, shuffling, decryption, etc.);
[0072] ● Auxiliary node P0: does not hold any data secret shares, and is responsible for performing local computing-intensive operations (such as key-value encoding matching, permutation rule generation), significantly reducing the communication overhead of multi-party interaction;
[0073] · Data owner: generates secret shares from the original data and distributes them to P1, P2, and then can go offline;
[0074] · Data analyst: submits a query request and obtains the final analysis result, and does not contact the original data and the intermediate computing process throughout the process.
[0075] This architecture supports multiple data owners, and P1 and P2 can be the data owners themselves or external entities, such as cloud service providers. In addition, to improve computational efficiency, the auxiliary party P0 is allowed to obtain certain additional information within a controlled range. The disclosure of these intermediate information does not involve specific original data and is negligible in practical applications, thus achieving a balance between privacy protection and execution efficiency.
[0076] Based on the above role division, the collaboration process of the system can be summarized into the following four steps (as Figure 1 ):
[0077] S1. Data distribution: After the data owners share the data with P1 and P2, they go offline, and P1 and P2 each hold a data share.
[0078] S2. Query trigger: The data analyst submits a query request to the system, and P1, P2, and P0 start a computing task.
[0079] S3. Collaborative computing: P1, P2, and P0 cooperate to complete secure computing through operations such as random coding and oblivious shuffling.
[0080] S4. Result return: P1 and P2 send the result shares to the data analyst, and the data analyst reconstructs the result.
[0081] Through role separation and task decoupling, this architecture transfers the high-cost oblivious operations (such as sorting and full-scale comparison) in traditional MPC to the local execution of the auxiliary node P0, reducing the number of multi-party interaction rounds. At the same time, P0 does not hold the data secret share and only completes the calculation based on the encrypted state coding, ensuring the privacy of the original data. This design not only supports the offline participation of data owners but also can adapt to high-latency network environments, providing basic support for the efficient and secure processing of large-scale data.
[0082] In the present invention, the collaborative computing stage mainly includes the following two core protocols, specifically including:
[0083] A. Secure connection protocol: Used for the efficient alignment of cross-table private data, solving the O(n 2 ) communication overhead problem in traditional solutions while maintaining a constant number of rounds;
[0084] B. Secure grouped aggregation protocol: Supports the encrypted state aggregation calculation of grouped key values, reducing the round complexity from
[0085] O(n) to O(logn).
[0086] The following will elaborate on the specific solutions of these two protocols in detail.
[0087] 2. Implementation scheme of the secure connection protocol
[0088] Efficient alignment of cross - institutional privacy data is achieved through dynamically generated random code matching and oblivious shuffling techniques. Figure 2 The specific implementation process is as follows:
[0089] (1) Data preparation: In the data distribution phase of the system, the data owner splits the key - value columns (such as user ID, transaction number) of the two tables (Table X and Table Y) to be joined into secret - shared shares and transmits them to computing nodes P1 and P2 respectively. For example, the key - value column x of Table X K is sharded into <x K >1 and <x K >2, and the key - value column y of Table Y K is sharded into <y K >1 and <y K >2. Computing nodes P1 and P2 generate a shared key k for subsequent encryption operations.
[0090] (2) Random code mapping: P1 and P2 perform secure code mapping on the key - values based on the shared key k using the LowMC block cipher algorithm. Specifically, for each key - value x K [i] of Table X, the computing node inputs its share <x K [i]> into the LowMC algorithm to generate the encrypted code e X [i]=LowMC k (<x K [i]>1 + <x K [i]>2). Similarly, the key - value columns of Table Y generate the code e Y . All the encoded values are transmitted to the auxiliary node P0, and P0 cannot reverse - deduce the original key - values from the codes, ensuring privacy security.
[0091] (3) Local matching and permutation rule generation: The auxiliary node P0 locally compares the encoded values e X and e Y of Table X and Table Y pair - by - pair, identifies all pairs of attributes with equal codes, and generates permutation rules π1 and π2. The design goal of the permutation rules is to align the matching rows to the same index position. For example, if the i - th row of Table X matches the j - th row of Table Y, then π1(i)=π2(j); the unmatched rows are uniformly randomly arranged at the end of the table.
[0092] (4) Composite permutation and secure shuffling: To hide the data alignment path, P1 and P2 perform an oblivious shuffle on Table X and Table Y respectively based on the matching permutations π1 and π2, each performing two random permutations and three re - randomization operations. Taking Table X as an example:
[0093] ο First random permutation: The participating parties P1 and P2 perform re - randomization on the share <x K > to generate <h>, and submit the share of P1 to P0. P0 generates a random permutation σ1 and synchronizes it to P2. The two parties jointly perform the σ1 permutation on the share and re-randomize it again to generate <c>;
[0094] ο Composite permutation generation: P0 calculates the composite permutation based on the target permutation π1 Guide P0 and P1 for the shares <c>Perform replacement and re-randomization again to obtain <g>, P0 sends the share to P1, and P1 and P2 pair with <g>Finally, perform one round of re-randomization to obtain <w>As a result of inadvertent shuffling. The entire process equivalently implements the original permutation π through the superposition of two permutations;
[0095] ο-fold randomization injection: After each permutation, P1 and P2 add random masks to the shares to ensure that no participant can infer data correlation through the intermediate state.
[0096] (5) Result generation and output: P1 and P2 will shuffle the share tables <x>and <y>Return to the data analyst through the secure channel. The data analyst reconstructs the table shares to obtain the joined result table.
[0097] Inner join: P1 and P2 retain all payloads (such as transaction amounts, timestamps) of the matching rows on both sides and return them to the data analyst as the joined result table.
[0098] Semi-join: P1 and P2 only retain the matching rows in the left table (Table X) and do not include the specific data of the right table (Table Y).
[0099] 3. Implementation Scheme of the Grouping Aggregation Protocol
[0100] The grouping aggregation protocol of the present invention realizes efficient private data statistics through dynamic coding alignment and parallel computing.
[0101] 3.1 Grouping Phase
[0102] The implementation scheme of the grouping protocol is similar to the joining protocol. It uses random coding and oblivious shuffling techniques to complete efficient grouping, ensuring that rows with the same key values are arranged adjacent to each other, while hiding the grouping logic to avoid revealing data distribution characteristics. The steps are as follows:
[0103] (1) Data preparation: The computing parties P1 and P2 hold the secret shares of the grouped key value column t of the data table T obtained in the system distribution phase and generate a shared key k. K
[0104] (2) Random coding mapping: Similar to the joining scheme, the three parties execute a random coding function based on the LowMC algorithm. P0 obtains the encoded value e of t. K K P0 dynamically generates a grouping permutation rule π according to the equality of e, making rows with the same encoding adjacent to each other, and ungrouped rows (such as invalid data or noise) are uniformly arranged at the end of the table. K
[0105] (3) Secure shuffle alignment grouping: To hide the grouping logic, P1 and P2 perform two random permutations. Through the same oblivious shuffling operation as the joining protocol, P0 and P1 and P2 jointly complete the rearrangement of table T based on π, ensuring that the secretly shared data is aligned in the same grouping order.
[0106] (4) Grouping result verification: P0 generates a grouping indication array d to mark the starting position of each group (such as d[i]=0 indicating that the i-th row is the starting point of a new group). d is secretly shared to P1 and P2 for subsequent aggregation operations.
[0107] 3.2 Aggregation Phase
[0108] For the grouped secret sharing table R, based on the Ladner-Fischer parallel prefix network, design a multi-level aggregation unit (AggNetUnit) to efficiently calculate aggregation values such as Sum, Avg, Max / Min. The specific implementation steps are as follows:
[0109] (2) Data preprocessing:
[0110] ο The grouped table R is aligned according to the grouping attribute <r K >, and P1 and P2 hold the secret shares of the aggregation attribute <r S >;
[0111] ο If the length of table R is not a power of 2, dummy data is supplemented (Sum is filled with 0, and Max / Min is supplemented with extreme values).
[0112] (3) Construction of the parallel prefix network (as Figure 3 shown):
[0113] ο Network structure: The grouped and aligned table is divided into multiple sub-blocks, and each sub-block independently calculates the local aggregation value. For example, for the Sum function, the sub-block calculates the local prefix sum; for the Max function, a local comparison tree is constructed. Then the local results are recursively merged;
[0114] ο Recursive merging: Starting from the bottom-level sub-blocks, the results are merged layer by layer upwards. For example, at the first layer, the aggregation values of adjacent two elements are calculated, and at the second layer, four elements are merged until the global result is generated at the top layer;
[0115] ο Definition of the basic unit: Given the grouping indicator array d ∈ {0,1} n and the aggregation attribute r S ∈ R n , we define AggNetUnit as Δ g:h =(a,b), which represents the aggregation information of the last group within the index range [g,h], where g,h ∈ {1,…,n}. In AggNetUnit Δ g:h =(a,b), the grouping information a ∈ {0,1} indicates whether all items within the range [g - 1,h] belong to the same group, and b ∈ R records the aggregation result of the last group within the range [g,h].
[0116] (4) Aggregation calculation logic: Build a parallel comparison tree (Max / Min) or a summation tree (Sum / Avg) within each group, and recursively select the extreme values; compare the ciphertext values through a secure comparison protocol to avoid plaintext leakage.
[0117] (5) Optimize Sum / Avg: For the Sum operation, calculate the prefix sum array u. P0's input is the permutation π generated based on the grouping result identification array d. The three parties perform an oblivious shuffle to shuffle the secret-shared prefix sum array u according to the permutation π, so that after the shuffle, the last row of each group in the prefix sum array is arranged in order at the beginning of the array. Extract the sum within the group by taking the difference between adjacent prefix sums: For the Avg operation, this optimization also applies.
[0118] (6) Result generation and output: P1 and P2 return the aggregated result share table to the data analyst through a secure channel, and the data analyst reconstructs the table shares.
[0119] 4. Privacy leakage control scheme
[0120] During the execution of the secure connection and grouped aggregation protocol, the auxiliary node P0 may indirectly infer sensitive data characteristics through statistical information (such as the number of matching rows, the number of groups). Therefore, the present invention proposes a dynamic privacy protection mechanism. Through differential privacy noise injection and data obfuscation techniques, the risk of statistical information leakage is reduced to a quantifiable and controllable differential privacy level. The specific implementation steps are as follows:
[0121] (1) Differential privacy noise injection (preprocessing stage)
[0122] Objective: Inject noise into the original key-value column before data sharing to ensure that the data processed subsequently is the noise-added result, meeting the (∈,δ)-differential privacy requirements.
[0123] Implementation steps:
[0124] . Noise generation: The data owner uses the Truncated Laplace Mechanism to generate noise. For example, add noise to the user ID column of table X to generate noise-added data with 1000 virtual IDs, changing the true number of matching rows from 100,000 rows to the noise-added result of 102,000 rows.
[0125] · Expand the encoding domain: To avoid key-value conflicts caused by virtual data, expand the original key-value encoding domain from 64 bits to 128 bits to ensure that the random encoding of virtual key values does not overlap with true key values.
[0126] (2) Dynamically identify the array to hide statistical information
[0127] Objective: Generate an identification array that hides statistical information based on the noise-added data to avoid directly exposing the number of matching rows or the grouping scale.
[0128] Implementation steps:
[0129] · Unique key connection: The auxiliary node P0 generates an array of matching row identifiers consisting of 0 / 1 (1 indicates matching, 0 indicates non - matching), rather than directly returning the number of matching rows. This array is secretly shared with the computing nodes P1 and P2 through secret sharing, and after joint decryption by the two, only the valid matching row data is extracted. For example, in cross - institutional user transaction analysis, P0 generates an identifier array containing 100,000 records (12,000 of which are matching marks), and after decryption by P1 and P2, only the rows marked as 1 are retained, hiding the true number of matches.
[0130] · Grouping aggregation: P0 generates an array of grouping boundary identifiers (marking intra - group continuity and inter - group separation points), and secretly shares the array with P1 and P2 through oblivious shuffling. For example, when medical cases are grouped by region, P0 marks the start and end positions of "Region A", and P1 and P2 complete the aggregation calculation without decrypting the entire array, avoiding revealing the grouping scale.
[0131] 5. Beneficial effects
[0132] The protocol of the present invention shows significant advantages in practical applications, especially in terms of communication efficiency, running time, scalability and adaptability, which have been significantly improved compared with the prior art (such as Secrecy and AHK+). The beneficial effects of the present invention are described in detail from these three dimensions below.
[0133] Substantial improvement in communication efficiency. When processing a table with 2 20 rows of data, the protocol of the present invention is significantly superior to Secrecy and AHK+ in terms of communication overhead. For unique key connection, the protocol of the present invention only requires 843.06 MB, while Secrecy and AHK+ require 184.55 GB and 3356.67 MB respectively. The solution of the present invention has 224 times less communication volume than Secrecy and 4 times less than AHK+. In terms of grouping aggregation, for the grouping Max / Min aggregation operation, AHK+ needs to transmit 3834.67 MB, and Secrecy is even as high as 60259.57 MB, while the protocol of the present invention only requires 2581.59 MB, reducing the communication burden. For the grouping summation operation, AHK+ needs 2014.67 MB, Secrecy needs 60896.05 MB, while the protocol of the present invention only requires 1763.18 MB, ensuring efficient calculation while reducing the communication volume. These results show that the solution of the present invention significantly reduces the communication overhead while ensuring the calculation efficiency, and is applicable to large - scale data processing scenarios.
[0134] Remarkable optimization of running time. When processing 2 20 When dealing with large-scale data, for the unique key join protocol, Secrecy takes 4358.4 seconds and 6311.6 seconds in LAN and WAN respectively, while the protocol of the present invention only takes 11.9 seconds (LAN) and 22.8 seconds (WAN), which is 366 times and 277 times faster than Secrecy in LAN and WAN environments respectively, and is also 9 - 19 times faster than AHK+. For the group-by-Max / Min operation, the execution time of Secrecy in LAN and WAN is 361.6 seconds and 997.4 seconds respectively, while the protocol of the present invention only takes 21.7 seconds (LAN) and 45.0 seconds (WAN), with an acceleration of about 22 times in the WAN environment. For the group-by-Sum operation, Secrecy takes 343.7 seconds (LAN) and 991.0 seconds (WAN), while the protocol of the present invention only takes 10.2 seconds (LAN) and 26.9 seconds (WAN), with an acceleration of about 37 times in the WAN environment. Figure 4 , 5 The comparison results shown in FIGS. 5 and 6 indicate that the solution of the present invention can effectively reduce the execution time and improve the query efficiency in different network environments, and has obvious application value in large-scale data processing.
[0135] Good scalability and adaptability. As the network latency increases, the growth trend of the time cost of the protocol of the present invention is much lower than that of Secrecy and AHK+. Since the protocol of the present invention adopts a low-round design, when the latency increases from 10 ms to 110 ms, the time cost of the unique key join protocol only increases by 15.8 s, while the time cost of AHK+ increases by 142.5 s. In addition, in the grouped Max / Min and grouped Sum operations, the time costs of the protocol of the present invention increase by 78.4 s and 32.9 s respectively, while the corresponding time costs of AHK+ increase by 167.7 s and 84.8 s, further verifying the high efficiency of the solution of the present invention. Therefore, compared with Secrecy and AHK+, the low-round protocol of the present invention shows better stability and adaptability in different network latency environments and can adapt to complex network environments.
[0136] Through the above analysis, it can be observed that although Secrecy and AHK+ perform well in dealing with small-scale data and low latency, their performance drops sharply as the data scale and latency increase, limiting their applicability in dealing with large-scale data and high latency. Compared with Secrecy and AHK+, although there is a certain leakage in the present invention, it is relatively stable and shows better scalability and adaptability under different data scales and network delays, and is more suitable for the needs of practical applications.
[0137] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising..." or "including..." does not exclude the existence of additional elements in the process, method, article or terminal device comprising the said element. In addition, in this text, "greater than", "less than", "exceeding", etc. are understood not to include the present number; "above", "below", "within", etc. are understood to include the present number.
[0138] Although the above-described embodiments have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the above description is only the embodiments of the present invention and does not limit the scope of patent protection of the present invention. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, is similarly included in the scope of patent protection of the present invention.< / y> < / x> < / w> < / g> < / g> < / c> < / c> < / h> < / w> < / g> < / g> < / c> < / c> < / y> < / x>
Claims
1. An efficient and secure database connection and grouping aggregation method based on a server-assisted architecture, characterized in that: The construction of the three-party server auxiliary architecture includes: Computer nodes P1 and P2 are responsible for storing the secret shares provided by the data owner and performing secure computing tasks; Auxiliary node P0: does not hold any data secret share and is responsible for performing local computation-intensive operations; Data owner: generates secret shares from the original data and distributes them to computer nodes P1 and P2, which can then be taken offline; Data analysts: submit query requests and obtain final analysis results without touching the original data or intermediate calculation process; The efficient and secure database connection and grouping aggregation method based on the server-assisted architecture specifically comprises the following steps: S1, data distribution: the data owner shares the data with computer nodes P1 and P2 and then goes offline. Computer nodes P1 and P2 hold their respective data shares. S2, query trigger: the data analyst submits a query request to the system, and the computer nodes P1, P2 and auxiliary node P0 start the computing task; S3, collaborative computing: Computer nodes P1, P2 and auxiliary node P0 collaborate to complete secure computing through random coding and inadvertent shuffling operations; S4. Result return: Computer nodes P1 and P2 send the result shares to the data analyst, who reconstructs the results.
2. The efficient and secure database connection and grouping aggregation method based on server-assisted architecture according to claim 1 is characterized in that: The step S3 collaborative computing includes the following two protocols: A. Secure Connection Protocol: It is used for efficient alignment of cross-table private data, solving the O(n) problem in traditional solutions while maintaining a constant round number. 2 ) Communication overhead problem; B. Group Aggregation Protocol: supports dense aggregation calculation of group key values, reducing the round complexity from O(n) to O(logn).
3. The efficient and secure database connection and grouping aggregation method based on server-assisted architecture according to claim 2 is characterized in that: The secure connection protocol is as follows: A1. Data preparation: In the data distribution phase of the system, the data owner splits the key value columns of the two tables X and Y to be connected into secret sharing shares and transmits them to computing nodes P1 and P2 respectively. Computing nodes P1 and P2 generate a shared key k for subsequent encryption operations. A2, Random code mapping: Based on the shared key k, computing nodes P1 and P2 use the LowMC block cipher algorithm to securely encode and map the key value. All the encoded values are transmitted to the auxiliary node P0. The auxiliary node P0 cannot reversely deduce the original key value through the encoding, ensuring privacy security; A3. Local matching and replacement rule generation: Auxiliary node P0 compares the encoding values e of table X and table Y locally one by one X and e Y , identify all attribute pairs with equal encodings and generate permutation rules π1 and π2. The permutation rules are designed to align matching rows to the same index position, and unmatched rows are randomly arranged at the end of the table; A4. Compound permutation and secure shuffle: To hide the data alignment path, computing nodes P1 and P2 perform an inadvertent shuffle on table X and table Y based on matching permutations π1 and π2, respectively, performing two random permutations and three re-randomization operations each time; A5. Result generation and output: Computing nodes P1 and P2 will shuffle the share table <x>and <y> The data is returned to the data analyst through a secure channel, and the data analyst reconstructs the table share to obtain the connection result table.< / y> < / x> 4. The efficient and secure database connection and grouping aggregation method based on server-assisted architecture according to claim 3 is characterized in that: Step A4: Compound permutation and secure shuffle: To hide the data alignment path, computing nodes P1 and P2 perform an inadvertent shuffle on table X and table Y based on matching permutations π1 and π2, respectively, performing two random permutations and three re-randomization operations each time. Take table X as an example: A41. First random replacement: Participants calculate the shares of nodes P1 and P2 <x K >Re-randomize <h>, and submit the shares of computing node P1 to auxiliary node P0. Auxiliary node P0 generates random permutation σ1 and synchronizes it to computing node P2. Both parties jointly perform σ1 permutation on the shares and re-randomize to generate <c> ;< / c> < / h> A42, composite permutation generation: auxiliary node P0 calculates the composite permutation based on the target permutation π1 Instruct the auxiliary node P0 and the computing node P1 to share <c>Permutate and re-randomize again to obtain <g>, the auxiliary node P0 sends the share to the computing node P1, and the computing nodes P1 and P2 <g>Finally, a round of re-randomization is performed to obtain <w> As a result of the inadvertent shuffling, the whole process is equivalent to realizing the original permutation π through the superposition of two permutations;< / w> < / g> < / g> < / c> A43. Re-randomization injection: After each permutation, computing nodes P1 and P2 add random masks to the shares to ensure that no participant can infer data association through intermediate states.
5. The efficient and secure database connection and grouping aggregation method based on server-assisted architecture according to claim 2 is characterized in that: The packet aggregation protocol includes a packetization phase and an aggregation phase, and the steps of the packetization phase are as follows: B1. Data preparation: Computing nodes P1 and P2 hold the grouping key column t of the data table T obtained during the system distribution phase K The secret share is used to generate a shared key k; B2, Random coding mapping: Same as the connection scheme, the three parties perform random coding functions based on the LowMC algorithm, and the auxiliary node P0 obtains t K The encoding value e K , the auxiliary node P0 according to e K The equality of is used to dynamically generate grouping permutation rules π, so that rows with the same code are arranged adjacently, and ungrouped rows are uniformly arranged at the end of the table; B3. Secure shuffle alignment grouping: To hide the grouping logic, computing nodes P1 and P2 perform two random permutations. Through the same inadvertent shuffle operation as the connection protocol, auxiliary node P0 and computing nodes P1 and P2 jointly complete the π-based rearrangement of table T to ensure that the secret shared data is aligned in the same grouping order; B5. Grouping result verification: The auxiliary node P0 generates a grouping indication array d to mark the starting position of each group. d is secretly shared with computing nodes P1 and P2 for subsequent aggregation operations. In the aggregation stage, for the grouped secret sharing table R, a multi-level aggregation unit is designed based on the Ladner-Fischer parallel prefix network to efficiently calculate the aggregation value. The steps are as follows: B6. Data preprocessing: οSecret sharing table R after grouping by group attribute <r K > Aligned, compute nodes P1 and P2 hold aggregated attributes <r S >'s secret share; If the length of the secret sharing table R is not a power of 2, add dummy data; B7. Parallel prefix network construction: οNetwork structure: The group-aligned table is divided into multiple sub-blocks, each sub-block independently calculates the local aggregation value, and then recursively merges the local results; ο Recursive merging: starting from the bottom sub-block, merge the results layer by layer upwards; οBasic unit definition: given a grouping indicator array d∈{0,1} n and the aggregation attribute r S ∈R n , define AggNetUnit as Δ g:h =(a,b), which represents the aggregation information of the last group in the index range [g,h], where g,h∈{1,…,n}, in AggNetUnitΔ g:h =(a,b), the grouping information a∈{0,1} indicates whether all items in the range [g-1,h] belong to the same group, and b∈R records the aggregation result of the last group in the range [g,h]; B8. Aggregate calculation logic: Construct a parallel comparison tree Max / Min or a summation tree Sum / Avg in each group and recursively select the extreme value; compare the secret values through a secure comparison protocol to avoid plaintext leakage; B9. Optimize Sum / Avg: For the Sum operation, the prefix sum array u is calculated, and the auxiliary node P0 inputs the permutation π generated based on the grouping result identification array d. The three parties perform an inadvertent shuffle to shuffle the secret shared prefix sum array u according to the permutation π, so that after the shuffle, the last row of each group in the prefix sum array is arranged in order at the beginning of the array, and the intra-group sum is extracted by the difference of adjacent prefix sums: For the Avg operation, this optimization also applies; B10. Result generation and output: Computing nodes P1 and P2 return the aggregated result share table to the data analyst through a secure channel, and the data analyst reconstructs the table share.
6. The efficient and secure database connection and grouping aggregation method based on server-assisted architecture according to claim 1 is characterized in that: The step S1 injects noise into the original key value column before data sharing to ensure that the subsequently processed data is a noised result and meets the (∈,δ)-differential privacy requirements, as follows: Noise generation: The data owner generates noise using a truncated Laplace mechanism; Extended encoding domain: To avoid key value conflicts caused by virtual data, the original key value encoding domain is extended from 64 bits to 128 bits to ensure that the random encoding of the virtual key value does not overlap with the real key value.
7. The efficient and secure database connection and grouping aggregation method based on server-assisted architecture according to claim 6 is characterized in that: It also includes generating a flag array based on the noisy data to hide the statistical information, avoiding directly exposing the number of matching rows or group size, as follows: Unique key connection: Auxiliary node P0 generates a matching row identification array consisting of 0 / 1: 1 indicates a match, 0 indicates no match, rather than directly returning the number of matching rows; the array is distributed to computing nodes P1 and P2 through secret sharing, and the two jointly decrypt and extract only valid matching row data; Group aggregation: Auxiliary node P0 generates a group boundary identification array, marking the intra-group continuity and inter-group separation points, and secretly shares the array to computing nodes P1 and P2 through inadvertent shuffling.