Full-secret-state distributed security aggregation calculation scheme in cloth-type scene

By building global indexes and local indexes in a distributed environment and splitting the aggregation computing task, the performance bottlenecks and difficult to balance safety and efficiency in the existing technology when processing large-scale data are solved, and efficient and secure data aggregation computing is achieved.

CN119939649APending Publication Date: 2025-05-06河钢数字技术股份有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411928224.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing fully dense distributed secure aggregation computing solution has performance bottlenecks when processing large-scale data, and it is difficult to balance security and efficiency, especially in distributed environments, where efficient and secure index structure and computing task splitting schemes are lacking.

Method used

A fully dense distributed secure aggregation calculation scheme in a distributed scenario is proposed. By constructing global indexes and local indexes in trusted master node TM and slave node #1, the aggregation calculation tasks are split, and encrypted aggregation calculation is performed in the TEE environment to ensure data privacy and the security of the calculation results.

Benefits of technology

It realizes efficient and secure data aggregation computing in a distributed environment, reduces network latency and bandwidth usage, improves computing efficiency, and ensures data privacy and integrity of computing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939649A_ABST
    Figure CN119939649A_ABST
Patent Text Reader

Abstract

The invention relates to a fully-encrypted distributed security aggregation computing scheme in a distributed scene, which comprises a client, the client is connected with a trusted master node (TM) through data transmission, the trusted master node (TM) is connected with a plurality of slave nodes # 1 through data transmission, the trusted master node (TM) comprises an REE and a TEE, the TEE internally comprises a query engine and a global index, and the query engine is connected with the slave nodes # 1 through data transmission. The slave node # 1 is composed of a security local index, a TEE query engine and a data set, and the TEE query engine is connected with the security local index and the data set through data transmission. According to the overall system provided by the embodiment of the invention, the system can realize safe aggregation calculation of data in a distributed environment, and meanwhile, the privacy and integrity of a calculation request are guaranteed; and support of multiple aggregation operators: a safe data transmission mechanism and an encryption aggregation scheme are studied and provided, sensitive information is ensured not to be leaked in the calculation process of different aggregation operators, the calculation performance is optimized, and sub-nodes needing to participate in calculation are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of fully-encrypted distributed secure aggregate computing, and in particular to a fully-encrypted distributed secure aggregate computing solution in a distributed scenario. Background Art

[0002] With the rapid development of Internet technology and the popularization of big data applications, individuals and institutions are facing increasing challenges in processing and analyzing data containing sensitive information. At the same time, data privacy and security issues have become key issues that need to be addressed. Traditional data processing models often require uploading data to external servers or cloud platforms, which undoubtedly increases the risk of data leakage or cyber attacks.

[0003] To meet this challenge, fully encrypted distributed secure aggregate computing technology has gradually become a key solution. The goal of this technology is to allow clients to effectively query and calculate distributed databases while protecting data privacy. The core idea of ​​fully encrypted is that during the entire data processing and calculation process, the original data is never exposed to any participant or external system. The data will be encrypted to ensure that only authorized clients can access and operate the data.

[0004] Distributed aggregate computing refers to the process of centrally processing and aggregating data stored in multiple nodes in a distributed system, aiming to calculate aggregate results that meet specific requirements (such as sum, average, maximum, minimum, etc.) and ensure that the privacy and security of data are not leaked during the calculation process. Especially in a distributed database environment, aggregate computing needs to consider factors such as data coordination, task scheduling, and load balancing between nodes.

[0005] The core issues involved in this technology include data encryption technology, secure communication protocols, collaborative computing mechanisms, identity authentication and authority management, etc., to ensure that data access and computing are strictly controlled during multi-party collaboration. In addition, while protecting data privacy, effective security measures must be taken to ensure the security of data transmission to prevent data from being stolen or tampered with during transmission and computing.

[0006] Fully encrypted distributed aggregate computing provides a privacy protection and secure computing solution for processing large-scale distributed data. By encrypting the data, participants can perform aggregate computing without decrypting the data, thus protecting sensitive information. This technology not only improves the efficiency of cooperation between multiple participating nodes, but also ensures the privacy of data and the accuracy of computing results, providing effective protection for data security and privacy protection in distributed computing.

[0007] Homomorphic encryption is an effective solution to the above problems. Existing homomorphic encryption technologies such as the Paillier cryptographic system can support data processing and aggregation in an encrypted state, and the calculation results can be obtained without exposing the data itself, thereby ensuring the privacy of the data. Frontier research has proposed a solution based on homomorphic encryption and Intel SGX technology. SGX aggregates multi-party data in a secure and isolated TEE, but each party's data remains private and will not be exposed to other parties or cloud computing service providers. Unfortunately, the above solution is aimed at a centralized cloud computing architecture, in which all data aggregation and computing tasks are concentrated in the cloud for processing. However, with the increase in data volume and the diversification of participants, a single centralized cloud computing architecture will face performance bottlenecks when processing large-scale data. The centralized cloud computing architecture transmits all raw data to the cloud for aggregation computing, which will introduce high network latency and bandwidth usage, especially when the data source is far away from the cloud data center. For example, data on remote devices or edge nodes needs to be uploaded to the cloud for processing for a long time, with high latency and large network traffic.

[0008] In addition to the above problems, existing solutions also have some significant limitations, the most important of which is that it is difficult to achieve an ideal balance between security and efficiency. On the one hand, these solutions often introduce complex encryption and authentication mechanisms to improve security, resulting in a significant increase in computing and storage overhead, making it difficult to meet the efficiency requirements in practical applications. On the other hand, some solutions that simplify security mechanisms in pursuit of efficiency are prone to security vulnerabilities and cannot provide reliable privacy protection and data integrity assurance in multi-node collaborative computing scenarios. In addition, existing methods generally lack a well-designed, efficient and secure index structure, which cannot quickly locate and access the required data, further affecting the overall performance of the system. More importantly, these solutions fail to fully utilize the characteristics of multiple nodes in a distributed environment, do not reasonably split the computing tasks and process them in parallel on multiple nodes, and thus cannot effectively improve the computing efficiency, resulting in the system showing obvious performance bottlenecks when processing large-scale data. In summary, these problems restrict the practical application of existing solutions in complex scenarios, and new technological breakthroughs are urgently needed to overcome these challenges. Therefore, we propose a fully encrypted distributed secure aggregation computing solution in a distributed scenario. Summary of the invention

[0009] The present application provides a fully encrypted distributed secure aggregate computing solution in a distributed scenario to solve the above-mentioned problems.

[0010] This application provides a fully encrypted distributed secure aggregation computing solution in a distributed scenario, including:

[0011] The client is connected to a trusted master node TM through data transmission, the trusted master node TM is connected to multiple slave nodes #1 through data transmission, the trusted master node TM includes REE and TEE, the TEE contains a query engine and a global index, the slave node #1 consists of a secure local index, a TEE query engine and a data set, the TEE query engine is connected to the secure local index and the data set through data transmission, and the query engine and the global index are interconnected through data transmission.

[0012] Preferably, the REE is specifically a processor chip capable of storing and processing data.

[0013] Preferably, the TEE represents a master node.

[0014] A fully encrypted distributed secure aggregation computing solution in a distributed scenario includes the following methods:

[0015] S1, global index construction;

[0016] S2, local index construction;

[0017] S3, aggregation calculation split;

[0018] S4. Aggregate calculation.

[0019] Preferably, the global index is the basis for aggregate calculation and split calculation. The global index is similar to a binary search tree structure and is used to divide calculations related to sensitive attributes. With the help of the global index and range sharding based on sensitive attributes, data can be quickly split and distributed to multiple SNs.

[0020] Preferably, the local index is constructed for each slave node SN i+1 , the sub-dataset stored is D i+1 , the value range of the sensitive attribute column is [S min , S max ], considering that TEE memory is relatively limited, it is necessary to reduce the amount of single data exchange between TEE and REE, so it is considered to convert the sub-dataset D i+1 Divide into several groups, save the min, max, sum, and count in each group as the secondary cache, and store them in the REE environment after encryption. The hit encrypted data group will be sent to the TEE environment to obtain the calculation result.

[0021] Preferably, when the aggregate computing splitting receives the aggregate computing request from the client, the trusted master node TM first needs to split the request and distribute it to multiple SNs. The parsing and splitting operation will be performed in the parsing engine of the TM. The operation takes the entire aggregate computing SQL statement as output and finally outputs a sensitive computing Formally, the parsing engine extracts encryption constraints related to sensitive attribute columns from SQL statements. Will Expressed as That is, find out where the sensitive attribute falls All data record IDs in it;

[0022] After sensitive calculation It will be sent to the TEE of TM, where the calculation is split and a sub-calculation set SQ is obtained. The module is built into the TEE of TM for the generation of sub-calculations. The algorithm uses the global index Ig and the encryption sensitive calculation As input, it takes a sub-computation set SQ as output, which will be distributed to SN for further calculation.

[0023] Preferably, the aggregate calculation comprises the following steps:

[0024] S41, the client will The calculation is sent to the TEE of the TM for parsing and splitting. After obtaining the sub-calculation set SQ, for each sub-calculation less than ID, r>∈SQ, the sub-calculation sq←r is sent to the SN numbered ID for sub-aggregation calculation;

[0025] S42, SN receives the sub-calculation sq, obtains the group number range that needs to be traversed according to the sensitive attribute range stored in itself and the range of sq, and initializes the minimum and maximum value results R at the same time;

[0026] S43, for those nodes that participate in the calculation of the entire node, query the first-level cache to obtain the minimum and maximum values ​​in the node, and maintain the calculation result R at the same time;

[0027] S44, for those groups in which the node partially participates in the calculation, but the entire group in the node participates in the calculation, query the secondary cache to obtain the minimum and maximum values ​​in the group, and maintain the calculation result R at the same time;

[0028] S45. For those groups whose partial data volume participates in the calculation within the node group, the data within the calculation range is sent to the TEE for traversal, and the calculation result R is maintained at the same time.

[0029] The above technical solution provided by the embodiment of the present application has the following advantages compared with the prior art:

[0030] The overall system provided by the embodiment of the present application includes secure index construction and computing request splitting, maintaining local encrypted indexes in the REE environment of each SN, and creating plaintext indexes in the TEE environment of the TM to achieve cross-node data access control. In order to adapt to the distributed computing mode, the computing request will be split into multiple sub-computations according to the characteristics of the task, and assigned to different SNs for processing. Each SN will perform encrypted aggregation calculations based on local data and computing requirements, and finally send the encrypted results back to the TM to ensure that the privacy of the data is protected and support efficient computing operations. In this way, the system can realize secure aggregation calculations of data in a distributed environment while ensuring the privacy and integrity of computing requests.

[0031] Support for multiple aggregation operators: In a distributed environment, aggregation calculations may cross nodes. To ensure that parallel calculations can be effectively performed when data is stored in a decentralized manner, and to ensure the security and consistency of all sub-results during the final aggregation, we need to research and provide secure data transmission mechanisms and encrypted aggregation solutions to ensure that the calculation process of different aggregation operators does not leak sensitive information, while optimizing computing performance and reducing the number of sub-nodes that need to participate in the calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0033] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0034] Figure 1 It is a system flow chart of the present invention. DETAILED DESCRIPTION

[0035] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0036] Various embodiments of the present application may exist in the form of a range. It should be understood that the description in the form of a range is only for convenience and simplicity and should not be understood as a hard limit to the scope of the present application; therefore, it should be considered that the range description has specifically disclosed all possible sub-ranges and single values ​​within the range. For example, it should be considered that the range description from 1 to 6 has specifically disclosed sub-ranges, such as from 1 to 3, from 1 to 4, from 1 to 5, from 2 to 4, from 2 to 6, from 3 to 6, etc., as well as single numbers within the range, such as 1, 2, 3, 4, 5 and 6, which are applicable regardless of the range. In addition, whenever a numerical range is indicated in the present application, it is meant to include any quoted numbers (fractions or integers) within the indicated range. Unless otherwise specified, various raw materials, reagents, instruments and equipment used in the present application, etc., can be purchased from the market or can be prepared by existing equipment.

[0037] In the present application, in the absence of any contrary description, the directional words used, such as "upper" and "lower", are specifically the directions of the drawings in the accompanying drawings. In addition, in the present application, the terms "include", "comprise", etc. refer to "including but not limited to". In the present application, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. In the present application, "and / or" describes the association relationship of the associated objects, indicating that three relationships may exist, for example, A and / or B, which can represent: A exists alone, A and B exist at the same time, and B exists alone. Wherein A, B can be singular or plural. In the present application, "at least one" refers to one or more, and "plural" refers to two or more. "At least one", "at least one of the following" or similar expressions refer to any combination of these items, including any combination of singular items or plural items. For example, "at least one of a, b, or c" or "at least one of a, b and c" can both mean: a, b, c, ab, i.e. a and b, ac, bc or abc, where a, b, c can be single or plural, respectively.

[0038] like Figure 1 As shown: This embodiment of the application provides a fully encrypted distributed secure aggregation computing solution in a distributed scenario, including:

[0039] The client is connected to a trusted master node TM through data transmission, the trusted master node TM is connected to multiple slave nodes #1 through data transmission, the trusted master node TM includes REE and TEE, the TEE contains a query engine and a global index, the slave node #1 consists of a secure local index, a TEE query engine and a data set, the TEE query engine is connected to the secure local index and the data set through data transmission, and the query engine and the global index are interconnected through data transmission.

[0040] The REE is specifically a processor chip capable of storing and processing data.

[0041] The TEE represents a master node.

[0042] A fully encrypted distributed secure aggregation computing solution in a distributed scenario includes the following methods:

[0043] S1, global index construction;

[0044] S2, local index construction;

[0045] S3, aggregation calculation split;

[0046] S4. Aggregate calculation.

[0047] The global index is the basis for aggregate calculation and split calculation. The global index is similar to a binary search tree structure and is used to divide calculations related to sensitive attributes. With the help of the global index and range sharding based on sensitive attributes, data can be quickly split and distributed to multiple SNs.

[0048] The local index is constructed for each slave node SN i+1 , the sub-dataset stored is D i+1 , the value range of the sensitive attribute column is [S min , S max ], considering that TEE memory is relatively limited, it is necessary to reduce the amount of single data exchange between TEE and REE, so it is considered to convert the sub-dataset D i+1 Divide into several groups, save the min, max, sum, and count in each group as the secondary cache, and store them in the REE environment after encryption. The hit encrypted data group will be sent to the TEE environment to obtain the calculation result.

[0049] When the trusted master node TM receives an aggregate computing request from a client, it first needs to split the request and distribute it to multiple SNs. The parsing and splitting operation will be performed in the parsing engine of the TM. This operation uses the entire aggregate computing SQL statement as output and finally outputs a sensitive computing Formally, the parsing engine extracts encryption constraints related to sensitive attribute columns from SQL statements. Will Expressed as That is, find out where the sensitive attribute falls All data record IDs in it;

[0050] After sensitive calculation It will be sent to the TEE of TM, where the calculation is split and a sub-calculation set SQ is obtained. The module is built into the TEE of TM for the generation of sub-calculations. The algorithm uses the global index Ig and the encryption sensitive calculation As input, it takes a sub-computation set SQ as output, which will be distributed to SN for further calculation.

[0051] The aggregation calculation includes the following steps:

[0052] S41, the client will The calculation is sent to the TEE of the TM for parsing and splitting. After obtaining the sub-calculation set SQ, for each sub-calculation less than ID, r>∈SQ, the sub-calculation sq←r is sent to the SN numbered ID for sub-aggregation calculation;

[0053] S42, SN receives the sub-calculation sq, obtains the group number range that needs to be traversed according to the sensitive attribute range stored in itself and the range of sq, and initializes the minimum and maximum value results R at the same time;

[0054] S43, for those nodes that participate in the calculation of the entire node, query the first-level cache to obtain the minimum and maximum values ​​in the node, and maintain the calculation result R at the same time;

[0055] S44, for those groups in which the node partially participates in the calculation, but the entire group in the node participates in the calculation, query the secondary cache to obtain the minimum and maximum values ​​in the group, and maintain the calculation result R at the same time;

[0056] S45. For those groups whose partial data volume participates in the calculation within the node group, the data within the calculation range is sent to the TEE for traversal, and the calculation result R is maintained at the same time.

[0057] When the maximum amount of data stored in each node is n / t, in the worst case, 2n / t data need to be traversed, which has high computing efficiency.

[0058] Specifically: the overall system can be divided into three stages: global index construction, local security index construction, and aggregate computing. The global index splits the data set into multiple sub-datasets and distributes them to the slave nodes. At the same time, the global index is maintained in the TEE of the master node. The local security index maintains the sharded data set information in the node, encrypts and virtualizes it and saves it in the REE environment. At the same time, the metadata information in each group is saved to speed up the calculation. Aggregate computing will split it into multiple sub-computing requests based on the global index in the TEE of the master node and distribute them to the corresponding slave nodes at the same time. The slave nodes then obtain local calculation results according to the local security index or metadata information, and finally send them back to the master node through the SSL channel. The master node further aggregates to obtain the final calculation results.

[0059] The above is only a specific implementation of the present application, so that those skilled in the art can understand or implement the present application. It will be apparent to those skilled in the art that various modifications to these embodiments are possible, and the general principles defined in the present application can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown in the present application, but will conform to the widest scope consistent with the principles and novel features applied for by the present application.

Claims

1. A fully encrypted distributed secure aggregation computing solution in a distributed scenario, characterized in that: include: The client is connected to a trusted master node TM through data transmission, the trusted master node TM is connected to multiple slave nodes #1 through data transmission, the trusted master node TM includes REE and TEE, the TEE contains a query engine and a global index, the slave node #1 consists of a secure local index, a TEE query engine and a data set, the TEE query engine is connected to the secure local index and the data set through data transmission, and the query engine and the global index are interconnected through data transmission.

2. According to claim 1, the fully encrypted distributed secure aggregation computing solution in a distributed scenario is characterized by: The REE is specifically a processor chip capable of storing and processing data.

3. According to claim 1, the fully encrypted distributed secure aggregation computing solution in a distributed scenario is characterized by: The TEE represents a master node.

4. A fully encrypted distributed secure aggregation computing solution in a distributed scenario, characterized in that: The following methods are included: S1, global index construction; S2, local index construction; S3, aggregation calculation split; S4. Aggregate calculation.

5. According to claim 3, the fully encrypted distributed secure aggregate computing solution in a distributed scenario is characterized by: The global index is the basis for aggregate calculation and split calculation. The global index is similar to a binary search tree structure and is used to divide calculations related to sensitive attributes. With the help of the global index and range sharding based on sensitive attributes, data can be quickly split and distributed to multiple SNs.

6. The fully secret distributed secure aggregate computing solution in a distributed scenario according to claim 3, characterized in that: The local index is constructed for each slave node SN i+1 , the sub-dataset stored is D i+1 , the value range of the sensitive attribute column is [S min , S max ], considering that TEE memory is relatively limited, it is necessary to reduce the amount of single data exchange between TEE and REE, so it is considered to convert the sub-dataset D i+1 Divide into several groups, save the min, max, sum, and count in each group as the secondary cache, and store them in the REE environment after encryption. The hit encrypted data group will be sent to the TEE environment to obtain the calculation result.

7. The fully secret distributed secure aggregate computing solution in a distributed scenario according to claim 3, characterized in that: When the trusted master node TM receives an aggregate computing request from a client, it first needs to split the request and distribute it to multiple SNs. The parsing and splitting operation will be performed in the parsing engine of the TM. This operation uses the entire aggregate computing SQL statement as output and finally outputs a sensitive computing Formally, the parsing engine extracts encryption constraints related to sensitive attribute columns from SQL statements. Will Expressed as That is, find out where the sensitive attribute falls All data record IDs in it; After sensitive calculation It will be sent to the TEE of TM, where the calculation is split and a sub-calculation set SQ is obtained. The module is built into the TEE of TM for the generation of sub-calculations. The algorithm uses the global index I g Encrypting sensitive computations As input, it takes a sub-computation set SQ as output, which will be distributed to SN for further calculation.

8. The fully secret distributed secure aggregate computing solution in a distributed scenario according to claim 7, characterized in that: The aggregation calculation includes the following steps: S41, the client will The calculation is sent to the TEE of the TM for parsing and splitting. After obtaining the sub-calculation set SQ, for each sub-calculation less than ID, r>∈SQ, the sub-calculation sq←r is sent to the SN numbered ID for sub-aggregation calculation; S42, SN receives the sub-calculation sq, obtains the group number range that needs to be traversed according to the sensitive attribute range stored in itself and the range of sq, and initializes the minimum and maximum value results R at the same time; S43, for those nodes that participate in the calculation of the entire node, query the first-level cache to obtain the minimum and maximum values ​​in the node, and maintain the calculation result R at the same time; S44, for those groups in which the node partially participates in the calculation, but the entire group in the node participates in the calculation, query the secondary cache to obtain the minimum and maximum values ​​in the group, and maintain the calculation result R at the same time; S45. For those groups whose partial data volume participates in the calculation within the node group, the data within the calculation range is sent to the TEE for traversal, and the calculation result R is maintained at the same time.

Citation Information

Cited By

  • Protein data encryption storage and operation method based on TEE server cluster

    CN121151140A