Safe verifiable multi-source data system supporting SQL (Structured Query Language) function

By designing a secure verified multi-source data system that supports SQL functions, and using new protocols and key generation algorithms to verify SQL query results, it solves the problem of unreliable aggregated query results in existing systems, realizes data integrity and accuracy, and improves the security and functionality of the database system.

CN120337273APending Publication Date: 2025-07-18ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510324827.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing joint security analysis system has insufficient verification mechanism for query results, especially in the query process involving aggregation function, users cannot effectively verify the accuracy of the calculation results, which affects the completeness of the data and the economic interests and reputation of the enterprise.

Method used

A secure and verified multi-source data system that supports SQL functions is designed. Through new protocols and key generation algorithms, the accuracy and completeness of the aggregate query results are ensured, including the key generation algorithm KeyGen, the encryption algorithm Enc and the decryption algorithm Dec, which supports the verification of aggregate functions such as sum(), avg(), count() and multiply().

Benefits of technology

It realizes reliability verification of SQL query results, ensures data integrity and accuracy, makes up for the lack of element cumulative operations in existing systems, improves the reliability of aggregation queries, and promotes the development of database technology in data security and privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337273A_ABST
    Figure CN120337273A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of cryptography, and discloses a secure verifiable multi-source data system supporting an SQL (Structured Query Language) function, called an SQL-svmDB (Secure Query Language-svmDB) system for short, entities related to the system comprise a plurality of database users, corresponding agents and a plurality of data owners; the system supports multi-table connection and complex aggregation operation. The invention provides a novel protocol, the protocol not only can verify the results of sum, avg and count functions in the SQL query in the aggregation query to ensure the integrity and accuracy of data, but also expands the protocol to support the multiplication operation, fills the blank that the existing SQL aggregation function in the security conjoint analysis mechanism lacks the multiplication operation on elements, and improves the security conjoint analysis mechanism. According to the method, the reliability of aggregation query can be improved, wider functions are introduced for a database system, and the development of a database technology in the aspects of data security and privacy can be promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of cryptography, and in particular relates to a secure and verifiable multi-source data system supporting SQL functions. Background Art

[0002] At present, the increasingly stringent global data compliance supervision has promoted the awakening of privacy protection awareness among data rights holders and data processing organizers, but at the same time, the data of each subsidiary and department has formed "islands" that are independent of each other and cannot be interconnected, limiting the effective circulation and collaboration of data. The security protection of data is not only to maintain privacy and corporate secrets, but also because once the data is improperly leaked or abused, its intrinsic value will be seriously depreciated, which will have a negative impact on the accuracy and fairness of data pricing. As the main data storage method, databases inevitably need to perform joint queries in order to achieve data integration and analysis across different databases or distributed systems, that is, query operations are performed simultaneously in multiple data sources to obtain joint results across multiple data sources. Therefore, while ensuring the privacy and security of multi-source data, balancing the contradiction between joint query needs and data circulation restrictions is currently a topic that needs to be solved urgently.

[0003] To address the above issues, multiple secure joint analysis techniques have emerged in academia and industry for realizing the joint query of multi-source data, such as Secrecy, Senate, Conclave, SMCQL, Scape, etc. The working entities of these systems include multiple database users U, one or more agents P, and multiple data owners O. When a database user U executes a query, it sends its query instruction to the agent P. The agent P generates a corresponding query plan according to the instruction and interacts with the corresponding data owner O. At the same time, it uses the secure multi-party computing technology as the backend analysis and computing engine to complete the secure multi-party computing of multi-party data. Finally, the agent P merges the query operation results and returns them to the user U. So far, some secure joint analysis systems have been open-sourced and shown certain practical value in practical applications. However, through research, it is found that they are too dependent on the intermediate agent and generally lack an effective verification mechanism for query results. Especially in the query process involving aggregation functions, users often cannot effectively verify the accuracy of the calculation results. This not only concerns the accuracy and integrity of data but also directly affects the economic interests and reputation of enterprises and organizations. Incorrect or unreliable data may lead enterprises to make decisions that are not conducive to their economic interests. For example, banks and financial institutions rely on accurate data to assess credit risks. Incorrect data may lead to lending to customers with poor credit, thus increasing the risk of bad debts. In other industries, such as retail and supply chain management, inaccurate data may lead to inventory backlogs or shortages, thus affecting sales and customer satisfaction. Therefore, the lack of such a verification mechanism will seriously affect users' trust in the system and limit its application in sensitive fields. Summary of the Invention

[0004] The object of the present invention is to provide a secure and verifiable multi-source data system that supports SQL functions to solve the above technical problems.

[0005] To solve the above technical problems, the specific technical solution of a secure and verifiable multi-source data system that supports SQL functions according to the present invention is as follows:

[0006] A secure and verifiable multi-source data system that supports SQL functions, abbreviated as the SQL-svmDB system. The entities involved in the system include multiple database users and their corresponding agents, as well as multiple data owners; the system supports multi-table joins and complex aggregation operations. The general format of the aggregation function query statement is as follows:

[0007] SELECT tablel.iteml, table2.item2,...tableN.itemN,

[0008] func(tablex.itemx,...:tablev.itemY)

[0009] FROM possessorA tablel, possessorB table2,....possessorN tableN JoINtable2ON tablel.xxx = table2.xxx,...WHERE..

[0010] Among them, table1, table2,..., tableN represent the names of the tables to be queried, item1, item2,..., itemN represent the names of the columns to be queried, and func() is an aggregation function, including sum(), avg(), count(), and the multiplication aggregation function multiply() customized by the SQL-svmDB system.

[0011] Furthermore, the sum() aggregation function and the multiply() function include addition function, multiplication function, statistical function, and average function.

[0012] Furthermore, the system involves multiple data owners S i , S i , S i ....S n , when performing a union query, it is necessary to calculate the data held by each data owner: for data owner S j , the data set to be analyzed is a specific column in the table, denoted as column vector (x 11 , x 12 ,..., x 1j )(j ∈ N T ); for S2, it is (x * , x 21 ,..., x 22 )(j ∈ N 2j ), and so on until the data vector from S T is (x n , x n1 ,..., x n2 )(j ∈ N nj ); for each x T (i ∈ {1, 2, 3,..., n}, j ∈ N ij ), user U and data owner S * respectively negotiate a set of key pairs; assume that user U and S1 jointly negotiate a key pair (p n , p i1 , p i2 ), user U and S2 jointly negotiate a key pair (q i1 , q i2 ), and S1 and S2 respectively send α i and βi Sent to the agent, where α i = x 1j p 11 + p 12 and β i = x 2j q 21 + q 22 The agent calculates where represents an addition or multiplication operation and sends it to the user.

[0013] Furthermore, in the addition function, data owners A and B respectively agree with user U on a set of secret values, which are (p1, p2) and (q1, q2) respectively, where |p1| > |q1|, and |p1| and |q1| are the bit lengths of p1 and q1 respectively; subsequently, A sends (ap1 + p2) to the agent, while B sends (bq1 + q2) to the agent. The agent calculates (ap1 + p2) + (bq1 + q2) and forwards it to user U; after the user receives the query result from the agent, subtract (p2 + q2) from it, 1. Divide by p1 to get the value of a; 2. Take the modulus of p1 and then divide by q1 to get the value of b, and add them together to get a + b; for the case involving multiple parties, each data owner can negotiate a set of secret values with the user, and the core principle is the same as that of the two-party case.

[0014] Furthermore, in the multiplication operation, data owners A and B respectively agree with user U on a set of secret values, which are (p1, p2) and (q1, q2) respectively, where |p1| > |q1| > |q2| > |p2|, and |p1|, |q1|, |q2|, and |p2| are the bit lengths of p1, q1, q2, and p2 respectively. A sends (ap1 + p2) to the agent, and B sends (bq1 + q2) to the agent. The agent calculates (ap1 + p2)(bq1 + q2), that is, abp1q1 + ap1q2 + bp2q1 + p2q2, and sends it to the user;

[0015] After the user receives it, directly divide by p1q1 to get the value of ab; during this process, if the agent tampers with the intermediate result, the user can first subtract the constant term p2q2 from the expression (abp1q1 + ap1q2 + bp2q1 + p2q2) received from the agent, and perform coefficient modulo operations on each term one by one starting from the first term. If the result is not zero, malicious behavior can be verified, and the same applies when there are multiple participating parties. Furthermore, the system includes a new protocol, denoted by "∏", which includes a key generation algorithm KeyGen, an encryption algorithm Enc, and a decryption algorithm Dec. The specific definition of the scheme ∏ = (KeyGen, Enc, Dec) is as follows:

[0016] (1) KeyGen(1 λ ) → (p i1 , p i2 ): This algorithm is a negotiation process between user U and data owner S i (i ∈ {1, 2,..., n}), generating a pair of keys (p i1 , p i2 ). The KeyGen() algorithm is jointly executed by user U and each data owner S i . By hashing the combination of two different pseudo-random number seeds (P Seed1 and P Seed1 ) and the i value, and applying the next_prime() function to the result to find the next prime number to obtain the values of p i1 and p i2 . Denote the above process as: p i1 = next prime (H(P Seed1 , i)), p i2 = next prime (H(P Seed2 , i));

[0017] (2) Enc(x i , state) → c state : This algorithm is executed by data owner S i . Its input includes the plaintext data x i and the state information state of S i . The plaintext data x i represents the data to be calculated queried locally by S i , while the state information state encapsulates a series of important information about S j , including the key pair (p i1 , p i2 );

[0018] (3) Dec(c, OP, state) → m / ⊥: This algorithm is executed by user U. Its input parameter c represents the data received from the proxy, that is, the result after the "OP" operation on c state . And OP is operation. In the case of incorrect decryption, the error flag "⊥" is returned. Whether the decryption is successful depends on multiple factors, including the correct operation OP, valid state information state, and the integrity of data c. m is the final calculation result, corresponding to γ j .

[0019] Further, it includes a query process, and the query process includes the following steps:

[0020] (1) Data owner S i Configures query permissions for its specific data. These permissions can be set to public query or restricted to specific user access. When a user obtains the data query permission, they can initiate a query;

[0021] (2) The user parses the original SQL query statement through the vproxy module of the client to generate a unilateral SQL statement adapted for execution by the data owner; the user then performs broadcast encryption on the parsed query fields and digitally signs each query statement one by one;

[0022] (3) The user sends the encrypted and signed SQL statements together with the original SQL statement to the intermediate proxy; the intermediate proxy is responsible for forwarding this information to the corresponding data owner;

[0023] (4) After receiving the SQL request, the data owner first verifies the signature. After successful verification, it decrypts the encrypted query fields to generate a complete query instruction; subsequently, the data owner executes the query operation;

[0024] (5) For data that does not require calculation during the query process, the data owner encrypts it with the public key and attaches a digital signature, while for data that requires calculation, it is encrypted using the ∏ protocol and sent to the proxy;

[0025] (6) The intermediate proxy is responsible for determining which data requires further calculation and performing the necessary arithmetic operations according to the protocol; for multi-table join operations, the proxy calls the PSI algorithm to implement;

[0026] (7) The intermediate proxy merges the encrypted query results from multiple parties;

[0027] (8) Finally, the intermediate proxy returns the merged result to the user;

[0028] After receiving the result, the user verifies the signature, then decrypts and verifies the query result.

[0029] A secure and verifiable multi-source data system supporting SQL functions of the present invention has the following advantages: The present invention proposes a new protocol that can not only verify the results of the "sum", "avg", "count" functions in SQL queries during aggregation queries, ensuring the integrity and accuracy of the data, but also extends the protocol to support multiplication operations, filling the gap in the lack of element multiplication operations in existing SQL aggregation functions in the secure joint analysis mechanism. It can not only improve the reliability of aggregation queries, but also introduce more extensive functions to the database system, and can also promote the development of database technology in terms of data security and privacy. Description of the Drawings

[0030] Figure 1Architecture diagram of the secure verifiable multi-source data system supporting SQL functions of the present invention;

[0031] Figure 2 Schematic diagram for the scheme description of the present invention;

[0032] Figure 3 System query flowchart of the present invention. Detailed implementation manners

[0033] To better understand the purpose, structure and function of the present invention, the following further describes in detail a secure verifiable multi-source data system supporting SQL functions of the present invention with reference to the accompanying drawings.

[0034] In this system, considering that the query mechanism based on the three-party computing model essentially outsources data, when the data owner outsources data to a third party, it is inevitable to consider the potential risk of malicious behavior. If two computing parties collude with each other, they may obtain the information of all data owners. According to current laws and regulations, data cannot cross domains, or can cross domains only after being processed. At the same time, from the perspective of industrial manufacturers, due to the possible difficulties in communication and adaptation among institutions, the role of the intermediate agent is particularly important. Therefore, the secure verifiable multi-source data system supporting SQL functions is further optimized on the basis of the intermediate agent query architecture, and a secure verifiable multi-source data system supporting SQL functions (SQL-secure verifiable multi-source DB), abbreviated as SQL-svmDB system, is designed. The specific system architecture is as Figure 1 shown as:

[0035] The entities involved in the SQL-svmDB system include multiple database users and their corresponding agents, as well as multiple data owners. Under the security assumption, we consider that the database users are honest and trustworthy, while the agents are regarded as malicious and the data owners are regarded as semi-honest.

[0036] First, the data owner authorizes their data and clarifies the query permissions of users. The data owner can set query permissions for their specific data, which can be that all users can query, or only specified users can query. Next, the user parses the original SQL statement through the vproxy on the client side. After parsing, a unilateral SQL statement that the corresponding data owner can directly use to execute is obtained. The user broadcasts and encrypts each field in these query statements and signs them one by one. Subsequently, the user sends these processed SQL statements and the original SQL statement together to the intermediate proxy, which then forwards them to the corresponding data owner. After receiving the SQL statement, the data owner performs signature verification. After successful verification, the data owner decrypts the query fields to obtain the complete query instruction and executes this instruction. After the data owner executes the query, the data that does not need to be calculated is encrypted with a public key and digitally signed, and the data that needs to be calculated is sent to the proxy after specific processing. The proxy determines which data needs to be calculated, performs the necessary calculations, and combines the query ciphertext results from multiple parties and returns them to the user. Finally, the user decrypts and verifies the query results after verifying the signature. When it comes to aggregation calculations such as "sum", "avg", "count", and "multiply", the user needs to perform specific calculations to obtain the query results to complete the entire query process.

[0037] The query structure of the SQL-svmDB system inherits the standard SQL syntax format with slight modifications and supports multi-table joins and complex aggregation operations. Aggregation functions are particularly important in secure multi-party joint analysis because they are often used to process sensitive data. The general format of the query statement is as follows (i.e., the original SQL statement input by the user):

[0038] SELECT tablel.iteml, table2.item2,...tableN.itemN,

[0039] func(tablex.itemx,...:tablev.itemY)

[0040] FROM possessorA tablel, possessorB table2,....possessorN tableN JoINtable2ON tablel.xxx=table2.xxx,...WHERE..

[0041] Among them, table1, table2, …, tableN represent the names of the tables to be queried, item1, item2, …, itemN represent the names of the columns to be queried, and func() is an aggregation function, including sum(), avg(), count(), and the multiplication aggregation function multiply() customized by the SQL-svmDB system, etc. Since the sum() function can support the functions of the avg() and count() functions, this invention focuses on introducing the scheme construction for the sum() aggregation function and the multiply() function.

[0042] ① Additive function

[0043] Taking two parties as an example, assume that there are existing data owners A and B, who hold table A and table B respectively. Table A stores the basic salary information of employees in a certain department, while table B stores the bonus information of employees in the same department. Now, a management user U wants to query the total income of all employees to evaluate the human resource cost of this department, that is, the sum of the basic salary and the bonus. Then the original SQL statement input by the user is as follows:

[0044] SELECT SUM(A.salary B.bonus)

[0045] FROM possessorA tableA A, possessorB tableB B

[0046] JOIN tableB B ON A.id = B.id;

[0047] ② Multiplication function

[0048] Still taking two parties as an example, there are two existing data owners A and B, who hold table A and table B respectively. A and B are two independent medical research institutions respectively. Table A stores the data information on the dosage of a specific drug (such as the number of times of use per month), while table B stores the data information on the side effects of the same drug (for example, the number of side effect events per thousand uses). A user U from a medical research institute wants to query the product of the dosage and the side effect incidence rate of the drug named "amoxicillin" for the risk assessment of this drug. Then the original SQL statement input by this user is as follows:

[0049] SELECT multiply(A.dosage, B.sideeffectincidence)

[0050] FROM possessorA tableA A, possessorB tableB B

[0051] JOIN tableB B ON A.drugname=B.drugname

[0052] WHERE tableA.drugname = 'Amoxicillin';

[0053] ③Statistical function

[0054] Here we take two parties as an example. Assume that the existing data owners A and B hold Table A and Table B respectively. Table A stores the basic salary information of employees in a department, while Table B stores the bonus information of employees in the department. Now a management user U wants to query the number of employees whose annual salary is less than 200,000 and whose overtime pay (a field in the bonus table) is greater than 19,000. The original SQL statement entered by the user is as follows:

[0055] SELECT count(A.employeeid,B.employeeid)

[0056] FROM possessorA tableA A, possessorB tableB B

[0057] JOIN tableB B ON A.employeeid=B.employeeid

[0058] where A.basicsalary<200000and B.bonus>19000and B.bonustype="overtime";

[0059] ④Average function

[0060] Here we take two parties as an example. Assume that the existing data owners A and B hold Table A and Table B respectively. Table A stores the basic salary information of employees in a certain department, while Table B stores the bonus information of employees in the department. Now a management user U wants to query the average value of the basic annual salary (the sum of the basic salary and the year-end bonus) of the employees. The original SQL statement entered by the user is as follows:

[0061] SELECT avg(A.basicsalary,B.bonus)

[0062] FROM possessorA tableA A, possessorB tableB B

[0063] JOIN tableB B 0N A.employeeid=B.employeeid

[0064] where B.bonustype = "Year - end bonus";

[0065] Scheme description and definition:

[0066] The SQL - svmDB scheme research involves multiple data owners S i , S i , S i ....S n , when performing a union query, it is necessary to calculate the data held by each data owner. Specifically, for data owner S j , the data set to be analyzed is a specific column in the table, expressed as a column vector (x 11 , x 12 ,..., x 1j )(j ∈ N T ); for S2, it is (x * , x 21 ,..., x 22 ,..., x 2j ), and so on, until the data vector from S T is (x n , x n1 ,..., x n2 ,..., x nj ). T .

[0067] For each x ij (i ∈ {1, 2, 3,..., n}, j ∈ N * ). The user U and the data owner S n respectively agree on a set of key pairs. Specifically, as Figure 2 shown, taking two parties as an example, the user U and S1 jointly negotiate a key pair (p i1 , p i2 ), the user U and S2 jointly agree on a key pair (q i1 , q i2 ), and S1 and S2 respectively send α i and β i to the proxy (where α i = x 1j p 11 + p 12 , β i = x 2j q 21 + q 22 ), and the proxy calculates (where represents an addition or multiplication operation), and sends it to the user.

[0068] Specifically, in the described addition function, data owners A and B respectively agree with user U on a set of secret values, namely (p1, p2) and (q1, q2), where |p1| > |q1| (|p1| and |q1| are the bit lengths of p1 and q1 respectively). Subsequently, to ensure the confidentiality of the query result, A sends (ap1 + p2) to the proxy, while B sends (bq1 + q2) to the proxy. The proxy calculates (ap1 + p2) + (bq1 + q2) and forwards it to user U.

[0069] After the user receives the query result from the proxy, subtract (p2 + q2) from it, 1. Divide by p1 to obtain the value of a; 2. Take the modulus of p1, then divide by q1 to obtain the value of b, and add them together to get a + b.

[0070] During this process, the proxy cannot access the real data. If the proxy has malicious intentions and attempts to tamper with the intermediate result, any tampering behavior can be regarded as adding a number c to the result of the expression (ap1 + p2) + (bq1 + q2). Then the user can subtract (p2 + q2) from the received result, and then take the modulus of p1 and q1 in turn. If the obtained result is not zero, it indicates that there is tampering, and the malicious behavior of the proxy is detected. This protocol effectively implements the encryption and authentication mechanisms only through the transmission of a single piece of data, ensuring data integrity and source authenticity. Even when the intermediate proxy is malicious, the reliability of the result can be verified.

[0071] For the case involving multiple parties, each data owner can negotiate a set of secret values with the user. The core principle is the same as that of the two - party case, so it will not be elaborated in detail here.

[0072] Similarly, in the multiplication operation, data owners A and B respectively agree with user U on a set of secret values, namely (p1, p2) and (q1, q2), where |p1| > |q1| > |q2| > |p2| (|q2|, |p2|, |q1|, |p1| are the bit lengths of q2, p2, q1, p1 respectively). A sends (ap1 + p2) to the proxy, and B sends (bq1 + q2) to the proxy. The proxy calculates (ap1 + p2)(bq1 + q2), that is, abp1q1 + ap1q2 + bp2q1 + p2q2, and sends it to the user.

[0073] After the user receives it, directly divide by p1q1 to obtain the value of ab. During this process, if the proxy tampers with the intermediate result, the user can first subtract the constant term p2q2 from the expression (abp1q1 + ap1q2 + bp2q1 + p2q2) received from the proxy, and perform coefficient modulus operations on each term one by one starting from the first term. If the result is not zero, the malicious behavior can be verified. The same applies to the case involving multiple participants.

[0074] To address the data verification flaws in current multi-source data security joint analysis, the present invention proposes a new protocol, denoted as "∏", which generally includes three parts: a key generation algorithm (KeyGen), an encryption algorithm (Enc), and a decryption algorithm (Dec). The specific definition of the scheme ∏=(KeyGen, Enc, Dec) is as follows:

[0075] (1) KeyGen(1 λ )→(p i1 , p i2 ): This algorithm is a negotiation process between user U and data owner S i (i∈{1, 2,..., n}), generating a set of key pairs (p i1 , p i2 ). The KeyGen() algorithm is jointly executed by user U and each data owner S i respectively. By performing a hash operation on the combination of two different pseudo-random number seeds (P Seed1 and P Seed1 ) and the i value, and applying the next_prime() function to the result to find the next prime number, the values of p i1 and p i2 are obtained. Denote the above process as: p i1 = next prime (H(P Seed1 , i)), p i2 = next prime (H(P Seed2 , i)).

[0076] (2) Enc(x i , state)→c state . This algorithm is executed by data owner S i , and its inputs include the plaintext data x i and the state information state of S i . The plaintext data x i represents the data to be calculated retrieved by S i locally, while the state information state encapsulates a series of important information about S j , including the key pair (p i1 , p i2 ).

[0077] (3) Dec(c, OP, state)→m / ⊥. This algorithm is executed by user U, and its input parameter c represents the data received from the proxy, that is, the result of c state after the "OP" operation, and OP is Operation. If the decryption fails, an error flag "⊥" is returned. The success of decryption depends on multiple factors, including the correct operation OP, the valid state information state, and the integrity of data c. m is the final calculation result, corresponding to γ mentioned above. j . The Dec() algorithm ensures that the system can not only securely store and transmit data, but also accurately and efficiently recover and calculate data.

[0078] During the design process of the entire solution, the query process is one of the core aspects, and its specific steps are crucial for the security and efficiency of the system. The following will elaborate on the specific query process in the SQL-svmDB solution, focusing on analyzing the operation steps at each stage, the execution details of the protocol, and the specific applications of various key technologies in the process, to help better understand the working mechanism and advantages of the entire solution. The specific query process is as Figure 3 shown:

[0079] (1) The data owner S j configures query permissions for its specific data, which can be set to public queries or restricted to specific user access. When a user obtains the data query permission, they can initiate a query.

[0080] (2) The user parses the original SQL query statement through the vproxy module of the client to generate a unilateral SQL statement adapted for the data owner to execute. The user then performs broadcast encryption on the parsed query fields and digitally signs each query statement one by one.

[0081] (3) The user sends the encrypted and signed SQL statements along with the original SQL statement to the intermediate proxy. The intermediate proxy is responsible for forwarding this information to the corresponding data owner.

[0082] (4) After receiving the SQL request, the data owner first verifies the signature. After successful verification, it decrypts the encrypted query fields to generate a complete query instruction. Subsequently, the data owner executes the query operation.

[0083] (5) For data that does not require calculation during the query process, the data owner encrypts it using the public key and attaches a digital signature, while for data that requires calculation, it is encrypted using the ∏ protocol and sent to the proxy.

[0084] (6) The intermediate proxy is responsible for determining which data requires further calculation and performing the necessary arithmetic operations according to the protocol (for multi-table join operations, the proxy calls the PSI algorithm to implement).

[0085] (7) The intermediate proxy merges the encrypted query results from multiple parties.

[0086] (8)Finally, the intermediate agent returns the merged result to the user.

[0087] After receiving the result, the user verifies the signature and then decrypts and verifies the query result.

[0088] It can be understood that the present invention is described by some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. In addition, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.

Claims

1. A secure and verifiable multi-source data system that supports SQL functions, abbreviated as the SQL-svmDB system, is characterized in that The entities involved in the system include multiple database users and their corresponding agents, as well as multiple data owners; the system supports multi-table joins and complex aggregation operations, and the general format of the aggregation function query statement is as follows: SELECT tablel.iteml,table2.item2,...tableN.itemN, func(tablex.itemx,...:tablev.itemY) FROM possessorA tablel,possessorB table2,....possessorN tableN JoIN table2 ON tablel.xxx=table2.xxx,...WHERE.. Among them, table1, table2, …, tableN represent the table names to be queried, item1, item2, …, itemN represent the column names to be queried, func() is an aggregation function, including sum(), avg(), count(), and the multiplication aggregation function multiply() defined by the SQL-svmDB system.

2. The secure and verifiable multi-source data system supporting SQL functions according to claim 1, wherein The sum() aggregation function and the multiply() function include addition, multiplication, statistical, and average functions.

3. The secure verifiable multi-source data system supporting SQL functions according to claim 1, characterized in that, The system involves multiple data owners \(S_1\) i , \(S_2\) i , \(S_3\) i .... \(S_n\) n . When performing a joint query, it is necessary to calculate the data held by each data owner: For data owner \(S_1\) i , the data set to be analyzed is a specific column in the table, expressed as a column vector \((x_1^1\) 11 , \(x_2^1\) 12 , …, \(x_m^1\) 1j ) T (\(j\in N\) * ); For \(S_2\), it is \((x_1^2\) 21 , \(x_2^2\) 22 , …, \(x_m^2\) 2j ) T , and so on until the data vector from \(S_n\) n is \((x_1^n\) n1 , \(x_2^n\) n2 , …, \(x_m^n\) nj ) T ; For each \(x_j^i\) ij (\(i\in\{1,2,3,\ldots,n\}, j\in N\) * ), user \(U\) and data owner \(S_i\) n respectively agree on a set of key pairs; Suppose user \(U\) and \(S_1\) jointly negotiate a key pair \((p_1\) i1 , \(p_2\) i2 ), user \(U\) and \(S_2\) jointly agree on a key pair \((q_1\) i1 , \(q_2\) i2 ), \(S_1\) and \(S_2\) respectively send \(\alpha_1\) i and \(\beta_1\) i to the proxy, where \(\alpha_1\) i = \(x_j^1p_1\) 1j + \(p_2\) 11 , \(\beta_1\) 12 = \(x_j^2q_1\) i + \(q_2\) 2j , and the proxy calculates \(\gamma\) 21 = 22 where i represents an addition or multiplication operation and sends it to the user. where represents an addition or multiplication operation and sends it to the user.

4. The secure and verifiable multi-source data system supporting SQL functions according to claim 3, characterized in that, In the addition function, data owners A and B respectively negotiate a set of secret values with user U, which are (p1, p2) and (q1, q2) respectively, where |p1| > |q1|, and |p1| and |q1| are the bit lengths of p1 and q1 respectively; Subsequently, A sends (ap1 + p2) to the agent, and at the same time B sends (bq1 + q2) to the agent. The agent calculates (ap1 + p2) + (bq1 + q2) and forwards it to user U; after the user receives the query result from the agent, subtract (p2 + q2) from it, 1. divide it by p1 to get the value of a; 2. take the modulus of p1, then divide it by q1 to get the value of b, and add them together to get a + b; for the case involving multiple parties, each data owner can negotiate a set of secret values with the user, and its core principle is the same as that of the two-party case.

5. The secure and verifiable multi-source data system supporting SQL functions according to claim 3, characterized in that In the multiplication operation, data owners A and B respectively negotiate a set of secret values with user U, which are (p1, p2) and (q1, q2) respectively, where |q1| > |q1| > |q2| > |p2|, and |p1|, |q1|, |q2|, |p2| are the bit lengths of p1, q1, q2, p2 respectively. A sends (ap1 + p2) to the agent, and B sends (bq1 + q2) to the agent. The agent calculates (ap1 + p2)(bq1 + q2), that is, abp1q1 + ap1q2 + bp2q1 + p2q2, and sends it to the user; After the user receives it, directly divide it by p1q1 to obtain the value of ab. In this process, if the proxy tampers with the intermediate result, the user can subtract the constant term p2q2 from the expression (abp1q1 + ap1q2 + bp2q1 + p2q2) received from the proxy, and perform coefficient modulo operations on each term one by one starting from the first term. If the result is not zero, malicious behavior can be verified. The same applies when multiple parties are involved.

6. The secure verifiable multi-source data system supporting SQL functions according to claim 3, characterized in that, The system includes a new protocol, denoted by "∏", which includes a key generation algorithm KeyGen, an encryption algorithm Enc, and a decryption algorithm Dec. The specific definition of the scheme ∏=(KeyGen, Enc, Dec) is as follows: (1)KeyGen(1 λ )→(p i1 ,p i2 ):This algorithm is a negotiation process between user U and data owner S i (i ∈ {1, 2, …, n}), generating a pair of keys (p i1 ,p i2 ). The KeyGen() algorithm is jointly executed by user U and each data owner S i respectively. By hashing the combination of two different pseudo-random number seeds (P Seed1 and P Seed1 ) and the value of i, and applying the next_prime() function to the result to find the next prime number to obtain the values of p i1 and p i2 . Denote the above process as: p i1 = next prime (H(P Seed1 , i)), p i2 = next prime (H(P Seed2 , i)); (2)Enc(x i ,state)→c state : This algorithm is executed by data owner S i . Its inputs include plaintext data x i and the state information state of S i . The plaintext data x i represents the data to be computed retrieved locally by S i , while the state information state encapsulates a series of important information about S i , including the key pair (p i1 , p i2 ); (3) Dec(c, OP, state) → m / ⊥: This algorithm is executed by user U. Its input parameter c represents the data received from the proxy, i.e., c state is the result after the operation of "OP", and OP is an operation. In the case of incorrect decryption, the error flag "⊥" is returned. Whether the decryption is successful depends on multiple factors, including the correct operation OP, the valid state information state, and the integrity of the data c. m is the final calculation result, corresponding to γ j .

7. The secure and verifiable multi-source data system supporting SQL functions according to claim 3, wherein It includes a query process, and the query process includes the following steps: (1)Data owner S i Configures query permissions for its specific data. These permissions can be set to public query or restricted to specific user access. When a user obtains the data query permission, they can initiate a query; (2) The user parses the original SQL query statement through the vproxy module of the client to generate a single-party SQL statement adapted for the data owner to execute. The user then performs broadcast encryption on the parsed query fields and performs digital signatures on each query statement one by one. (3) The user sends the encrypted and signed SQL statements together with the original SQL statement to the intermediate proxy. The intermediate proxy is responsible for forwarding this information to the corresponding data owner. (4) After receiving the SQL request, the data owner first verifies the signature. After passing the verification, it decrypts the encrypted query fields to generate a complete query instruction. Subsequently, the data owner executes the query operation. (5) For the data that does not need to be calculated during the query process, the data owner encrypts it with the public key and attaches a digital signature, while for the data that needs to be calculated, it is encrypted using the ∏ protocol and sent to the proxy. (6) The intermediate proxy is responsible for determining which data needs further calculation and performing the necessary arithmetic operations according to the protocol. For multi-table join operations, the proxy calls the PSI algorithm to implement. (7) The intermediate proxy merges the encrypted query results from multiple parties. (8) Finally, the intermediate proxy returns the merged result to the user. After receiving the result, the user verifies the signature, and then decrypts and verifies the query result.