Data processing method, product, equipment and computer readable storage medium

By blocking and reconstructing ciphertext status data in a multi-server hiding query system, replacing multiplication operations to reduce communication between servers, the problems of high communication overhead and security risks are solved, and the effect of reducing computing and communication overhead is achieved.

CN120068158AActive Publication Date: 2025-05-30INSPUR (BEIJING) ELECTRONICS INFORMATION IND CO LTD

Patent Information

Application Number
CN202510546604.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-05-30
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

In a multi-server hidden query scenario based on secret sharing, frequent interactions between managers need to be made, resulting in high communication overhead and security risks.

Method used

By blocking and reconstructing ciphertext status data between multiple servers of the query system, and replacing the secretly shared number multiplication operation instead of multiplication operation, the frequency of communication between servers is reduced.

Benefits of technology

It reduces the computing and communication overhead of the client, reduces the number of interactions between the servers, and improves the security and efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068158A_ABST
    Figure CN120068158A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method, a product, equipment and a computer readable storage medium, and relates to the technical field of privacy computing, a plurality of servers in a query system perform block division on ciphertext state data of each round of iteration reconstruction, and perform inner product operation on each data block and a query vector share of a secret sharing state to obtain a query vector share of a secret sharing state; according to the technical scheme, the secret sharing multiplication operation is replaced by the secret sharing multiplication operation, so that the problems of high communication overhead and security risk caused by the secret sharing multiplication operation are solved, and the technical effects of reducing the communication frequency between the servers and reducing the calculation and communication overhead of the clients are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a data processing method, product, device, and computer-readable storage medium. Background Art

[0002] The Private Information Retrieval (PIR) technology allows a user to retrieve target information from a remote database without revealing the query intention through cryptographic protocols (such as homomorphic encryption and secret sharing), solving the problems of user privacy leakage and data ownership conflict in traditional query modes. However, in the application of multi-server private information retrieval scenarios based on secret sharing, frequent interactions are usually required between management parties, and auxiliary data generated by a trusted third party is used to perform multiplication operations of secret sharing, resulting in high communication overhead and security risks.

[0003] Therefore, how to provide a solution to the above technical problems is an issue that those skilled in the art need to solve currently. Summary of the Invention

[0004] The present invention provides a data processing method, product, device, and computer-readable storage medium to at least solve the problems of high communication overhead and security risks in related technologies.

[0005] The present invention provides a data processing method, which is applied to any server in a query system. Multiple servers each store secret sharing shares of numerical data of target information. The data processing method of any server includes:

[0006] Obtain query information sent by a client, where the query information includes query vector secret shares in each round of multiple rounds of iteration;

[0007] Divide the ciphertext state data of the current round of iteration into at least one data block, perform inner product operations on the at least one data block and the query vector secret shares of the current round of iteration respectively to generate multiple encrypted intermediate data, and send the multiple encrypted intermediate data to other servers, and reconstruct the ciphertext state data of the next round of iteration based on the multiple encrypted intermediate data generated by each server. Repeat this step until the iteration termination condition is met; the ciphertext state data of the first round of iteration is obtained based on the secret sharing shares of each server;

[0008] Return the ciphertext state data when the iteration termination condition is met to the client, so that the client can decrypt and reconstruct the numerical data of the target information.

[0009] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above data processing method are implemented.

[0010] The present invention also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of the above data processing method when executing the computer program.

[0011] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program implements the steps of the above data processing method when executed by a processor.

[0012] Through the present invention, since the multiple servers in the query system respectively block the ciphertext state data reconstructed in each round of iteration, and perform inner product operations on each data block with the query vector shares in the secret sharing state, replacing the multiplication operation of secret sharing with the scalar multiplication operation of secret sharing, the problems of high communication overhead and security risks brought by the multiplication operation of secret sharing are solved, achieving the technical effects of reducing the communication frequency between servers and reducing the calculation and communication overhead of the client. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In order to more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0014] Figure 1 It is a schematic structural diagram of a query system provided by an embodiment of the present invention;

[0015] Figure 2 It is a flowchart of a data processing method provided by an embodiment of the present invention;

[0016] Figure 3 It is a schematic diagram of the underlying principle of block calculation and main round reduction provided by an embodiment of the present invention;

[0017] Figure 4 It is a schematic structural diagram of a data processing system provided by an embodiment of the present invention;

[0018] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention;

[0019] Figure 6 It is a schematic structural diagram of a computer-readable storage medium provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0021] It should be noted that in the description of the present invention, the terms "include", "comprise" or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present invention are used to distinguish similar objects and not to describe a particular order or sequence.

[0022] In order to enable those skilled in the art in the technical field to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0023] To facilitate the understanding of the data processing method provided in this embodiment, first, the query system applicable to this data processing method will be described. Refer to Figure 1As shown in the figure, the query system shown in this embodiment includes t servers, denoted as Server 1, ……, Server t respectively. The query system also includes a data provider and a client. The data provider stores the encrypted data in different servers respectively. The client sends encrypted query information to the servers, and the servers return encrypted target information to the client. Among them, the data provider is the owner or management unit of sensitive data (i.e., the data to be protected in this embodiment, and sensitive data is data that is not expected to be made public). Limited by its own storage resources or computing resources, the data provider entrusts the sensitive data to multiple third-party cloud platforms (i.e., servers). To facilitate subsequent numerical operations on the data, the data provider needs to numerically process the sensitive data (such as text data). Specifically, the data provider first encodes the sensitive data into a digital sequence using source coding technology, segments the digital sequence into equal-length segments, and performs numerical conversion on each segment one by one. Further, to avoid leakage of sensitive data, before uploading the numerical data to the server, the data provider performs some form of cryptographic technology processing on the sensitive data, such as using homomorphic encryption or secret sharing and other technologies for processing, and finally obtains the secret sharing shares of the numerical data in this embodiment. The server is a hosting platform for sensitive data, with a large amount of storage, computing, and bandwidth resources, and is responsible for responding to the information query work of the client. A single server does not master the original numerical data, but the data processed by cryptographic technology. After receiving the information query request, according to a series of predefined calculation rules, it cooperates with other servers to perform a series of calculations and feedbacks the calculation results to the client. The client is the user of sensitive data, with limited computing and bandwidth resources. It generates a series of data based on the index or keyword of the target information and sends a query request to the server. In addition to obtaining the target information, it requires that the query conditions (the index or keyword of the target information) are not leaked. Of course, the identities of the participating objects in the above query system can be cross-overlapped. For example, the data provider can also assume the role of the server at the same time.

[0024] An embodiment of the present invention provides a data processing method. The data processing method in this embodiment can be implemented by each of the above servers. Refer to Figure 2 and, in combination with the execution process of the data processing method, describe the data processing method in detail.

[0025] A data processing method provided in this embodiment includes:

[0026] S101: Obtain the query information sent by the client, where the query information includes the query vector secret shares in each round of multiple rounds of iteration;

[0027] In this embodiment, the client sets the length of the query vector, generates query information according to the index of the target information and the scale of the database, and sends it to each server, so that each server can perform subsequent operations according to the query information it receives, that is, the secret share of the query vector in each round of multiple rounds of iteration it receives.

[0028] It can be understood that the query vector is split into multiple independent parts through secret sharing technology, and each part is a secret share of the query vector. Each server only holds one of the shares. Exemplarily, the query vector is split into multiple shares using secret sharing technology , and distributed to t servers. It can be understood that a single share does not disclose any information about the original query vector. Multiple servers need to cooperate to combine their respective shares to reconstruct the original query vector.

[0029] S102: Divide the ciphertext state data of the current round of iteration into at least one data block, perform inner product operations on the at least one data block and the secret share of the query vector of the current round of iteration respectively to generate multiple encrypted intermediate data, and send them to other servers, and reconstruct the ciphertext state data of the next round of iteration based on the multiple encrypted intermediate data generated by each server. Repeat this step until the iteration termination condition is met; the ciphertext state data of the first round of iteration is obtained based on the secret share of each server.

[0030] S103: Return the ciphertext state data when the iteration termination condition is met to the client, so that the client can decrypt it to reconstruct the numerical data of the target information.

[0031] It can be understood that to avoid the leakage of sensitive data, before uploading sensitive data to the server, the data party numerically processes the sensitive data to obtain numerical data, and splits the numerical data into multiple independent parts through secret sharing technology. Each part is a secret share, and multiple secret shares are sent to multiple servers one by one. In this embodiment, for the first round of iteration, each server performs a homomorphic encryption operation on the locally stored secret share to obtain initial encrypted data, and then sends the initial encrypted data calculated by itself to other servers. Each server reconstructs the ciphertext state data used in the first round of iteration based on the initial encrypted data calculated locally and the initial encrypted data sent by other servers. The ciphertext state data in this embodiment is specifically the numerical data of the ciphertext state.

[0032] For each server, in each iteration process, the ciphertext state data corresponding to the current iteration is divided into multiple data blocks. The inner product operation is performed on each data block and the query vector secret share of the current iteration held by itself (the addition and multiplication involved are the homomorphic addition and homomorphic scalar multiplication operations in the homomorphic encryption algorithm), generating multiple encrypted intermediate data. The multiple encrypted intermediate data are sent to other servers so that each server can reconstruct the ciphertext state data in the next iteration based on the encrypted intermediate data generated by itself and the encrypted intermediate data sent by other servers. The above iteration process is repeated until the iteration termination condition is met. It can be understood that after each iteration, the amount of ciphertext state data decreases. Exemplarily, assume the amount of ciphertext state data , the block length , then in the first round, it is divided into 4 blocks, in the second round, it is divided into 2 blocks, and finally 1 block. In this embodiment, the query vector and the data are encrypted throughout the process. The server cannot know the user's intention or the original content. By performing inner product on blocks step by step, the target data is gradually focused, redundant calculations are reduced, and secret sharing prevents single point of failure. Losing some shares does not affect data recovery.

[0033] In this embodiment, the iteration termination condition includes but is not limited to stopping when the ciphertext state data is reduced to a single block, or reaching the maximum number of iterations. The server returns the ciphertext state data that meets the iteration termination condition to the client, and the client decrypts it to obtain the target information.

[0034] Suppose there are three servers in total, namely the first server S1, the second server S2, and the third server S3. For the first server S1, the other servers are the second server S2 and the third server S3. For the second server S2, the other servers are the first server S1 and the third server S3. For the third server S3, the other servers are the second server S2 and the first server S1.

[0035] It can be seen that in this embodiment, the multiple servers in the query system respectively divide the ciphertext state data reconstructed in each iteration into blocks, and perform inner product operations on each data block and the query vector shares in the secret sharing state, replacing the multiplication operation of secret sharing with the scalar multiplication operation of secret sharing, eliminating the dependence on a trusted third party, reducing the communication frequency between servers, reducing the number of data blocks after each iteration, reducing the processing volume in subsequent rounds, and the client only needs to send a small amount of query information to obtain the target information, significantly reducing the calculation and communication overhead of the client. In addition, the ciphertext operation throughout the process ensures the privacy of the query content, intermediate data, and results.

[0036] In an exemplary embodiment, the process of dividing the ciphertext state data of the current iteration into at least one data block includes:

[0037] Determine the query vector length;

[0038] Divide the ciphertext state data of the current round of iteration into at least one data block according to the length of the query vector, and the length of the data block is equal to the length of the query vector.

[0039] In this embodiment, first determine the length l of the query vector. The server divides the ciphertext state data into groups of every l, obtaining at least one data block, aiming to facilitate the subsequent inner product operation with the secret share of the query vector.

[0040] In an exemplary embodiment, the process of dividing the ciphertext state data of the current round of iteration into at least one data block includes:

[0041] Divide the ciphertext state data of the current round of iteration into at least one data block;

[0042] Determine whether there is a target data block with a length less than the length of the query vector among the at least one data block;

[0043] If so, add a preset supplementary value to the target data block so that the length of the supplemented target data block is equal to the length of the query vector.

[0044] In this embodiment, the target data block may be the tail data block after division in any round of iteration. Determine whether the length of the tail data block is less than the length of the query vector. The preset supplementary value can be 0. The server can expand the actual block length to l by padding with 0, so as to facilitate the subsequent inner product operation with the secret share of the query vector it holds.

[0045] In an exemplary embodiment, the process of performing inner product operations on at least one data block and the secret share of the query vector of the current round of iteration respectively, generating multiple encrypted intermediate data and sending them to other servers includes:

[0046] For each data block, perform an inner product calculation on the data block and the secret share of the query vector of the current round of iteration to obtain an inner product calculation result;

[0047] Use the preset public key to perform a first encryption calculation on multiple inner product calculation results respectively to obtain multiple encrypted intermediate data. It can be understood that the first encryption calculation can be approximate encryption or semi-encryption, etc.

[0048] Among them, the process of using the preset public key to perform a first encryption calculation on multiple inner product calculation results respectively to obtain multiple encrypted intermediate data includes: for each inner product calculation result, determine the current random element, and use the preset public key and the current random element to perform a first encryption calculation on the inner product calculation result to obtain intermediate encrypted data.

[0049] In an exemplary embodiment, the process of reconstructing the ciphertext state data for the next round of iteration based on multiple encrypted intermediate data generated by each server includes: performing a homomorphic addition operation on the multiple encrypted intermediate data generated by each server to obtain the ciphertext state data for the next round of iteration; the process of obtaining the ciphertext state data for the first round of iteration based on the secret sharing shares of each server includes: performing a second encryption calculation on the secret sharing shares of each server using a preset public key to generate multiple initial encrypted data; performing a homomorphic addition operation on the multiple initial encrypted data generated by each server to obtain the ciphertext state data for the first round of iteration. Among them, the second encryption calculation can specifically be a complete homomorphic encryption operation.

[0050] Specifically, in non-first iterations, perform a homomorphic addition operation on the multiple encrypted intermediate data generated by each server to obtain the ciphertext state data for the next round of iteration. In the first iteration, perform a second encryption calculation on the secret sharing shares of each server using a preset public key to generate multiple initial encrypted data. Specifically, perform a homomorphic addition operation on the multiple initial encrypted data generated by each server to obtain the ciphertext state data for the first round of iteration.

[0051] In an exemplary embodiment, the server receives query information from the client, completes a series of operations in an interactive and collaborative manner according to predefined calculation rules, and finally feeds back the calculation result to the client.

[0052] Specifically, the server first calculates the sequence , where, , , , represents the number of data in the 0th round of iteration, represents the number of data in the 1st round of iteration, represents the number of data in the last round of iteration, represents the total number of data before iteration, represents the number of data in the i-th round of iteration, represents taking the floor of the corresponding element, and l represents the length of the query vector.

[0053] For , the u-th server selects a random element , and encrypts the held share using the Paillier encryption algorithm, . Among them, the symbol represents the set consisting of all elements in the ring that are relatively prime to n. Exemplarily, the u-th server calculates: .

[0054] Among them, represents the u-th server under the control of the public key pk, the share it holds The result data after performing the Paillier homomorphic encryption operation. To ensure the correctness of the final result, take the smallest positive integer in the ring equivalence class for calculation.

[0055] Each server collaborates to calculate , that is, the numerical data for reconstructing the ciphertext state (i.e., the ciphertext state data in this embodiment). Specifically, for , the u-th server sends the result of the homomorphic encryption operation of the share it holds to other servers and receives the result of the homomorphic encryption operation of the share held by other servers sent by other servers (i.e., the w-th server ), where the value of w is the share , and . Then, for , the server performs the following calculation locally . .

[0056] Among them, the symbol represents the homomorphic addition operation in the Paillier encryption algorithm, and the symbol represents the addition operation on the ring , represents the result data after performing the Paillier homomorphic encryption operation on the -th data in the 0-th round of iteration under the control of the public key pk, represents the result data after performing the Paillier homomorphic encryption operation on the share of the -th data held by the 1st server under the control of the public key pk, represents the result data after performing the Paillier homomorphic encryption operation on the share of the -th data held by the 2nd server under the control of the public key pk, represents the result data after performing the Paillier homomorphic encryption operation on the share of the -th data held by the t-th server under the control of the public key pk, ​​​

[0057] The server collaboratively calculates the target information in a way of block calculation and round-by-round reduction. Specifically, in each round of iterative calculation, the server first divides the numerical data in ciphertext state into blocks of every l, and performs an inner product operation on each block and the query vector share of the current round of iteration (the addition and multiplication involved are the homomorphic addition and homomorphic scalar multiplication operations in the homomorphic encryption algorithm). Then, a homomorphic encryption operation is performed on the result of the inner product operation again, and the corresponding numerical data is collaboratively reconstructed again. Formally, for , each server loops through the following steps. The following takes the calculation process executed by the u-th server in the i-th round of iteration as an example for illustration. For , the u-th server performs the following calculations locally:

[0058] ;

[0059] ;

[0060] ;

[0061] where represents the homomorphic scalar multiplication operation in the Paillier encryption algorithm. For , , . Among them, represents the -th data in the i-th iteration, represents the -th data in the (i - 1)-th iteration, represents the 0-th component of the query vector in the (i - 1)-th round of iteration, represents the -th data in the (i - 1)-th round of iteration, represents the 1-st component of the query vector in the (i - 1)-th round of iteration, represents the -th data in the (i - 1)-th round of iteration, represents the last component of the query vector in the (i - 1)-th round of iteration. represents the last data in the i-th iteration, represents the -th data in the (i - 1)-th round of iteration, represents the -th data in the (i - 1)-th round of iteration, represents the -th data in the (i - 1)-th iteration, represents the a component

[0062] For , the u-th server randomly selects a random element and performs the following calculations locally:

[0063] ;

[0064] where is the u-th share of the i-th data in the i-th iteration The Paillier homomorphic ciphertext obtained by the second encryption calculation based on a certain random element (denoted as ), performs the first encryption calculation on , that is, calculates to obtain The Paillier homomorphic ciphertext encrypted by the second encryption based on the random element . Therefore, it is still denoted as , . The symbol mod represents the modulo operation. For example represents the remainder of a divided by b .

[0065] Each server cooperates to calculate . For , the u-th server sends to other servers (such as the w-th server , and ), and receives the share from other servers . The u-th server calculates locally:

[0066] .

[0067] where represents the result data after performing the Paillier homomorphic encryption operation on the i-th data in the i-th iteration under the control of the public key pk , is the result data after performing the Paillier homomorphic encryption operation on the u-th share of the i-th data in the i-th iteration under the control of the public key pk The i-th .

[0068] After completing the above loop operation, the number of data values of the ciphertext state is . For , the u-th server calculates locally ;

[0069] Among them, , represents the $u$-th share of the 0-th data in the -th round of iteration, represents the 0-th data in the -th round of iteration, represents the $u$-th share of the 0-th component of the query vector in the -th round of iteration, represents the 1-st data in the -th round of iteration, represents the $u$-th share of the 1-st component of the query vector in the -th round of iteration, represents the -th data in the -th round of iteration, represents the $u$-th share of the -th component of the query vector in the -th round of iteration. Sending as the final ciphertext state data when the iteration termination condition is satisfied to the client, it can be understood that represents the secret share based on Paillier homomorphic addition, represents the secret share of addition on the ring .

[0070] In an exemplary embodiment, the preset public key is the public key in the key pair generated by two prime numbers determined according to the length information of the digital sequence segment corresponding to the numerical data of the target information by the client according to the homomorphic encryption algorithm.

[0071] In this embodiment, it is assumed that the index of the target information queried by the client in the database is , and the length of the query vector is set to $l$. The client secretly selects two large prime numbers according to the digital sequence segmentation length of the data party, constructs a key pair based on the key generation algorithm of the homomorphic encryption algorithm, and then sends the public key to each server, and the private key is secretly stored locally. According to the information such as the target information index and the database scale, the client generates query information in the way of additive secret sharing and sends it to each server.

[0072] Specifically, the client generates a key pair ($pk$, $sk$) of the Paillier homomorphic encryption algorithm, sends the public key $pk$ to each server, and the private key $sk$ is secretly stored locally. Exemplarily, the client randomly selects two distinct large prime numbers $p$, $q$ (the bit lengths of $p$ and $q$ are both not less than 512), such that , and , where represents the greatest common divisor of. Calculate , where represents the least common multiple. Select such that there exists an integer a such that , for example . Among them, the symbol represents the set consisting of all elements relatively prime to in the ring . Set the function to calculate . The public key is denoted as , and the private key is denoted as , where g represents the first public key, n represents the second public key, represents the first private key, represents the second private key.

[0073] In an exemplary embodiment, the data processing method further includes:

[0074] Determine the query vector length, the index of the target information, and the database scale through the client;

[0075] Generate query information based on the query vector length, the index of the target information, and the database scale.

[0076] Among them, the process of generating query information based on the query vector length, the index of the target information, and the database scale includes:

[0077] Calculate the absolute index sequence of the target information in each round of iteration based on the query vector length, the index of the target information, and the database scale;

[0078] Calculate the relative index sequence according to the absolute index sequence and the query vector length;

[0079] Construct a query vector sequence using the relative index sequence; the query vector sequence includes the query vectors in each round of iteration;

[0080] For the query vector in each round of iteration, randomly split the query vector to obtain multiple query vector secret shares, and send the multiple query vector secret shares to multiple servers respectively.

[0081] Specifically, the client calculates the absolute index sequence , where: , , , , represents the iteration round number label, similar to i, and the only difference between the two is the value range.

[0082] represents rounding up, that is, greater than or equal to The smallest integer, denotes the round-down operation, i.e., the largest integer less than or equal to . The absolute index in this embodiment refers to the absolute position of the target information in the current round of iterative data sequence when the server performs iterative calculations round by round.

[0083] The client calculates the relative index sequence , where:

[0084] .

[0085] The symbol mod represents the modulo operation. The relative index in this embodiment refers to the relative position of the target information in the current data sequence block in the current round of iteration when the server performs iterative calculations round by round.

[0086] The client constructs the query vector sequence such that the th component of is 1 and the remaining components are 0. Among them, Figure 3 denotes the query vector used by the server in the Figure 3 round of iterative calculation, as shown in The client generates the query vector sequence for additive secret sharing on the ring , , ,

[0087] ​​​​​​​​​​​​​​​It can be understood that the client distributes the secret shares of the query vector sequence to different servers, so that a single server cannot obtain the complete query vector information. This effectively prevents the server from reverse-inferring the sensitive information of the client, such as query intent, data pattern, etc., thereby enhancing the privacy protection of the client. During the distributed computing process, the query vector sequence exists in each server in the form of secret shares, avoiding the transmission and storage of data in plaintext between multiple nodes. Even if a certain server is compromised, the attacker cannot directly obtain the complete data information, reducing the risk of data leakage. Only the necessary intermediate results or summary information need to be exchanged between servers, without transmitting the complete data, reducing the communication volume within the system.

[0088] In an exemplary embodiment, the data processing method further includes:

[0089] Encoding the numerical data of the target information into a digital sequence;

[0090] Segmenting the digital sequence into equal-length segments, and converting each segment into numerical data one by one. Performing secret sharing on each converted numerical data to obtain multiple secret sharing shares; the number of secret sharing shares is equal to the number of servers;

[0091] Distributing the multiple secret sharing shares to multiple servers.

[0092] In this embodiment, it is assumed that the data party has pieces of sensitive data (i.e., the data to be protected in this embodiment), and there are t non-colluding servers , non-colluding means that in a distributed system or cryptographic protocol, it is assumed that each server (server node) will not share the sensitive information it holds to ensure the security and privacy of the system. The data party uses source coding technology to encode the data into rows of digital sequences, and then segment the digital sequences into equal lengths and perform numerical conversion on each segment one by one. The data party performs additive secret sharing on the numerical data and distributes the secret sharing shares to each server through a secret channel to complete the distributed storage of sensitive data.

[0093] Specifically, the data party uses source coding technology to encode the pieces of sensitive data it holds to obtain rows of digital sequences . Exemplarily, the Huffman coding technique can be adopted to encode sensitive data into a uniquely decodable bit sequence, and the character coding set is made public. Huffman coding is a technique for encoding data based on the frequency of character occurrences, that is, the higher the frequency of a character, the shorter the corresponding bit string, so as to achieve the purpose of information compression. The character coding set refers to the correspondence between characters and their coding results that includes all characters in the sensitive data, such as , where indicates that the coding result of the capital letter A is 01.

[0094] The data party selects a sufficiently large positive integer m (such as m ≥ 1024). For the digital sequence , the data party makes all digital sequences of equal length and the sequence length divisible by m by padding a number of digits at the end (for example, when using Huffman coding, padding 0 or 1), and then segments with consecutive m-bit digital segments as the basic unit. For each digital segment of digital sequences, the data party converts it into an element on a specific ring to obtain a data matrix. Exemplarily, assuming a certain bit segment is , then the corresponding element on the ring is . In this way, the data party obtains a data matrix M with the number of rows being on the ring . The ring represents an algebraic structure formed by elements under addition and multiplication operations. Among them, represents the set of all elements that are congruent to modulo , which is called an equivalence class of modulo . The addition operation on the ring is defined as , and the multiplication operation is defined as .

[0095] For each column element of the data matrix M, denoted as , represents the 0th element of the current column, represents the first element of the current column, represents the last element of the current column. The data party performs additive secret sharing of these elements among t servers. Exemplarily, for , the data party randomly selects t elements for the th column such that , and then secretly sends to the u-th server It can be understood that in each cycle of the solution of the present invention, one column of the data matrix M is processed each time, and each column has elements.

[0096] It can be understood that the numerical data of the target information is split into multiple secret shares and stored in different servers respectively. Even if one of the servers is compromised, the attacker cannot obtain the complete sensitive data. Since the servers do not share the sensitive information they hold, even if multiple servers are attacked, as long as not all servers are compromised simultaneously and cooperate to leak data, the security of the data can still be guaranteed. Multiple servers can calculate and process the data shares they hold simultaneously, so as to achieve parallel computing and improve the efficiency and speed of data processing. During the data processing, the servers only need to exchange the necessary intermediate results or summary information, rather than transmitting the complete data. This reduces the communication overhead within the system and improves the overall data processing efficiency. In a large-scale distributed system, the reduction of communication overhead further improves the system performance.

[0097] In an exemplary embodiment, the process of the client reconstructing the numerical data of the target information after decryption includes:

[0098] The client reconstructs the numerically encrypted data after homomorphic encryption according to the ciphertext status data returned by each server;

[0099] Use the private key in the key pair to decrypt the homomorphic ciphertext of the numerically encrypted data after homomorphic encryption to obtain the numerical data of the target information, and convert the numerical data of the target information into a digital sequence, and convert it into a digital sequence;

[0100] Convert the digital sequence into the numerical data of the target information.

[0101] In this embodiment, the client receives the result data (i.e., the encrypted status data when the iteration termination condition is satisfied) fed back by each server, restores the intermediate result of the numerically encrypted data of the target information through the secret reconstruction algorithm, and then obtains the intermediate result of the numerical data through homomorphic decryption. Perform modulo operation on the intermediate result of the numerical data to obtain the numerical data, and obtain the numerical data of all bit segments of the target information through multiple loops. Arrange all the numerical data in a predetermined order, convert it into a bit sequence, and then restore the complete target information according to the character encoding set.

[0102] Specifically, the client calculates , where represents the ciphertext status data returned by the 1st server, represents the ciphertext status data returned by the 2nd server, represents the ciphertext status data returned by the tth server, It represents the result data after performing the homomorphic addition operation in the Paillier encryption algorithm on the ciphertext state data returned by t servers. For perform the Paillier homomorphic decryption operation, that is, calculate , to obtain the intermediate result of the numerical data . Then, calculate , to obtain the numerical data . It represents performing the Paillier homomorphic decryption operation on under the control of the private key sk.

[0103] Based on the query information and the key pair, the server and the client perform the above series of steps on each column element of the data matrix M respectively, and all the numerical data of the target information bit sequence can be obtained. The client arranges all the numerical data in a predetermined order, converts it into a bit sequence, and then performs an inverse transformation according to the character encoding set to restore the complete target information.

[0104] In summary, in this embodiment, to reduce the communication overhead of the client, a method of block calculation and round-by-round reduction is proposed to extract the target information from the database. Specifically, the client calculates the relative index of the target information in each round based on the target information index and the database scale, and constructs the query vector for each round accordingly. To hide the target information index, the client distributes the query vector to each server in the form of additive secret. Each server divides the numerical data sequence in ciphertext state according to the length of the query vector, and performs an inner product operation with the query vector of the current round iteration to reduce the database scale while retaining the relevant data of the target information. Repeat the above steps until the database scale is reduced to 1, and finally extract the relevant data of the target information. To reduce the communication overhead of the server, it is proposed to use homomorphic encryption technology to encrypt the numerical data in the secret sharing state, and then reconstruct the numerical data in ciphertext state accordingly, and perform an inner product operation with the query vector in the secret sharing state, so as to replace the multiplication operation in secret sharing with the multiplication operation in secret sharing, reduce the interaction times between servers, reduce the communication overhead, and at the same time maintain the privacy of the numerical data. To support the private query of multiple types of data, a numerical conversion scheme for multiple types of data based on source coding is proposed. Specifically, the data party represents the sensitive data as a digital sequence based on source coding technology, divides the digital sequence into equal-length segments, and converts each segment into an element on a specific ring, that is, realizes the numerical conversion of the sensitive data.

[0105] The beneficial effects are as follows. Through the application of the homomorphic encryption algorithm, the execution process of the algorithm no longer depends on a trusted third party, reducing security risks and effectively reducing the communication overhead of the server. By extracting target information from the database through block calculation and round-by-round reduction, the client only needs to send a small amount of query information to obtain the target information, significantly reducing the calculation and communication overhead of the client. Under the assumption that all servers do not collude, it is difficult for the server to obtain the target information index of the client through the query information, protecting the privacy of the client. The client cannot obtain other sensitive data of the data party other than the target information based on the feedback result of the server, protecting the privacy of the data party.

[0106] From the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0107] Referring to Figure 4 , the embodiments of the present invention further provide a data processing system, including:

[0108] An acquisition module 11, configured to acquire query information sent by a client, where the query information includes query vector secret shares in each round of multiple rounds of iteration;

[0109] A processing module 12, configured to divide the ciphertext state data of the current round of iteration into at least one data block, perform inner product operations on the at least one data block and the query vector secret shares of the current round of iteration respectively to generate multiple encrypted intermediate data, and send the multiple encrypted intermediate data to other servers, and reconstruct the ciphertext state data of the next round of iteration based on the multiple encrypted intermediate data generated by each server. Repeat this step until the iteration termination condition is met; the ciphertext state data of the first round of iteration is obtained based on the secret sharing shares of each server;

[0110] A feedback module 13, configured to return the ciphertext state data when the iteration termination condition is met to the client, so that the client can decrypt and reconstruct the numerical data of the target information.

[0111] In an exemplary embodiment, the process of dividing the ciphertext state data of the current round of iteration into at least one data block includes:

[0112] Determine the length of the query vector;

[0113] Divide the ciphertext state data of the current round of iteration into at least one data block according to the length of the query vector, and the length of the data block is equal to the length of the query vector.

[0114] In an exemplary embodiment, the process of dividing the ciphertext state data of the current round of iteration into at least one data block according to the length of the query vector includes:

[0115] Divide the ciphertext state data of the current round of iteration into at least one data block according to the length of the query vector;

[0116] Determine whether there is a target data block with a length less than the length of the query vector in at least one data block;

[0117] If so, add a preset supplementary value to the target data block so that the length of the supplemented target data block is equal to the length of the query vector.

[0118] In an exemplary embodiment, the process of performing an inner product operation on at least one data block and the query vector secret share of the current round of iteration to generate multiple encrypted intermediate data and then sending them to other servers includes:

[0119] For each data block, perform an inner product calculation on the data block and the query vector secret share of the current round of iteration to obtain an inner product calculation result;

[0120] Use a preset public key to perform a first encryption calculation on multiple inner product calculation results to obtain multiple encrypted intermediate data.

[0121] In an exemplary embodiment, the process of using a preset public key to perform a first encryption calculation on multiple inner product calculation results to obtain multiple encrypted intermediate data includes:

[0122] For each inner product calculation result, determine the current random element, and use the preset public key and the current random element to perform a first encryption calculation on the inner product calculation result to obtain intermediate encrypted data.

[0123] In an exemplary embodiment, the process of reconstructing the ciphertext state data of the next round of iteration based on multiple encrypted intermediate data generated by each server includes:

[0124] Perform a homomorphic addition operation on multiple encrypted intermediate data generated by each server to obtain the ciphertext state data of the next round of iteration;

[0125] The process of obtaining the ciphertext state data of the first round of iteration based on the secret sharing shares of each server includes:

[0126] Use a preset public key to perform a second encryption calculation on the secret sharing shares of each server to generate multiple initial encrypted data;

[0127] Perform a homomorphic addition operation on multiple initial encrypted data generated by each server to obtain the ciphertext state data of the first round of iteration.

[0128] In an exemplary embodiment, the preset public key is the public key in the key pair generated by the client according to the length information of the digital sequence segment corresponding to the numerical data of the target information according to the homomorphic encryption algorithm.

[0129] In an exemplary embodiment, the process of the client reconstructing the numerical data of the target information after decryption includes:

[0130] The client reconstructs the numerically encrypted data according to the ciphertext status data returned by each server;

[0131] Use the private key in the key pair to decrypt the homomorphic ciphertext of the numerically encrypted data to obtain the numerical data of the target information, and convert the numerical data of the target information into a digital sequence;

[0132] Convert the digital sequence into the numerical data of the target information.

[0133] In an exemplary embodiment, the iteration termination conditions include that the number of iterations reaches a preset number and / or the number of data blocks of the ciphertext status data is reduced to a preset number.

[0134] In an exemplary embodiment, the data processing system is further configured to:

[0135] Encode the numerical data of the target information into a digital sequence;

[0136] Perform equal-length segmentation on the digital sequence, and convert each segment into numerical data one by one. Perform secret sharing on each converted numerical data to obtain multiple secret sharing shares; the number of secret sharing shares is equal to the number of servers;

[0137] Distribute multiple secret sharing shares to multiple servers.

[0138] In an exemplary embodiment, the data processing system is further configured to:

[0139] Determine the query vector length, the index of the target information, and the database size through the client;

[0140] Generate query information based on the query vector length, the index of the target information, and the database size.

[0141] In an exemplary embodiment, the process of generating query information based on the query vector length, the index of the target information, and the database size includes:

[0142] Calculate the absolute index sequence of the target information in each round of iteration based on the query vector length, the index of the target information, and the database size;

[0143] Calculate the relative index sequence according to the absolute index sequence and the query vector length;

[0144] Construct a query vector sequence using the relative index sequence; the query vector sequence includes the query vectors in each round of iteration;

[0145] For the query vector in each iteration, the query vector is randomly split to obtain multiple query vector secret shares, and the multiple query vector secret shares are respectively sent to multiple servers.

[0146] For the description of the features in the corresponding embodiments of the data processing system in this embodiment, reference can be made to the relevant descriptions in the corresponding embodiments of the data processing method, which will not be elaborated here one by one.

[0147] Please refer to Figure 5 , an embodiment of the present invention further provides an electronic device, including:

[0148] A memory 21 for storing computer programs;

[0149] A processor 22 for implementing the steps of any of the above data processing methods when executing the computer program.

[0150] The electronic device further includes:

[0151] An input interface 23 connected to the processor 22 via a communication bus 26, for obtaining externally imported computer programs, parameters, and instructions, and storing them in the memory 21 under the control of the processor 22. The input interface can be connected to an input device to receive parameters or instructions manually input by the user. The input device can be a touch layer covered on the display screen, or a button, trackball, or touchpad provided on the terminal housing.

[0152] A display unit 24 connected to the processor 22 via a communication bus 26, for displaying the data sent by the processor 22. The display unit can be a liquid crystal display screen or an electronic ink display screen, etc.

[0153] A network port 25 connected to the processor 22 via a communication bus 26, for communicating and connecting with external terminal devices. The communication technology used for this communication connection can be a wired communication technology or a wireless communication technology, such as Mobile High-Definition Link technology, Universal Serial Bus, High-Definition Multimedia Interface, Wi-Fi technology, Bluetooth communication technology, Low Energy Bluetooth communication technology, communication technology based on IEEE802.11s, etc.

[0154] Please refer to Figure 6 , an embodiment of the present invention further provides a computer-readable storage medium 30, on which a computer program 31 is stored, and when the computer program 31 is executed by a processor, the steps of any of the above data processing methods are implemented.

[0155] In an exemplary embodiment, the above-mentioned computer-readable storage medium 30 may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store the computer program 31.

[0156] An embodiment of the present invention also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of any one of the above data processing methods are implemented.

[0157] An embodiment of the present invention also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any one of the above data processing method embodiments are implemented.

[0158] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0159] The above has introduced in detail a data processing method, product, electronic device, and computer-readable storage medium provided by the present invention. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A data processing method, characterized in that: Applied to any server of the query system, a plurality of the servers each store a secret sharing share of the numerical data of the target information, and a data processing method of any server includes: Obtaining query information sent by the client, the query information including a secret share of a query vector for each round in multiple iterations; Divide the ciphertext state data of the current iteration into at least one data block, perform inner product operation on at least one data block and the secret share of the query vector of the current iteration, generate multiple encrypted intermediate data and send them to other servers, and reconstruct the ciphertext state data of the next iteration based on the multiple encrypted intermediate data generated by each server, and repeat this step until the iteration termination condition is met; the ciphertext state data of the first iteration is obtained based on the secret sharing share of each server; The ciphertext state data when the iteration termination condition is satisfied is returned to the client so that the client can reconstruct the numerical data of the target information after decryption.

2. The data processing method according to claim 1, characterized in that: The process of dividing the ciphertext state data of the current iteration into at least one data block includes: Determine the query vector length; The ciphertext state data of the current iteration is divided into at least one data block according to the length of the query vector, and the length of the data block is equal to the length of the query vector.

3. The data processing method according to claim 2, characterized in that: The process of dividing the ciphertext state data of the current iteration into at least one data block according to the query vector length includes: Dividing the ciphertext state data of the current iteration into at least one data block according to the query vector length; Determine whether there is a target data block in at least one of the data blocks whose length is less than the length of the query vector; If yes, a preset supplementary value is added to the target data block so that the length of the supplemented target data block is equal to the length of the query vector.

4. The data processing method according to claim 1, characterized in that: The process of performing inner product operations on at least one of the data blocks and the query vector secret share of the current iteration to generate a plurality of encrypted intermediate data and then sending the generated encrypted intermediate data to the other server ends includes: For each of the data blocks, performing an inner product calculation on the data block and the secret share of the query vector of the current iteration to obtain an inner product calculation result; A preset public key is used to perform a first encryption calculation on the plurality of inner product calculation results to obtain a plurality of encrypted intermediate data.

5. The data processing method according to claim 4, characterized in that: The process of performing a first encryption calculation on the plurality of inner product calculation results using a preset public key to obtain a plurality of encrypted intermediate data includes: For each of the inner product calculation results, a current random element is determined, and a first encryption calculation is performed on the inner product calculation result using a preset public key and the current random element to obtain intermediate encrypted data.

6. The data processing method according to claim 4, characterized in that: The process of reconstructing the ciphertext state data of the next iteration based on the plurality of encrypted intermediate data generated by each of the server ends includes: Performing homomorphic addition operation on the plurality of encrypted intermediate data generated by each of the servers to obtain ciphertext state data for the next round of iteration; The process of obtaining the ciphertext state data of the first round of iteration based on the secret sharing shares of each of the servers includes: Using the preset public key to perform a second encryption calculation on the secret sharing share of each server to generate a plurality of initial encrypted data; A homomorphic addition operation is performed on the multiple initial encrypted data generated by each of the servers to obtain the ciphertext state data of the first round of iteration.

7. The data processing method according to claim 4, characterized in that: The preset public key is a public key in a key pair generated by the client according to a homomorphic encryption algorithm using two prime numbers determined according to length information of a digital sequence segment corresponding to the digitized data of the target information.

8. The data processing method according to claim 7, characterized in that: The process of reconstructing the numerical data of the target information after decryption by the client includes: The client reconstructs the homomorphically encrypted numerical data according to the ciphertext state data returned by each of the servers; The homomorphic ciphertext of the homomorphically encrypted digitized data is decrypted using the private key in the key pair to obtain the digitized data of the target information, and the digitized data of the target information is converted into a digital sequence.

9. The data processing method according to claim 1, characterized in that: The iteration termination condition includes that the number of iterations reaches a preset number and / or the number of data blocks of the ciphertext state data is reduced to a preset number.

10. The data processing method according to claim 1, characterized in that: The data processing method further includes: Encoding the data to be protected into a digital sequence; Divide the digital sequence into segments of equal length, and convert the segments into numerical data one by one, and perform secret sharing on each of the converted numerical data to obtain a plurality of secret sharing shares; the number of the secret sharing shares is equal to the number of the server ends; Distribute the plurality of secret sharing shares to the plurality of service terminals.

11. The data processing method according to any one of claims 1 to 10, characterized in that: The data processing method further includes: Determining the query vector length, the index of target information and the database size through the client; The query information is generated based on the query vector length, the index of the target information, and the database size.

12. The data processing method according to claim 11, characterized in that: The process of generating the query information based on the query vector length, the index of the target information and the database size includes: Calculate the absolute index sequence of the target information in each round of iteration based on the query vector length, the index of the target information and the database size; Calculate a relative index sequence according to the absolute index sequence and the query vector length; Constructing a query vector sequence using the relative index sequence; the query vector sequence includes query vectors in each round of iteration; For the query vector in each round of iteration, the query vector is randomly split to obtain a plurality of query vector secret shares, and the plurality of query vector secret shares are sent to the plurality of server ends respectively.

13. A computer program product, characterized in that The invention comprises a computer program / instruction, which, when executed by a processor, implements the steps of the data processing method according to any one of claims 1 to 12.

14. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data processing method according to any one of claims 1 to 12 when executing the computer program.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data processing method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • Cipher state database query method and device based on copy secret sharing

    CN115455488A

  • Decryption method, related device and storage medium

    CN115589281A

  • GPU acceleration method and system for homomorphic encryption technology

    CN119740252A

  • Privacy information retrieval method, system and equipment and storage medium

    CN119760780A

  • Secure Substring Search to Filter Encrypted Data

    US20190220620A1

Cited By

  • Data hiding query method and electronic equipment

    CN121637570A