System and method for protecting data privacy and storage medium
By combining a fully homomorphic cryptosystem and a secure two-party computation protocol, the security and efficiency issues of the k-means algorithm in privacy protection are solved, realizing a clustering algorithm that does not leak data in distributed computing, thus improving the security and efficiency of data privacy protection.
Patent Information
- Application Number
- CN202510967916.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-11-11
AI Technical Summary
Existing k-means clustering algorithms have security issues regarding privacy protection. During the calculation process, cluster center or merging information can be easily leaked, leading to the leakage of information related to the original data. Furthermore, the current technology is relatively inefficient.
Employing a fully homomorphic cryptosystem and a secure two-party computation protocol with unintended transmission, the system outsources computation by dividing the data into two parts and sending them to two non-colluding servers. It uses a privacy-preserving k-means algorithm for clustering, and the servers generate cluster labels through an interactive protocol and return them to the data provider without revealing the original data.
It achieves the protection of data privacy and confidentiality during the clustering process, without disclosing any intermediate calculation results, improving computational efficiency and network transmission efficiency, and enhancing security.
Smart Images

Figure CN120930177A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data privacy technology, and in particular to a system, method and storage medium for protecting data privacy. Background Technology
[0002] Clustering algorithms are widely used unsupervised machine learning methods that divide data in a dataset into multiple categories based on similarity, making it easier to observe the characteristics of different categories. The k-means algorithm is one such partition-based clustering algorithm. Traditional plaintext clustering algorithms require collecting all plaintext data before computation. With increasing emphasis on data privacy and confidentiality, when the data to be clustered is scattered across different entities, the transmission and collection of plaintext data becomes difficult. Therefore, it is necessary to design clustering algorithms that protect data privacy.
[0003] Current privacy protection methods for the k-means algorithm have security issues. They may require disclosing the intermediate result of cluster centers during the calculation process, or disclosing merging information when merging partial results between different computing servers during distributed computing. These data leaks can lead to varying degrees of leakage of original data-related information. Summary of the Invention
[0004] This application provides a system, method, and storage medium for protecting data privacy, in order to solve problems such as data privacy leakage that are easily caused in related technologies.
[0005] The first aspect of this application provides a system for protecting data privacy, comprising: a first server, a second server, and at least one data provider; for each data provider, the data provider sends first secret-shared data of privacy data to the first server, and the data provider sends second secret-shared data of privacy data to the second server; the first server processes the first secret-shared data, the second server processes the second secret-shared data, and after the first server and the second server jointly process the data, they obtain complete clustering labels, and send the complete clustering labels to all data providers.
[0006] Optionally, both the first and second servers are equipped with a privacy-preserving k-means algorithm. The first and second servers input a vector dataset in secret-sharing form, which includes either first or second secret-sharing data. The execution flow of the privacy-preserving k-means algorithm includes: initializing cluster labels; performing clustering iterations based on the vector dataset; if the current iteration is the first iteration, selecting multiple data points from the vector dataset as cluster centers; the first server randomly generates the indices of the selected data points and sends these indices to the second server; otherwise, calculating cluster centers based on the number of data points belonging to the target cluster; the first and second servers calculate a distance matrix based on the vector dataset and the distance centers; calculate new cluster labels based on the distance matrix; share the new cluster labels with each other; and generate complete cluster labels based on their own cluster labels and the other server's cluster labels. If the new cluster labels match the old cluster labels, the clustering iteration ends; if the new cluster labels do not match the old cluster labels, the old cluster labels are updated using the new cluster labels, and the clustering iteration continues.
[0007] Optionally, the first server and the second server each invoke the squared distance calculation algorithm to calculate the distance matrix. The execution flow of the squared distance calculation algorithm includes: the second server encodes its own cluster centers and the first secret shared data into a polynomial, encrypts the polynomial to obtain encrypted ciphertext, and sends the encrypted ciphertext to the first server; the first server performs homomorphic operations based on its own cluster centers, the second secret shared data, and the encrypted ciphertext to obtain intermediate ciphertext of the squared distance, and sends the intermediate ciphertext of the squared distance to the second server; the second server decrypts the intermediate ciphertext of the squared distance to obtain the target matrix, and calculates the first distance matrix based on the target matrix and the first secret shared data; the first server calculates the second distance matrix based on the random matrix and the second secret shared data; and the distance matrix is determined based on the first distance matrix and the second distance matrix.
[0008] Optionally, the first server and the second server each invoke an extreme value algorithm to calculate the cluster labels. The calculation formula for the extreme value algorithm is as follows:
[0009] <l i >←arg min j∈[k] ( <D ij >);
[0010] Among them, l i For clustering labels, D ij Let be the distance matrix, j be the cluster center, and k be the cluster.
[0011] Optionally, the privacy-preserving k-means algorithm is constructed based on a fully homomorphic cryptosystem of the RLWE (Ring Learning With Errors) problem and a secure two-party computation protocol based on unintentional transmission.
[0012] Optionally, the fully homomorphic cryptosystem is defined by the polynomial degree, the plaintext modulus, and the ciphertext modulus. The fully homomorphic system also supports negative cyclic shift, LWE (Learning With Errors) extraction, LWE ciphertext packing, and RLWE ciphertext packing.
[0013] Optionally, the secure two-party computation protocol includes an extremum algorithm and a division algorithm. The extremum algorithm is used to calculate the maximum or minimum value and the index corresponding to the maximum or minimum value based on the given arithmetic secret-sharing data. The division algorithm is used to calculate the integer quotient of the secret-sharing data and the target public divisor, and the calculation result is in secret-sharing form.
[0014] Optionally, the data provider's processing flow includes: representing the privacy data as fixed-point numbers and encoding it into integers; splitting the integers into first secret-sharing data and second secret-sharing data.
[0015] The second aspect of this application provides a method for protecting data privacy, applied to a data privacy protection system as described in the above embodiments. The method includes: for each data provider, the data provider sends first secret-shared data of privacy data to a first server, and the data provider sends second secret-shared data of privacy data to a second server; the first server processes the first secret-shared data, the second server processes the second secret-shared data, and after the first server and the second server jointly process the data, they obtain complete clustering labels, and send the complete clustering labels to all data providers.
[0016] A third aspect of this application provides a computer-readable storage medium having a computer program or instructions stored thereon, which are executed by a processor to perform the data privacy protection method as described in the above embodiments.
[0017] Therefore, this application has at least the following beneficial effects:
[0018] This application constructs a system for protecting data privacy. In this system, the data provider divides its privacy-preserving data into two parts and sends them separately to two independent servers for outsourced computation. The two servers, through an interactive privacy computation protocol, generate clustering labels for the data provided by the data provider and return the results to their respective data providers. The two servers cannot access the original privacy-preserving data, thus protecting the privacy and confidentiality of the data. This solves the technical problems in related technologies that easily lead to data privacy leaks.
[0019] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0020] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0021] Figure 1 This is a schematic diagram of a system for protecting data privacy according to an embodiment of this application;
[0022] Figure 2 This is a schematic diagram of data clustering labels provided according to embodiments of this application;
[0023] Figure 3 This is a flowchart of a method for protecting data privacy provided according to an embodiment of this application. Detailed Implementation
[0024] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0025] Before describing the solution of this application, let me first introduce some related content.
[0026] The k-means algorithm is a partition-based clustering algorithm. Based on the number of clusters specified by the user, it initially randomly assigns a corresponding number of data points as cluster centers and assigns all data points to the cluster containing the nearest cluster center. The algorithm then iterates through multiple rounds. In each iteration, the average center of all data points in each cluster is calculated and used as the new cluster center. Then, all data points are again assigned to the cluster containing the nearest cluster center, completing one iteration update. The algorithm terminates after the user-specified number of iterations or when the data point classification labels no longer change, and the current clustering result is taken as the final result.
[0027] Current privacy protection methods for the k-means algorithm suffer from efficiency and security issues. On the security front, some methods require disclosing the intermediate result of cluster centers during computation, or leaking merging information when merging partial results between different computing servers in distributed computing. These data leaks lead to varying degrees of leakage of information related to the original data. On the efficiency front, some methods typically use a single cryptographic technique, such as the computationally inefficient Paillier semi-homomorphic encryption system based on the discrete logarithm problem, resulting in long privacy computation protocols. Other methods rely solely on inadvertent transmission, leading to extremely high network transmission rounds and total data volume.
[0028] To this end, this application proposes to combine a fully homomorphic cryptosystem based on the RLWE problem and a secure two-party computation protocol based on unintentional transmission to jointly construct a data privacy-protecting system that does not disclose any intermediate computation information, so as to realize the computation of clustering algorithms.
[0029] The following description, with reference to the accompanying drawings, outlines a system, method, and storage medium for protecting data privacy according to embodiments of this application. Addressing the security issues inherent in current clustering algorithm privacy protection methods mentioned in the background, such as the need to disclose intermediate results like cluster centers during computation or the leakage of merging information when merging partial results between different computing servers in distributed computing, this application provides a system for protecting data privacy. In this method, the data provider divides its privacy data into two parts and sends them to two non-colluding servers for outsourced computation. The two servers, through an interactive privacy computation protocol, generate clustering result labels for the data provided by the data provider and return the results to their respective data providers. The two servers cannot access the original privacy data, thus protecting data privacy and confidentiality. This solves the problem of data privacy leakage inherent in related technologies.
[0030] Specifically, Figure 1 This is a schematic diagram of a system for protecting data privacy provided in an embodiment of this application.
[0031] like Figure 1 As shown, the data privacy protection system 10 includes: a first server 11, a second server 12, and at least one data provider 13.
[0032] Among them, for each data provider 13, the data provider 13 sends the first secret sharing data of the private data to the first server 11, and the data provider 11 sends the second secret sharing data of the private data to the second server 12. The first server 11 processes the first secret sharing data, and the second server 12 processes the second secret sharing data. After the first server 11 and the second server 12 jointly process the data, a complete clustering label is obtained, and the complete clustering label is sent to all data providers 13.
[0033] It can be understood that the embodiment of this application constructs a system 10 for protecting data privacy. The data provider 13 divides the private data into two parts and sends them to two non-colluding servers for outsourcing calculation respectively. The two servers generate a clustering result label for the data provided by the data provider 13 through an interactive privacy computing protocol, and return the result to each data provider. The two servers cannot obtain the original private data, so as to protect the privacy and confidentiality of the data.
[0034] Further, in the embodiment of this application, the processing flow of the data provider 13 includes: representing the private data as a fixed-point number and encoding it as an integer; splitting the integer into the first secret sharing data and the second secret sharing data.
[0035] It can be understood that the data provider 13 in the embodiment of this application processes the private data to generate data in the form of secret sharing held by the two corresponding servers.
[0036] Assume that the first server 11 is P0 and the second server 12 is P1. Each data provider D i All its vector data First, represent and encode it as an integer in fixed-point number Then split it into the form of secret sharing and send it to the two computing servers respectively. That is, for any x ij , data provider D i Randomly generate a vector <x ij >0 = r ij , calculate <x ij >1 = x ik -r ij , and then send <x ij >0 to P0 and <x ij >1 to P1.
[0037] Specifically, this application uses fixed-point numbers to represent decimals and adopts the secret sharing technology as one of the ways to store the ciphertext of intermediate results. For an integer 0 ≤ x < t, take The two parties P0 and P1 participating in the calculation respectively hold Denoted as the two parties holding secret sharing <x> t When t is explicitly stated in the context, this superscript is omitted and it is directly written as t. <x>For brevity, the value access space is represented using... <x>∈S means x∈S. For a decimal... This invention uses fixed-point number representation to encode it as an integer. Let the precision of the fixed-point number be φ, then the encoding is... It is worth noting that when a fixed-point number is multiplied once, its precision doubles to 2φ.
[0038] Furthermore, in this embodiment, both the first server 11 and the second server 12 are equipped with a privacy-preserving k-means algorithm. The first server 11 and the second server 12 are input with a vector dataset in secret-sharing form, which includes either first secret-sharing data or second secret-sharing data. The execution flow of the privacy-preserving k-means algorithm includes: initializing cluster labels, performing clustering iterations based on the vector dataset, wherein if the current iteration is the first iteration, multiple data points are selected from the vector dataset as cluster centers, and the first server 11 randomly generates the selected cluster centers. The first server 11 and the second server 12 send the data point index to the second server 12; otherwise, they calculate the cluster center based on the number of data points belonging to the target cluster. The first server 11 and the second server 12 calculate the distance matrix based on the vector dataset and the cluster center, calculate the new cluster label based on the distance matrix, and share the new cluster label with the other server. They generate a complete cluster label based on their own cluster label and the other server's cluster label. If the new cluster label is consistent with the old cluster label, the clustering iteration ends. If the new cluster label is inconsistent with the old cluster label, the old cluster label is updated using the new cluster label, and the clustering iteration continues.
[0039] It is understood that both the first server 11 and the second server 12 in this application embodiment are equipped with a privacy-preserving k-means algorithm. The input is a vector dataset in the form of secret sharing. In the execution process of the privacy-preserving k-means algorithm, multiple iterations are performed to obtain the clustering labels of the privacy data.
[0040] Furthermore, in this embodiment of the application, the first server 11 and the second server 12 each invoke an extreme value algorithm to calculate the clustering labels. The calculation formula for the extreme value algorithm is as follows:
[0041]
[0042] Among them, l i For clustering labels, D ij Let be the distance matrix, j be the cluster center, and k be the cluster.
[0043] Specifically, in the privacy-preserving k-means algorithm, the maximum number of iterations T and the number of clusters k ≥ 2 are given in advance. Two computing servers P0 and P1 are input with a vector dataset in a secret-sharing format. The algorithm outputs cluster labels l i ∈[k], i∈[n]. The algorithm execution process is as follows.
[0044] 1. Initialize l′ i ←0, i∈[n].
[0045] 2. Perform T iterations. In the t-th iteration:
[0046] (1) If t = 0, from Randomly select k data points as cluster centers <c0> , <c1>,…, <c k-1 When selecting a data point, P0 can randomly generate the index of the selected data point and send it to P1; otherwise, calculate... Where |{i:l′ i =j}| indicates based on label l′ i Determine the number of data points belonging to cluster j.
[0047] (2) Both parties call the squared distance calculation algorithm, and the two sets of secret sharing vectors are input as { <x i >} i∈[n] ,{ <c j >} j∈[k] Output the distance matrix of secret sharing. <d>.
[0048] (3) Both parties call the extreme value algorithm to calculate for i∈[n]:
[0049] <l i >←arg min j∈[k] ( <D ij >);
[0050] Both parties publicly disclosed their secret sharing <l i >0, <l i >1, obtain the complete cluster label l i .
[0051] (4) If for all i∈[n] l i =l′ i The algorithm terminated prematurely, l i This is the final result.
[0052] (5) Update l′ i ←k i ,i∈[n].
[0053] 3. If the iteration process does not end prematurely, the cluster label obtained from the last calculation will be... i This is the final result.
[0054] Furthermore, in the embodiments of this application, the privacy-preserving k-means algorithm is constructed based on a fully homomorphic cryptosystem of the RLWE problem and a secure two-party computation protocol based on unintended transmission.
[0055] It is understood that the privacy-preserving k-means algorithm in this application is constructed based on the fully homomorphic cryptosystem of the RLWE problem and the secure two-party computation protocol based on unintentional transmission. Homomorphic encryption assists the efficient linear algorithm for distance calculation, while secure multi-party computation is used to handle nonlinear calculations such as division and comparison for calculating cluster centers. The two technologies complement each other, reducing transmission overhead and computation time overhead respectively.
[0056] Compared to previous designs that used a single computing technology, this application further improves the time efficiency and network transmission efficiency of the algorithm interaction protocol execution, which is conducive to the further acceptance, deployment and promotion of this privacy protection algorithm in practice.
[0057] Furthermore, in the embodiments of this application, the fully homomorphic cryptosystem is defined by the polynomial degree, the plaintext modulus, and the ciphertext modulus. The fully homomorphic system also supports negative cyclic shift, LWE extraction, LWE ciphertext packing, and RLWE ciphertext packing.
[0058] Specifically, this application implements homomorphic encryption through a fully homomorphic cryptosystem.
[0059] The homomorphic encryption system used in this application must support operations on integer polynomial rings, and the BFV cryptosystem can be adopted. This cryptosystem is defined by three parameters: the polynomial degree N. p The plaintext modulus is t, and the ciphertext modulus is q. Its plaintext space is a polynomial ring. That is, the plaintext is less than N. p A polynomial of degree n with integer coefficients, whose coefficients are modulo t. The ciphertext space is... in That is, the ciphertext consists of two polynomials with integer coefficients, and these two polynomials are less than N. p The coefficient is modulo q. The encryption of plaintext a is abbreviated as: Where sk is the encryption key. When the key is explicitly specified in the context, it is omitted and simply abbreviated as sk. This fully homomorphic cryptosystem supports addition and multiplication of ciphertext. Specifically, let the decryption algorithm be Dec(a,sk) = Dec(a), and this cryptosystem satisfies:
[0060] Dec(Enc(a)+Enc(b))=a+b;
[0061] Dec(Enc(a)·b)=ab.
[0062] In addition, the cryptographic system also supports the following operations:
[0063] 1. Negative cyclic shift. For polynomial rings. Let element a in the array be denoted as . For the case where k < 0, due to the properties of the polynomial ring, we can define Shift(a,k) = Shift(k + 2N). p In a cryptographic system, suppose the ciphertext c consists of two polynomials, i.e., c = (a, b), defined as follows:
[0064] Shift(c,k)=(Shift(a,k),Shift(b,k));
[0065] The cryptographic system guarantees that Dec(Shift(c,k)) = Shift(Dec(c),k).
[0066] 2. LWE Extraction. In cryptographic systems, ciphertext is usually RLWE ciphertext. LWE ciphertext can be extracted from RLWE ciphertext. LWE ciphertext only encrypts one coefficient of the original plaintext polynomial. The extraction algorithm is denoted as Extract(c,i). Let [a] i Let represent the coefficient of the i-th term in polynomial a, and DecLWE(c) represent the decryption method for LWE ciphertext. Then, the LWE extraction methods are:
[0067] DecLWE(Extract(c,i))=[Dec(c)] i ;
[0068] In addition, LWE ciphertext supports addition of integers, i.e. have:
[0069] DecLWE(c+a)=DecLWE(c)+a.
[0070] 3. LWE encrypted packaging. For k≤log2N p It can be packaged. LWE ciphertext l0,…,l n-1 To become a single RLWE ciphertext c = PackLWEs(n,(l0,…,l n-1 )),satisfy:
[0071]
[0072] 4. RLWE encrypted packaging. For k≤log2N p It can be packaged. RLWE ciphertext c0,…,c n-1 To become a separate RLWE ciphertext. Let... Then we have:
[0073]
[0074] Furthermore, in the embodiments of this application, the secure two-party computation protocol includes an extreme value algorithm and a division algorithm. The extreme value algorithm is used to calculate the corresponding maximum or minimum value and the index corresponding to the maximum or minimum value based on the given arithmetic secret sharing data. The division algorithm is used to calculate the integer quotient of the secret sharing data and the target public divisor, and the calculation result is in the form of secret sharing.
[0075] Specifically, the fundamental operators for secure two-party computation based on inadvertent transmission in this application all use secret sharing as the input and output of the algorithm. The fundamental operators used in this application include the following two types:
[0076] 1. Extreme value. max( <a0>,…, n-1 >) Generate the maximum value shared by a given set of n arithmetic secrets. arg max( 0,…, n-1 >) Generates the index corresponding to the maximum value among n arithmetic secrets. min(·) and arg min(·) can be defined similarly.
[0077] 2. Division. Divide ( d) Computational secret sharing t Given a divisor d and a publicly known divisor d, the result is secretly shared.
[0078] Further, in this embodiment, the first server 11 and the second server 12 each call the squared distance calculation algorithm to calculate the distance matrix. The execution flow of the squared distance calculation algorithm includes: the first server 11 encodes the cluster centers and the first secret sharing data into a polynomial, encrypts the polynomial to obtain encrypted ciphertext, and sends the encrypted ciphertext to the second server 12; the second server 12 performs homomorphic operations based on the cluster centers, the second secret sharing data, and the encrypted ciphertext to obtain the squared distance intermediate ciphertext, and sends the squared distance intermediate ciphertext to the first server 11; the first server 11 decrypts the squared distance intermediate ciphertext to obtain the target matrix, and calculates the first distance matrix based on the target matrix and the first secret sharing data; the second server 12 calculates the second distance matrix based on the random matrix and the second secret sharing data; and the distance matrix is determined based on the first distance matrix and the second distance matrix.
[0079] The squared distance calculation algorithm of this application takes two sets of vectors in ciphertext form as input, uses homomorphic encryption to calculate the square of the Euclidean distance between any pair of vectors in the vector space, forms a squared distance matrix, and returns it to both parties in ciphertext form as the algorithm output.
[0080] Specifically, the squared distance calculation algorithm of this application includes the following steps:
[0081] The input to the squared distance calculation algorithm is two sets of d-dimensional vectors secretly shared by participants P0 and P1. <p i >} i∈[n] ,{ j >} j∈[m] ,in The output is the squared distance matrix of the secret sharing. Define two functions to encode a vector as a polynomial: for a vector p = (p0, p1, ..., p... d-1 ), λ≥d, then:
[0082]
[0083] The specific process of the algorithm is as follows.
[0084] 1. The smallest power of two such that λ≥d, let be...
[0085] 2. For participant P P For P∈{0,1}, for i∈[n], j∈[m], calculate:
[0086]
[0087] 3. P1 encrypts and sends. Give it to P0.
[0088] 4. Calculate P0 for i∈[n], j∈[m]:
[0089]
[0090]
[0091] Note that qqPk is calculated here. j At that time, it is necessary to send the same LWE ciphertext qq j Packaging is performed using μ times.
[0092] 5. P0 for I∈[n ′ ]calculate:
[0093]
[0094] Where, when i≥n, we can take These are the LWE ciphertext with 0 encryption, the zero vector, and the RLWE ciphertext with 0 encryption, respectively.
[0095] 6. P0 for I∈[n ′ Calculate for j∈[m]:
[0096]
[0097] cr Ij ←ppPk I +qqPk j -2·Shift(pqPk Ij ,1-λ).
[0098] 7. Take P0 randomly generates a matrix And encode its row vector, that is, for k∈[K], calculate:
[0099]
[0100] 8. Calculate P0 for k∈[K]:
[0101]
[0102] crDense k ←PackRLWEs(λ,crGroup k )-s k .
[0103] 9. P0 will store all ciphertext {crDense} k } k∈[K] Send to P1.
[0104] 10. Decrypt P1 to obtain K items of length N. p A vector, denoted as a matrix
[0105] 11. Calculate the results of P0 and P1 respectively in the following ways. <d> 0, <d>1:
[0106] For i∈[n], j∈[m], let:
[0107] v=(i modμ)·λ+((Im+j)modλ),
[0108] calculate:
[0109]
[0110] In addition, it should be noted that the computing server and the data provider can overlap. When there are more than two data providers, two can be selected from all the data providers as the servers used for computing.
[0111] In summary, this application designs a k-means algorithm computation scheme for data privacy, allowing two or more data providers to outsource their data in encrypted form to two non-colluding servers. These two servers generate clustering result labels for all data points through an interactive privacy computation protocol. Apart from the final clustering labels generated by the algorithm, the original data and intermediate computation results of any data provider are not disclosed to other data providers or the two computation servers during this process, thus protecting data privacy and confidentiality.
[0112] This application presents a privacy-preserving k-means clustering scheme applied to a data outsourcing scenario involving two non-colluding computing servers and multiple data providers. The computing servers and data providers can overlap; either two providers can be selected from among all data providers, or two additional non-colluding participants can be requested to provide computing services. The computing servers are designed as P0 and P1, and each data provider is D. i All of its vector data First, it is encoded as an integer using fixed-point representation. Then, it is split into secret-sharing forms and sent to the two computing servers respectively; that is, for any x... ij Data provider D i Randomly generated vectors <x ij >0=r ij ,calculate <x ij >1=x ij -r ij Then <x ij >0 is sent to P0. <x ij >1 can be sent to P1. After this splitting and sending is completed, the two computing servers can assume that they hold the union of vector data from all data providers in the form of a secret sharing, and can apply the privacy-preserving k-means algorithm described above. After the two computing servers P0 and P1 complete the clustering algorithm and obtain the clustering labels, one of them can send all the clustering labels back to all participants.
[0113] The following specific embodiment describes the process of clustering data while protecting data privacy.
[0114] Each data holder (equivalent to a data provider) possesses a subset of dataset S1, along with two dedicated servers for privacy-preserving computation. Dataset S1 is a commonly used dataset for testing clustering algorithms, containing 5000 2D vectors arranged in 15 Gaussian clusters, each containing 300 to 350 data points. Before computation begins, all data holders convert their data into a fixed-point representation using the method described above, secretly share it, and send it to the two computation servers. The fixed-point precision used is φ = 16, and the modulus of the arithmetic secret sharing is t = 2. 59 After receiving the data in the secret-shared format, the two servers establish a connection and initialize the homomorphic encryption system parameters. P1 generates the encryption key, and P0 performs homomorphic computation, consistent with the squared distance calculation algorithm. The homomorphic encryption system parameters are set to N. p =8192, t=2 59 ,q≈2 180 After initialization, both parties perform the privacy-preserving k-means algorithm as described in the invention to obtain clustering label results. Since the clustering label results are publicly available to both computing servers, the process ends after either computing server returns the clustering label results to all data holders. The visualization results of the clustering labels on the S1 dataset are as follows.< / d> < / d> Figure 2 As shown, dots of different colors represent different clusters.
[0115] The data privacy protection scheme proposed in this application specifically includes:
[0116] 1. This application allows the calculation of clustering results using complete sets of vector data to be clustered from multiple data providers, while protecting the original data information of each data provider from being disclosed to any other participating parties. This application facilitates the centralized use of similar data originally scattered across multiple parties for unsupervised machine learning data analysis, improves the value of data utilization, and contributes to the development and promotion of privacy protection technologies.
[0117] 2. During the calculation process, apart from the clustering label results of each iteration, no intermediate calculation results such as the distance information between data points are disclosed. Compared with existing works that require the disclosure of information such as the distance information between points and the coordinates of the cluster centers, the scheme proposed in this application greatly enhances the data confidentiality of the algorithm and can provide higher security.
[0118] 3. This application comprehensively employs multiple cryptographic tools and techniques, including homomorphic encryption and secure multi-party computation based on unintentional transmission. Homomorphic encryption assists in the efficient linear algorithm for distance calculation, while secure multi-party computation is used to handle nonlinear calculations such as division and comparison for calculating cluster centers. The two techniques complement each other, reducing both transmission overhead and computation time overhead. Compared to previous designs that used a single computation technique, this application further improves the time efficiency of the algorithm interaction protocol execution and the network transmission efficiency, which is conducive to the further acceptance, deployment, and promotion of this privacy-preserving algorithm in practice.
[0119] According to the data privacy protection system proposed in the embodiments of this application, the data provider divides the privacy data into two parts and sends them to two non-colluding servers for outsourced computation. The two servers generate clustering result labels for the data provided by the data provider through an interactive privacy computation protocol and return the results to each data provider. The two servers cannot obtain the original privacy data from each other, thereby achieving the protection of data privacy and confidentiality.
[0120] Next, with reference to the accompanying drawings, a method for protecting data privacy according to embodiments of this application is described.
[0121] Figure 3 This is a flowchart of a method for protecting data privacy according to an embodiment of this application.
[0122] like Figure 3 As shown, this method for protecting data privacy includes the following steps:
[0123] In step S101, for each data provider, the data provider sends the first secret sharing data of privacy data to the first server, and the data provider sends the second secret sharing data of privacy data to the second server.
[0124] In step S102, the first server processes the first secret-shared data, the second server processes the second secret-shared data, and the first and second servers work together to process the data to obtain complete cluster labels, which are then sent to all data providers.
[0125] It should be noted that the foregoing explanation of the system embodiment for protecting data privacy also applies to the method for protecting data privacy in this embodiment, and will not be repeated here.
[0126] According to the data privacy protection method proposed in the embodiments of this application, the data provider divides the privacy data into two parts and sends them to two non-colluding servers for outsourced computation. The two servers generate clustering result labels for the data provided by the data provider through an interactive privacy computation protocol and return the results to each data provider. The two servers cannot obtain the original privacy data from each other, thereby achieving the protection of data privacy and confidentiality.
[0127] This application also provides a computer-readable storage medium storing a computer program or instructions thereon, which, when executed by a processor, implements the above-described method for protecting data privacy.
[0128] In the description of this specification, the references to "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0129] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0130] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0131] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or more of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (FPGAs), field-programmable gate arrays (FPGAs), etc.
[0132] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments. < / d> < / c0> < / x> < / x> < / x>
Claims
1. A system for protecting data privacy, characterized in that, include: A first server, a second server, and at least one data provider; For each data provider, the data provider sends a first secret sharing data of privacy data to the first server, and the data provider sends a second secret sharing data of privacy data to the second server. The first server processes the first secret sharing data, and the second server processes the second secret sharing data. After the first server and the second server process the data together, they obtain a complete cluster label, and send the complete cluster label to all the data providers.
2. The system for protecting data privacy according to claim 1, characterized in that, Both the first server and the second server are equipped with a privacy-preserving k-means algorithm. The first server and the second server input a vector dataset in secret-sharing form, which includes either first secret-sharing data or second secret-sharing data. The execution flow of the privacy-preserving k-means algorithm includes: Initialize cluster labels and perform clustering iterations based on the vector dataset. If the current iteration is the first iteration, select multiple data points from the vector dataset as cluster centers. The first server randomly generates the index of the selected data points and sends the index to the second server. Otherwise, calculate the cluster centers based on the number of data points whose cluster labels belong to the target cluster. The first server and the second server calculate a distance matrix based on the vector dataset and the cluster centers, calculate new cluster labels based on the distance matrix, share the new cluster labels with each other's servers, and generate complete cluster labels based on their own cluster labels and each other's cluster labels. If the new cluster label is the same as the old cluster label, the clustering iteration ends. If the new cluster label is different from the old cluster label, the old cluster label is updated using the new cluster label, and the clustering iteration continues.
3. The system for protecting data privacy according to claim 2, characterized in that, The first server and the second server each invoke the squared distance calculation algorithm to calculate the distance matrix. The execution flow of the squared distance calculation algorithm includes: The second server encodes the first secret sharing data into a polynomial based on its own cluster center and the first secret sharing data, encrypts the polynomial to obtain encrypted ciphertext, and sends the encrypted ciphertext to the first server. The first server performs homomorphic operations based on its own cluster center, the second secret sharing data, and the encrypted ciphertext to obtain the square distance intermediate ciphertext, and then sends the square distance intermediate ciphertext to the second server; The second server decrypts the ciphertext of the squared distance to obtain the target matrix, and calculates the first distance matrix based on the target matrix and the first secret sharing data; The first server calculates the second distance matrix based on the random matrix and the second secret sharing data; The distance matrix is determined based on the first distance matrix and the second distance matrix.
4. The system for protecting data privacy according to claim 2, characterized in that, The first server and the second server each invoke an extreme value algorithm to calculate cluster labels. The calculation formula for the extreme value algorithm is as follows: <l i >←argmin j∈[k] (<D ij >); Among them, l i For clustering labels, D ij Let be the distance matrix, j be the cluster center, and k be the cluster.
5. The system for protecting data privacy according to claim 1, characterized in that, The privacy-preserving k-means algorithm is constructed based on a fully homomorphic cryptosystem of the RLWE problem and a secure two-party computation protocol based on unintended transmission.
6. The system for protecting data privacy according to claim 5, characterized in that, The fully homomorphic cryptosystem is defined by the polynomial degree, plaintext modulus, and ciphertext modulus. The fully homomorphic system also supports negative cyclic shift, LWE extraction, LWE ciphertext packing, and RLWE ciphertext packing.
7. The system for protecting data privacy according to claim 5, characterized in that, The secure two-party computation protocol includes an extremum algorithm and a division algorithm, wherein, The extreme value algorithm is used to calculate the corresponding maximum or minimum value and the index corresponding to the maximum or minimum value based on the given arithmetic secret sharing data; The division algorithm is used to calculate the integer quotient of the secret-shared data and the target public divisor, and the calculation result is in the form of secret sharing.
8. The system for protecting data privacy according to claim 1, characterized in that, The data provider's processing flow includes: The privacy data is represented as a fixed-point number and encoded as an integer; The integer is split into the first secret sharing data and the second secret sharing data.
9. A method for protecting data privacy, characterized in that, The method is applied to the data privacy protection system according to any one of claims 1-8, wherein the method includes: For each of the data providers, the data provider sends a first secret sharing data of privacy data to the first server, and the data provider sends a second secret sharing data of privacy data to the second server; The first server processes the first secret-shared data, the second server processes the second secret-shared data, and after the first server and the second server process the data together, they obtain complete cluster labels, which are then sent to all the data providers.
10. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, The computer program or instructions are executed by a processor to implement the method for protecting data privacy as described in claim 9.