Privacy protection k-means method for distributed data
In distributed data processing, the server negotiates the initial centroid and user local encryption calculation method, combined with homomorphic encryption and pseudo-random value schemes, the problem of the distributed k-means algorithm leaking intermediate data at each iteration is solved, and efficient and secure distributed data clustering analysis is achieved.
Patent Information
- Application Number
- CN202510263610.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-20
AI Technical Summary
The existing distributed k-means algorithm leaks intermediate data at each iteration, causing user privacy and data security to be compromised.
By negotiating the initial centroid between servers and performing encrypted global centroid calculations locally on the user, combining homomorphic encryption and pseudo-random value schemes, safe calculation and encrypted transmission of global centroids are realized to avoid intermediate data leakage.
It realizes the clustering analysis of distributed data without leaking user privacy data, ensuring data security and computing efficiency, and achieving the same accurate clustering results as plaintext k-means.
Smart Images

Figure CN120180494A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data security, and particularly to a privacy-preserving k-means method for distributed data. Background Art
[0002] Traditional centralized data analysis methods face many limitations: on the one hand, it is difficult for them to efficiently process data across regions and platforms; on the other hand, there is a risk of leaking sensitive information during the data integration process. To address these challenges, a distributed privacy-preserving k-means clustering method has emerged.
[0003] The distributed privacy-preserving k-means clustering technology can complete in-depth mining of user behavior patterns, preference characteristics, etc. without directly accessing or transferring the original data. This helps enterprises improve service quality and optimize the user experience, and enables more scientific and reasonable decision-making in many important fields such as public security, healthcare, and education reform.
[0004] While the distributed privacy-preserving k-means clustering technology brings convenience to our lives, it also brings new challenges. Existing distributed data privacy-preserving k-means clustering methods are divided into distributed algorithms and outsourcing algorithms. However, both of these algorithms face different challenges. The outsourcing algorithm involves multiple users uploading encrypted data, executing the algorithm on one or more servers, and finally returning the clustering results to all parties. The outsourcing algorithm well solves the privacy problem of data mining for distributed data and reduces the pressure on users in terms of computing power and storage. However, outsourcing data to a cloud server deprives the customer of direct control over their data, inevitably bringing some new problems. The transmission of encrypted data brings a huge communication burden to the cloud, and existing outsourcing schemes have the problem of low availability of encrypted data, and a large amount of computational overhead is required for clustering. The distributed algorithm requires all users to cooperate in execution. Users calculate gradients locally in plaintext, encrypt them, and then the server aggregates them into intermediate clustering centers and returns them to the users, and finally a clustering model is trained. The distributed algorithm has a smaller computational overhead and communication volume than the outsourcing algorithm. However, existing distributed algorithms leak some information during each iteration, such as leaking the number of each class, leaking the sum of data, and leaking intermediate clustering centers. Although significant progress has been made in existing research, it is still a challenge to meet simultaneous security, efficiency, and practicality. Summary of the Invention
[0005] The purpose of the present invention is to provide a privacy-preserving k-means method for distributed data, aiming to solve the technical problem that existing distributed k-means algorithms leak intermediate data during each iteration.
[0006] To achieve the above object, the present invention provides a privacy-preserving k-means method for distributed data, comprising the following steps:
[0007] Step 1: The server negotiates the initial centroid, processes it and sends it to the user;
[0008] Step 2: The user performs the nearest clustering calculation locally based on the encrypted global centroid or initial centroid of the previous round;
[0009] Step 3: The local centroid and the number of samples in each cluster are scrambled and encrypted and then transmitted to different servers for secure aggregation;
[0010] Step 4: The two servers interact to complete the calculation and encryption of the global centroid, and send it to the user for iteration;
[0011] Step 5: When the sum of the distances between the local cluster centroid and the local cluster centroid of the previous iteration no longer changes or changes very little, the user sends a termination message; when the number of termination messages exceeds a certain threshold or reaches the iteration limit, the iteration stops; the server calculates the final centroid and sends it to all users.
[0012] Optionally, the specific method of global centroid encryption in step 2 is as follows:
[0013] The distance between the sample point and the centroid is calculated using the Euclidean distance. i With the centroid c j The distance d ij Subtract the sample point x i With the centroid c j' The distance d ij' , expand the formula into If d ij -d ij' <0, then x i Closer to c j , otherwise x i Closer to c j' , the user learns the centroids of all and c j -c j' , repeat k-1 times to find the nearest centroid of the sample, which is and c j -c j' Adding positive random perturbations does not affect the results.
[0014] Optionally, during the secure aggregation in step 3, the user uploads the private data plus the mask generated from the seed, encrypts them separately using the public keys of server TS1 and server TS2, and sends them to TS2 and TS1. The user sends the seed to the corresponding server using the paired key. The server reconstructs the mask using the randomly shared seed by the user, aggregates the private data, and eliminates the mask through the addition property of the Paillier homomorphic algorithm.
[0015] Optionally, during the calculation of the global centroid in step 4, for the two servers TS1 and TS2, TS1 generates a random vector r1 and a random number r3, and TS2 generates a random number r2; TS1 calculates where E2(·) is the ciphertext encrypted using the homomorphic public key generated by server TS2, and s is the sum of samples for each class; by the homomorphic property TS1 sends E2(r1s) to TS2, and TS2 decrypts it using the private key sk2 to get D2(E2(r1s)) = r1s; similarly, TS2 also performs related calculations m is the number of samples for each class, and sends E1(r2m) to TS1; TS1 decrypts it using the private key sk1 to get D1(E1(r2m)) = r2m; then TS1 calculates and sends it to TS2, and TS2 calculates to get TS2 encrypts and and then sends the encryption result and to TS1. After that, the random vector r1 and the random number r3 can be eliminated through the homomorphic encryption property, and finally E2(c) and E2(c 2 ) are obtained, where E2(c T c)←E2(sum(c 2 )).;
[0016] TS1 can calculate and E2(c j -c j' ) in the ciphertext state; then add a random number to hide the intermediate data and calculate and send it to TS2. After TS2 decrypts it and adds the random number r 2jj' to get where j∈(1,...,k), j'∈(2,...,k), j<j', r 1jj' , r 2jj' ≥0;
[0017] Server TS2 sends r 1jj' r 2jj' (c j -cj' ) Sent to each user, where r1jj' is a random number generated by server TS1, and r 2jj' is a random number generated by server TS2, and r 1jj' , r 2jj' ≥ 0.
[0018] Optionally, in step 5, when the sum of the distances between the local clustering centroids and the local clustering centroids in the previous iteration no longer changes or changes very little, the formula is expressed as
[0019]
[0020] where t is the number of iterations, c is the centroid, k is the number of clusters, and ε is a preset value.
[0021] The present invention provides a privacy-preserving k-means method for distributed data. First, two servers negotiate the initial centroids and send them to the users; the users locally perform the nearest clustering calculation, scramble and encrypt the local centroids and the number of samples in each cluster, and then send them to different servers for aggregation; the two servers cooperate to calculate the global clustering center under the premise of protecting privacy, and send it to the users for iteration; when the preset number of iterations is reached or the abort request exceeds the threshold, the process terminates, and the server calculates the final clustering center and distributes it. The present invention solves the problem that the existing distributed k-means algorithm leaks intermediate data during each iteration. The method of the present invention neither leaks the user's private data, nor leaks the global centroid and the number of samples in each cluster to the server and the user, and can achieve the same accurate clustering result as the plaintext k-means without leaking any private data. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.
[0023] Figure 1 It is a schematic flowchart of a privacy-preserving k-means method for distributed data according to the present invention.
[0024] Figure 2 It is a schematic execution composition diagram of a specific embodiment of a privacy-preserving k-means method for distributed data according to the present invention.
[0025] Figure 3 It is a schematic flowchart of the data aggregation algorithm according to the present invention.
[0026] Figure 4It is a schematic flow diagram of the algorithm for calculating the global clustering center of the present invention.
[0027] Figure 5 It is a schematic comparison diagram between the present method and the existing classification in a specific embodiment of the present invention.
[0028] Figure 6 It is a schematic comparison diagram of the running time of the present method and the existing privacy-preserving k-means scheme on different data sets in a specific embodiment of the present invention.
[0029] Figure 7 It is a schematic comparison diagram of the communication volume of the present method and the existing privacy-preserving k-means scheme on different data sets in a specific embodiment of the present invention.
[0030] Figure 8 It is a schematic comparison diagram of the running time of the present method and the existing privacy-preserving k-means scheme under different parameters in a specific embodiment of the present invention.
[0031] Figure 9 It is a schematic comparison diagram of the communication volume of the present method and the existing privacy-preserving k-means scheme under different parameters in a specific embodiment of the present invention. Detailed implementation manners
[0032] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as limiting the present invention.
[0033] The present invention provides a privacy-preserving k-means method for distributed data, including the following steps:
[0034] Step 1: The server negotiates the initial centroids, processes them, and sends them to the users.
[0035] Step 2: The users perform the nearest clustering calculation locally according to the encrypted global centroids or the initial centroids of the previous round.
[0036] Step 3: The local centroids and the number of samples in each cluster are scrambled and encrypted, and then sent to different servers for secure aggregation.
[0037] Step 4: Through the interaction of two servers, the calculation and encryption of the global centroids are completed, and then sent to the users for iteration.
[0038] Step 5: When the sum of the distances between the local clustering centroids and the local clustering centroids of the previous iteration no longer changes or changes very little, the users send termination messages; when the number of termination messages exceeds a certain threshold or reaches the iteration limit, the iteration stops; the server calculates the final centroids and sends them to all users.
[0039] The execution process is as follows Figure 1 shown. First, the servers (TS1 and TS2) negotiate the initial centroid, process it, and send it to the user. The user executes the nearest cluster search algorithm locally and uses the data aggregation algorithm to scramble and encrypt the local gradient and the number of clusters, and then sends them to different servers for aggregation. Without compromising the user's privacy, the two servers calculate the global cluster center, obtain the scrambled global centroid, and send it to the user for continued iteration. When the iteration limit is reached or the number of user requests to abort exceeds the threshold, the iteration terminates. After the iteration terminates, the server outputs the final cluster center and sends the final cluster center to each user.
[0040] Specifically, the input data for clustering is distributed among many users, consisting of n users DP1 ··· DP n Each user DPi has its own m private samples X i ={x i1 , x i2 ,..., x im}, where x ij is composed of a u-dimensional vector describing the sample characteristics. The combined data of all users forms a data set {x1, x2,..., x n}. The goal of the present invention is to establish mutual privacy protection between the server and users DP1 ··· DP n when calculating the k best cluster centers of the combined data {x1, x2,..., x n}, and neither the server nor the users can infer any private information about other users and the intermediate data generated by the calculation except for the final result of the k-means clustering.
[0041] The following further describes the specific embodiments and the algorithms in the process of step execution
[0042] The implementation of the method in this embodiment consists of two parts: system setting and privacy protection cluster centroid update. The system setting is used to set the encryption key and generate the initial centroid. User DP i and servers TS1 and TS2 respectively establish pairwise keys ek i,1 , ek i,2 , and then the two servers respectively generate their own public and private key pairs for encryption, send the public keys to other users and servers, and finally the two servers jointly generate the initial centroid, scramble it, and distribute it to the users.
[0043] 1. Nearest cluster search algorithm
[0044] In step 2, the user is assigned to the nearest center within the iteration. Let d ij be the distance between the i-th sample x i and the j-th cluster center cj The distance between, that is
[0045] d ij =(x i -c j ) T (x i -c j ). (1)
[0046] Now consider the distance from x i to clusters c j and c j' There is
[0047]
[0048] If the server sends and c j -c j' to the user, the user calculates d ij -d ij' If d ij -d ij' < 0, then x i is closer to c j , otherwise x i is closer to c j' This process is repeated k - 1 times until x i finds the nearest cluster center among all k clusters. However, the cluster centers need to be kept secret from the server and the user (how the cluster centers are kept secret from the server will be introduced in the third stage). Therefore, we can and c j -c j' add random perturbations
[0049]
[0050] where ρ(j,j') > 0 is a random number. There are
[0051] Since ρ(j,j') > 0, adding the random number has no effect on how close x i is to the cluster center. After obtaining this information, each user can determine the nearest center based on their own private data. The specific process is as follows: The user's sample x l First calculate
[0052]
[0053] S (1,2) is related to ρ (1,2) (c1 - c2). If S (1,2) < 0, the sample x lCloser to the centroid c1, otherwise the sample x l is closer to the centroid c2. Next, use the centroids and c j to represent the two centroids, where If the sample x l is closer to the centroid then the next step is to calculate If the sample x l is closer to the centroid c j then the next step is to calculate S (j,j+1) . This process is repeated k - 1 times to find the nearest center of the samples.
[0054] User DP i After each sample is classified, the user can obtain the list of clustered samples shown in Table 1.
[0055]
[0056]
[0057] Table 1 User DP i The samples contained in the class in User DP are then used to calculate User DP i the sum of the samples in each class and S i =(s i,1 ,..., s i,k ), where
[0058]
[0059] the number of samples m i =(m i,1 ,..., m i,k ) in each class in User DP i, where
[0060] m i,j =|C i,j | (5)
[0061] which is used to calculate the global centroid.
[0062] 2. Data aggregation algorithm (secure aggregation step of Step 3)
[0063] To protect the intermediate data S i =(s i,1 ,..., s i,k ), m i =(m i,1 ,..., m i,k ) of User DP i and the aggregation results s j , m jDuring the aggregation process, it cannot be obtained by other users and the server. A scheme based on homomorphic encryption supplemented by pseudo-random values is proposed to aggregate privacy data.
[0064] The algorithm is first performed by user DP i to randomly extract a seed b i , which is sent to the PRG (pseudo-random number generator), and then a mask matrix P of the same size as the privacy data matrix is generated i = PRG(b i ). Finally, the privacy data vector and the mask vector are added together to obtain
[0065] S' i = S i + P i . (6)
[0066] For m i , the same method is used for processing. User DP i uses the seed e i to generate the mask vector p i , which is added to m i to obtain
[0067] m' i = m i + p i . (7)
[0068] Then, DP i uses pk2 pk1 to encrypt S' i and m′ i into E2(S' i ) and E1(m' i ) respectively, and then sends them to TS1 and TS2 respectively. As Figure 3 shown.
[0069] After the sending is completed, the server TS1 and TS2 aggregate the received E2(S' i ), ∈(1,...,n) and E1(m' i ), ∈(1,...,n) to obtain
[0070] E2(S') = E2(S'1)·...·E2(S' n ) = E2(S'1 +... + S' n ) (8)
[0071] and
[0072] E1(m') = E1(m'1)·...·E1(m' n ) = E1(m'1 +... + m' n ) (9)
[0073] And send an aggregation request to all users.
[0074] After receiving the aggregation request from the server, each online user DP i Uploads the seed b i And the seed e i To the corresponding server via the pairwise key. The servers TS1 and TS2 respectively send the seeds b i , e i Into the PRG to recover the mask. Using the additive homomorphism of Paillier encryption, the mask is removed, and the servers TS1 and TS2 respectively obtain E2(S) and E1(m). Since the server TS1 does not have sk2 and the server TS2 does not have sk1, the servers TS1 and TS2 cannot obtain the intermediate values S and m.
[0075] 3. Global centroid secure calculation algorithm
[0076] Recalculating the centroid of each class according to the average value is a crucial step in the k-means algorithm. However, homomorphic encryption does not support division operations, and this method uses Algorithm 3 to solve this problem. The problem is as follows: The server TS1 has the encrypted data E2(S) = (E2(s1),..., E2(s k )), and the server TS2 has the encrypted data E1(m) = (E1(m1),..., E1(m k )). A privacy-preserving secure protocol is needed to perform the following functions
[0077]
[0078] Where j ∈ (1,..., k), j' ∈ (2,..., k), j > j', and ρ j,j' Is a pair of positive random numbers that are kept secret from both users and servers.
[0079] Without knowing the private data S, m, ρ, and the centroid c, the server obtains And ρ j,j' (c j - c j' ). The algorithm is as Figure 4 Shown.
[0080] Figure 4 Shows the centroid negotiation calculation process of the servers TS1 and TS2 for one centroid (the same for other centroids). First, TS1 and TS2 respectively generate random numbers. TS1 generates a random vector r1 and a random number r3, and TS2 generates a random number r2. TS1 calculates By the homomorphic property TS1 sends E2(r1s) to TS2, and TS2 decrypts it using the private key sk2 to obtain D2(E2(r1s)) = r1s. Similarly, TS2 also performs relevant calculations and sends E1(r2m) to TS1. TS1 decrypts it using the private key sk1 to obtain D1(E1(r2m)) = r2m. Then TS1 calculates r2r3m and sends it to TS2, and TS2 calculates to obtain TS2 encrypts and and then sends the encryption result and to TS1. After that, through the property of homomorphic encryption, the random vector r1 and the random number r3 can be eliminated, and finally E2(c) and E2(c 2 ) are obtained. Among them, E2(c T c) ← E2(sum(c 2 )).
[0081] The calculation of other centroids is the same as Figure 4 , and TS1 can calculate and E2(c j -c j' ) in the ciphertext state. Then add random numbers to hide the intermediate data and calculate and send it to TS2. After TS2 decrypts it, add the random number r2jj' to obtain r 1jj' r 2jj' (c j -c j' ). Among them, j ∈ (1,..., k), j' ∈ (2,..., k), j < j', r 1jj' , r 2jj' ≥ 0.
[0082] The server TS2 sends r 1jj' r 2jj' (c j -c j' ) to each user, where r1jj' is the random number generated by the server TS1, r 2jj' is the random number generated by the server TS2, and r 1jj' , r 2jj' ≥ 0.
[0083] 4. Output the final centroid algorithm
[0084] Repeatedly execute the above algorithms 1, 2, and 3 until there is almost no change or no change in the clustering process. At the end of each iteration, the user needs to compare the newly obtained clustering center with the clustering center of the previous iteration. User Dp iWhen executing the first stage, when the sum of the distances of the cluster centers does not change or changes very little
[0085]
[0086] where s i,j is the cluster center of the j-th class in the latest iteration, and s i,j ' is the cluster center of the j-th class in the previous iteration. User Dp i uploads the abort information. When the number of abort information received by the server is greater than η max the fourth stage is executed. First, TS1 and TS2 are generated respectively. TS1 generates a random vector r1 and a random number r3, and TS2 generates a random number r2. TS1 calculates By the homomorphic property TS1 sends E2(r1s) to TS2, and TS2 decrypts it using the private key sk2 to get D2(E2(r1s)) = r1s. Similarly, TS2 also performs related calculations sends E1(r2m) to TS1. TS1 decrypts it using the private key sk1 to get D1(E1(r2m)) = r2m. Then TS1 calculates r2r3m and sends it to TS2, and TS2 calculates to get TS2 sends to TS1. TS1 eliminates the perturbation to obtain the centroid c and sends it to TS2 and all user DPs.
[0087] The clustering effect can be seen in Figure 5 , Figure 5 By dimensionality reduction and visualization, the clustering performance of DTK-means and K-means on plaintext data was evaluated using the MNIST [dataset]. To speed up the processing, only the first 5000 samples of the MNIST dataset were selected. Principal component analysis (PCA) reduced the high-dimensional feature space of the MNIST dataset to two dimensions for visualization while retaining the variance as much as possible. The number of clusters was set to 10, corresponding to the ten different classes in the MNIST dataset. Scatter plots were used to visualize each sample in the two-dimensional PCA space. The points were colored according to their assigned clusters, providing an intuitive color-coded representation of the clustering results. The cluster centroids were highlighted with different markers to clearly represent the centroids of each cluster. Figure 5 The clustering results were compared. As can be seen from the right subplot, compared with the K-means algorithm, the DTK-means algorithm does not cause any samples to be misclassified. This result shows that the present invention can protect sensitive information without sacrificing clustering performance.
[0088] To further evaluate the performance of the method of the present invention (hereinafter referred to as: DTK-means), a comprehensive comparison of the performance of DTK-means with two existing privacy-preserving k-means schemes is carried out in the present invention. The following are two other privacy-preserving k-means schemes (details are as follows) for comparison:
[0089] 1) HE-based distributed method "Mutual Privacy Preserving k-Means Clustering in Social Participatory Sensing": This research focuses on keeping the clustering centers confidential from users, designs an algorithm for finding the nearest cluster, while the clustering centers are kept confidential from users. The cloud uses homomorphic encryption to update the clustering centers, hereinafter referred to as Wen Kai.
[0090] 2) Outsourcing method based on A-SS and PC-HE "Efficient privacy-preserving outsourced k-means clustering on distributed data": The implementation of this algorithm largely relies on secret sharing technology, supplemented by homomorphic encryption. This algorithm first allows each user to submit their data to the server, then the server performs k-means clustering on the combined data without compromising data privacy, and finally outputs the clustering centers. Hereinafter referred to as Wen Guo.
[0091] Table 2 lists the comparison results of DTK-means and these two schemes in terms of security features. The participating parties can be one server and at least two users (n≥2) or two servers and at least two users (n≥2). The present invention and Wen Kai belong to distributed algorithms, and the security features of the present invention are comprehensively superior to those of Wen Kai, and are equivalent to the security attributes of Wen Guo which belongs to an outsourcing algorithm.
[0092] Table 3 lists the computational cost and communication cost of DTK-means and other privacy-preserving k-means schemes for one iteration. The computational cost is represented by the number of executions of basic operations, including encryption (E), decryption (D), modular operation (P), and multiplication (×). Some items with lower costs are omitted in Table 3. It can be seen that the theoretical computational cost and communication cost of Wen Guo are much greater than those of the present invention and Wen Kai. The reason is that the k-means clustering of Wen Guo is carried out under ciphertext, and a series of ciphertext calculations such as secure distance, secure comparison, and secure minimum are required in the k-means clustering process. While the k-means calculations of the present invention and Wen Kai are carried out in the plaintext state without the need for complex ciphertext calculations. The computational cost of Wen Kai is not much different from that of the present invention. The reason is that the encryption scheme adopted by the present invention requires the user to perform one Paillier encryption, and the server needs to perform multiple Paillier decryptions. In Wen Kai, the intermediate data only needs to be encrypted once by the user and decrypted once by the server, but additional homomorphic addition calculations are required, and the computational amount is comparable to that of the DTK-means algorithm. The communication cost of the present invention's scheme is the lowest among the three schemes because it only requires the user to upload the intermediate data to the server, and the server performs the calculation, transmission, and return. In contrast, in Wen Kai, the user not only needs to transmit data to the server but also needs to transmit data between users to cover the data. Wen Guo requires the plaintext data to be completely uploaded after encryption, and multiple transmissions are required for each subsequent calculation. Therefore, the communication complexity of the present invention's scheme is less than that of Wen Kai and Wen Guo.
[0093] From the comparison results, the scheme of the present invention has the highest efficiency. The communication cost of the present invention is the lowest, and the scheme of the present invention is the only one that supports dynamically changing users during model training. Generally speaking, compared with other privacy-preserving k-means schemes, the present invention reaches a higher or the same security level and is superior to the other two schemes in terms of overall performance.
[0094]
[0095]
[0096] Table 2 Comparison of Different Privacy-Preserving Schemes
[0097]
[0098] Table 3 Comparison of Computational Complexity and Communication Complexity of Different Schemes
[0099] * m number of samples, u number of features, k number of clusters, n number of users, l bit length of the squared Euclidean distance
[0100] * N is the modulus of the Paillier cryptosystem, and L(N) is the bit length of N.
[0101] The present invention uses three datasets for computational efficiency and communication volume experiments. The first one is the famous Iris [dataset]. The second dataset is the Wine [dataset]. The dataset used in the last experiment is the KEGG metabolic reaction network [dataset]. To ensure the reliability of experimental comparison, all three schemes are run in the same experimental environment.
[0102] (1) Computational efficiency: Figure 6 As can be seen, for the present invention and the Wen Kai scheme, as the number of iteration rounds increases, the training time increases linearly. The scheme of the present invention shows the highest efficiency, and the Wen Guo scheme has the longest computational time. The experimental results are consistent with the theoretical analysis.
[0103] (2) Communication efficiency: Figure 7 It shows that for the present invention and the Wen Kai scheme, as the number of iteration rounds increases, the training time increases linearly. The present invention has the lowest communication volume, and the Wen Guo scheme has the highest communication volume.
[0104] The present invention conducts some experiments on three algorithms using different parameters to evaluate the performance of these algorithms. These parameters determine the computational and communication costs of the algorithms. The parameters include the number of sample points (m), the number of features (u), and the number of clusters (k). In the experiment, the present invention keeps one parameter unchanged and studies the influence of the changes of the other two parameters on the computational and communication costs. The KEGG metabolic reaction network dataset is used in the experiment. Figure 6 and Figure 7 show the running time and communication cost for one iteration using different parameters. The results show that the running time and communication cost increase linearly with the increase of any one parameter, but at different rates.
[0105] Figure 8 a and Figure 9 a shows that when the number of sample points (m) increases, the growth rate of the computational and communication costs of DTK-means and the Wen Kai scheme is slower than that of the Wen Guo scheme, and the communication cost of the latter almost remains unchanged. This is because the number of sample points (m) has less influence on the computational complexity of DTK-means and the Wen Kai scheme. In addition, the communication cost does not involve the number of sample points (m), so it almost remains unchanged. On the contrary, in the Wen Guo scheme, almost every item in the computational and communication costs includes the number of sample points (m). Therefore, the costs related to the Wen Guo scheme increase at the same rate as the number of sample points (m).
[0106] Figure 8 b and Figure 9b It shows that as the number of features (u) increases, the computational and communication costs of DTK-means and Wen Kai's scheme increase at the same rate as the number of features (u). This is because the computational and communication costs of DTK-means and Wen Kai's scheme are mainly determined by the number of features (u). The computational and communication costs of Wen Guo's scheme increase relatively slowly. In terms of computational cost, encryption (E) and decryption (D) account for a large part of the total cost. Given that m >> u, the impact of the number of features (u) on the overall computational cost is much smaller. Therefore, the increase in the number of features (u) will not significantly affect the computational cost of Wen Guo's scheme. Regarding communication cost, the transmission of ciphertext accounts for most of the cost. Since the number of features (u) is much smaller than the number of sample points (m), the increase in features has a negligible impact on communication cost. Therefore, the growth of communication cost is also relatively slow.
[0107] Figure 8 c and Figure 9 c shows that for DTK-means and Wen Kai's scheme, an increase in the number of clusters (k) leads to a faster growth in computational and communication costs. However, these costs are still significantly lower than those of Wen Guo's scheme. This is usually due to m >> k2. The growth rate of the computational and communication costs of Wen Guo is consistent with the increase in the number of clusters (k), which is consistent with the theoretical analysis.
[0108] From the above analysis, it can be concluded that the experimental results are consistent with the theoretical analysis. The increase in the number of features (u) has the greatest impact on DTK-means and Wen Kai's [8] scheme, while the increase in the number of clusters (k) has the most significant impact on Wen Guo's scheme.
[0109] In summary, compared with the current methods, the present invention has the following beneficial effects:
[0110] (1) DTK-means has higher security than traditional distributed algorithms, neither leaking users' private data nor revealing the global clustering centers and the number of classes to the server and users.
[0111] (2) The present invention can achieve clustering results as accurate as those of plaintext k-means without leaking any private data.
[0112] (3) Compared with those that also achieve a high level of security, the present invention has lower computational and communication costs.
[0113] The above-disclosed are only one or more preferred embodiments of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.
Claims
1. A privacy-preserving k-means method for distributed data, characterized in that: The following steps are involved: Step 1: The server negotiates the initial centroid, processes it and sends it to the user; Step 2: The user performs the nearest clustering calculation locally based on the encrypted global centroid or initial centroid of the previous round; Step 3: The local centroid and the number of samples in each cluster are scrambled and encrypted and then transmitted to different servers for secure aggregation; Step 4: The two servers interact to complete the calculation and encryption of the global centroid, and send it to the user for iteration; Step 5: When the sum of the distances between the local cluster centroid and the local cluster centroid of the previous iteration no longer changes or changes very little, the user sends a termination message; when the number of termination messages exceeds a certain threshold or reaches the iteration limit, the iteration stops; the server calculates the final centroid and sends it to all users.
2. The privacy-preserving k-means method for distributed data according to claim 1, characterized in that: The specific method of global centroid encryption in step 2 is as follows: The distance between the sample point and the centroid is calculated using the Euclidean distance. i With the centroid c j The distance d ij Subtract the sample point x i With the centroid c j' The distance d ij' , expand the formula into If d ij -d ij' <0, then x i Closer to c j , otherwise x i Closer to c j' , the user learns the centroids of all and c j -c j' , repeat k-1 times to find the nearest centroid of the sample, which is and c j -c j' Adding positive random perturbations does not affect the results.
3. The privacy-preserving k-means method for distributed data according to claim 2, characterized in that: In the process of secure aggregation in step 3, the user uploads the private data plus the mask generated by the seed, and then encrypts it using the public keys of server TS1 and server TS2 respectively, and sends it to TS2 and TS1. The user uses the pairwise key to send the seed to the corresponding server, and the server uses the random seed shared by the users to reconstruct the mask. Through the addition characteristics of the Paillier homomorphic algorithm, the private data is aggregated and the mask is eliminated.
4. The privacy-preserving k-means method for distributed data according to claim 3, characterized in that: Step 4 During the calculation of the global centroid, two servers TS1 and TS2, TS1 generates a random vector r1 and a random number r3, and TS2 generates a random number r2; TS1 calculates Where E2(·) is the ciphertext encrypted with the homomorphic public key generated by server TS2, s is the sample sum of each class; according to the homomorphic property TS1 sends E2(r1s) to TS2, which decrypts it using private key sk2 to obtain D2(E2(r1s))=r1s. Similarly, TS2 also performs related calculations. m is the number of samples in each class, E1(r2m) is sent to TS1; TS1 uses the private key sk1 to decrypt and obtain D1(E1(r2m))=r2m; then TS1 calculates r2r3m and sends it to TS2, TS2 calculates get TS2 Pair and Encrypt, and then the encrypted result and Send it to TS1, and then the random vector r1 and the random number r3 can be eliminated through the homomorphic encryption property, and finally E2(c) and E2(c 2 ), where E2(c T c)←E2(sum(c 2 )).; TS1 can be calculated in ciphertext state and E2(c j -c j' ); then add random numbers to hide the intermediate data and calculate Send to TS2, TS2 decrypts and adds random number r 2jj' get where j∈(1,...,k),j'∈(2,...,k),j<j',r 1jj' , r 2jj' ≥0; Server TS2 will r 1jj' r 2jj' (c j -c j' ) is sent to each user, where r1jj' is a random number generated by server TS1, r 2jj' is a random number generated by server TS2, and r 1jj' , r 2jj' ≥0.
5. The privacy-preserving k-means method for distributed data according to claim 4, characterized in that: In step 5, when the sum of the distances between the local cluster centroid and the local cluster centroid of the previous iteration does not change or changes very little, the formula is expressed as Where t is the number of iterations, c is the centroid, k is the number of clusters, and ε is the preset value.