Privacy-preserving outsourcing method based on K-means clustering based on linear transformation

By adopting the K-means clustering privacy protection outsourcing method based on linear transformation in the cloud computing environment, the problem of K-means clustering data privacy protection in cloud computing is solved, and efficient, secure, accurate and verifiable clustering results are achieved, computing overhead is reduced and fraudulent behavior of cloud servers is detected.

CN116170138BActive Publication Date: 2025-05-09HARBIN ENG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310131205.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-17
Publication Date
2025-05-09
Estimated Expiration
2043-02-17

AI Technical Summary

Technical Problem

It is difficult for the existing technology to achieve privacy protection of K-means clustering in a cloud computing environment. It not only ensures the security and accuracy of data, but also avoids excessive computing overhead and fraudulent behaviors of cloud servers.

Method used

The K-means clustering privacy protection outsourcing method based on linear transformation is adopted to ensure the privacy and integrity of the data during transmission and processing through steps such as key generation, data encryption, cloud computing, result verification and data decryption. The method uses permutation, orthogonal transformation, and random number operations to encrypt the data and detects spoofing behavior of the cloud server through verification steps.

Benefits of technology

It realizes 100% accuracy, security, efficiency and verifiability of K-means clustering data in a cloud computing environment, reduces computing overhead, and can effectively detect fraudulent behavior of cloud servers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116170138B_ABST
    Figure CN116170138B_ABST
Patent Text Reader

Abstract

The present invention discloses a privacy-preserving outsourcing method for K-means clustering based on linear transformation. The method comprises the following steps: 1. The data owner randomly generates a key using a key generation algorithm; 2. The data owner permutes the index order and attribute order corresponding to each record in D to obtain D', and the data owner uses the key to transform D' into D” and sends it to the cloud; 3. The cloud executes the K-means mean clustering task and returns the K-means clustering result and the centroid of each cluster to the data owner; 4. The data owner verifies the clustering result; 5. After verifying the clustering result returned by the cloud successfully, the data owner restores the index order corresponding to each record in D” through π1 to obtain the true K-means clustering result. This method can achieve 100% accuracy, security, efficiency, and verifiability through an efficient linear transformation technique.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of privacy protection, and relates to a clustering privacy protection method, and specifically to a privacy protection outsourcing method of K-means clustering based on linear transformation. Background Art

[0002] Cloud computing is a user-centric computing service. In recent years, with the continuous development of cloud computing technology and the increasing number of cloud service providers, more and more users and enterprises have chosen cloud computing services. As one of the important applications of cloud computing, outsourcing computing technology has also become a hot topic of concern. In a cloud computing environment, user terminals with limited computing power and storage resources can outsource complex computing tasks to cloud servers for processing, and enjoy the endless computing and storage resources provided by the cloud computing platform in a pay-as-you-go manner. This new computing model reduces the burden of personal computing and avoids users' large investments in local hardware and software and maintenance. Users can remotely store data in the cloud for processing and enjoy high-quality applications and services in the cloud on demand. With the rapid development of cloud computing technology, outsourcing computing makes it possible to accelerate cluster analysis.

[0003] Clustering is one of the main tasks in exploratory data mining and statistical data analysis, and is widely used in healthcare, social networks, image analysis, pattern recognition and other fields. At the same time, with the rapid growth of big data, data mining and analysis are also facing challenges in terms of the scale, variety and speed of data clustering. In order to efficiently manage large-scale datasets and support their clustering, public cloud infrastructure is a major consideration in terms of performance and economy. However, the use of public cloud services inevitably introduces privacy issues. This is not only because many of the data involved in data mining applications are sensitive in nature, such as personal health information, localized data, financial data, etc., but also because the public cloud is an open environment operated by an external third party. For example, a promising trend in predicting personal disease risk is to cluster the health records of existing patients. Therefore, when outsourcing sensitive datasets to public clouds for clustering, appropriate privacy protection mechanisms must be given.

[0004] At present, privacy-preserving computing technologies for K-means clustering are mainly divided into two categories. One category is represented by differential privacy protection methods, which are characterized by high efficiency, but poor security and low accuracy. The other category is represented by homomorphic encryption, which has the characteristics of strong security, high accuracy and low efficiency. In addition, current research work basically ignores the dishonesty of the cloud and does not take verification measures for the clustering results returned by the cloud.

[0005] Therefore, it is necessary to invent a privacy-preserving outsourcing method for K-means clustering that can simultaneously achieve security, efficiency, accuracy, and verifiability. Summary of the invention

[0006] The purpose of the present invention is to provide a privacy protection outsourcing method based on linear transformation K-means clustering, which saves computing resources, has low communication costs, has high accuracy, and can detect cloud deception behavior. The method can achieve 100% accuracy, security, efficiency and verifiability through efficient linear transformation technology.

[0007] The objective of the present invention is achieved through the following technical solutions:

[0008] A privacy protection outsourcing method based on linear transformation K-means clustering includes the following steps:

[0009] Step 1: Key generation: The data owner has a large-scale original data set D, and the number of data samples is n, and the sample dimension is m. The data owner uses the key generation algorithm to randomly generate a key sk = (π 1 ,π 2 ,Q,L,c), where π 1 , π 2 are two random permutation functions, Q is an orthogonal real matrix generated by Householder transformation, L is a real matrix with the same row vectors, and c is a random real number (c>0);

[0010] Step 2: Data encryption: After generating the key sk, the data owner first uses π 1 , π 2 The index order and attribute order corresponding to each record in the original data set D are permuted to obtain the permuted data set D'. Next, the data owner uses the key (Q, L, c) to convert D' into an encrypted data set D" and sends D" to the cloud.

[0011] Step 3, cloud computing: After receiving the encrypted data set D", the cloud performs the K-means mean clustering task. After the computing task is completed, the cloud returns the K-means clustering results and the centroid of each cluster to the data owner;

[0012] Step 4: Result verification: The data owner verifies the clustering results returned from the cloud;

[0013] Step 5: Data decryption: After verifying the clustering results returned by the cloud, the data owner can use π 1 Restore the index order corresponding to each record in the encrypted data set D" to obtain the true K-means clustering result.

[0014] Compared with the prior art, the present invention has the following advantages:

[0015] 1. The present invention can achieve data privacy protection during the K-means clustering task on the cloud server, with low computational overhead, without affecting the accuracy of clustering, and can detect cloud deception with a probability close to 1.

[0016] 2. The present invention has accuracy and privacy: The encryption method designed by the present invention first uses two permutations to disrupt the order of each data index and the order of attributes of the data set, and then ensures the invariance of the distance while encrypting the data through orthogonal transformation and translation transformation, and finally ensures the privacy of the distance by multiplying the random number operation. Therefore, the clustering result of the encrypted data is completely consistent with the clustering result of the original data, and the cloud cannot obtain any sensitive information of the data owner during the data input and output process and the calculation process.

[0017] 3. The present invention is highly efficient: By using the method designed by the present invention, the computational overhead is reduced from O(ktmn) to (kmn), where k represents the number of clusters, t represents the number of iterations, n represents the number of samples, and m represents the dimension of each sample. Among them, the computational overhead of the encryption stage is O(mn), and the stage with the largest computational overhead is the verification stage - O(kmn), so the present invention proposes three improved verification methods according to different situations. A large number of experiments have also proved the high efficiency of the method designed by the present invention.

[0018] 4. The present invention is verifiable: In order for the data owner to detect the cheating behavior of the cloud with a probability close to 1, three points need to be met: (1) the clustering has converged, (2) the centroid of each cluster is correct, and (3) each data is assigned to the nearest cluster, so it can be satisfied by performing one iteration. The present invention can detect the cheating behavior of the cloud with a probability close to 1. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 The system model diagram of the privacy protection outsourcing method based on K-means clustering based on linear transformation.

[0020] Figure 2 The overall flow chart of the privacy protection outsourcing method of K-means clustering based on linear transformation.

[0021] Figure 3 This is a diagram showing that the data in each cluster is concentrated in the improved verification scheme 1.

[0022] Figure 4 This is a diagram showing the data dispersion in each cluster in the improved verification scheme 1.

[0023] Figure 5 This is a diagram showing the data dispersion in each cluster in the improved verification scheme 1.

[0024] Figure 6Comparison of time costs at each stage when dimension m = 100 and number of clusters k = 10.

[0025] Figure 7 The time cost comparison between using the method of the present invention and not using the method of the present invention when the dimension m=100 and the number of clusters k=10 is shown in FIG.

[0026] Figure 8 The comparison of the rates of the data owners of the method of the present invention when the dimensions are different.

[0027] Fig. 9 The comparison of the rates of the data owners of the method of the present invention when the number of clusters is different.

[0028] Fig.10 Comparison of the cloud speed of the method of the present invention when the dimensions are different.

[0029] Fig.11 The comparison of the cloud rate of the method of the present invention when the number of clusters is different.

[0030] Fig.12 Comparison of the time cost between the verification method and the improved verification scheme 1.

[0031] Fig.13 Comparison of the time cost of the verification method with the improved verification schemes 2 and 3.

[0032] Fig.14 Comparison of the time cost of the data owner between the method proposed in the present invention and the known method PIPC

[0033] Fig.15 Comparison of the time cost in the cloud between the method proposed in the present invention and the conventional method PIPC. DETAILED DESCRIPTION

[0034] The technical solution of the present invention is further described below in conjunction with the accompanying drawings, but is not limited thereto. Any modification or equivalent replacement of the technical solution of the present invention without departing from the spirit and scope of the technical solution of the present invention should be included in the protection scope of the present invention.

[0035] Figure 1The figure shows the system model diagram of the privacy protection outsourcing method of K-means clustering based on linear transformation, which includes two parts: the data owner with limited computing resources and the cloud server with abundant computing resources. The purpose of the data owner is to realize the K-means clustering task of computing large-scale data sets with the help of the cloud server. In order to protect the privacy of the original data set, the data owner first generates a key, encrypts the original data set with the key, and then sends the encrypted data set to the cloud. After receiving the sent data set, the cloud performs the clustering task and returns the calculation result to the data owner. After receiving the result returned by the cloud, the data owner first verifies the correctness of the returned result. If it is incorrect, it is invalid; if it is correct, the user end decrypts it to obtain the final result.

[0036] Figure 2 The figure shows the overall flow chart of the privacy protection outsourcing method of K-means clustering based on linear transformation, which includes the following steps:

[0037] Step 1: Key generation: The data owner has a large-scale original data set D, and the number of data samples is n, and the sample dimension is m. The data owner uses the key generation algorithm to randomly generate a key sk = (π 1 ,π 2 ,Q,L,c). The specific steps include:

[0038] Step 1.1: The data owner generates π based on the number of data samples n and the sample dimension m 1 , π 2 , the specific form is as follows:

[0039]

[0040] Step 1.2: The data owner generates an m-dimensional orthogonal real number matrix Q through the Householder transformation in the QR decomposition, which specifically includes the following steps:

[0041] Step 1.2.1, generate m random real numbers q 1 ,q 2 ,…,q m ;

[0042] Step 1.2.2: Construct a column vector q=(q 1 ,q 2 ,…,q m ) T ;

[0043] Step 1.2.3, construct the Householder transformation matrix, the formula is as follows:

[0044]

[0045] Where I is the identity matrix;

[0046] Step 1.3: Generate an n×m real number matrix L. The specific steps are as follows:

[0047] Step 1.3.1, generate m random real numbers l 1 , l 2 ,…,l m ;

[0048] Step 1.3.2: Construct an m-dimensional column vector l = (l 1 , l 2 ,…,l m ) T ;

[0049] Step 1.3.3, construct an n×m matrix L = (l, l, ..., l) T ;

[0050] Step 1.4, generate a random real number c (c>0).

[0051] Step 2: Data encryption:

[0052] Step 2.1: The data owner first uses π 1 , π 2 The index order and attribute order corresponding to each record in the original data set D are permuted to obtain the permuted data set D'. The formula is as follows:

[0053]

[0054] Step 2.2: The data owner uses the key (Q, L, c) to encrypt D' into an encrypted data set D'. The specific formula is as follows:

[0055] D" = c x (D' x QL).

[0056] Step 3, cloud computing: After receiving the encrypted data set D", the cloud performs the K-means mean clustering task. After the computing task is completed, the cloud returns the K-means clustering results and the centroid of each cluster to the data owner.

[0057] Step 4: Result verification: After receiving the clustering results returned from the cloud, the data owner verifies the returned results. The specific steps include:

[0058] Step 4.1: The data owner calculates whether the centroid of each cluster is consistent with the centroid returned by the cloud based on the clustering results returned by the cloud;

[0059] Step 4.2: If the returned results are consistent, the data owner continues to calculate the distance between each data point and all centroids to verify whether each data point is assigned to the nearest cluster; if the returned results are inconsistent, it means that the clustering results returned by the cloud are wrong;

[0060] Step 4.3: If the verification passes, return TRUE; otherwise, return FALSE.

[0061] Step 5: Data decryption: After verifying the clustering results returned by the cloud, the data owner can use π 1 Restore the index order corresponding to each record in the encrypted data set D" to obtain the true K-means clustering result. The formula is as follows:

[0062]

[0063] In addition, in order to improve the efficiency of verification, the present invention also provides three improved result verification schemes for verifying the correctness of the returned results. The specific verification schemes are as follows:

[0064] Verification solution 1:

[0065] The geometry-based verification scheme is applicable to two-dimensional or three-dimensional data. The specific steps are as follows:

[0066] (1) The data owner calculates whether the centroid of each cluster is consistent with the centroid returned by the cloud based on the clustering results returned by the cloud.

[0067] (2) If the returned results are consistent, continue to verify whether each data point is assigned to the nearest cluster. The clustering results are displayed on a two-dimensional plane. The nearest cluster center is obtained by calculating the distance between each cluster center and the centroid of other clusters. A circle is drawn with half the distance between the cluster center and the nearest cluster center as the radius and the cluster center as the center. If the data points of each cluster are relatively concentrated, then these circles will contain all the points in the encrypted data set D”, such as Figure 3 As shown. If the points inside each circle are consistent with the points in the clusters in the results returned by the cloud, it means that each point is assigned to the nearest cluster. If the data points of each cluster are scattered, there will be some data points scattered outside the circle, such as Figure 4 As shown. However, if the points inside the circle correspond to the results returned by the cloud, it means that these data points inside the circle are assigned to the nearest cluster. Continue to verify whether the data points outside the circle are assigned to the nearest cluster by finding the second closest centroid to each centroid and drawing a circle with a radius of half the distance between the two centroids.

[0068] (3) For data points outside the circle, after the second circle is drawn, see if they are inside the corresponding circle. If they are inside the corresponding circle, it means that the data point belongs to this cluster, and check whether the result is consistent with the result returned by the cloud. If the data point falls at the intersection of two or more circles, calculate the distance between these data points and the centroid of the intersection circle respectively, find the cluster with the closest distance, and check the result returned by the cloud, such as Figure 5 As shown. Keep expanding the radius and repeat the process of drawing circles until all data points have been checked. If there are still a small number of data points outside all circles at the end, then for each of these data points, calculate their distance to the k centroids to find the nearest cluster, and check the results returned by the cloud against these data points. If the results returned by the cloud are consistent with the verified results, the results returned by the cloud are correct.

[0069] Verification solution 2:

[0070] In addition to returning the clustering results and the centroid of each cluster, the cloud also runs the MDS dimensionality reduction algorithm on the encrypted data set D” and returns the reduced matrix Z of size m×2. If the data point has a large number of dimensions, the specific verification steps are as follows:

[0071] (1) The data owner calculates whether the centroid of each cluster is consistent with the centroid returned by the cloud based on the clustering results returned by the cloud;

[0072] (2) If the returned results are consistent, the data owner continues to calculate the distance between each data point and all centroids in the reduced-dimensional data set Z through the MDS algorithm to verify whether each data point is assigned to the nearest cluster; if the returned results are inconsistent, it means that the clustering results returned by the cloud are wrong;

[0073] (3) If the verification passes, return TRUE; otherwise, return FALSE.

[0074] Verification solution 3:

[0075] If the data owner has used some preprocessing methods on the data set D in advance to achieve good clustering results, the error sum of squares SSE and silhouette coefficient method can be used to verify whether the cloud performs the task honestly. The specific steps are as follows:

[0076] (1) In step 3, during the cloud computing process, the cloud calculates the average distance a of each data point to other data points in its cluster and the minimum value b of the average distance of the data point to all data points in other clusters based on the clustering results and returns them to the data owner;

[0077] (2) The data owner calculates the silhouette coefficient S and SSE of each data based on the returned results. The formula is as follows:

[0078]

[0079]

[0080] Among them, k represents the number of clusters, C i (i=1,2,...,k) represents the i-th cluster, x represents all samples, u i (i=1,2,...,k) represents the i-th centroid.

[0081] (3) If s ≥ 0 and SSE → 0, it means that the cloud returns the correct result.

[0082] Table 1 gives the experimental results of the time cost of each stage when the dimension m=100 and the number of clusters k=10.

[0083] Table 1

[0084] Number of clusters k Dimension m Data <![CDATA[t KeyGen ]]> <![CDATA[t Encrypt ]]> <![CDATA[t Decrypt ]]> <![CDATA[t Verify ]]> <![CDATA[t Do ]]> 10 100 5000 0.000521 0.011504 0.000065 0.024359 0.036449 10 100 10000 0.001181 0.020890 0.000105 0.043012 0.065188 10 100 15000 0.001307 0.033188 0.000120 0.071761 0.106376 10 100 20000 0.001838 0.042524 0.000290 0.10055 0.145202 10 100 25000 0.002265 0.050958 0.000327 0.11527 0.168820

[0085] In the experiment, for the convenience of description, we use t KeyGen ,t Encrypt ,t cloud ,t Verify ,t Decrypt It represents the time cost of five stages: key generation, data encryption, cloud computing, result verification, and data decryption. original and t Do They represent the time cost of performing the K-means clustering task without using the present invention and the time cost of performing the clustering task using the method proposed by the present invention. Specifically, t Do =t KeyGen +t Encrypt +t Verify +t Decrypt Therefore, t original / t Do reflects the rate of the data owner, t original / t cloud reflects the rate in the cloud. If the ratio t original / t cloud When the value is around 1, it means that the method proposed in this invention will not add extra computing burden to the cloud. In the experiment, we selected dimension m = 100 and cluster number k = 10 as the benchmark experiment and conducted 5 sets of experiments. The size of data n started from 5000 and increased to 25000 in units of 5000. The experimental results are all in seconds. In order to better display the experimental results in the table, we visualize the experimental data, as shown in the figure. Figure 6 As shown. Figure 6 It can be seen that t KeyGen and t Encrypt The smallest and slow growing, t VerifyThe largest and fastest growing, which indicates that the verification phase has the highest cost.

[0086] Table 2 shows the experimental results of the time cost of using the method of the present invention and not using the method of the present invention when the dimension m=100 and the number of clusters k=10.

[0087] Table 2

[0088] Number of clusters k Dimension m Data <![CDATA[t Do ]]> <![CDATA[t cloud ]]> <![CDATA[t original ]]> <![CDATA[t original / t Do ]]> <![CDATA[t original / t cloud ]]> 10 100 5000 0.036449 1.07355 1.388768 38.102 1.295 10 100 10000 0.065188 3.146684 3.573628 54.820 1.136 10 100 15000 0.106376 7.324363 8.200740 77.092 1.120 10 100 20000 0.145202 18.437218 16.457947 113.345 0.893 10 100 25000 0.168820 20.752455 20.219925 20.219925 0.974

[0089] As the data n increases from 5000 to 25000, t original From 1.39 seconds to 20.22 seconds, t Do never exceeds 0.17 seconds, and t Do The ratio of increased from the initial 38 to 119, proving that the method proposed in this invention is applicable to large-scale data sets. In order to better display the experimental results in the table, we visualize the experimental data, as shown in Figure 7 As shown. Figure 7 It can be seen that as the data size n increases, t Do Slow growth, t original Rapidly growing, and t Do Always much smaller than t original , which shows that the method proposed in this invention can indeed reduce the computing overhead of the data owner.

[0090] Figure 8 and Fig. 9 A comparison of the data owner rates of the method of the present invention is given when the dimensionality and the number of clusters are different.

[0091] For different dimensions m and cluster numbers k, as the data size n increases, the speedup ratio of the data owner is original / t Do This shows that the present invention is effective for different dimensions, cluster numbers and sample numbers, and can significantly reduce the computational cost of the data owner.

[0092] Fig.10 and Fig.11 The speed comparison of the cloud side of the method of the present invention is given when the dimension and the number of clusters are different. It can be seen from the figure that the speedup ratio of the cloud side is t original / t cloud It is close to 1, which proves that the present invention does not increase the computational overhead of the cloud server.

[0093] Fig.12 and Fig.13The time cost comparison of the verification method and the improved verification schemes 1, 2, and 3 is given. As can be seen from the figure, the time cost of the three improved verification methods proposed by the present invention is less than the time cost of the original verification algorithm. This shows that the three improved verification methods proposed by the present invention are more efficient.

[0094] Fig.14 and Fig.15 A comparison of the time cost of the data owner of the method proposed in the present invention and the known method PIPC is given. It can be clearly seen from the figure that the time cost of the method proposed in the present invention is less than that of the known method PIPC (Pipc: Privacy and integrity-preserving clustering analysis for load profiling in smart grids. IEEE Internet of Things Journal, 9(13): 10851–10861, 2022.) both in the data owner and in the cloud.

Claims

1. A privacy protection outsourcing method based on linear transformation K-means clustering, characterized by The method comprises the following steps: Step 1, key generation: The data owner has a large-scale original data set D, and the number of data samples is n, and the sample dimension is m. The data owner uses the key generation algorithm to randomly generate a key sk = (π1, π2, Q, L, c), where π1 and π2 are two random permutation functions, Q is an orthogonal real number matrix generated by Householder transformation, L is a real number matrix with the same row vector, c is a random real number, c>0; Step 2, data encryption: After generating the key sk, the data owner first uses π1 and π2 to permute the index order and attribute order corresponding to each record in the original data set D to obtain the permuted data set D'. Next, the data owner uses the key (Q, L, c) to convert D' into an encrypted data set D" and sends D" to the cloud; Step 3, cloud computing: After receiving the encrypted data set D", the cloud performs the K-means mean clustering task. After the computing task is completed, the cloud returns the K-means clustering results and the centroid of each cluster to the data owner; Step 4: Result verification: The data owner verifies the clustering results returned from the cloud; Step 5, data decryption: After successfully verifying the clustering results returned by the cloud, the data owner restores the index order corresponding to each record in the encrypted data set D” through π1 to obtain the true K-means clustering results.

2. The privacy protection outsourcing method based on K-means clustering based on linear transformation according to claim 1 is characterized in that The step 1 specifically includes the following steps: Step 1.1, the data owner generates π1 and π2 according to the number of data samples n and the sample dimension m. The specific form is as follows: Step 1.2, the data owner generates an m-dimensional orthogonal real number square matrix Q through the Householder transformation in QR decomposition; Step 1.3, generate an n×m real number matrix L; Step 1.4: Generate a random real number c.

3. The privacy protection outsourcing method based on K-means clustering based on linear transformation according to claim 2 is characterized in that The step 1.2 specifically includes the following steps: Step 1.2.

1. Generate m random real numbers q1, q2, ..., q m ; Step 1.2.2: Construct a column vector q = (q1, q2, ..., q m ) T ; Step 1.2.3, construct the Householder transformation matrix, the formula is as follows: Where I is the identity matrix.

4. The privacy protection outsourcing method based on K-means clustering based on linear transformation according to claim 2 is characterized in that The specific steps of step 1.3 are as follows: Step 1.3.

1. Generate m random real numbers l1, l2, ..., l m ; Step 1.3.2: Construct an m-dimensional column vector l = (l1, l2, ..., l m ) T ; Step 1.3.3, construct an n×m matrix L = (l, l, ..., l) T .

5. The privacy protection outsourcing method based on linear transformation K-means clustering according to claim 1 is characterized in that In step 2, the formula for permuting the data set D' is as follows: Where, i = 1, 2, ..., n, j = 1, 2, ..., m; The formula for encrypting the dataset D" is as follows: D" = c x (D' x QL).

6. The privacy protection outsourcing method based on K-means clustering based on linear transformation according to claim 1 is characterized in that The step 4 specifically includes the following steps: Step 4.1: The data owner calculates whether the centroid of each cluster is consistent with the centroid returned by the cloud based on the clustering results returned by the cloud; Step 4.2: If the returned results are consistent, the data owner continues to calculate the distance between each data point and all centroids to verify whether each data point is assigned to the nearest cluster; If the returned results are inconsistent, it means that the clustering results returned by the cloud are wrong; Step 4.3: If the verification passes, return TRUE; otherwise, return FALSE.

7. The privacy protection outsourcing method based on K-means clustering based on linear transformation according to claim 1 is characterized in that In step 4, if the data point is two-dimensional data, the following steps are specifically included: Step 4.1: The data owner calculates whether the centroid of each cluster is consistent with the centroid returned by the cloud based on the clustering results returned by the cloud; Step 4.2: If the returned results are consistent, continue to verify whether each data point is assigned to the nearest cluster; The clustering results are displayed on a two-dimensional plane. The nearest cluster center is obtained by calculating the distance between each cluster center and the centroid of other clusters. A circle is drawn with the cluster center as the center, with half the distance between the cluster center and the nearest cluster center as the radius. If the data points of each cluster are relatively concentrated, then these circles will contain all the points in the encrypted data set D”; if the points in each circle are consistent with the points in the cluster in the result returned by the cloud, it means that each point is assigned to the nearest cluster; if the data points of each cluster are scattered, then there will be some data points scattered outside the circle, but for the points in the circle, if they correspond to the results returned by the cloud, it means that these data points in the circle are assigned to the nearest cluster; continue to verify whether the data points outside the circle are assigned to the nearest cluster, by finding the second closest centroid to each centroid and drawing a circle with half the distance between the two centroids as the radius; Step 4.3: For data points outside the circle, after the second circle is drawn, check whether they are inside the corresponding circle. If they are inside the corresponding circle, it means that the data point belongs to this cluster, and check whether the result is consistent with the result returned by the cloud; if the data point falls at the intersection of two or more circles, calculate the distance between these data points and the centroid of the intersection circle respectively, find the cluster with the closest distance, and check the result returned by the cloud; keep expanding the radius and repeat the process of drawing circles until all data points have been checked; If there are still a small number of data points outside all circles at the end, then for each of these data points, calculate their distances to the k centroids respectively to find the nearest cluster, and check the results returned by the cloud against these data points; if the results returned by the cloud are consistent with the verified results, the results returned by the cloud are correct.

8. The privacy protection outsourcing method based on K-means clustering based on linear transformation according to claim 1 is characterized in that In step 4, if the data point dimensions are large, the following steps are specifically included: Step 4.1: The data owner calculates whether the centroid of each cluster is consistent with the centroid returned by the cloud based on the clustering results returned by the cloud; Step 4.2: If the returned results are consistent, the data owner continues to calculate the distance between each data point and all centroids in the reduced-dimensional data set Z through the MDS algorithm to verify whether each data point is assigned to the nearest cluster. If the returned results are inconsistent, it means that the clustering results returned by the cloud are wrong; Step 4.3: If the verification passes, return TRUE; otherwise, return FALSE.

9. The privacy protection outsourcing method based on K-means clustering based on linear transformation according to claim 1, characterized in that In step 4, if the data owner uses a data preprocessing method before performing K-means clustering to achieve a good clustering effect, the following steps are specifically included: Step 4.1: In the cloud computing process of step 3, the cloud calculates the average distance a of each data point to other data points in its cluster and the minimum value b of the average distance of the data point to all data points in other clusters according to the clustering results and returns them to the data owner; Step 4.2: The data owner calculates the silhouette coefficient S and SSE of each data according to the returned results; Step 4.3: If s ≥ 0 and SSE → 0, it means that the cloud returns the correct result.

10. The privacy protection outsourcing method based on linear transformation K-means clustering according to claim 9, characterized in that The silhouette coefficient S and SSE are formulated as follows: Among them, k represents the number of clusters, C i represents the i-th cluster, x represents all samples, u i represents the i-th centroid, i=1,2,...,k.

Citation Information

Patent Citations

  • K-means clustering method and system with privacy protection function

    CN107145791A

  • Verifiable multi-party k-means federated learning method with privacy protection function

    CN112487481A