Efficient privacy protection KNN classification method

By adopting Euclidean triplets and divide-and-control bubble protocols in the privacy protection KNN classification method, the challenges of existing methods in communication overhead and computing efficiency are solved, and efficient and secure KNN classification calculations are achieved.

CN120217140APending Publication Date: 2025-06-27FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510271926.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing privacy protection KNN classification methods have great challenges in communication overhead and computing efficiency, especially on large-scale data sets, which lead to a significant reduction in computing efficiency and a higher number of communication rounds.

Method used

The zero-online communication distance calculation method based on Euclid triple and the efficient nearest neighbor search method based on the divide-and-control bubble bubbling protocol are adopted to reduce communication overhead and improve computing efficiency.

Benefits of technology

It significantly reduces online communication overhead, improves computing efficiency, and enables KNN classification to be efficiently executed under the protection of multi-party privacy, and is suitable for large-scale data sets and distributed scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217140A_ABST
    Figure CN120217140A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of cyberspace security, provides an efficient privacy protection KNN classification method, and aims to solve the problem that an existing privacy protection KNN classification method is high in communication overhead in an online stage. The communication overhead is optimized by adopting the following two key technologies: (1) a new Euclidean triple is designed, so that online communication is not needed when the safe Euclidean square distance is calculated; and (2) a divide-and-conquer bubbling protocol is provided, and the number of communication rounds required for selecting k nearest neighbors in the KNN classification process is greatly reduced. The framework supports cooperation of a plurality of data owners, ensures data privacy security, and is suitable for large-scale data sets and real-time application scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of privacy computing and machine learning, and particularly relates to an efficient privacy-preserving KNN (K-Nearest Neighbors) classification method, which is applicable to scenarios such as medical diagnosis and recommendation systems that require cross-institutional collaborative training and are data-sensitive. Background Art

[0002] In recent years, with the rapid development of big data and artificial intelligence technologies, the K-Nearest Neighbors (KNN) classification method [1][2] has been widely used in many fields, including medical diagnosis, financial risk control, recommendation systems, and bioinformatics, due to its intuitiveness, non-parametric nature, and high classification accuracy. However, the high accuracy of KNN classification depends on a large amount of high-quality data. In practical applications, this data is often distributed among different institutions and organizations and cannot be directly shared due to privacy protection regulations (such as GDPR [3], etc.) and security policies, thus restricting the performance and applicability of the KNN classification model.

[0003] To solve this problem, in recent years, researchers have proposed various privacy-preserving KNN classification methods [4][5][6][7][8], mainly based on secure multi-party computing technologies such as homomorphic encryption, secret sharing, and garbled circuits, enabling multiple data owners to jointly perform KNN classification without revealing the original data. However, the existing privacy-preserving KNN classification frameworks still face significant challenges in terms of communication overhead and computational efficiency, mainly reflected in the following two aspects:

[0004] 1. The communication overhead of secure Euclidean distance calculation is too large:

[0005] In the process of KNN classification, calculating the distance between the query sample and all samples in the database is a core step. Existing methods usually use secure multi-party computing protocols for privacy-preserving Euclidean distance calculation, such as secure multiplication protocols based on secret sharing or addition and multiplication operations based on homomorphic encryption. However, these methods require a large amount of online communication for each distance calculation, such as the exchange of secure shared intermediate variables or the homomorphic calculation of encrypted data, resulting in a significant reduction in computational efficiency, especially on large-scale datasets.

[0006] 2. The number of communication rounds for nearest neighbor search is too many:

[0007] When selecting the nearest k neighbors, existing methods usually adopt a sequential comparison selection algorithm. For example, some methods use a sequential bubble protocol to complete k-nearest neighbor search within O(kn) communication rounds. However, since these methods need to compare and exchange data one by one, the number of communication rounds is high, and their efficiency is limited in a wide area network environment, unable to meet the privacy computing requirements of low latency and high throughput.

[0008] References:

[0009] [1] Mucherino A, Papajorgji P J, Pardalos P M, et al. K-nearest neighbor classification[J]. Data mining in agriculture, 2009: 83-106;

[0010] [2] Kataria A, Singh M D. A review of data classification using k-nearest neighbour algorithm[J]. International Journal of Emerging Technology and Advanced Engineering, 2013, 3(6): 354-360;

[0011] [3] Voigt P, Von dem Bussche A. The eu general data protection regulation(gdpr)[J]. A Practical Guide, 1st Ed., Cham: Springer International Publishing, 2017, 10(3152676): 10-5555;

[0012] [4] Sun M, Yang R. An efficient secure k nearest neighbor classification protocol with high-dimensional features[J]. International Journal of Intelligent Systems, 2020, 35(11): 1791-1813;

[0013] [5]Li Z, Wang H, Zhang S, et al. SecKNN: FSS-Based Secure Multi-Party KNN Classification Under General Distance Functions[J]. IEEE Transactions on Information Forensics and Security, 2023, 19: 1326-1341;

[0014] [6]Wu W, Parampalli U, Liu J, et al. Privacy preserving k-nearest neighbor classification over encrypted database in outsourced cloud environments[J]. World Wide Web, 2019, 22: 101-123;

[0015] [7]Wong W K, Cheung D W, Kao B, et al. Secure kNN computation on encrypted databases[C] / / Proceedings of the 2009 ACM SIGMOD International Conference on Management of data. 2009: 139-152;

[0016] [8]Cui N, Yang X, Wang B, et al. SVkNN: Efficient secure and verifiable k-nearest neighbor query on the cloud platform[C] / / 2020 IEEE 36th International Conference on Data Engineering(ICDE). IEEE, 2020: 253-264。 Summary of the Invention

[0017] The objective of the present invention is to provide an efficient privacy - preserving KNN classification method that is efficient, secure, and has low communication overhead. This method can ensure that multiple data owners jointly complete the KNN classification calculation without revealing data privacy. The present invention proposes a brand - new optimization scheme, including a zero - online - communication distance calculation method based on Euclidean triples and an efficient nearest - neighbor search method based on the divide - and - conquer bubble protocol, thereby reducing communication overhead, improving calculation efficiency, and enabling the KNN classification to be efficiently executed under multi - party privacy protection.

[0018] An efficient privacy - preserving KNN (K - Nearest Neighbor) classification method provided by the present invention is as follows:

[0019] (1) Offline data sharing stage

[0020] The data owner encrypts its local data using a secret - sharing protocol and distributes the encrypted data to two or more computing parties. After this stage, the data of the data owner will be stored in encrypted form among the computing parties, thus ensuring data security.

[0021] (2) Online KNN classification calculation stage

[0022] (2.1) After step (1) is completed, use the Euclidean triple method to achieve secure Euclidean square - distance calculation without online communication.

[0023] (2.2) Use the divide - and - conquer bubble protocol to reduce the number of communication rounds required for nearest - neighbor selection from O(kn) to O(k log n), where n is the number of samples in the dataset and k is the number of selected nearest neighbors; specifically: in each iteration, first divide the sample distances into two equal - length lists (if odd, perform an additional comparison and exchange), and then perform secure comparison and exchange operations based on secret sharing on the elements in the two lists to ensure that smaller elements are exchanged to the previous list; then repeat the above process for the previous list, and finally, the nearest - neighbor selection can be completed within O(k log n) communication rounds.

[0024] (3) Result reconstruction stage

[0025] The computing party returns the encrypted classification result to the user, and the user recovers the final KNN classification result through the secret - sharing protocol. In this process, the computing party will not obtain the plaintext information of the classification result, ensuring data privacy security.

[0026] In the present invention, the zero - online - communication distance calculation based on Euclidean triples in step (2.1), taking two computing parties as an example, and the rest can be analogized, is as follows:

[0027] (2.1.1) Euclidean triples are pre-generated in the offline data sharing phase: The computing party P0 and the computing party P1 pre-generate multiple Euclidean triples Where: represents is additively secret shared with the computing party;

[0028] (2.1.2) Calculate the distance in the online classification calculation phase: The database sample and the query sample After being secretly shared using masks, that is, P0 holds and U i , and V, that is, P1 holds and and and Calculate the squared distance of the secret sharing value by the following method:

[0029] P0 calculates:

[0030] P1 calculates:

[0031] In this way, the secure Euclidean squared distance calculation completely depends on the randomly generated numbers in the offline phase and local operations, without additional communication in the online phase, and can significantly improve the calculation efficiency in the scenario of large-scale data sets.

[0032] In the present invention, the efficient nearest neighbor search based on the divide-and-conquer bubble protocol described in step (2.2) is specifically as follows:

[0033] (2.2.1) Partitioning: Assume that the length of the input distance list in the form of secret sharing and its corresponding label list is l, and the list is divided into two lists, each list having a size of

[0034] (2.2.2) Parallel comparison and exchange: The elements in the two lists are executed one by one for secure comparison and exchange operations to ensure that the smaller distance values are exchanged to the front; if the list length l is odd, an additional comparison and exchange is performed between the last element and the first element;

[0035] (2.2.3) Iterative recursion: Update the list length to Repeat the above process for the reduced list until the list length is reduced to 1;

[0036] (2.2.4) Obtain and delete the first element, and repeat steps (2.2.1) to (2.2.3): The first three steps bubble the minimum value to the first position of the list; take out this element and delete it, and then continue to repeat the above steps k - 1 times to obtain the remaining k - 1 nearest neighbors.

[0037] In the present invention, the KNN classification calculation is applicable to a distributed environment of multiple data owners, and specifically includes the following optimizations:

[0038] (1) During the data sharing phase, the data owners adopt a secret sharing protocol, so that the computing party cannot directly access the original data;

[0039] (2) During the KNN classification process, each computing party only calculates based on the encrypted data to ensure that the classification process meets the privacy protection requirements;

[0040] (3) After the classification calculation is completed, the computing party adopts a secret sharing recovery mechanism to securely return the final classification result to the user without leaking additional information.

[0041] In the present invention, the method is applicable to different data distribution modes, including:

[0042] (1) Horizontal data distribution, that is, different data owners hold data with the same feature set but different samples. After adopting the secret sharing mechanism for data sharing, KNN calculation is performed;

[0043] (2) Vertical data distribution, that is, different data owners hold data with the same samples but different features. After aligning the data through the private set intersection protocol, secret sharing is performed, and then KNN calculation is performed.

[0044] The beneficial effects of the present invention are as follows:

[0045] (1) Greatly reduce the online communication overhead: The present invention realizes a distance calculation scheme with zero online communication by pre - generating Euclidean triples, and adopts a divide - and - conquer bubble protocol in the nearest neighbor search, optimizing the original sequential bubble operation into parallel processing, effectively reducing the number of communication rounds. Therefore, in large - scale data sets or distributed scenarios, the impact of network bandwidth and communication delay can be significantly reduced.

[0046] (2) Support multi - party collaboration and ensure data privacy: Through the secret sharing protocol, each data owner can complete the joint calculation without exposing the original data, and the entire KNN classification process is always carried out in ciphertext. The computing party cannot obtain the data plaintext or intermediate results, avoiding the risk of sensitive information leakage.

[0047] (3) Adapt to multiple data distribution patterns: The present invention is compatible with horizontal and vertical data distributions and is applicable when multiple parties have different data types or scales; if different data owners hold the same samples but different features, the features can be aligned based on the private set intersection protocol before performing classification, further expanding the application scenarios.

[0048] (4) Simple and easy-to-expand calculation process: The Euclidean triples and divide-and-conquer bubble sorting adopted can be combined with existing secure comparison protocols and secret sharing frameworks, having good scalability and portability. Description of the Drawings

[0049] Figure 1 It is a schematic diagram of the architecture of the present invention. Detailed Embodiment

[0050] The present invention will be further described below with reference to the embodiments and the accompanying drawings.

[0051] Embodiment 1: Taking cross-hospital medical diagnosis as a specific application scenario, combined with Figure 1 the shown architecture, the efficient privacy-preserving KNN classification method of the present invention will be elaborated in detail.

[0052] 1. Scenario Setup

[0053] Suppose there are multiple hospitals (data owners) that respectively hold the medical records of different patients. The data of each hospital includes several features (such as age, blood pressure, blood sugar, test indicators, etc.) and corresponding disease labels. Due to concerns about patient privacy, the hospitals cannot directly share the original data. Now a certain patient (query user) hopes to use this cross-hospital collaborative KNN model to determine the type of disease he / she has, but is worried about the leakage of personal sensitive information during the transmission or calculation process.

[0054] 2. Offline Data Sharing Stage

[0055] Multiple hospitals perform secret sharing on their respective data: Each hospital randomly splits its patient data (including feature vectors and labels) into secret shares and sends them to two or more computing parties respectively. In this way, each computing party can only obtain encrypted information and cannot restore the complete features or diagnostic labels of any patient.

[0056] Pre-generate Euclidean triples: In the offline stage, the computing parties cooperate to generate random numbers and use these random numbers to construct preprocessed Euclidean triples. This process has nothing to do with the actual patient queries and can be completed during idle periods or at uniformly arranged time points to accelerate the subsequent distance calculation.

[0057] 3. Online KNN Classification Calculation Stage

[0058] (1) Secret sharing of query data: When a patient initiates a query, their feature vectors (such as age, blood pressure, test indicators, etc.) are also split through secret sharing and sent to the computing party, so that each party holds different secret shares.

[0059] (2) Euclidean distance calculation (zero online communication)

[0060] Since the patient data of each hospital and the features of the query user have been distributed to the computing party in ciphertext, the computing party can complete the distance calculation locally using the Euclidean triples generated in the offline phase.

[0061] Compared with traditional secure multiplication protocols, this scheme does not require additional online communication during the calculation process, which can greatly reduce network bandwidth and time overhead.

[0062] (3) Nearest neighbor search based on divide-and-conquer bubble protocol

[0063] The computing party performs a divide-and-conquer bubble operation on all distances: first, the distance list is divided into two parts, and then the elements in the two lists are compared and exchanged in parallel one by one.

[0064] Through iterative comparison and exchange, the smaller distance values will gradually "bubble" to the front of the list.

[0065] Finally, the computing party obtains the ciphertext indexes and ciphertext labels of the top k nearest neighbor samples of the query user.

[0066] (4) Generation of classification results: The computing party performs a secure voting or counting operation (such as majority voting) on the labels of these k neighbors to generate a ciphertext representation of the final classification result.

[0067] 4. Result reconstruction phase

[0068] The computing party returns the classification result in ciphertext form to the query user, and the user decrypts it with the secret sharing shares they hold to obtain the plaintext disease prediction category. Since each computing party only holds incomplete encrypted shares of the labels, they will not obtain any diagnostic results or other sensitive information of the patient during this process.

[0069] It should be noted that in the above scenario, if the hospital has sufficient computing power and bandwidth resources, the computing party can be directly served by the hospital.

Claims

1. An efficient privacy-preserving KNN classification method, characterized by The specific steps are as follows: (1) Offline data sharing stage The data owner uses a secret sharing protocol to encrypt its local data and distribute the encrypted data to two or more computing parties. After this stage is completed, the data owner's data will be stored in a secret form between the computing parties to ensure data security; (2) Online KNN classification calculation stage (2.1) After step (1) is completed, the Euclidean triple method is used to achieve secure Euclidean square distance calculation without the need for online communication; (2.2) A divide-and-conquer bubble protocol is used to reduce the number of communication rounds required for nearest neighbor selection from O(kn) to O(k logn), where n is the number of samples in the data set and k is the number of selected neighbors. Specifically, in each iteration, the sample distance is first divided into two lists of equal length. If it is an odd number, an additional comparison and exchange is performed. Then, a secure comparison and exchange operation based on secret sharing is performed on the elements in the two lists to ensure that the smaller element is exchanged to the previous list. Then, the above process is repeated for the previous list, and the nearest neighbor selection can be completed within O(k log n) rounds of communication. (3) Result reconstruction stage The computing party returns the classification results in encrypted form to the user, and the user recovers the final KNN classification results through the secret sharing recovery protocol; During this stage, the computing party will not obtain the plaintext information of the classification results, ensuring data privacy and security.

2. The method according to claim 1, characterized in that The zero online communication distance calculation based on the Euclidean triple described in step (2.1) is as follows, taking two computing parties as an example, and the rest are analogous: (2.1.1) Pre-generate Euclidean triples in the offline data sharing phase: Computing parties P0 and P1 pre-generate multiple Euclidean triples in: represent The addition secret is shared with the computing party; (2.1.2) Distance calculation in online classification calculation phase: Database samples and query sample After being shared with the masked secret, the computing party, P0, holds and U i , and V, that is, the computing party P1 holds and and and The squared distance is calculated by The secret sharing value of: P0 calculation: P1 calculation:

3. The method according to claim 1, characterized in that The efficient nearest neighbor search based on the divide-and-conquer bubble protocol in step (2.2) is specifically as follows: (2.2.1) Partitioning: Assume that the length of the distance list in the form of input secret sharing and its corresponding label list is l, divide the list into two lists of equal size, each of which is (2.2.2) Parallelized comparison and exchange: Perform safe comparison and exchange operations on the elements in the two lists one by one to ensure that the smaller distance value is exchanged to the previous list; If the length l is an odd number, the last element is additionally compared and exchanged with the first element; (2.2.3) Iterative recursion: Update the list length to Repeat the above process for the reduced list until the length of the list is reduced to 1; (2.2.4) Get and delete the first element, and repeat steps (2.2.1) to (2.2.3): The first three steps bubble the minimum value to the first position of the list; take out the element and delete it, then continue to repeat the above steps k-1 times to get the remaining k-1 nearest neighbors.

4. The method according to claim 1, characterized in that The KNN classification calculation is applicable to the distributed environment of multiple data owners, and specifically includes the following optimizations: (1) The data owner adopts a secret sharing protocol during the data sharing phase, so that the computing party cannot directly access the original data; (2) During the KNN classification process, each computing party performs calculations only based on encrypted data to ensure that the classification process meets privacy protection requirements; (3) After the classification calculation is completed, the computing party uses a secret sharing recovery mechanism to securely return the final classification result to the user without leaking additional information.

5. The method according to claim 1, characterized in that The method is applicable to different data distribution modes, including: (1) Horizontal data distribution, that is, different data owners hold data with the same feature set but different samples, and use a secret sharing mechanism to share data before performing KNN calculations; (2) Vertical data distribution, that is, different data owners hold data of the same sample but different features. After aligning the data through the privacy set intersection protocol, secret sharing is performed and then KNN calculation is performed.