Multi-party manifold learning method based on secret sharing

By designing an efficient and secure top-k algorithm and a full-source shortest path algorithm based on secret sharing technology, the security and efficiency problems of manifold learning in existing technologies are solved, achieving efficient and secure computation in multi-party manifold learning, reducing communication complexity and protecting data privacy.

CN120825288BActive Publication Date: 2025-11-18NANKAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511328188.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-17
Publication Date
2025-11-18
Estimated Expiration
2045-09-17

AI Technical Summary

Technical Problem

Existing technologies lack secure and effective k-nearest neighbor algorithms and all-source shortest path algorithms to support manifold learning. Furthermore, existing solutions struggle to balance performance and security, or suffer from high computational complexity, resulting in excessive communication volume and round-robin complexity, making them difficult to scale.

Method used

We design efficient and secure top-k algorithms and full-source shortest path algorithms based on secret sharing technology. By utilizing the idea of ​​opening after shuffling and matrix symmetry, combined with batch communication technology, we reduce the number of communication rounds and secret sharing operations. We also employ the low-round-count Floyd-Warshall algorithm and secure multidimensional scaling method to ensure that data privacy is not compromised.

Benefits of technology

It achieves efficient and secure k-nearest neighbor and all-source shortest path computation in multi-party manifold learning, significantly reduces communication complexity, ensures that the computing server cannot infer the data owner's data information, and improves computational efficiency and privacy protection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120825288B_ABST
    Figure CN120825288B_ABST
Patent Text Reader

Abstract

The application belongs to the field of privacy protection multi-party data cooperation research, and particularly relates to a multi-party manifold learning method based on secret sharing. The method comprises the following steps: multiple data owners first distribute the secret sharing shares of the respective data sets to three computing servers; each computing server interactively calculates the k-nearest neighbor distance matrix of all sample pairs on the secret shares of the data sets based on a novel secure top-k algorithm; then, each computing server interactively calculates the distance of the all-source shortest path on the secret shares of the k-nearest neighbor distance matrix by using a low-round Floyd-Warshall algorithm; finally, each computing server performs secure multi-dimensional scaling on the secret shared shortest path distance matrix to obtain the low-dimensional embedding of the data set, and any two servers send the secret shares of the low-dimensional embedding to the designated user to restore the plaintext locally. The application reduces the number of distance calculations and proposes an efficient secure top-k and all-source shortest path algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention pertains to the research of privacy-preserving multi-party data collaboration, specifically involving a multi-party manifold learning method based on secret sharing. Background Technology

[0002] In the era of big data, the exponential growth of data presents both significant challenges and opportunities for information extraction and decision-making. Data mining involves developing algorithms and technologies to discover meaningful patterns, relationships, and insights from raw datasets. Data held by a single institution or enterprise often suffers from limited volume and narrow information dimensions, making it difficult to meet the demands of complex applications for large samples and diverse information. Utilizing data from multiple institutions or enterprises can often improve the effectiveness of data mining. However, as more and more countries emphasize the protection of data privacy, regulations such as the General Data Protection Regulation (GDPR) have been introduced, and giants like Facebook and Uber have been fined heavily for data violations. To mitigate risks, many institutions have become extremely cautious when sharing data, exacerbating the "data silo" phenomenon. This prevents the effective integration and utilization of highly valuable multi-party data, limiting the further release of data value. Therefore, exploring secure and efficient data collaboration mechanisms while ensuring data privacy and security has become a top priority.

[0003] Secret sharing is one of the main technical approaches to solving the problem of secure multi-party data collaboration. It provides cryptographically provable security and has attracted widespread attention from academia and industry due to its practicality. Secret sharing schemes typically share input data in secret shares among multiple participants, who interactively compute the required functions. It can efficiently implement a wide range of functions, including but not limited to mixed-mode operations in the arithmetic and Boolean domains such as addition, subtraction, multiplication, division, equality, and comparison; complex scientific operations such as exponentiation, square roots, and trigonometric functions; and set operations such as shuffling, permutation, and sorting. Due to the rich function support and high performance of secret sharing, numerous studies have proposed multi-party manifold learning methods based on secret sharing, including principal component analysis, classification, clustering, regression, neural networks, and SQL-like analysis. However, there is currently a lack of support for manifold learning tasks, with isomap being a representative method. Isomap excels at capturing the inherent nonlinear structure of data, a crucial ability in fields where data exhibits intrinsic nonlinear relationships. Therefore, manifold learning is widely used in nonlinear dimensionality reduction, denoising, and data visualization in fields such as image and speech processing, bioinformatics, and financial analysis.

[0004] The main problem with current research is the lack of secure and efficient k-nearest neighbor (kNN) and all-source shortest path (USS) algorithms to support manifold learning. Current secure kNN computation schemes suffer from various issues. Some schemes sacrifice security for performance, weakening the security model; some calculate approximate results; and some have high preprocessing costs. Current secure shortest path computation schemes are based on Dijkstra's algorithm, modifying the original to guarantee randomness, but this increases the algorithm's complexity, leading to high communication and round complexity in secret-sharing implementations and poor scalability. Therefore, designing secure and efficient kNN and USS algorithms based on secret-sharing technology, and subsequently constructing a secret-sharing multi-party manifold learning method, is a pressing need. Summary of the Invention

[0005] The purpose of this invention is twofold. First, it proposes efficient and secure top-k algorithms and full-source shortest path algorithms based on secret sharing technology, providing an efficient solution for multi-party manifold learning methods that allow multiple data owners to jointly perform secret sharing computations using their respective datasets. Second, it protects the privacy of data owners' private datasets, ensuring that during computation, the computation server cannot deduce any information about the data owner's dataset from its perspective, including communication messages, memory accesses, and keys.

[0006] To design an efficient and secure top-k algorithm, this paper constructs a method based on the "shuffle and reveal" approach, building upon the plaintext fast selection algorithm. This involves appending a method to the end of each input vector element. In Bit values ​​are used to ensure that each element in the input vector is unique. Then, a secure shuffling protocol is used to randomly shuffle the input vector, ensuring that the comparison results in subsequent partitioning operations are independent of the original input vector and can be securely disclosed. This allows data exchange operations to be easily completed locally on the server. To design an efficient and secure full-source shortest path algorithm, this invention expands the two nested loops of the original Floyd-Warshall algorithm (i.e., the Floyd algorithm) and uses batch communication technology combined with data dependencies, thereby significantly reducing the number of communication rounds. Simultaneously, by leveraging the characteristics of Isomap, the number of secret sharing operations is effectively reduced.

[0007] To achieve the above objectives, the present invention adopts the following technical solution:

[0008] A multi-party manifold learning method based on secret sharing includes the following steps:

[0009] Step 1: The data owners encrypt and shard their respective datasets using a secret sharing protocol, and then upload the sharded dataset secret shares to three computing servers respectively.

[0010] Step 2: Each computing server calculates the secret share of the k-nearest neighbor distance matrix for all sample pairs on the secret share of the dataset, using Euclidean distance as the distance metric.

[0011] Step 3: Each computing server uses the low-round-number Floyd-Warshall algorithm on the secretly shared k-nearest neighbor distance matrix to calculate the distance of the shortest path across all sources. The calculation results are stored in the secretly shared shortest path distance matrix.

[0012] Step 4: Each computing server performs secure multidimensional scaling on the shortest path distance matrix of the secret sharing to obtain the low-dimensional embedding of the dataset. Any two computing servers send the secret share of the low-dimensional embedding to the designated user to restore it to plaintext locally.

[0013] In a further optimization of this technical solution, the number of data owners in step 1 is not limited, and the data owner datasets are secretly shared with the computing server using a three-party semi-honest replication secret sharing scheme.

[0014] In a further optimization of this technical solution, the secure k-nearest neighbor calculation in step 2 utilizes matrix symmetry to halve the number of secret sharing operations in the Euclidean distance calculation.

[0015] In a further optimization of this technical solution, the safe k-nearest neighbor calculation in step 2 uses the safe top-k algorithm to find the minimum value in each row of the distance matrix. Each distance corresponds to a given sample. Find the nearest neighbors and set the distances of the remaining non-neighbors to 1. .

[0016] This technical solution is further optimized so that the low-round-count Floyd-Warshall algorithm in step 3 does not reveal the number of edges. By expanding the two nested loops within the algorithm, the operations are divided into three batches according to data dependencies, and communication within each batch is parallelized, significantly reducing the number of communication rounds. Given the input distance matrix, express The distance value of the location, let These are pointers to the inner two loops, This is a specific value of the outermost loop pointer. To determine the size of the distance matrix, the three batches are defined as follows: 1) Batch 1 includes all Secret sharing operation at time, i.e. ,2) Batch 2 includes (1) all The secret sharing operation at time, that is, for all (2) All The secret sharing operation at time, that is, for all 3) Batch 3 includes all The secret sharing operation at time, that is, for all .

[0017] This technical solution is further optimized by step 2, which is based on the idea of ​​"shuffling and opening". An efficient and secure top-k algorithm is designed on the basis of the plaintext fast selection algorithm.

[0018] Further optimization of this technical solution is that the low-round-count Floyd-Warshall algorithm in step 3 reduces the number of secret sharing operations based on the characteristics of Isomap: 1) Operations to update the diagonal elements of the matrix can be omitted, including the diagonal operations in batch 1 and batch 3; 2) Operations in batch 2 can be omitted; 3) Operations in the lower triangular part of the matrix in batch 3 can be omitted.

[0019] In a further optimization of this technical solution, the safe top-k algorithm in step 2 appends a parameter to the end of each element of the input vector. In The bit values ​​ensure that the elements of the input vector are unique, and the input vector is randomly shuffled using a secure shuffling protocol. The comparison results in subsequent partitioning operations are independent of the input vector, allowing for secure opening and facilitating data exchange locally on the server.

[0020] This technical solution is further optimized so that the safe top-k algorithm in step 2 is equal to the position of the benchmark value. or The probability of the algorithm stopping after each partition doubles, thus reducing the average number of partitions.

[0021] Unlike current technologies, this invention proposes a multi-party manifold learning method based on secret sharing. This method does not limit the number of data owners; instead, data owners outsource their datasets to a computation server through secret sharing. The computation server interactively performs computations on the secret shares, ensuring high security, as the server cannot deduce any information related to the data owner's private data. This method halves the number of secret sharing operations in Euclidean distance and distance squared calculations using matrix symmetry. Based on the "shuffle and open" idea, an efficient and secure top-k algorithm is designed to support k-nearest neighbor computation, building upon the plaintext fast selection algorithm. A low-round-count Floyd-Warshall algorithm on secret sharing is designed for full-source shortest path computation. The algorithm does not reveal the number of edges; by expanding two nested loops within the algorithm and employing batch communication techniques based on data dependencies, the number of communication rounds is significantly reduced, and the number of secret sharing operations is reduced based on the characteristics of Isomap. In secure multidimensional scaling operations, a parallel Jacobi method is used to decompose matrix features, and a secure top-k algorithm is used to extract the largest eigenvalues ​​and their corresponding eigenvectors. Attached Figure Description

[0022] Figure 1 This is the overall flowchart of this solution;

[0023] Figure 2 This is the system architecture diagram of this solution;

[0024] Figure 3 This is a diagram showing the steps of the safe top-k algorithm;

[0025] Figure 4 This is a step diagram of the low-round-count Floyd-Warshall algorithm;

[0026] Figure 5 This is a flowchart of the secure multidimensional scaling algorithm steps. Detailed Implementation

[0027] The technical content, structural features, objectives, and effects of this technical solution are further explained below with reference to specific embodiments and accompanying drawings.

[0028] It should be emphasized that, in order to avoid unnecessary details from interfering with the understanding of this application, the accompanying drawings only show the structures and / or processing steps closely related to the solution of this application, and other details unrelated to this application are omitted.

[0029] This invention proposes a multi-party manifold learning method based on secret sharing, employing a setup of three computing servers and multiple users. For the overall process of the solution, please refer to [link / reference]. Figure 1 As shown. Specifically, we focus on a commonly used manifold learning algorithm, isomap. First, the data owners negotiate predefined parameters, including the number of neighbors. and dimensionality reduction Each dataset is secretly shared with three computing servers using a replication secret-sharing scheme. Each computing server calculates the Euclidean distance between all sample pairs based on its secret share of the complete dataset, using secret-sharing multiplication and square root operations. A secure top-k algorithm, constructed based on secret-sharing shuffling, secret-sharing comparison, and secret share opening operations, finds the distance between all sample pairs. The nearest neighbor distance matrix is ​​obtained by taking the distances of the nearest neighbors and setting the distances of the remaining non-neighbor samples as maxima. Then, the distance matrix of the all-source shortest path is calculated on the secret share of the nearest neighbor distance matrix using the low-round-number Floyd-Warshall algorithm. Finally, based on the secure parallel Jacobi algorithm and the secure top-k algorithm, the distance matrix of the all-source shortest path is securely multi-dimensionally scaled to obtain the secret share of the low-dimensional embedding of the dataset. Any two computing servers send the secret share of the low-dimensional embedding to the designated user, who opens it locally as a plaintext result.

[0030] For the architecture of this invention, please refer to [link / reference]. Figure 2 As shown, the data owner and computing service are deconstructed. There is no limit to the number of data owners, but there are three computing servers. The computing servers do not collude, but may attempt to infer the data owner's dataset from their own perspective. Each data owner sends a secret share of their dataset to the three computing servers. The servers interactively compute the low-dimensional embedding of the dataset on the secret share, and finally send the secret share to a designated user to open it as a plaintext low-dimensional embedding. This method can provide accurate computation results to the designated user. Simultaneously, the computing servers cannot infer any information about the data owner's dataset from their perspective regarding communication messages, memory accesses, and keys, thus providing a high level of privacy protection.

[0031] The following examples illustrate this in detail.

[0032] Step 1: Multiple data owners select three untrusted computing servers that do not collude, and use a replication secret sharing scheme to secretly share their respective datasets with the servers: arbitrary variables It will be divided into three random parts And is defined as a form of secret sharing. ,server Holding secret shares ,server Holding secret shares ,server Holding secret shares The secret share of any server cannot be restored to its original value.

[0033] Step 2: Three computing servers interactively perform secure k-nearest neighbor computation on a secretly shared input dataset. Let... The number of samples in the dataset. For sample dimensions, The number of neighbors is defined by the user. First, the distance matrix is ​​initialized with zeros. It interactively computes the Euclidean distance between all sample pairs based on secret-shared multiplication and square-root operations. Utilizing the symmetry of the distance matrix, the distance computation operation is reduced by half. It is known from the... The sample to the first The distance from the nth sample is equal to the distance from the nth sample. The sample to the first The distance between samples is calculated only. The distance in the upper triangle is calculated, and the matrix is ​​added to its transpose to determine the position. Values ​​assigned to positions .

[0034] Use a safe top-k algorithm to find the minimum value in each row of the matrix. Each distance corresponds to a given sample. Find the nearest neighbors and set the distances of the remaining non-neighbors to 1. Specifically, for each row of the matrix, the permutation generated by the safe top-k algorithm will minimize... Arrange the distances to the front Position, then back The distance to the location is set to By applying the corresponding permutation in reverse order to each row, the distances are rearranged back to their original positions. Finally, based on the minimum operation on secret sharing, the distance matrix is ​​symmetric by setting the value at a given position to the minimum of the values ​​at its transpose.

[0035] Based on the idea of ​​"shuffling and opening," a safe top-k algorithm is designed on the basis of the original fast selection algorithm. The algorithm requires that the elements of the input vector are unique; otherwise, it can be solved by appending to the end of each element. The values ​​in the middle are used to implement this, and the additional values ​​include Bits. The input vector is randomly shuffled at the beginning of the algorithm, so that the comparison results in subsequent partitions are independent of the input vector, allowing them to be safely opened. Furthermore, data exchange can be easily performed locally. For detailed steps of the safe top-k algorithm, see [link to algorithm]. Figure 3 .vector Record the permutations between the input and output vectors. The computation server first uses a secure shuffling protocol. An interactive implementation of random shuffling of a secret-shared vector is performed. After shuffling, the vector is repeatedly partitioned within a shrinking range. The partitioning operation is based on comparison operations on the secret share and opening operations on the secret-shared value. The partitioning operation randomly selects a pivot value, moves elements smaller than the pivot value to the left, and elements larger than the pivot value to the right, and then calculates the position of the pivot value. The element swaps are performed locally by the computation server based on the results of the open comparisons.

[0036] In the standard fast selection algorithm, when the position of the reference value is equal to... When the time comes, the algorithm stops because its goal is to find the first... Small or First The largest element, in this case, is the element that is found to be the first... Smaller elements. Although this stopping condition also applies to searching for preceding elements... The elements, that is, the smallest or the largest. However, this invention proposes a better stopping condition: when the position of the reference value is equal to... or When the partition stops, the condition is valid because when When it is the baseline value, for any and ,have Therefore, for any They all This means This is itself a valid top-k result. Because the randomly selected pivot position allows the algorithm to stop in two cases: when the pivot position equals... or The algorithm stops when the partitioning occurs, and the probability of the algorithm stopping after each partition doubles, so the average number of partitions is expected to decrease.

[0037] The average time complexity of the original fast selection algorithm is Therefore, the number of secret sharing comparisons in the secure top-k algorithm of this invention is... The corresponding communication complexity is Bits. Similar to quicksort, the average number of partitions is... Since the communication in the secret sharing comparison operation of each partition is performed in parallel, the round complexity of the secure top-k algorithm is O(n log n). The cost of a safe shuffle is... Bit and The cost of wheels is being concealed.

[0038] Step 3: Each computing server uses a round-efficient, secure Floyd-Warshall algorithm on the secretly shared k-nearest neighbor distance matrix to calculate the distance of the shortest path across all sources. The calculation results are stored in the secretly shared shortest path distance matrix. For detailed steps of the round-efficient, secure Floyd-Warshall algorithm proposed in this invention, please refer to... Figure 4 The derivation process is as follows. For those containing Input matrix of points The original Floyd-Warshall algorithm uses the following three nested loops.

[0039] for :

[0040] for :

[0041] for :

[0042] .

[0043] This structure led to The circuit depth is too high. Since the number of communication rounds in a secret-sharing protocol is proportional to the circuit depth, this overhead is impractical. This invention first expands the two loops within the Floyd-Warshall algorithm, analyzes its data dependencies, and batch processes the communication of independent secret-sharing operations, significantly reducing the number of communication rounds. For For each specific value, the operation will be divided into the following three batches:

[0044] Batch 1: Includes all Secret sharing operation at time, i.e. .

[0045] Batch 2: Includes (1) all The secret sharing operation at time, that is, for all (2) All The secret sharing operation at time, that is, for all .

[0046] Batch 3: Includes all The secret sharing operation at time, that is, for all .

[0047] Data dependencies can be observed: (1) Operations in batch 1 do not depend on the calculation results of other batches; (2) Operations in batch 2 depend only on batch 1. The calculation results; (3) The operations in batch 3 depend only on batch 2: and The calculation results. Therefore, as long as batch 1, batch 2, and batch 3 are executed sequentially, the operations in each batch can be executed in parallel. This reduces the circuit depth of the Floyd-Warshall algorithm to .

[0048] The Floyd-Warshall algorithm is further optimized by reducing the number of secret sharing operations based on the characteristics of the manifold learning algorithm Isomap.

[0049] Optimization 1: In Isomap, due to the non-negativity of the distance metric, all distances in the nearest neighbor matrix are non-negative; specifically, the distance from any sample to itself is zero. Therefore, the distance of the shortest path from any sample to itself is zero. Thus, (1) the operation of updating the diagonal elements of the matrix can be saved, including the diagonal operations in batches 1 and 3, since their values ​​are known to be zero; (2) batch 2 can be saved because... The calculation will not change the result.

[0050] Optimization 2: The output of the nearest neighbor algorithm is an undirected graph, meaning the nearest neighbor matrix is ​​symmetric. Clearly, the Floyd-Warshall algorithm preserves the matrix symmetry, allowing the elimination of operations in the lower triangle in batch 3. Operations are performed only in the upper triangle, and the results are copied to the corresponding transpose positions in the lower triangle.

[0051] Compared to the original Floyd-Warshall algorithm, the optimized algorithm reduces the number of secret sharing operations from Reduce to .

[0052] Step 4: Each computing server secretly shares the shortest path distance matrix. Perform secure multidimensional scaling to obtain low-dimensional embeddings of the dataset. For detailed steps of the secure multidimensional scaling algorithm, please refer to [link to algorithm documentation]. Figure 5 Multidimensional scaling first calculates the shortest path distance matrix. The square matrix of distances is then transformed into an inner product matrix by applying bicentering. ,

[0053] The squared distance matrix is ​​defined as , in express The square of the distance at location, It is a central matrix, defined as ,in It is the identity matrix. Multidimensional scaling of the matrix. Perform eigenvalue decomposition and calculate dimensional embedding, where The dimension after dimensionality reduction: in It is by A diagonal matrix composed of the largest eigenvalues It is a matrix composed of corresponding eigenvectors. It is calculated Dimensional embedding.

[0054] The secure version of multidimensional scaling performs a secret-shared distance squaring operation only in the upper triangle of the matrix and uses the secure parallel Jacobi protocol. This method calculates the eigenvalues ​​and eigenvectors of the matrix. Experiments show that this method is more efficient than QR-based methods in secret-sharing environments. In this example, the secure top-k algorithm will minimize... The values ​​are arranged in the first position of the vector. The goal of secure multidimensional scaling is to find [a specific location]. The largest eigenvalue can be transformed into finding the largest eigenvalue using a safe top-k algorithm. The smallest eigenvalue, at which point the vector finally... The feature values ​​at each position correspond to The largest eigenvalue.

[0055] Finally, any two computing servers will embed the low dimension. The secret share is sent to the designated user, and the result is restored locally to plaintext.

[0056] The invention was validated experimentally using the Wisconsin breast cancer dataset and synthetic datasets. The scheme was run on three LEGION REN9000K-34IAZ compute servers, each equipped with a 12th Gen Intel(R) Core(TM) i9-12900KF 3.20 GHz CPU and 32GB of RAM. Computation was performed on a single thread, and the servers were connected via a 5Gbps bandwidth network with a 0.2ms RTT. The semi-honest three-party replication secret-sharing scheme does not use a preprocessing paradigm, so the end-to-end runtime of the scheme was measured experimentally. In all experiments, the ring size of the secret share was set to 128 bits.

[0057] First, the time performance of the secure top-k and full-source shortest path algorithms was tested separately. In both tests, all input data was randomly synthesized. Data distribution did not affect computational or communication overhead because all algorithms were randomized. Regardless of whether the input was real or random, the computation server performed calculations on a uniformly random secret share.

[0058] The secure top-k algorithm of this invention is compared with the three-way radix sort algorithm, which is considered the most practical solution currently available. The test comparison results are shown in Table 1. The running time (in seconds) of the secure top-k algorithm and the three-way radix sort algorithm using synthetic datasets with different input sizes (from 200,000 to 1,000,000) shows that the secure top-k algorithm of this invention is 9 to 13.6 times faster than the three-way radix sort algorithm. The secure top-k algorithm can complete the top-k calculation for 200,000 secretly shared input data in less than 0.8 seconds, indicating that for smaller datasets, the top-k calculation is almost instantaneous.

[0059] Table 1. Runtime (seconds) of the secure top-k algorithm and three-way radix sort using a synthetic dataset.

[0060]

[0061] The low-round-count Floyd algorithm of this invention is compared with the state-of-the-art secure Dijkstra algorithm in calculating the shortest path across all sources. For a fair comparison, the secure Dijkstra algorithm is also instantiated using a three-party replication secret-sharing scheme, and its secure shuffling protocol is replaced with a state-of-the-art three-party protocol. The test comparison results are shown in Table 2, which displays the running time (in seconds) of the secure Floyd algorithm and the secure Dijkstra algorithm using synthetic datasets with different data volumes. The low-round-count Floyd algorithm of this invention significantly improves the efficiency of calculating the shortest path across all sources, being 684.9 to 1818.5 times faster than the secure Dijkstra algorithm. This performance improvement is mainly due to the reduction in communication overhead, especially the reduction in the number of communication rounds.

[0062] Table 2. Running time (seconds) of the safe Floyd algorithm and the safe Dijkstra algorithm using synthetic datasets.

[0063]

[0064] Finally, assuming the dataset is jointly owned by multiple data owners, the time performance of the secret-shared multi-party manifold learning algorithm Isomap was tested using the Wisconsin breast cancer dataset. With 20 nearest neighbors, the algorithm reduced the original 30-dimensional dataset to 5 dimensions. The baseline scheme compared to this invention uses the secure Dijkstra algorithm to compute the full-source shortest path. The test comparison results are shown in Table 3, which shows the running time (in seconds) of this scheme and the baseline scheme for manifold learning on the Wisconsin breast cancer dataset with different data volumes. The running speed of this invention is approximately 11.1 to 28.2 times faster than the baseline scheme. The performance improvement becomes more significant as the number of samples increases.

[0065] Table 3. Runtime (seconds) of this protocol and baseline protocol on the Wisconsin breast cancer dataset.

[0066]

[0067] The advantages of this invention are as follows: 1. Data owners outsource their datasets to the computing server in a secret sharing manner, without limiting the number of data owners; 2. It has high security, as the computing server cannot deduce any information related to the data owner's private data; 3. It halves the number of secret sharing operations in Euclidean distance and distance squared calculations by utilizing matrix symmetry; 4. It provides accurate and secure top-k and k-nearest neighbor calculations, with the secure top-k algorithm being at least 9 times faster than three-way radix sort; 5. It employs a low-round-count Floyd-Warshall algorithm for secret sharing, using batch communication technology based on data dependencies to achieve a low number of communication rounds, and utilizing the characteristics of Isomap to reduce the number of secret sharing operations, without revealing the number of edges.

[0068] It should be noted that the relational terms such as "first" and "second" mentioned in this document are used only to distinguish different entities or operations and do not necessarily require or imply an actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," and any other variations thereof are intended to indicate non-exclusive inclusion, meaning that the process, method, article, or terminal device may include other elements not explicitly listed, or elements inherent to the process, method, article, or terminal device itself, in addition to the listed elements. Unless further specified, elements defined by "comprising..." or "including..." do not exclude the presence of other elements in a process containing these elements.

[0069] In addition, the terms “more than,” “greater than,” and “less than” in this article should be understood as excluding the number itself; while the terms “above,” “below,” and “within” should be understood as including the number itself.

[0070] Although the embodiments described above have been detailed, those skilled in the art, after understanding the basic inventive concept of the present invention, can make various changes and improvements to these embodiments. Therefore, the above content is only a few specific embodiments of the present invention and should not be construed as limiting the scope of patent protection of the present invention. Any modifications to the equivalent structure or equivalent process based on the content of this specification and drawings, or direct or indirect applications of the principles of the present invention in other related technical fields, should be covered within the scope of patent protection of the present invention.

Claims

1. A multi-party manifold learning method based on secret sharing, characterized in that, Includes the following steps: Step 1: The data owners encrypt and shard their respective datasets using a secret sharing protocol, and then upload the sharded dataset secret shares to three computing servers respectively. Step 2: Each computing server calculates the secret share of the k-nearest neighbor distance matrix for all sample pairs on the secret share of the dataset, using Euclidean distance as the distance metric. Step 2 calculates the k-nearest neighbor distance matrix using a safe top-k algorithm to find the smallest k-nearest neighbor in each row of the distance matrix. Each distance corresponds to a given sample. Find the nearest neighbors and set the distances of the remaining non-neighbors to 1. ; Step 2 of the safe top-k algorithm appends to the end of each element of the input vector. In The bit values ​​ensure that the elements of the input vector are unique, and the input vector is randomly shuffled using a secure shuffling protocol. The comparison results in subsequent partitioning operations are independent of the input vector. Step 3: Each computing server uses a low-round-number Floyd algorithm on the secretly shared k-nearest neighbor distance matrix to calculate the distance of the shortest path across all sources. The calculation results are stored in the secretly shared shortest path distance matrix. In step 3, the low-round-count Floyd algorithm does not reveal the number of edges. It expands the two nested loops within the algorithm and divides the operations into three batches based on data dependencies. Communication within each batch is parallelized, allowing... Given the input distance matrix, express The distance value of the location, let These are pointers to the inner two loops, This is a specific value of the outermost loop pointer. The size of the distance matrix is ​​defined for the three batches as follows: 1) Batch 1: Includes all Secret sharing operation at time, i.e. ,2) Batch 2: including (1) all The secret sharing operation at time, that is, for all (2) All The secret sharing operation at time, that is, for all 3) Batch 3: Includes all The secret sharing operation at time, that is, for all ; Step 4: Each computing server performs secure multidimensional scaling on the shortest path distance matrix of the secret sharing to obtain the low-dimensional embedding of the dataset. Any two computing servers send the secret share of the low-dimensional embedding to the designated user to restore it to plaintext locally.

2. The multi-way manifold learning method based on secret sharing as described in claim 1, characterized in that, In step 1, the number of data owners is unlimited, and the data owner datasets are secretly shared with each computing server using a three-party semi-honest replication secret sharing scheme.

3. The multi-party manifold learning method based on secret sharing as described in claim 2, characterized in that, The specific details of the three-party semi-honest replication secret sharing scheme are as follows: arbitrary variables It will be divided into three random parts And is defined as a form of secret sharing. ,server Holding secret shares ,server Holding secret shares ,server Holding secret shares .

4. The multi-party manifold learning method based on secret sharing as described in claim 1, characterized in that, The k-nearest neighbor distance matrix calculation in step 2 utilizes matrix symmetry to halve the number of secret sharing operations in the Euclidean distance calculation.

5. The multi-party manifold learning method based on secret sharing as described in claim 1, characterized in that, The safe top-k algorithm in step 2 is equal to the position of the benchmark value. or The probability of the algorithm stopping after each partition doubles.

6. The multi-party manifold learning method based on secret sharing as described in claim 1, characterized in that, In step 3, the low-round-number Floyd algorithm reduces the number of secret sharing operations based on the characteristics of equidistant mapping: 1) Operations updating the diagonal elements of the matrix are omitted, including the diagonal operations in batch 1 and batch 3; 2) Operations in batch 2 are omitted; 3) Operations in the lower triangular part of the matrix in batch 3 are omitted.

7. The multi-way manifold learning method based on secret sharing as described in claim 1, characterized in that, The secure multidimensional scaling steps in step 4 are as follows: Multidimensional scaling first calculates the shortest path distance matrix. The square matrix is ​​then transformed into an inner product matrix by applying bicentering. , The squared distance matrix is ​​defined as ,in express The square of the distance at location, It is a central matrix, defined as ,in It is an identity matrix, a multidimensional scaling matrix. Perform eigenvalue decomposition and calculate dimensional embedding, where The dimension after dimensionality reduction: in It is by A diagonal matrix composed of the largest eigenvalues It is a matrix composed of corresponding eigenvectors. It is calculated Dimensional embedding.

Citation Information

Patent Citations

  • An on-line monitoring method of packed column flooding state based on integrated manifold learning

    CN109214268A

  • Unsupervised nonlinear adaptive manifold learning method

    CN109961088A