FSS-based privacy protection mean shift clustering method and device

Through the FSS-based privacy protection mean offset clustering method, data is divided and uploaded to an edge server that does not collude with each other. The offline stage is used to generate keys and random values, and iterative calculations and mode hiding are performed in combination with security protocols, which solves the problems of large amount of privacy protection calculations and data leakage in the existing technology, and achieves efficient privacy protection and data security.

CN120034327APending Publication Date: 2025-05-23JOINT WARFARE COLLEGE NAT DEFENSE UNIV OF THE CHINESE PEOPLES LIBERATION ARMY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510205309.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When existing clustering methods deal with privacy protection in multi-data owner scenarios, there are challenges such as large computing volume and relying on homomorphic encryption or auxiliary information, making it difficult to effectively protect the privacy of sensitive data.

Method used

The FSS-based privacy protection mean offset clustering method is adopted. By dividing the data into two copies and uploading them to edge servers that do not collude with each other, the FSS key and random values are generated in the offline stage, combining secure sampling, secure selection and secure negative index protocols, iterative calculations and mode hiding are performed, and clustering results in secret sharing are output.

Benefits of technology

It greatly reduces communication overhead and computing complexity in multi-data owner scenarios, while effectively protecting data privacy, preventing access mode leakage, and meeting privacy needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034327A_ABST
    Figure CN120034327A_ABST
Patent Text Reader

Abstract

The invention relates to a privacy protection mean shift clustering method and device based on FSS. The method comprises the steps that a data owner divides original data into two parts and uploads the two parts to two edge servers which are not communicated with each other, a trusted third party generates an FSS secret key and a random value in an offline stage and distributes the FSS secret key and the random value to the edge servers, and in an online stage, the edge servers randomly select seed points from shared data through a secure sampling protocol and send the seed points to the edge servers. Iteratively calculating a mean shift vector of each seed point, calculating a Gaussian kernel weight by using a security negative index protocol, deleting a repetitive mode through a security selection protocol, inserting a virtual mode to hide a real clustering number, distributing a nearest mode label for each data point, and outputting a clustering result in a secret sharing form; and the data user merges the clustering results to obtain a final clustering label. By adopting the method, a secure data transmission mode can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing and privacy protection, and particularly to a privacy-preserving mean shift clustering method and apparatus based on FSS. Background Art

[0002] Clustering is an important unsupervised learning method and is widely used in many fields such as image processing, network traffic analysis, and finance. However, in practical applications, clustering usually involves sensitive data of multiple data owners, such as medical data, financial data, etc., and data privacy protection has become a key issue. Traditional clustering methods such as the K-means algorithm have limitations, such as requiring the number of clusters to be predefined and being only effective for convex clusters; the density-based mean shift algorithm can handle clusters of any shape, but has a large computational cost. Most existing privacy-preserving clustering schemes rely on homomorphic encryption or secret sharing techniques, but still face challenges such as high computational overhead or reliance on auxiliary information. Summary of the Invention

[0003] Based on this, it is necessary to provide a privacy-preserving mean shift clustering method and apparatus based on FSS for the above technical problems.

[0004] A privacy-preserving mean shift clustering method based on FSS, the method includes:

[0005] The data owner divides the original data into two parts and uploads them to two non-colluding edge servers respectively;

[0006] The trusted third party generates FSS keys and random values in the offline phase and distributes them to the edge servers;

[0007] In the online phase, the edge server randomly selects seed points from the shared data through a secure sampling protocol, iteratively calculates the mean shift vector of each seed point, calculates the Gaussian kernel weight using a secure negative exponent protocol, deletes duplicate patterns through a secure selection protocol, inserts virtual patterns to hide the true number of clusters, assigns the nearest pattern label to each data point, and outputs the clustering result in the form of secret sharing; the secure negative exponent protocol is a negative exponent protocol based on piecewise polynomials and the least squares method;

[0008] The data user merges the clustering results to obtain the final clustering labels.

[0009] In one embodiment, the negative exponent protocol based on piecewise polynomials and the least squares method is:

[0010] Divide the domain of the secure negative exponent protocol into a preset number of intervals; among them, dense interval division is adopted in the area with large gradient changes, and sparse interval division is adopted in the area with gentle gradient changes;

[0011] For each of the intervals, a quadratic polynomial is used to approximate the negative exponential function e -x ; where the quadratic polynomial is expressed as:

[0012] nExp(x)=α i,2 x 2 +α i,1 x+α i,0

[0013] α i,2 , α i,1 and α i,0 It is obtained by querying the polynomial coefficient table calculated in advance by the least square method;

[0014] Use the distributed comparison function to determine the interval to which the input data x belongs, and return the secret sharing result of the corresponding polynomial;

[0015] When the input data x is hidden by a random mask, mask-hidden polynomial coefficients are generated.

[0016] In one embodiment, the security selection protocol includes:

[0017] Determine the selection function as:

[0018]

[0019] Among them, a represents the virtual mode, d represents the candidate mode;

[0020] Using r in1 To hide x, use r in2 Hide the virtual mode a and the candidate mode d, and construct the offset function as:

[0021]

[0022] Comparison using distributed comparison functions and The output is selected based on the comparison result.

[0023] In one embodiment, the safe sampling protocol includes:

[0024] Offline stage:

[0025] A trusted third party generates a random permutation π i and a random vector a i , and calculate b 0 +b 1 =π 0 (π 1 (a 0 )+a 1 ), and a random vector r of length m; where i∈{0,1};

[0026] The trusted third party will randomly arrange π i , random vector a i and random vector r are sent to edge server s 1 and edge servers 2 ;

[0027] Online stage:

[0028] Edge Servers 1 Will self-sustaining data [x] 0 With a random vector a 0 Add and get the mask data [x′] 0 , and the first mask data [x′] 0 Send to edge servers 2 ; where [x'] 0 =[x] 0 +a 0

[0029] Edge Servers 2 According to the random arrangement π 1 and mask data [x′] 0 For self-sustaining data[x] 1 Perform a shuffle operation to obtain the first obfuscated data x′, and compare the first obfuscated data x′ with the random vector a 1 Add, get the second obfuscated data x″, set [z] 1 =-b 1 , and sends the second obfuscated data x″ to the edge server s 1 ; where x′=π 1 ([x′] 0 +[x] 1 ), x'=x'+a 1 ;

[0030] Edge Servers 1 Using the random sequence π 0 Shuffle the second obfuscated data x″ to obtain [z] 0 =π 0 (x″)-b 0 ;

[0031] Edge Servers 1 and edge servers 2 According to the random vector r from the processed data [z] 1 and [z] 0 Select the corresponding sample points and output [y] i =[z r ] i .

[0032] In one of the embodiments, it further includes: an offline phase:

[0033] A trusted third party generates a random vector r for each seed point f,j , and the random vector r f,j Secret sharing is r f,j,0 and r f,j,1 To edge servers 1 and edge servers 2 , and generate the key k for the secure negative exponential protocol f,j,0 and k f,j,1 , and the random vector r f,j,0 and r f,j,1 , key k f,j,0 and k f,j,1 Send to edge servers 1 and edge servers 2 ;

[0034] Online stage:

[0035] Edge Servers 1 and edge servers 2 m seed points are randomly selected from the shared data through a secure sampling protocol; for each selected seed point, the edge server s 1 and edge servers 2 Calculate the square distances to other points respectively;

[0036] According to the square distance, the Gaussian kernel is calculated using a safe negative exponential function to obtain a Gaussian kernel value. The mean shift vector is calculated based on the Gaussian kernel value and the seed point position is updated. When the change of the seed point is lower than a specified threshold, the seed point converges and the converged seed point is inserted into the pattern list.

[0037] In one of the embodiments, it further includes: an offline phase:

[0038] A trusted third party generates a random vector r for masking in1 and a random vector r in2 , and generate a key for the secure selection protocol, the random vector r in1 , random vector r in2 And the key is sent to the edge server s 1 and edge servers 2 ;

[0039] Online stage:

[0040] Edge Servers 1 and edge servers 2 Create an empty pattern list;

[0041] For each candidate pattern, if the pattern list is empty, add the candidate pattern to the pattern list;

[0042] If the pattern list is not empty, the squared Euclidean distance between the current candidate pattern and the candidate patterns in the pattern list is calculated to obtain a distance matrix, the distance matrix is ​​summed to determine the minimum distance; the minimum distance and the current candidate pattern are masked, and the safe selection protocol is used to compare the masked minimum distance with the threshold. If the masked minimum distance exceeds the threshold, the current candidate pattern is inserted into the pattern list, otherwise, a virtual pattern is inserted. After all candidate patterns are processed, the pruned pattern list is output;

[0043] Edge Servers 1 and edge servers 2 Create an empty list of cluster labels;

[0044] For each data point, calculate the squared Euclidean distance between it and each pattern in the pruned pattern list, and calculate the distance matrix. Sum the distance matrix and find the index corresponding to the minimum distance. Assign the label of the nearest pattern corresponding to the index to the current data point and output the clustering result in the form of secret sharing.

[0045] The above-mentioned FSS-based privacy-preserving mean shift clustering method and device, first, the data owner divides the original data into two parts and uploads them to two non-colluding edge servers respectively, to ensure that a single share cannot leak any original data information; secondly, FSS keys and random numbers are pre-generated in the offline stage, and complex cryptographic operations are transferred to the non-real-time link, which greatly reduces the communication overhead in the online stage; further, in the pattern deletion stage, the actual number of clusters is hidden by inserting virtual patterns, and the shuffling and masking technology of the secure sampling protocol is combined to prevent access pattern leakage, thereby meeting the privacy needs of multiple data owners in the scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 It is a flowchart of a privacy-preserving mean-shift clustering method based on FSS in one embodiment;

[0047] Figure 2 It is a flowchart of a privacy-preserving mean-shift clustering device based on FSS in one embodiment;

[0048] Figure 3 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0049] To make the objectives, technical solutions and advantages of this application more clear and understandable, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely used to explain this application and are not intended to limit this application.

[0050] This invention is applied to the following application environment: The privacy-preserving mean-shift clustering framework includes four types of entities as follows:

[0051] Data Owners:

[0052] Own sensitive data x and cannot directly share the original data due to privacy restrictions.

[0053] Split the data into two secret sharing shares x 0 and x 1 , and upload them to two non-colluding edge servers s 1 and edge server s 2 .

[0054] Edge Servers, s 1 and s 2 ):

[0055] Receive the sharing shares x 0 and x 1 from the data owner, and collaboratively execute security protocols (such as secure negative exponent calculation, mean-shift iteration, etc.).

[0056] During the processing, the original data or intermediate results cannot be inferred to ensure privacy.

[0057] Trusted Dealer, T:

[0058] Generate and distribute random values and function secret sharing (FSS) keys for the edge servers in the offline phase to support secure calculations (such as comparison, multiplication, etc.) in the online phase.

[0059] Data User, V:

[0060] The authorizing party (which may be the data owner or other third party), by merging the result shares y 0 and y 1 returned by the edge servers, reconstructs the complete clustering result for subsequent analysis tasks.

[0061] 1. Data splitting and uploading:

[0062] The data owner splits the original data x into x 0 and x 1 , satisfying x 0 + x1 =x.

[0063] x 0 and x 1 Upload to edge servers separately 1 and edge servers 2 .

[0064] 2. Offline stage preparation:

[0065] The trusted third party generates a random mask and FSS key and distributes them to the edge server s 1 and edge servers 2 .

[0066] The security protocol in the online phase does not require additional key negotiation, thus reducing real-time communication overhead.

[0067] 3. Online secure computing:

[0068] Edge Servers 1 and edge servers 2 Collaboratively perform privacy-preserving computations (e.g., secure negative exponentiation, secure selection protocols, etc.) using pre-generated keys and masks.

[0069] 4. Result reconstruction and use:

[0070] Data users from edge servers 1 and edge servers 2 Get the secret sharing share y of the clustering result 0 and 1 , after merging, we get the plaintext result y=y 0 +y 1 .

[0071] Users can only access the final cluster labels and cannot trace back to the original data or intermediate calculation processes.

[0072] 5. Privacy and security mechanisms

[0073] Data segmentation and secret sharing:

[0074] Single shared datax 0 and x 1 It does not contain complete information and requires collaboration between both parties to restore it, preventing single point leakage.

[0075] FSS and masking technology:

[0076] The calculation process is encrypted by FSS key and random mask to ensure that the intermediate results (such as kernel weights and mode positions) are always confidential.

[0077] Non-collusion assumption:

[0078] Edge Servers1 and the edge server s 2 Do not collude with each other, and any server cannot independently infer valid information.

[0079] In one of the embodiments, as Figure 1 shown, a privacy-preserving mean shift clustering method based on FSS is provided, including the following steps:

[0080] Step 102, the data owner divides the original data into two parts and uploads them to two non-colluding edge servers respectively.

[0081] Step 104, the trusted third party generates FSS keys and random values in the offline phase and distributes them to the edge servers.

[0082] Step 106, in the online phase, the edge server randomly selects seed points from the shared data through a secure sampling protocol, iteratively calculates the mean shift vector of each seed point, calculates the Gaussian kernel weight using the secure negative exponential protocol, deletes duplicate patterns through the secure selection protocol, inserts virtual patterns to hide the true number of clusters, assigns the nearest pattern label to each data point, and outputs the clustering result in the form of secret sharing.

[0083] The secure negative exponential protocol is a negative exponential protocol based on piecewise polynomials and the least squares method.

[0084] Step 108, the data user merges the clustering results to obtain the final clustering labels.

[0085] For the above privacy-preserving mean shift clustering method based on FSS, first, the data owner divides the original data into two parts and uploads them to two non-colluding edge servers respectively to ensure that a single share cannot disclose any original data information. Secondly, by pre-generating FSS keys and random numbers in the offline phase, complex cryptographic operations are transferred to the non-real-time link, greatly reducing the communication overhead in the online phase. Further, in the pattern deletion phase, the true number of clusters is hidden by inserting virtual patterns, and the shuffling and masking techniques of the secure sampling protocol are combined to prevent the access pattern from being leaked, meeting the privacy requirements in the multi-data owner scenario.

[0086] In one of the embodiments, the negative exponential protocol based on piecewise polynomials and the least squares method is as follows:

[0087] Divide the domain of the secure negative exponential protocol into a preset number of intervals; among them, dense interval division is adopted in the area with large gradient changes, and sparse interval division is adopted in the area with gentle gradient changes. For each of the intervals, a quadratic polynomial is used to approximate the negative exponential function e -x ; where the quadratic polynomial is expressed as:

[0088] nExp(x) = α i,2x 2 +α i,1 x+α i,0

[0089] α i,2 , α i,1 and α i,0 It is obtained by querying the polynomial coefficient table calculated in advance by the least squares method; using the distributed comparison function to determine the interval to which the input data x belongs, and returning the secret sharing result of the corresponding polynomial; when the input data x is hidden by a random mask, the mask-hidden polynomial coefficients are generated.

[0090] In this embodiment, a balance between efficiency and security is achieved in negative exponential calculations through dynamic piecewise polynomial approximation and FSS privacy protection mechanism: first, the input interval is adaptively divided according to the gradient change, and the quadratic polynomial coefficients are pre-calculated in combination with the least squares method, which not only ensures accuracy but also avoids iterative calculation overhead; secondly, the distributed comparison function (DCF) is used to securely determine the interval to which the input belongs, and mask coefficients are generated to prevent the leakage of original data and polynomial information, thereby realizing non-interactive privacy computing.

[0091] In one embodiment, the security selection protocol includes:

[0092] Determine the selection function as:

[0093]

[0094] Among them, a represents the virtual mode, d represents the candidate mode; r in1 To hide x, use r in2 Hide the virtual mode a and the candidate mode d, and construct the offset function as:

[0095]

[0096] Comparison using distributed comparison functions and The output is selected based on the comparison result.

[0097] In one embodiment, the safe sampling protocol includes:

[0098] Offline stage:

[0099] A trusted third party generates a random permutation π i and a random vector a i , and calculate b 0 +b 1 =π 0 (π 1 (a 0 )+a 1), and a random vector r of length m; where i∈{0,1}; the trusted third party will randomly arrange π i , random vector a i and random vector r are sent to edge server s 1 and edge servers 2 ;

[0100] Online stage:

[0101] Edge Servers 1 Will self-sustaining data [x] 0 With a random vector a 0 Add and get the mask data [x′] 0 , and the first mask data [x′] 0 Send to edge servers 2 ; where [x'] 0 =[x] 0 +a 0 .

[0102] Edge Servers 2 According to the random arrangement π 1 and mask data [x′] 0 For self-sustaining data[x] 1 Perform a shuffle operation to obtain the first obfuscated data x′, and compare the first obfuscated data x′ with the random vector a 1 Add, get the second obfuscated data x″, set [z] 1 =-b 1 , and sends the second obfuscated data x″ to the edge server s 1 ; where x′=π 1 ([x′] 0 +[x] 1 ), x'=x'+a 1 .

[0103] Edge Servers 1 Using the random sequence π 0 Shuffle the second obfuscated data x″ to obtain [z] 0 =π 0 (x″)-b 0 .

[0104] Edge Servers 1 and edge servers 2 According to the random vector r from the processed data [z] 1 and [z] 0 Select the corresponding sample points and output [y] i =[z r ] i .

[0105] In one embodiment, the steps of implementing the edge server to randomly select seed points from shared data through a secure sampling protocol, iteratively calculate the mean drift vector of each seed point, and calculate the Gaussian kernel weight using a secure negative exponential protocol include:

[0106] Offline stage:

[0107] A trusted third party generates a random vector r for each seed point f,j , and the random vector r f,j Secret sharing is r f,j,0 and r f,j,1 To edge servers 1 and edge servers 2 , and generate the key k for the secure negative exponential protocol f,j,0 and k f,j,1 , and the random vector r f,j,0 and r f,j,1 , key k f,j,0 and k f,j,1 Send to edge servers 1 and edge servers 2 ;

[0108] Online stage:

[0109] Edge Servers 1 and edge servers 2 m seed points are randomly selected from the shared data through a secure sampling protocol; for each selected seed point, the edge server s 1 and edge servers 2 The square distances between the points and other points are calculated respectively; according to the square distances, the Gaussian kernel is calculated using the safe negative exponential function to obtain the Gaussian kernel value; the mean shift vector is calculated according to the Gaussian kernel value and the seed point position is updated; when the change of the seed point is lower than the specified threshold, the seed point is determined to be converged, and the converged seed point is inserted into the pattern list.

[0110] In one embodiment, the steps of deleting repeated patterns through a secure selection protocol, inserting virtual patterns to hide the actual number of clusters, assigning a nearest pattern label to each data point, and outputting a clustering result in a secret sharing form include:

[0111] Offline stage:

[0112] A trusted third party generates a random vector r for masking in1 and a random vector r in2 , and generate a key for the secure selection protocol, the random vector r in1 , random vector r in2 And the key is sent to the edge server s 1and edge servers 2 ;

[0113] Online stage:

[0114] Edge Servers 1 and edge servers 2 Create an empty pattern list;

[0115] For each candidate pattern, if the pattern list is empty, add the candidate pattern to the pattern list;

[0116] If the pattern list is not empty, the squared Euclidean distance between the current candidate pattern and the candidate patterns in the pattern list is calculated to obtain a distance matrix, the distance matrix is ​​summed to determine the minimum distance; the minimum distance and the current candidate pattern are masked, and the safe selection protocol is used to compare the masked minimum distance with the threshold. If the masked minimum distance exceeds the threshold, the current candidate pattern is inserted into the pattern list, otherwise, a virtual pattern is inserted. After all candidate patterns are processed, the pruned pattern list is output; the edge server s 1 and edge servers 2 Create an empty cluster label list; for each data point, calculate the squared Euclidean distance between it and the patterns in the pruned pattern list, and calculate the distance matrix, sum the distance matrix, and find the index corresponding to the minimum distance, assign the label of the nearest pattern corresponding to the index to the current data point, and output the clustering result in secret sharing form.

[0117] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0118] In one embodiment, if Figure 2 As shown, a privacy-preserving mean shift clustering device based on FSS is proposed, and the device includes:

[0119] The data uploading module 202 is used for the data owner to split the original data into two parts and upload them to two edge servers that are not in collusion with each other;

[0120] An offline module 204, used for a trusted third party to generate FSS keys and random values ​​in an offline phase and distribute them to the edge server;

[0121] The online module 206 is used for the online stage, where the edge server randomly selects seed points from the shared data through a secure sampling protocol, iteratively calculates the mean drift vector of each seed point, calculates the Gaussian kernel weight using a secure negative exponential protocol, deletes repeated patterns through a secure selection protocol, inserts virtual patterns to hide the actual number of clusters, assigns a nearest pattern label to each data point, and outputs a clustering result in a secret sharing form; the secure negative exponential protocol is a negative exponential protocol based on piecewise polynomials and least squares method;

[0122] The clustering module 208 is used for the data user to merge the clustering results to obtain the final clustering label.

[0123] In one embodiment, the negative exponential protocol based on piecewise polynomial and least squares method is:

[0124] The domain of the secure negative exponential protocol is divided into a preset number of intervals; dense interval division is used in areas with large gradient changes, and sparse interval division is used in areas with gentle gradient changes;

[0125] For each of the intervals, a quadratic polynomial is used to approximate the negative exponential function e -x ; where the quadratic polynomial is expressed as:

[0126] nExp(x)=α i,2 x 2 +α i,1 x+α i,0

[0127] α i,2 , α i,1 and α i,0 It is obtained by querying the polynomial coefficient table calculated in advance by the least square method;

[0128] Use the distributed comparison function to determine the interval to which the input data x belongs, and return the secret sharing result of the corresponding polynomial;

[0129] When the input data x is hidden by a random mask, mask-hidden polynomial coefficients are generated.

[0130] In one embodiment, the security selection protocol includes:

[0131] Determine the selection function as:

[0132]

[0133] Among them, a represents the virtual mode, d represents the candidate mode;

[0134] Using r in1 To hide x, use r in2 Hide the virtual mode a and the candidate mode d, and construct the offset function as:

[0135]

[0136] Comparison using distributed comparison functions and The output is selected based on the comparison result.

[0137] In one embodiment, the safe sampling protocol includes:

[0138] Offline stage:

[0139] A trusted third party generates a random permutation π i and a random vector a i , and calculate b 0 +b 1 =π 0 (π 1 (a 0 )+a 1 ), and a random vector r of length m; where i∈{0,1};

[0140] The trusted third party will randomly arrange π i , random vector a i and random vector r are sent to edge server s 1 and edge servers 2 ;

[0141] Online stage:

[0142] Edge Servers 1 Will self-sustaining data [x] 0 With a random vector a 0 Add and get the mask data [x′] 0 , and the first mask data [x′] 0 Send to edge servers 2 ; where [x'] 0 =[x] 0 +a 0 ;

[0143] Edge Servers 2 According to the random arrangement π 1 and mask data [x′] 0 For self-sustaining data[x] 1 Perform a shuffle operation to obtain the first obfuscated data x′, and compare the first obfuscated data x′ with the random vector a 1 Add, get the second obfuscated data x″, set [z] 1=-b 1 , and sends the second obfuscated data x″ to the edge server s 1 ; where x′=π 1 ([x′] 0 +[x] 1 ), x'=x'+a 1 ;

[0144] Edge Servers 1 Using the random sequence π 0 Shuffle the second obfuscated data x″ to obtain [z] 0 =π 0 (x″)-b 0 ;

[0145] Edge Servers 1 and edge servers 2 According to the random vector r from the processed data [z] 1 and [z] 0 Select the corresponding sample points and output [y] i =[z r ] i .

[0146] In one embodiment, the online module 206 is also used in the offline phase:

[0147] A trusted third party generates a random vector r for each seed point f,j , and the random vector r f,j Secret sharing is r f,j,0 and r f,j,1 To edge servers 1 and edge servers 2 , and generate the key k for the secure negative exponential protocol f,j,0 and k f,j,1 , and the random vector r f,j,0 and r f,j,1 , key k f,j,0 and k f,j,1 Send to edge servers 1 and edge servers 2 ;

[0148] Online stage:

[0149] Edge Servers 1 and edge servers 2 m seed points are randomly selected from the shared data through a secure sampling protocol; for each selected seed point, the edge server s 1 and edge servers 2 Calculate the square distances to other points respectively;

[0150] According to the square distance, the Gaussian kernel is calculated using a safe negative exponential function to obtain a Gaussian kernel value. The mean shift vector is calculated based on the Gaussian kernel value and the seed point position is updated. When the change of the seed point is lower than a specified threshold, the seed point converges and the converged seed point is inserted into the pattern list.

[0151] In one embodiment, the online module 206 is also used in the offline phase:

[0152] A trusted third party generates a random vector r for masking in1 and a random vector r in2 , and generate a key for the secure selection protocol, the random vector r in1 , random vector r in2 And the key is sent to the edge server s 1 and edge servers 2 ;

[0153] Online stage:

[0154] Edge Servers 1 and edge servers 2 Create an empty pattern list;

[0155] For each candidate pattern, if the pattern list is empty, add the candidate pattern to the pattern list;

[0156] If the pattern list is not empty, the squared Euclidean distance between the current candidate pattern and the candidate patterns in the pattern list is calculated to obtain a distance matrix, the distance matrix is ​​summed to determine the minimum distance; the minimum distance and the current candidate pattern are masked, and the safe selection protocol is used to compare the masked minimum distance with the threshold. If the masked minimum distance exceeds the threshold, the current candidate pattern is inserted into the pattern list, otherwise, a virtual pattern is inserted. After all candidate patterns are processed, the pruned pattern list is output;

[0157] Edge Servers 1 and edge servers 2 Create an empty list of cluster labels;

[0158] For each data point, calculate the squared Euclidean distance between it and each pattern in the pruned pattern list, and calculate the distance matrix. Sum the distance matrix and find the index corresponding to the minimum distance. Assign the label of the nearest pattern corresponding to the index to the current data point and output the clustering result in the form of secret sharing.

[0159] For the specific definition of the FSS-based privacy-preserving mean shift clustering device, please refer to the definition of the FSS-based privacy-preserving mean shift clustering method above, which will not be repeated here. Each module in the above-mentioned FSS-based privacy-preserving mean shift clustering device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0160] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a multi-modal personalized health management program generation method based on a large model is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0161] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0162] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method in the above embodiment when executing the computer program.

[0163] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method in the above embodiment are implemented.

[0164] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0165] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0166] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A privacy-preserving mean-shift clustering method based on FSS, characterized in that: The method comprises: The data owner splits the original data into two parts and uploads them to two independent edge servers respectively; The trusted third party generates the FSS key and random value in the offline phase and distributes them to the edge server; In the online stage, the edge server randomly selects seed points from the shared data through a secure sampling protocol, iteratively calculates the mean drift vector of each seed point, calculates the Gaussian kernel weight using a secure negative exponential protocol, deletes repeated patterns through a secure selection protocol, inserts virtual patterns to hide the actual number of clusters, assigns the nearest pattern label to each data point, and outputs the clustering result in the form of secret sharing; the secure negative exponential protocol is a negative exponential protocol based on piecewise polynomials and least squares method; Data users merge the clustering results to obtain the final cluster labels.

2. The method according to claim 1, characterized in that: The negative exponential protocol based on piecewise polynomials and least squares method is: The domain of the secure negative exponential protocol is divided into a preset number of intervals; dense interval division is used in areas with large gradient changes, and sparse interval division is used in areas with gentle gradient changes; For each of the intervals, a quadratic polynomial is used to approximate the negative exponential function e -x ; where the quadratic polynomial is expressed as: nExp(x)=α i,2 x 2 +a i,1 x+a i,0 α i,2 , α i,1 and α i,0 It is obtained by querying the polynomial coefficient table calculated in advance by the least square method; Use the distributed comparison function to determine the interval to which the input data x belongs, and return the secret sharing result of the corresponding polynomial; When the input data x is hidden by a random mask, mask-hidden polynomial coefficients are generated.

3. The method according to claim 2, characterized in that The security selection protocol includes: Determine the selection function as: Among them, a represents the virtual mode, d represents the candidate mode; Using r in1 To hide x, use r in2 Hide the virtual mode a and the candidate mode d, and construct the offset function as: Comparison using distributed comparison functions and The output is selected based on the comparison result.

4. The method according to claim 2, characterized in that: Safe sampling protocols include: Offline stage: A trusted third party generates a random permutation π i and a random vector a i , and calculate b0+b1=π0(π1(a0)+a1), and a random vector r of length m; where i∈{0,1}; The trusted third party will randomly arrange π i , random vector a i and the random vector r are sent to edge server s1 and edge server s2 respectively; Online stage: The edge server s1 adds the self-supporting data [x]0 to the random vector a0 to obtain the mask data [x′]0, and sends the first mask data [x′]0 to the edge server s2; wherein, [x′]0=[x]0+a0; The edge server s2 performs a shuffle operation on the self-sustaining data [x]1 according to the random arrangement π1 and the mask data [x′]0 to obtain the first obfuscated data x′, adds the first obfuscated data x′ to the random vector a1 to obtain the second obfuscated data x″, sets [z]1=-b1, and sends the second obfuscated data x″ to the edge server s1; wherein x′=π1([x′]0+[x]1), x″=x'+a1; The edge server s1 shuffles the second obfuscated data x″ using the random sequence π0 to obtain [z]0=π0(x″)-b0; Edge servers s1 and s2 select corresponding sample points from the processed data [z]1 and [z]0 according to the random vector r and output [y] i =[z r ] i .

5. The method according to any one of claims 1 to 4, characterized in that: The edge server randomly selects seed points from the shared data through a secure sampling protocol, iteratively calculates the mean drift vector of each seed point, and uses a secure negative exponential protocol to calculate the Gaussian kernel weights, including: Offline stage: A trusted third party generates a random vector r for each seed point f,j , and the random vector r f,j Secret sharing is r f,j,0 and r f,j,1 to edge servers s1 and s2, and generate the key k for the secure negative exponential protocol f,j,0 and k f,j,1 , and the random vector r f,j,0 and r f,j,1 , key k f,j,0 and k f,j,1 Send to edge server s1 and edge server s2; Online stage: Edge servers s1 and s2 randomly select m seed points from the shared data through a secure sampling protocol; for each selected seed point, edge servers s1 and s2 calculate the square distance between it and other points respectively; According to the square distance, the Gaussian kernel is calculated using a safe negative exponential function to obtain a Gaussian kernel value. The mean shift vector is calculated based on the Gaussian kernel value and the seed point position is updated. When the change of the seed point is lower than a specified threshold, the seed point converges and the converged seed point is inserted into the pattern list.

6. The method according to any one of claims 1 to 4, characterized in that: The secure selection protocol is used to remove duplicate patterns, insert virtual patterns to hide the true number of clusters, assign the nearest pattern label to each data point, and output the clustering results in the form of secret sharing, including: Offline stage: A trusted third party generates a random vector r for masking in1 and a random vector r in2 , and generate a key for the secure selection protocol, the random vector r in1 , random vector r in2 And the key is sent to edge server s1 and edge server s2; Online stage: Edge servers s1 and s2 create empty pattern lists; For each candidate pattern, if the pattern list is empty, add the candidate pattern to the pattern list; If the pattern list is not empty, the squared Euclidean distance between the current candidate pattern and the candidate patterns in the pattern list is calculated to obtain a distance matrix, the distance matrix is ​​summed to determine the minimum distance; the minimum distance and the current candidate pattern are masked, and the safe selection protocol is used to compare the masked minimum distance with the threshold. If the masked minimum distance exceeds the threshold, the current candidate pattern is inserted into the pattern list, otherwise, a virtual pattern is inserted. After all candidate patterns are processed, the pruned pattern list is output; Edge servers s1 and s2 create empty clustering label lists; For each data point, calculate the squared Euclidean distance between it and each pattern in the pruned pattern list, and calculate the distance matrix. Sum the distance matrix and find the index corresponding to the minimum distance. Assign the label of the nearest pattern corresponding to the index to the current data point and output the clustering result in the form of secret sharing.

7. A privacy-preserving mean-shift clustering device based on FSS, characterized in that: The device comprises: The data upload module is used by the data owner to split the original data into two parts and upload them to two independent edge servers respectively; An offline module, used for a trusted third party to generate FSS keys and random values ​​in an offline phase and distribute them to the edge server; An online module is used in the online stage. The edge server randomly selects seed points from the shared data through a secure sampling protocol, iteratively calculates the mean drift vector of each seed point, calculates the Gaussian kernel weight using a secure negative exponential protocol, deletes repeated patterns through a secure selection protocol, inserts virtual patterns to hide the actual number of clusters, assigns the nearest pattern label to each data point, and outputs the clustering result in the form of secret sharing; the secure negative exponential protocol is a negative exponential protocol based on piecewise polynomials and least squares method; The clustering module is used by data users to merge clustering results to obtain the final clustering label.

8. The device according to claim 7, characterized in that: The negative exponential protocol based on piecewise polynomials and least squares method is: The domain of the secure negative exponential protocol is divided into a preset number of intervals; dense interval division is used in areas with large gradient changes, and sparse interval division is used in areas with gentle gradient changes; For each of the intervals, a quadratic polynomial is used to approximate the negative exponential function e -x ; where the quadratic polynomial is expressed as: nExp(x)=α i,2 x 2 +a i,1 x+a i,0 α i,2 , α i,1 and α i,0 It is obtained by querying the polynomial coefficient table calculated in advance by the least square method; Use the distributed comparison function to determine the interval to which the input data x belongs, and return the secret sharing result of the corresponding polynomial; When the input data x is hidden by a random mask, mask-hidden polynomial coefficients are generated.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.