Feature selection method based on secure multi-party computing

Through secure multi-party computing technology, the data set is outsourced to the server for feature selection, which solves the problems of high computing complexity and privacy leakage, realizes efficient feature selection under privacy protection, and improves the accuracy and reliability of data analysis.

CN120354444APending Publication Date: 2025-07-22NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510310053.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the prior art, the feature selection process has high computational complexity and cannot effectively protect data privacy, which poses a risk of privacy leakage.

Method used

Through secure multi-party computing technology, the data set is split and outsourced to the server for calculation. The communication channels between servers and secure multi-party computing technology are used to calculate feature scores and construct a single-hot matrix for feature selection to avoid directly exposing the original data.

Benefits of technology

It realizes effective feature selection without leaking data privacy, reduces computing complexity, improves the accuracy and reliability of data analysis, and complies with the requirements of laws and regulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354444A_ABST
    Figure CN120354444A_ABST
Patent Text Reader

Abstract

The invention provides a feature selection method based on secure multi-party computing, which is technically characterized in that the execution of a computing task is completed by a server by using a secure multi-party computing technology and adopting a data outsourcing form, so that the resource consumption of a user side is reduced, the feature score of a data set feature is obtained by using various secure multi-party computing technologies, and the user experience is improved. And further, a comparison vector is obtained through cyclic comparison, and a one-hot matrix is obtained by adopting a lookup table technology, so that a selected feature matrix is obtained, the calculation complexity of a feature selection process is reduced, feature selection under privacy protection is realized, and the problem that the feature selection process has relatively high complexity is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of computer vision and machine learning, and particularly to a feature selection method based on secure multi-party computation. Background Art

[0002] The purpose of feature selection is to select the most useful part of features from the original dataset, remove redundant or irrelevant features, so as to improve the performance and generalization ability of the model. Feature selection can not only simplify the model, reduce the computational cost, but also help improve the accuracy and interpretability of the model. Among them, the filtering method is one of the three commonly used feature selection methods. It scores features through statistical indicators such as analysis of variance, correlation coefficient, chi-square test, etc., and then selects features according to the scores.

[0003] Secure Multi-Party Computation (SMPC) is a cryptographic technology aimed at protecting the privacy of data of all participating parties, ensuring that each party does not need to disclose its own input data during the process of joint computation. It allows multiple participants to perform joint computations securely without trusting each other, and obtain the computation results without exposing their private data. Secure multi-party computation based on secret sharing mainly includes the following key steps: (1) Data splitting: The participating party encrypts or splits its own private data into several parts and distributes these shares to other participating parties. (2) Protocol execution: The participating parties perform computations according to a pre-determined protocol, which usually uses technologies such as homomorphic encryption, secret sharing, and garbled circuits to ensure that each participating party cannot know the specific input of other parties during the computation process. (3) Result synthesis: After the computation is completed, each party combines the computation results it holds to obtain the final computation result, while ensuring that no party can recover the original data of other parties from it. However, in the prior art, it is necessary to centralize all data for processing, which has a significant risk of privacy leakage. Summary of the Invention

[0004] Each exemplary embodiment of this application provides a feature selection method based on secure multi-party computation, so as to at least have the technical effect of preventing personal sensitive data from being leaked during the feature selection process.

[0005] According to one aspect of this application, each exemplary embodiment of this application provides a feature selection method based on secure multi-party computation, and the method includes the following steps:

[0006] S1, establish a communication channel between servers. The user locally splits their respective datasets into a feature set and a label set, performs one-hot encoding on the label set. After determining the number of outsourcing servers, the user establishes a communication channel with the servers, and distributes the feature set and the label set to the servers in the form of secret sharing;

[0007] S2. Calculate the shares of the feature scores held by each server. Based on the share technology and combined with the secure multi-party computing technology, the server calculates the feature scores of each feature.

[0008] S3. Based on the feature scores, the server performs a cyclic comparison of the feature scores of each feature to obtain a comparison vector. Among them, the comparison vector is represented as a one-hot vector, including t 1s, corresponding to the positions of the features with the top t highest feature scores.

[0009] S4. The server constructs a one-hot matrix based on the comparison vector and performs a multiplication calculation between the original features and the one-hot matrix to obtain the selected feature matrix.

[0010] The present application has the following beneficial effects:

[0011] Through the secure multi-party computing technology, each participating party can jointly perform feature selection without exposing the original data. This not only prevents the leakage of personal sensitive data during the feature selection process but also ensures that all parties can carry out effective data collaboration on the premise of compliance and security, thereby improving the accuracy and reliability of the analysis and meeting the requirements of laws and regulations at the same time. Privacy-preserving feature selection is of great significance for making full use of cross-institutional data resources while protecting data privacy, and for improving the efficiency and credibility of data analysis.

[0012] The present invention mainly aims at the data privacy security and calculation efficiency problems in feature selection, and designs a feature selection method based on secure multi-party computing. The user distributes the data set to the server in the form of secret sharing in advance, which can transfer the calculation overhead to the server with greater computing resources. A network communication channel is established between the servers, and the subsequent calculation tasks are borne. The servers interact and calculate, and obtain the feature score vector according to multiple secure multi-party computing primitives. According to the feature score vector, the server constructs a one-hot matrix, thereby completing the selection of the original feature matrix through the one-hot matrix. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0014] Figure 1 It is a schematic diagram of the overall process involved in the method of the present invention;

[0015] Figure 2 It is a schematic diagram of the calculation process of the feature scores involved in the method of the present invention;

[0016] Figure 3 It is a schematic diagram of the one-hot matrix for calculating the feature scores involved in the method of the present invention. Specific embodiments

[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the preferred embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments.

[0018] All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.

[0019] The present invention is proposed in view of the fact that the existing feature selection methods cannot provide effective privacy protection for data sets and the computational complexity of the feature selection process is high. The present invention outsources user data in the form of secret sharing, and a server with more computing resources completes all computing tasks. The server adopts secure multi-party computing technology to calculate the feature values of each feature. As Figures 1 to 2 shown, the present invention provides a feature selection method based on secure multi-party computing, and the method includes the following steps:

[0020] S1. Establish a communication channel between servers. The user locally splits their respective data sets into a feature set and a label set, and performs one-hot encoding on the label set. After determining the number of outsourcing servers, the user establishes a communication channel with the servers, and distributes the feature set and the label set to the servers in the form of secret sharing;

[0021] S2. Calculate the shares of the feature scores held by each server. The server calculates the feature scores of each feature based on the share technology in combination with the secure multi-party computing technology;

[0022] S3. The server circularly compares the feature scores of each feature based on the feature scores to obtain a comparison vector. Among them, the comparison vector is represented as a one-hot vector, including t 1s, corresponding to the positions of the features with the top t highest feature scores;

[0023] S4. The server constructs a one-hot matrix based on the comparison vector, and multiplies the original features by the one-hot matrix to obtain the selected feature matrix.

[0024] It should be noted that the current feature selection scheme cannot achieve efficient privacy protection for data, and there are problems such as a large number of loop iterations and high complexity in the process of constructing a one-hot matrix based on feature scores to achieve feature selection. Therefore, a feature selection method based on secure multi-party computation is proposed. In the prior art, there already exists a way to construct a one-hot matrix based on feature scores. This way consists of two nested loops. The outer loop obtains the index value of the highest value in the current feature vector, and the inner loop generates a one-hot vector through equality judgment, with only one position being 1, and its index value corresponding to the index value of the highest value in the feature vector during this loop. However, such a nested loop method has a large computational complexity. To achieve feature selection under privacy protection and at the same time reduce the computational complexity, the present invention utilizes various secure multi-party computation primitives. First, the data is securely outsourced, and the server calculates the feature scores under secret sharing. After obtaining the feature score vector, a comparison vector is obtained through loop comparison, and then a one-hot matrix is further constructed to achieve feature selection.

[0025] In one embodiment, before the user starts the computing task in S1, the local data is securely sent to the server in the form of secret sharing, where the data of all users needs to ensure the same data dimension, data features, and data labels.

[0026] In one embodiment, the secure multi-party computation technology in S2 includes oblivious transfer, garbled circuits, and lookup tables for auxiliary computation.

[0027] It should be noted that in the feature selection process, secure multi-party computation technology is used to ensure the privacy and security of data. Its characteristic is that the computational process of feature selection is all based on the secure multi-party computation primitive - secret sharing to achieve, and on this basis, other secure multi-party computation primitives such as oblivious transfer, garbled circuits, lookup tables, etc. are used for auxiliary computation.

[0028] In one embodiment, before the computing task in S3 starts, the user enters the offline phase.

[0029] It should be noted that in the present invention, the server has relatively large computing resources. After the user starts the computing task, the user can go offline and does not need to participate in the execution of the computing task.

[0030] In one embodiment, S1 specifically includes:

[0031] S11, select two high-performance computing service devices S1 and S2, and build a network communication channel. A number of data holders, called users C1,…,C n , each establish a network communication channel with S1 and S2. Users C1,…,C n each hold data feature sets D1,…,D n , data label sets L1,…,Ln , where each data set has the same number of samples M, the same number of features K, the number of categories L, and the label set is in the form of one-hot encoding, forming a matrix of size M×L;

[0032] S12, users C1,…,C n D1,…,D n Split into and All users will Send to S1, Sent to S2; the user uses secret sharing technology to send the label set L1,…,L n Split into and and will Send to S1, Send to S2;

[0033] S13, users jointly set the threshold value of each data feature, which is f1,…,f K ,After completing the above operations, the user enters the offline stage.

[0034] refer to Figure 3 , Figure 3 Represents the process of constructing a one-hot matrix based on feature scores. After obtaining the feature scores, it is necessary to select the corresponding features from the original feature matrix based on the highest t feature scores. The process of looping to determine the highest value in the current vector and then selecting the corresponding feature has a large loop cost. The user first generates a comparison vector through loop comparison, and obtains a one-hot vector through the lookup table technology. The elements in the one-hot vector corresponding to the highest α eigenvalues are 1, otherwise they are 0. Therefore, there is no need to explicitly derive the size relationship between the first α eigenvalues. It is only necessary to obtain the index of the first α feature scores, which corresponds to the position of the one-hot vector elements.

[0035] Accordingly, in one embodiment, the process from S2 to S4 includes the step of constructing a one-hot matrix according to the feature scores, which specifically includes:

[0036] S1, the servers hold the share of feature score vectors and jointly discuss and select a threshold value t, which represents the number of features that need to be selected from the original feature set.

[0037] S2, the server jointly compares the numerical values of each element in the feature score vector one by one, that is, for the i-th element in the feature score vector, compare it with the i+1-th element. If the i-th element is greater than or equal to the i+1-th element, the result is set to 1, otherwise it is 0. After comparing all elements one by one, a comparison vector u is obtained, which represents the size comparison results of all elements with other elements.

[0038] S3. Use the generated comparison vector u as the input to the lookup table to obtain a one-hot vector v with a length of K. If the i-th position of the one-hot vector v is 1, it means that the score of the corresponding feature is among the top α of all feature scores; otherwise, its feature score is lower and not among the top α.

[0039] S4. According to the one-hot vector v, expand it column-wise to obtain a matrix W. The expansion rule is v[i] = W[i, i], that is, the i-th element of the one-hot vector is the element at the i-th row and i-th column of the matrix W. The other positions on W are 0.

[0040] It should be noted that this approach can effectively reduce the number of loops for constructing the one-hot matrix from feature scores. That is, through the above technical solution, the feature selection process only requires a small number of loop operations.

[0041] In one embodiment, the processes of S2 to S4 include the step of constructing a one-hot matrix according to feature scores, which specifically includes:

[0042] S1. When the server receives the data shares, first perform a comparison operation on the dataset shares and the threshold value to generate a one-hot matrix C of size M×K T ; divide each feature into two subsets. In one subset, the feature values of all entries for this feature are greater than or equal to the threshold value, and in the other subset, the feature values of all entries for the feature are less than the threshold value. Perform column-wise addition on the one-hot matrix C T to obtain a vector a, which gives the number of times each feature in all samples is greater than or equal to the threshold value. Correspondingly, subtract a from the total number of samples M to obtain the number of times each feature in all samples is less than the threshold value, denoted by b.

[0043] S2. Calculate the number of categories in each subset; the server uses the one-hot matrix C T and the category matrix L to generate a one-hot matrix A of size K×L, where each element represents the number of samples in which the feature value of the k-th feature is greater than or equal to the threshold value and belongs to the l-th category; subtract each element in A from the a value to obtain a matrix A', which represents the number of samples in which the feature value of the k-th feature is less than the threshold value and belongs to the l-th category; correspondingly, by performing a negation operation on the one-hot matrix C T to obtain the calculation result for each feature in all samples that is less than the threshold value, and then perform the same calculation with the category matrix L to generate a one-hot matrix B, where each element represents the number of samples in which the feature value of the k-th feature is less than the threshold value and belongs to the l-th category. Subtract each element in B from the b value to obtain a matrix B', which represents the number of samples in which the feature value of the k-th feature is less than the threshold value and does not belong to the l-th category.

[0044] S3. For matrices A and A', calculate the ratio p of each element to a, and then calculate the information entropy p ln p of each feature; for matrices B and B', calculate the ratio p' of each element to b, and then calculate the information entropy p' ln p' of each feature.

[0045] S4. Calculate the proportion of each subset in the entire data set to obtain the weight w, and calculate the product of the weight and the information entropy, so as to obtain the mutual information value of each feature and form a feature score vector.

[0046] S5. Sort according to the feature values in the feature score vector, select the highest α feature values, and return their corresponding index values; and select these features from the original data feature set to form a new data feature set.

[0047] It should be noted that in this embodiment, user data is outsourced to achieve the availability of large-scale data calculation, and the feature selection process uses secure multi-party computing technology to ensure data privacy and security, so as to solve the data privacy and security problems and computational complexity problems in the feature selection process of the data set, and provide an effective balance for improving the availability and privacy of the data set in the calculation process.

[0048] The above is only the preferred embodiment of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

[0049] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.

[0050] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A feature selection method based on secure multi-party computation, characterized in that, The method comprises the following steps: S1, build a communication channel between servers. Users split their data sets into feature sets and label sets locally, and perform one-hot encoding on the label sets. After determining the number of outsourced servers, users build a communication channel with the servers and distribute the feature sets and label sets to the servers in the form of secret sharing. S2, calculates the share of feature scores held by each server. The server calculates the feature score of each feature based on the share technology and the secure multi-party computing technology; S3, the server cyclically compares the feature scores of each feature based on the feature scores to obtain a comparison vector, wherein the comparison vector is represented as a one-hot vector containing t 1s corresponding to the positions of the features with the highest feature scores before t. S4, the server constructs a one-hot matrix based on the comparison vector, and multiplies the original feature by the one-hot matrix to obtain a selected feature matrix.

2. The feature selection method based on secure multi-party computation according to claim 1, wherein In S1, before the computing task starts, the user securely sends local data to the server in the form of secret sharing, where all users' data must have the same data dimension, data features and data labels.

3. The feature selection method based on secure multi-party computation according to claim 1, wherein The secure multi-party computing technology of S2 includes oblivious transmission, obfuscated circuits, and lookup tables to assist in computing.

4. A feature selection method based on secure multi-party computation according to claim 1, characterized in that The user enters the offline phase before the S3 computing task starts.

5. A feature selection method based on secure multi-party computation according to claim 1, characterized in that, The S1 specifically includes: S11. Select two high-performance computing service devices S1 and S2 to establish a network communication channel. A number of data holders are called users C1, …, C n , each of which establishes a network communication channel with S1 and S2. Users C1, …, C n each hold data feature sets D1, …, D n , data label sets L1, …, L n , where each data set has the same number of samples M, the same number of features M, the number of classes is L, and the label set is in the form of one-hot encoding, forming a matrix of size M×L; S12, users C1, …, C n Split D1, …, D n into and All users will send to S1, and send to S2; users split tag sets L1, …, L n into and and will send to S1, and send to S2; S13. The user jointly sets the threshold values for each data feature, which are f1, …, f K . After completing the above operations, the user enters the offline phase.

6. The feature selection method based on secure multi-party computation according to claim 1, wherein The process from S2 to S4 includes the step of constructing a one-hot matrix according to the feature scores, which specifically includes: The servers hold a share of the feature score vector and jointly negotiate to select a threshold value t, which represents the number of features that need to be selected from the original feature set. The server jointly compares the numerical values of each element in the feature score vector one by one, that is, for the i-th element in the feature score vector, compare it with the i+1-th element. If the i-th element is greater than or equal to the i+1-th element, the result is set to 1, otherwise it is 0. After comparing all elements one by one, a comparison vector u is obtained, which represents the size comparison results of all elements with other elements. The generated comparison vector u is used as the input of the lookup table to obtain a one-hot vector v with a length of K. If the i-th position of the one-hot vector v is 1, it means that the score of the corresponding feature is in the first α positions of all feature scores. Otherwise, its feature score is low and is not before the first α positions. According to the one-hot vector v, expand it column-wise to get the matrix W. The expansion rule is v[i]=W[i,i], that is, the i-th element of the one-hot vector is the element on the i-th row and i-th column of the matrix W. The other positions on W are 0.

7. A feature selection method based on secure multi-party computation according to claim 1, characterized in that, The process from S2 to S4 specifically includes: After the server receives the data shares, it first performs a comparison operation on the data set shares and the threshold value to generate a one-hot matrix C of size M×K T ; The features are divided into two subsets. In one subset, the feature values of this feature for all entries are greater than or equal to the threshold value, and in the other subset, the feature values of the features for all entries are less than the threshold value. For the one-hot matrix C T Perform column-wise addition to obtain the vector a, and thus obtain the number of each feature greater than or equal to the threshold value in all samples. Correspondingly, subtract a from the total number of samples M to obtain the number of each feature less than the threshold value in all samples, which is represented by b; Calculate the number of categories in each subset; the server generates a one-hot matrix A of size K×L through the one-hot matrix C T and the category matrix L, where each element represents the number of samples in which the eigenvalue of the k-th feature is greater than or equal to the threshold value and belongs to the l-th category; subtract each element in A from the a value to obtain the matrix A’, which represents the number of samples in which the eigenvalue of the k-th feature is less than the threshold value and belongs to the l-th category; correspondingly, by taking the negation operation on the one-hot matrix C T to obtain the calculation result of each feature less than the threshold value for all samples, and then perform the same calculation with the category matrix L to generate a one-hot matrix B, where each element represents the number of samples in which the eigenvalue of the k-th feature is less than the threshold value and belongs to the l-th category, subtract each element in B from the b value to obtain the matrix B’, which represents the number of samples in which the eigenvalue of the k-th feature is less than the threshold value and does not belong to the l-th category; For matrices A and A', the ratio p of each element to a is calculated respectively, and the information entropy pln p of each feature is calculated accordingly; for matrices B and B', the ratio p' of each element to b is calculated respectively, and the information entropy p'ln p' of each feature is calculated accordingly; Calculate the proportion of each subset to the entire data set to obtain the weight w, calculate the product of the weight and information entropy, and thus obtain the mutual information value of each feature to form a feature score vector; Sort according to the eigenvalues in the feature score vector, select the highest α eigenvalues, and return their corresponding index values; and select these features from the original data feature set to form a new data feature set.